News & Research
The latest AI research and news with real-world stakes. Each item is sourced, dated, summarized in plain English and tagged by impact area, and checked against its source before it appears.
Kind
Impact area
5672 items
- ResearchWhite Rose Research Online (University of Leeds, The University of Sheffield, University of York)2026-06-09QCP
Structural Causal World Models for Safety Assurance of AI-based Autonomy · Zou, Jie, STEFANAKOS, IOANNIS, Shahbeigi Roudposhti, Sepeedeh et al.
This paper introduces Structural Causal World Models (SCWMs), a formal framework grounded in structural causal models to support safety assurance of AI-based autonomous systems. SCWMs provide interpretable, machine-verifiable representations that unify symbolic constraints, probabilistic uncertainty, and causal dependencies, enabling traceable hazard analysis, safety requirement derivation, and run-time monitoring. The methodology is domain-agnostic and is illustrated through autonomous driving examples, aiming to close the semantic gap in defining safety requirements for complex AI systems. The work contributes to reducing uncertainty in safety assurance by providing a basis for causal hazard and risk analysis and verification of probabilistic guarantees.
- ResearcharXiv2026-06-08QP
CCBENCH: Assessing LLM Cultural Competence via Implicitly Signaled Norms using Health Queries · Vasudha Varadarajan, Akhila Yerukola, Mona T. Diab et al.
CCBENCH is a new benchmarking framework that evaluates how well large language models (LLMs) adapt to users' implicitly signaled cultural values, rather than relying on static demographic labels. Using a health-query case study (CCBENCH-Health), the authors constructed 60 theoretically grounded personas spanning six cultures and assessed five leading LLMs across 3,120 unique interactions. Results show that even the best-performing models give culturally appropriate responses only 20–30% of the time, with chain-of-thought prompting yielding only modest 3–5% average gains. A persistent asymmetry is found where models perform better when personas deviate from cultural norms than when they follow them—most starkly in the Afghan context (average 8.8%)—suggesting models default to built-in biases rather than adapting to cultural cues.
- ResearcharXiv2026-06-08EQ
SafeGEO: Understanding Generative Engine Optimization Risks in Recommendation Agents · Qianfeng Wen, Yifan Simon Liu, Xin Liu et al.
SafeGEO examines how Generative Engine Optimization (GEO) — the practice of rewriting web content to boost visibility in AI-generated outputs — can be weaponized in recommendation agents to make flawed products appear better supported than they are. The authors build an evaluation suite with 22 GEO attack variants across 600 recommendation cases and find that such attacks increase the rate at which flawed products enter recommendation sets by up to 83.2%. They also test agent-side defenses like defensive prompting and structured evidence checks, which reduce harmful promotion by up to 39.2%, but cannot fully close the gap to baseline performance without any GEO attacks. The findings highlight a meaningful and unresolved vulnerability in AI-powered recommendation systems where seller-controlled content can systematically mislead AI agents.
- ResearcharXiv2026-06-08EQ
Less Context, Better Agents: Efficient Context Engineering for Long-Horizon Tool-Using LLM Agents · Abhilasha Lodha, Mahsa Pahlavikhah Varnosfaderani, Abir Chakraborty et al.
This paper investigates how to manage context efficiently for LLM-based agents handling enterprise workflows, specifically automated expense itemization in Microsoft Dynamics 365 Finance and Operations. The authors evaluate four GPT-5 configurations on a 50-task hotel expense benchmark, finding that pruning context to the last 5 tool call/response pairs plus automated summarization achieves 91.6% complete itemization and 99.64% average amount itemized, while reducing token usage and runtime compared to retaining full conversation history. Full-context retention achieved only 71.0% completion at nearly 1.5 million tokens and over 14 hours, whereas the pruning-plus-summarization approach used roughly 553,000 tokens and under 6 hours. The results demonstrate that selective retention of recent tool interactions combined with compact summarization improves both reliability and efficiency for enterprise tool-use agent workflows.
- ResearcharXiv2026-06-08CP
Local Is Not a Sufficient Privacy Boundary: Governing OS-Integrated On-Device AI · Jonghyun Chung, Sanket Badhe
This paper argues that running AI on a local device does not, by itself, constitute a meaningful privacy boundary, because on-device assistants can still aggregate sensitive data from email, calendars, files, and screenshots, persist derived state, invoke tools, and route requests to cloud infrastructure. The authors develop an OS-centered privacy framework that treats privacy as an institutional accountability problem, specifying a threat model, a six-part risk taxonomy, privacy-by-architecture controls, and a four-level audit rubric. They apply the rubric to Apple Intelligence/Foundation Models, Android AICore/Gemini Nano, and Microsoft Recall using publicly available documentation. The work has direct implications for how regulators, platform vendors, and auditors should govern on-device AI, emphasizing constrained information flow, bounded authority, visible user control, and auditable governance across the OS lifecycle.
- ResearcharXiv2026-06-08QP
Gaming AI-Assisted Peer Reviews Poses New Risks to the Scientific Community · Lin Li, Qi Zhang, Xander Davies et al.
This paper demonstrates that AI-assisted peer review systems are vulnerable to a simple, low-cost adversarial attack: superficially rephrasing a manuscript's abstract without changing its scientific content. The strongest attack achieves a success rate of about 38% in improving AI review outcomes—rising above 50% when the original AI review suggests rejection—increasing acceptance ratings by up to +1.31 points on a 10-point scale for Gemini 3 Flash reviewers and boosting scores on criteria like soundness, significance, and perceived contribution. The attack costs roughly $1 and 5 minutes per submission and is difficult to distinguish from ordinary editing, meaning authors may be incentivized to optimize for AI judgment rather than scientific merit. The authors argue that AI review tools should not be treated as neutral evaluators in high-stakes peer review without systematic robustness testing, transparent safeguards, and careful human oversight.
- ResearcharXiv2026-06-08Q
Invisible to humans, visible to machines: a preregistered audit of Unicode fidelity across four biomedical bibliographic APIs · Przemysław Czuma
This preregistered audit tests whether four major biomedical bibliographic APIs (PubMed E-utilities, Crossref, OpenAlex, Semantic Scholar) faithfully reproduce Unicode characters from published abstracts, using PubMed Central JATS XML as ground truth across a random sample of 4,000 articles. The study finds two systematic, near-total character losses: PubMed preserves typographic punctuation in only 0.6% of eligible abstracts, and OpenAlex preserves special whitespace in 0% of cases, while mathematical symbols and Greek letters are preserved at over 95% fidelity across all APIs. Additionally, Crossref returns no abstract at all for 24.6% of papers, with Elsevier and ACS showing 0% coverage. These findings matter because silently degraded text directly undermines the quality of biomedical LLM training corpora, scientometric analyses, and any corpus-based research that assumes API-returned text matches the published source.
- ResearcharXiv2026-06-08QC
A Controlled Audit of Pretraining Contamination in Public Medical Vision-Language Benchmarks · Bruce Changlong Xu, Lan Wu, Alexander Ryu
This paper audits whether publicly available medical vision-language benchmarks (SLAKE-En, PathVQA, VQA-RAD, and OmniMedVQA) may have been present in the pretraining data of open vision-language models, which would inflate reported accuracy. Using four detection methods—image-side near-neighbour overlap, canonical-order exchangeability, tail enrichment, and cross-model overlap—the authors find measurable image-source overlap on SLAKE-En (up to 19.8% of images flagged) and text-side signals on SLAKE-En and OmniMedVQA, though manual review suggests distributional rather than exact pixel-level duplication. Critically, some detector families (Min-K%++ tail enrichment and cross-model top-K overlap) prove unreliable as standalone contamination signals on small cohorts, as a model without plausible medical-VQA exposure (BLIP-2) reproduces apparent positive signals. The findings raise important concerns about the validity of benchmark evaluations for medical AI and highlight the need for more robust contamination detection methods.
- ResearcharXiv2026-06-08QP
Deployment-Time Memorization in Foundation-Model Agents · Lei, Chen, Guilin Zhang et al.
This paper examines how memory-design choices in long-lived AI agents—such as summarization aggressiveness, retrieval breadth, and deletion mode—jointly affect personalization utility, privacy risk, and deletion fidelity. The authors introduce two metrics (Personalization Recall and Adversarial Extraction Rate) plus a Forgetting Residue Score, finding on the LongMemEval benchmark that key-fact summarization reduces canary extraction by 76% on Gemma 3 12B and 64% on GPT-4o-mini while preserving nearly all personalization recall. However, the same compression creates a deletion-fidelity failure: raw-only deletion leaves derived summary copies recoverable in roughly 20% of cases, and only full-pipeline purge or tombstone redaction eliminates residue entirely. The work argues that persistent agent memory must be treated as a first-class memorization mechanism requiring evaluation of what agents can recall, what is extractable, and what can be truly erased.
- ResearcharXiv2026-06-08Q
BenSyc: Benchmarking Conversational Sycophancy and Human Alignment in LLMs for Bengali Contexts · Kazi Noshin, Sajib Acharjee Dip, Ranat Das Prangon et al.
BenSyc introduces the first benchmark for studying conversational sycophancy in Bengali social contexts, built from 11,840 Reddit posts and 170,000 comments from communities across Bangladesh and West Bengal. The benchmark uses human-validated binary labels and a five-level taxonomy (Invalidation, Neutral, Support, Validation, and Escalation) to assess whether LLM responses shift from balanced support toward excessive validation or escalatory alignment. Evaluating more than 15 open and proprietary LLMs, the best system achieves only 61.8 Macro-F1 on binary detection and 61.7 Macro-F1 on five-class classification, with many models frequently generating strongly validating or escalatory responses in emotionally charged situations. The findings underscore the need for culturally grounded multilingual benchmarks to evaluate socially aligned conversational AI.
- ResearcharXiv2026-06-08W
The Empirically Grounded Adaptive Virtual Patient for Psychotherapy Training: Disclosure That Responds to Therapist Micro-Skills · Angela Chen, Siwei Jin, Catherine Bao et al.
This paper introduces the Adaptive Virtual Patient (AVP), a simulated patient system for training psychotherapy micro-skills such as empathic responding and exploratory probing. Unlike fixed-script systems or unconstrained LLMs, the AVP uses a structural equation model fit to nearly 2,000 hours of real psychotherapy transcripts to dynamically adjust patient disclosure levels—from guarded to fully open—based on trainee behavior each turn, with an LLM generating utterances conditioned on those disclosure levels. In an evaluation with 20 clinicians and trainees across 80 sessions, AVP disclosure rose appropriately in response to therapist empathy and exploration, while a prompt-only baseline stayed flat. The system offers a scalable, empirically grounded tool for psychotherapy training that gives realistic, adaptive feedback on trainee skill.
- ResearcharXiv2026-06-08QP
Principled Uncertainty in Clinical AI: End-to-End Bayesian Modelling and Algorithmic Equity Auditing Across Multimodal Patient Data · Oladimeji Anthonio, Dimeji Abdulsobur Olawuyi, Oloruntoba Ajayi et al.
This paper presents a Bayesian deep learning framework for clinical AI that quantifies uncertainty in predictions made from multimodal patient data, and then uses those uncertainty estimates as a formal measure of algorithmic equity. The architecture combines variational encoders, precision-weighted fusion, and a decomposed uncertainty output separating aleatoric from epistemic uncertainty, achieving an Expected Calibration Error of 0.096 on 1,000 simulated patients. Equity audits show that epistemic uncertainty systematically flags underserved groups: primary/rural facility patients exhibit a 15.3% uncertainty equity gap (p < 0.001, effect size = 0.698), low-SES patients a 6.8% gap, and elderly patients a 3.9% gap, while no significant sex-based disparity is found. The findings argue that calibrated uncertainty is an actionable equity signal relevant to clinical deployment decisions, not merely a technical modeling property.
- ResearcharXiv2026-06-08EP
Powering the Future of AI: Navigating the Trade-offs for Europe's Energy Transition and Net-Zero Goals · Mohammad Hemmati, Gbemi Oluleye, Vassilis M. Charitopoulos
Using a spatially explicit optimisation model of Europe across 21 AI growth scenarios, this study quantifies how the rapid expansion of AI-driven hyperscale data centres could add 73–723 TWh of extra electricity demand by 2050, risking cumulative CO2 emissions overshoots of 67–181 MtCO2 between 2030 and 2050. The analysis finds that after 2030, AI infrastructure location will be driven more by firm power availability and system flexibility than by clean energy abundance, with moderate scenarios requiring 200 additional hours of firm generation and increasing levelised cost of energy by 35 EUR/MWh in key hubs. Existing infrastructure would need at least 70 GW of additional capacity even under pessimistic scenarios, rising to 226 GW under managed growth pathways, while improved efficiency could significantly reduce capacity needs and system peaks. The paper concludes that although net-zero targets for 2050 may still be achievable, intermediate-year emission risks are substantial and EU carbon-neutral goals could be compromised without policy adaptation to accelerating digital transformation.
- ResearcharXiv2026-06-08QP
Clinically Grounded Privacy Evaluation of Medical LMs · Sasha Ronaghi, Sana Tonekaboni, Lena Stempfle et al.
This paper introduces a clinically grounded framework for evaluating privacy leakage in medical language models (LMs), moving beyond simple training-text recovery to assess realistic threat scenarios. Testing on an LM pretrained on 378,000 clinical notes, the researchers find that routine encounter metadata—such as a patient's name, date of birth, and visit date—can elicit high rates of verbatim memorization and sensitive-diagnosis recovery (e.g., AUROC 0.91 for abortion, 0.81 for HIV). The study also cautions that exact-match memorization metrics can overstate risk, as 36% of memorized tokens reflect templated documentation rather than uniquely identifying content. These findings highlight serious privacy risks from training on longitudinal clinical data and offer a practical evaluation framework relevant to health data governance and policy.
- ResearcharXiv2026-06-08QP
PRISM: Recovering Instruction Sets from Language Model Activations · Gilad Gressel, Rahul Pankajakshan, Julia Diament et al.
PRISM is a new method for recovering the full set of active instructions steering a large language model (LLM) agent by decoding its internal hidden states. The system, called an activation-conditioned interpreter, is trained using judge-guided reinforcement learning (GRPO) to produce a faithful bullet list of instructions, constraints, prohibitions, and subgoals active at inference time. Across benign, constrained, prompt-injection, and hidden-objective scenarios, PRISM outperforms existing activation-to-language baselines, with the largest gains on security-relevant objectives. This matters for AI monitoring and safety: it gives auditors a principled way to detect when models are being steered by hidden or injected instructions they were not intended to follow.
- ResearcharXiv2026-06-08EQ
AI Scientists Are Only as Good as Their Evidence: A Stratified Ablation of Proprietary Data and Reasoning Skills in Drug-Asset Valuation · Yinan Wang
This paper tests whether AI agent capability in drug-asset valuation is primarily limited by reasoning quality or by the evidence the agent can access. Using a controlled three-arm ablation, the authors find that adding reasoning scaffolds (structured tools, valuation playbooks, verifiers) improves calibration and audit discipline, but fails to overcome a factual ceiling imposed by evidence gaps—non-proprietary agents recover only 25–38% of a curated competitive gold record. Adding a proprietary corpus (Noah AI) raises gold-record recovery to 96% and drives 'completeness-aware decision utility' to 7.43 versus 1.76–2.57 for web-only or scaffold-enhanced agents. The key finding is that proprietary evidence sets the upper bound of what an AI scientist can know and therefore decide, which has direct implications for how enterprises deploy AI in high-stakes scientific decisions.
- ResearcharXiv2026-06-08Q
Deterministic Integrity Gates for LLM-Assisted Clinical Manuscript Preparation: An Auditable Biomedical Informatics Architecture · Yoojin Nam, Jinhoon Jeong, Namkug Kim
This paper presents MedSci Skills, an open-source architecture for verifying AI-generated clinical manuscripts by pairing LLM-based generation with deterministic integrity gates. The system uses 43 modular skills and a 21-detector deterministic tier to catch fabricated citations, numerical drift from source tables, and unmet reporting-guideline items (STARD, PRISMA, STROBE) before a manuscript advances to the next stage. In a seeded-defect test, the deterministic gates detected all 27 injected defects with no false positives, while a single-prompt LLM reviewer caught only 11, missing defects hidden in code, bibliography, and style. The architecture produces an auditable, re-executable verification trail intended to support human oversight of LLM-assisted manuscript preparation.
- ResearcharXiv2026-06-08Q
LLM-Orchestrated Conformance Checking in Stroke Care Without Computer-Interpretable Guidelines · Giorgio Leonardi, Stefania Montani, Manuel Striani et al.
This paper presents a modular LLM-based framework that performs conformance checking in healthcare — assessing whether patient care pathways follow clinical guidelines — without requiring formal Computer-Interpretable Guidelines (CIGs). Applied to stroke care at Alessandria Hospital's neurological ward, the system automatically extracts patient traces from clinical discharge letters, derives normative rules from unstructured guideline text, translates them into executable scripts, and computes a Trace Conformance Indicator. Evaluated against 50 rules and hundreds of patient traces, the framework found that more than 86% of traces were conformant, demonstrating both the feasibility of LLM-orchestrated compliance analysis and a high level of adherence to stroke care guidelines at that institution. The approach is significant because it removes a key barrier to real-world clinical guideline compliance monitoring by eliminating the need for pre-existing formal guideline representations.
- ResearcharXiv2026-06-08CP
AI Assurance in UK Defence: Challenges in Operationalising JSP 936 · Callum Cockburn, Sam Farrow
This report analyzes the practical difficulties of implementing JSP 936 Part 1, the UK Ministry of Defence's AI assurance directive, through a structured interpretive review. It identifies eight thematic challenge areas—including adequacy of evidence and argument, human interaction management, operational environment definition, systems-of-systems integration, AI performance assessment, safety and security analysis, ethicality measurement, and managing AI's inherent complexity. The authors argue that while JSP 936 provides a useful governance basis, significant unresolved technical, organisational, and assurance questions remain due to the socio-technical nature of AI-enabled systems and uncertainty in real-world deployment. The report calls for further methods, guidance, and organisational capability to support ambitious, safe, and responsible AI adoption across UK Defence.
- ResearcharXiv2026-06-08QP
Pretrained, Frozen, Still Leaking: Auditing Cross-Encoder Attribute Transfer in EEG Foundation Models · Jianwei Tai
This paper audits privacy risks in EEG foundation models (BIOT, LaBraM, and EEGPT) by evaluating their frozen embeddings across four endpoints simultaneously: raw reconstruction, membership inference, identity linkage, and differential privacy on downstream heads. The key finding is that a single attribute decoder trained on one encoder can transfer to other encoders via a linear bridge, leaking spectral attributes even when individual single-endpoint audits appear to pass, with subject-disjoint confidence interval lower bounds of at least 0.081 across all six cross-encoder directions. The authors introduce an audit-endpoint disagreement score (AEDS), which is positive across all tested datasets (p<0.001), while standard defenses including DP-SGD and LiRA membership inference (AUC 0.50–0.70) fail to close the attribute leakage channel. The framework provides a joint release decision rule that can block model releases based on cross-encoder attribute leakage rather than relying on scattered single-endpoint defenses.
- ResearcharXiv2026-06-08QP
Culturally-Adapted Red-Teaming Across East and Southeast Asian Contexts: A Methodological and Comparative Analysis · Hyeji Choi, Yongtaek Lim, Minwoo Kim
This paper investigates a critical gap in multilingual safety testing of large language models (LLMs): the common practice of directly translating English safety benchmarks into other languages fails to capture culturally embedded threats, social norms, and legal frameworks. The authors built paired direct-translation and culturally-adapted datasets for Korean, Japanese, Thai, and Khmer, finding that culturally-adapted prompts produced higher Attack Success Rates across all 16 language-model combinations (mean +9.3 percentage points), while direct translation underestimated risk in 44 of 48 category-language combinations. Cultural Realism scores for direct-translated prompts averaged only 0.17 out of 3.0, versus up to 2.51 for culturally-adapted prompts, confirming that translation-only approaches produce inputs that diverge systematically from real-world multicultural settings. The findings make a strong methodological case that valid LLM safety evaluation requires culture-specific benchmark adaptation, not merely linguistic conversion.
- ResearcharXiv2026-06-08WE
Autonomous Incident Resolution at Hyperscale: An Agentic AI Architecture for Network Operations · Arun Malik
This paper presents a deployed multi-agent AI architecture designed to autonomously detect, diagnose, and remediate network incidents in large-scale cloud infrastructure without human intervention. The system uses hierarchical agent decomposition, runbook-derived knowledge, and progressive autonomy with layered safety controls including authorization checks and rollback mechanisms. Deployed in production at a major cloud provider, it achieves autonomous resolution rates exceeding 90% for common incident categories. The work demonstrates that agentic AI can meaningfully reduce reliance on human-driven incident response at hyperscale while preserving operational safety guarantees.
- ResearcharXiv2026-06-08EQ
Context Rot in AI-Assisted Software Development: Repurposing Documentation Consistency for AI Configuration Artifacts · Christoph Treude, Sebastian Baltes
This paper introduces 'context rot,' the phenomenon where AI coding assistant configuration files (such as CLAUDE.md, AGENTS.md, and .cursorrules) become stale as software evolves, causing the persistent context guiding AI tool behavior to diverge from the actual codebase. The authors argue that decades of software documentation consistency research provides an immediate toolbox for detecting this problem, and they present a research roadmap connecting existing approaches to this new setting. As preliminary evidence, applying an existing README/wiki consistency checker to a statistically representative sample of 356 repositories found stale code element references in 23.0% of repositories, demonstrating that traditional documentation consistency tools can already surface context rot. The findings matter because undetected context rot may silently degrade the quality of AI-assisted code generation across development teams.
- ResearcharXiv2026-06-08QP
Context-Fractured Decomposition Attacks on Tool-Using LLM Agents: Exploiting Artifact Provenance Gaps · Xiaofeng Lin, Yukai Yang, Daniel Guo et al.
This paper introduces Context-Fractured Decomposition (CFD), a family of jailbreak attacks targeting tool-using LLM agents that operate across multiple steps and artifact states. Unlike existing multi-turn jailbreaks such as Crescendo and Tree of Attacks—which assume a single contiguous conversation—CFD exploits 'provenance gaps,' where harmful behavior is elicited later in a pipeline by combining individually innocuous tool actions and artifacts from earlier interactions. The authors demonstrate that CFD improves jailbreak success rates by up to 28.3 percentage points over state-of-the-art baselines, even against strong single-turn judges, and propose provenance lineage tagging as a mitigation direction. The findings reveal a critical security gap in real-world agent deployments where enforcement is fragmented across tools, modules, and time.
- ResearcharXiv2026-06-08Q
Decoy-Calibrated Failure Audits for Language Models · Vyzantinos Repantis, Ameya Gawde, Harshvardhan Singh
This paper introduces Janus, a statistical auditing procedure for identifying where language models fail and whether those failure patterns are real or artifacts of selection bias. Janus scores candidate error explanations (called descriptors) by their error-rate lift and then compares them against fake 'decoy' descriptors with the same frequencies but random assignments, confirming only those that beat the decoy baseline and replicate on held-out data. In controlled experiments on multi-table lookup tasks and two public benchmarks (MuSiQue and LongBench v2), Janus avoids false positives that uncalibrated methods report — for example, reducing 20 reported descriptors to zero confirmed findings on LongBench v2. The method matters for quality assurance because it provides a principled way to separate proposing failure explanations from credibly reporting them, reducing the risk that auditors mistake noise for real model weaknesses.