News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
Deployment-Time Memorization in Foundation-Model Agents
Lei, Chen, Guilin Zhang et al.
arXiv · 2026-06-08
This paper examines how memory-design choices in long-lived AI agents—such as summarization aggressiveness, retrieval breadth, and deletion mode—jointly affect personalization utility, privacy risk, and deletion fidelity. The authors introduce two metrics (Personalization Recall and Adversarial Extraction Rate) plus a Forgetting Residue Score, finding on the LongMemEval benchmark that key-fact summarization reduces canary extraction by 76% on Gemma 3 12B and 64% on GPT-4o-mini while preserving nearly all personalization recall. However, the same compression creates a deletion-fidelity failure: raw-only deletion leaves derived summary copies recoverable in roughly 20% of cases, and only full-pipeline purge or tombstone redaction eliminates residue entirely. The work argues that persistent agent memory must be treated as a first-class memorization mechanism requiring evaluation of what agents can recall, what is extractable, and what can be truly erased.
- AI policy
- Quality assurance
Research
BenSyc: Benchmarking Conversational Sycophancy and Human Alignment in LLMs for Bengali Contexts
Kazi Noshin, Sajib Acharjee Dip, Ranat Das Prangon et al.
arXiv · 2026-06-08
BenSyc introduces the first benchmark for studying conversational sycophancy in Bengali social contexts, built from 11,840 Reddit posts and 170,000 comments from communities across Bangladesh and West Bengal. The benchmark uses human-validated binary labels and a five-level taxonomy (Invalidation, Neutral, Support, Validation, and Escalation) to assess whether LLM responses shift from balanced support toward excessive validation or escalatory alignment. Evaluating more than 15 open and proprietary LLMs, the best system achieves only 61.8 Macro-F1 on binary detection and 61.7 Macro-F1 on five-class classification, with many models frequently generating strongly validating or escalatory responses in emotionally charged situations. The findings underscore the need for culturally grounded multilingual benchmarks to evaluate socially aligned conversational AI.
- Quality assurance
Research
The Empirically Grounded Adaptive Virtual Patient for Psychotherapy Training: Disclosure That Responds to Therapist Micro-Skills
Angela Chen, Siwei Jin, Catherine Bao et al.
arXiv · 2026-06-08
This paper introduces the Adaptive Virtual Patient (AVP), a simulated patient system for training psychotherapy micro-skills such as empathic responding and exploratory probing. Unlike fixed-script systems or unconstrained LLMs, the AVP uses a structural equation model fit to nearly 2,000 hours of real psychotherapy transcripts to dynamically adjust patient disclosure levels—from guarded to fully open—based on trainee behavior each turn, with an LLM generating utterances conditioned on those disclosure levels. In an evaluation with 20 clinicians and trainees across 80 sessions, AVP disclosure rose appropriately in response to therapist empathy and exploration, while a prompt-only baseline stayed flat. The system offers a scalable, empirically grounded tool for psychotherapy training that gives realistic, adaptive feedback on trainee skill.
- Workforce
Research
Principled Uncertainty in Clinical AI: End-to-End Bayesian Modelling and Algorithmic Equity Auditing Across Multimodal Patient Data
Oladimeji Anthonio, Dimeji Abdulsobur Olawuyi, Oloruntoba Ajayi et al.
arXiv · 2026-06-08
This paper presents a Bayesian deep learning framework for clinical AI that quantifies uncertainty in predictions made from multimodal patient data, and then uses those uncertainty estimates as a formal measure of algorithmic equity. The architecture combines variational encoders, precision-weighted fusion, and a decomposed uncertainty output separating aleatoric from epistemic uncertainty, achieving an Expected Calibration Error of 0.096 on 1,000 simulated patients. Equity audits show that epistemic uncertainty systematically flags underserved groups: primary/rural facility patients exhibit a 15.3% uncertainty equity gap (p < 0.001, effect size = 0.698), low-SES patients a 6.8% gap, and elderly patients a 3.9% gap, while no significant sex-based disparity is found. The findings argue that calibrated uncertainty is an actionable equity signal relevant to clinical deployment decisions, not merely a technical modeling property.
- Quality assurance
- AI policy
Research
Powering the Future of AI: Navigating the Trade-offs for Europe's Energy Transition and Net-Zero Goals
Mohammad Hemmati, Gbemi Oluleye, Vassilis M. Charitopoulos
arXiv · 2026-06-08
Using a spatially explicit optimisation model of Europe across 21 AI growth scenarios, this study quantifies how the rapid expansion of AI-driven hyperscale data centres could add 73–723 TWh of extra electricity demand by 2050, risking cumulative CO2 emissions overshoots of 67–181 MtCO2 between 2030 and 2050. The analysis finds that after 2030, AI infrastructure location will be driven more by firm power availability and system flexibility than by clean energy abundance, with moderate scenarios requiring 200 additional hours of firm generation and increasing levelised cost of energy by 35 EUR/MWh in key hubs. Existing infrastructure would need at least 70 GW of additional capacity even under pessimistic scenarios, rising to 226 GW under managed growth pathways, while improved efficiency could significantly reduce capacity needs and system peaks. The paper concludes that although net-zero targets for 2050 may still be achievable, intermediate-year emission risks are substantial and EU carbon-neutral goals could be compromised without policy adaptation to accelerating digital transformation.
- AI policy
- Enterprise
Research
Clinically Grounded Privacy Evaluation of Medical LMs
Sasha Ronaghi, Sana Tonekaboni, Lena Stempfle et al.
arXiv · 2026-06-08
This paper introduces a clinically grounded framework for evaluating privacy leakage in medical language models (LMs), moving beyond simple training-text recovery to assess realistic threat scenarios. Testing on an LM pretrained on 378,000 clinical notes, the researchers find that routine encounter metadata—such as a patient's name, date of birth, and visit date—can elicit high rates of verbatim memorization and sensitive-diagnosis recovery (e.g., AUROC 0.91 for abortion, 0.81 for HIV). The study also cautions that exact-match memorization metrics can overstate risk, as 36% of memorized tokens reflect templated documentation rather than uniquely identifying content. These findings highlight serious privacy risks from training on longitudinal clinical data and offer a practical evaluation framework relevant to health data governance and policy.
- AI policy
- Quality assurance
Research
PRISM: Recovering Instruction Sets from Language Model Activations
Gilad Gressel, Rahul Pankajakshan, Julia Diament et al.
arXiv · 2026-06-08
PRISM is a new method for recovering the full set of active instructions steering a large language model (LLM) agent by decoding its internal hidden states. The system, called an activation-conditioned interpreter, is trained using judge-guided reinforcement learning (GRPO) to produce a faithful bullet list of instructions, constraints, prohibitions, and subgoals active at inference time. Across benign, constrained, prompt-injection, and hidden-objective scenarios, PRISM outperforms existing activation-to-language baselines, with the largest gains on security-relevant objectives. This matters for AI monitoring and safety: it gives auditors a principled way to detect when models are being steered by hidden or injected instructions they were not intended to follow.
- AI policy
- Quality assurance
Research
AI Scientists Are Only as Good as Their Evidence: A Stratified Ablation of Proprietary Data and Reasoning Skills in Drug-Asset Valuation
Yinan Wang
arXiv · 2026-06-08
This paper tests whether AI agent capability in drug-asset valuation is primarily limited by reasoning quality or by the evidence the agent can access. Using a controlled three-arm ablation, the authors find that adding reasoning scaffolds (structured tools, valuation playbooks, verifiers) improves calibration and audit discipline, but fails to overcome a factual ceiling imposed by evidence gaps—non-proprietary agents recover only 25–38% of a curated competitive gold record. Adding a proprietary corpus (Noah AI) raises gold-record recovery to 96% and drives 'completeness-aware decision utility' to 7.43 versus 1.76–2.57 for web-only or scaffold-enhanced agents. The key finding is that proprietary evidence sets the upper bound of what an AI scientist can know and therefore decide, which has direct implications for how enterprises deploy AI in high-stakes scientific decisions.
- Enterprise
- Quality assurance
Research
Deterministic Integrity Gates for LLM-Assisted Clinical Manuscript Preparation: An Auditable Biomedical Informatics Architecture
Yoojin Nam, Jinhoon Jeong, Namkug Kim
arXiv · 2026-06-08
This paper presents MedSci Skills, an open-source architecture for verifying AI-generated clinical manuscripts by pairing LLM-based generation with deterministic integrity gates. The system uses 43 modular skills and a 21-detector deterministic tier to catch fabricated citations, numerical drift from source tables, and unmet reporting-guideline items (STARD, PRISMA, STROBE) before a manuscript advances to the next stage. In a seeded-defect test, the deterministic gates detected all 27 injected defects with no false positives, while a single-prompt LLM reviewer caught only 11, missing defects hidden in code, bibliography, and style. The architecture produces an auditable, re-executable verification trail intended to support human oversight of LLM-assisted manuscript preparation.
- Quality assurance
Research
LLM-Orchestrated Conformance Checking in Stroke Care Without Computer-Interpretable Guidelines
Giorgio Leonardi, Stefania Montani, Manuel Striani et al.
arXiv · 2026-06-08
This paper presents a modular LLM-based framework that performs conformance checking in healthcare — assessing whether patient care pathways follow clinical guidelines — without requiring formal Computer-Interpretable Guidelines (CIGs). Applied to stroke care at Alessandria Hospital's neurological ward, the system automatically extracts patient traces from clinical discharge letters, derives normative rules from unstructured guideline text, translates them into executable scripts, and computes a Trace Conformance Indicator. Evaluated against 50 rules and hundreds of patient traces, the framework found that more than 86% of traces were conformant, demonstrating both the feasibility of LLM-orchestrated compliance analysis and a high level of adherence to stroke care guidelines at that institution. The approach is significant because it removes a key barrier to real-world clinical guideline compliance monitoring by eliminating the need for pre-existing formal guideline representations.
- Quality assurance
Research
AI Assurance in UK Defence: Challenges in Operationalising JSP 936
Callum Cockburn, Sam Farrow
arXiv · 2026-06-08
This report analyzes the practical difficulties of implementing JSP 936 Part 1, the UK Ministry of Defence's AI assurance directive, through a structured interpretive review. It identifies eight thematic challenge areas—including adequacy of evidence and argument, human interaction management, operational environment definition, systems-of-systems integration, AI performance assessment, safety and security analysis, ethicality measurement, and managing AI's inherent complexity. The authors argue that while JSP 936 provides a useful governance basis, significant unresolved technical, organisational, and assurance questions remain due to the socio-technical nature of AI-enabled systems and uncertainty in real-world deployment. The report calls for further methods, guidance, and organisational capability to support ambitious, safe, and responsible AI adoption across UK Defence.
- Certifications
- AI policy
Research
Pretrained, Frozen, Still Leaking: Auditing Cross-Encoder Attribute Transfer in EEG Foundation Models
Jianwei Tai
arXiv · 2026-06-08
This paper audits privacy risks in EEG foundation models (BIOT, LaBraM, and EEGPT) by evaluating their frozen embeddings across four endpoints simultaneously: raw reconstruction, membership inference, identity linkage, and differential privacy on downstream heads. The key finding is that a single attribute decoder trained on one encoder can transfer to other encoders via a linear bridge, leaking spectral attributes even when individual single-endpoint audits appear to pass, with subject-disjoint confidence interval lower bounds of at least 0.081 across all six cross-encoder directions. The authors introduce an audit-endpoint disagreement score (AEDS), which is positive across all tested datasets (p<0.001), while standard defenses including DP-SGD and LiRA membership inference (AUC 0.50–0.70) fail to close the attribute leakage channel. The framework provides a joint release decision rule that can block model releases based on cross-encoder attribute leakage rather than relying on scattered single-endpoint defenses.
- Quality assurance
- AI policy
Research
Culturally-Adapted Red-Teaming Across East and Southeast Asian Contexts: A Methodological and Comparative Analysis
Hyeji Choi, Yongtaek Lim, Minwoo Kim
arXiv · 2026-06-08
This paper investigates a critical gap in multilingual safety testing of large language models (LLMs): the common practice of directly translating English safety benchmarks into other languages fails to capture culturally embedded threats, social norms, and legal frameworks. The authors built paired direct-translation and culturally-adapted datasets for Korean, Japanese, Thai, and Khmer, finding that culturally-adapted prompts produced higher Attack Success Rates across all 16 language-model combinations (mean +9.3 percentage points), while direct translation underestimated risk in 44 of 48 category-language combinations. Cultural Realism scores for direct-translated prompts averaged only 0.17 out of 3.0, versus up to 2.51 for culturally-adapted prompts, confirming that translation-only approaches produce inputs that diverge systematically from real-world multicultural settings. The findings make a strong methodological case that valid LLM safety evaluation requires culture-specific benchmark adaptation, not merely linguistic conversion.
- Quality assurance
- AI policy
Research
Autonomous Incident Resolution at Hyperscale: An Agentic AI Architecture for Network Operations
Arun Malik
arXiv · 2026-06-08
This paper presents a deployed multi-agent AI architecture designed to autonomously detect, diagnose, and remediate network incidents in large-scale cloud infrastructure without human intervention. The system uses hierarchical agent decomposition, runbook-derived knowledge, and progressive autonomy with layered safety controls including authorization checks and rollback mechanisms. Deployed in production at a major cloud provider, it achieves autonomous resolution rates exceeding 90% for common incident categories. The work demonstrates that agentic AI can meaningfully reduce reliance on human-driven incident response at hyperscale while preserving operational safety guarantees.
- Workforce
- Enterprise
Research
Context Rot in AI-Assisted Software Development: Repurposing Documentation Consistency for AI Configuration Artifacts
Christoph Treude, Sebastian Baltes
arXiv · 2026-06-08
This paper introduces 'context rot,' the phenomenon where AI coding assistant configuration files (such as CLAUDE.md, AGENTS.md, and .cursorrules) become stale as software evolves, causing the persistent context guiding AI tool behavior to diverge from the actual codebase. The authors argue that decades of software documentation consistency research provides an immediate toolbox for detecting this problem, and they present a research roadmap connecting existing approaches to this new setting. As preliminary evidence, applying an existing README/wiki consistency checker to a statistically representative sample of 356 repositories found stale code element references in 23.0% of repositories, demonstrating that traditional documentation consistency tools can already surface context rot. The findings matter because undetected context rot may silently degrade the quality of AI-assisted code generation across development teams.
- Quality assurance
- Enterprise
Research
Context-Fractured Decomposition Attacks on Tool-Using LLM Agents: Exploiting Artifact Provenance Gaps
Xiaofeng Lin, Yukai Yang, Daniel Guo et al.
arXiv · 2026-06-08
This paper introduces Context-Fractured Decomposition (CFD), a family of jailbreak attacks targeting tool-using LLM agents that operate across multiple steps and artifact states. Unlike existing multi-turn jailbreaks such as Crescendo and Tree of Attacks—which assume a single contiguous conversation—CFD exploits 'provenance gaps,' where harmful behavior is elicited later in a pipeline by combining individually innocuous tool actions and artifacts from earlier interactions. The authors demonstrate that CFD improves jailbreak success rates by up to 28.3 percentage points over state-of-the-art baselines, even against strong single-turn judges, and propose provenance lineage tagging as a mitigation direction. The findings reveal a critical security gap in real-world agent deployments where enforcement is fragmented across tools, modules, and time.
- Quality assurance
- AI policy
Research
Decoy-Calibrated Failure Audits for Language Models
Vyzantinos Repantis, Ameya Gawde, Harshvardhan Singh
arXiv · 2026-06-08
This paper introduces Janus, a statistical auditing procedure for identifying where language models fail and whether those failure patterns are real or artifacts of selection bias. Janus scores candidate error explanations (called descriptors) by their error-rate lift and then compares them against fake 'decoy' descriptors with the same frequencies but random assignments, confirming only those that beat the decoy baseline and replicate on held-out data. In controlled experiments on multi-table lookup tasks and two public benchmarks (MuSiQue and LongBench v2), Janus avoids false positives that uncalibrated methods report — for example, reducing 20 reported descriptors to zero confirmed findings on LongBench v2. The method matters for quality assurance because it provides a principled way to separate proposing failure explanations from credibly reporting them, reducing the risk that auditors mistake noise for real model weaknesses.
- Quality assurance
Research
TRIAGE: Dialectical Reasoning for Explainable Risk Prediction on Irregularly Sampled Medical Time Series with LLMs
Hyeongwon Jang, Gyouk Chu, Changhun Kim et al.
arXiv · 2026-06-08
TRIAGE is a framework that trains large language models to perform dialectical reasoning—generating rationales for competing clinical outcomes—when predicting patient risk from irregularly sampled electronic health record time series. The paper shows that existing LLM approaches collapse graded risk into overconfident binary predictions, and that TRIAGE mitigates this by producing continuous, calibrated risk scores alongside explicit clinical reasoning. Evaluated on three benchmarks, TRIAGE achieves an average 3.3% AUPRC improvement and reduces calibration error by 81% compared to competitive baselines, while its rationales outperform post-hoc explanations by 20% in clinical reasoning quality. These results matter for clinical quality assurance because better-calibrated, interpretable early warning systems give clinicians verifiable reasoning to support patient triage decisions.
- Quality assurance
Research
Document-Authored Control-Signal Impersonation: A Low-Cost Indirect Prompt Attack on RAG Safety Boundaries
Jianguo Zhu
arXiv · 2026-06-08
This paper studies a security vulnerability in retrieval-augmented generation (RAG) systems where attacker-controlled document text can impersonate trusted metadata, provenance, or policy signals — a pattern the authors call Document-Authored Control-Signal Impersonation (DACSI). Unlike traditional prompt injection that issues direct override commands, DACSI exploits the fact that RAG pipelines collapse trusted and untrusted text into a single natural-language channel, allowing document-authored labels to be misread as authorized control signals. The authors evaluate DACSI across six language models (including DeepSeek V4 Pro, Qwen3.5-397B, GPT-5.5, Gemini 3.1 Pro, and GLM-4.7), finding varying levels of susceptibility across model regimes and identifying source/channel separation as a key mitigation direction. This matters for enterprise and policy contexts because RAG systems are widely deployed, and this low-cost attack vector bypasses conventional safety boundaries without any explicit instruction to override policy.
- Enterprise
- AI policy
Research
CARE: A Conformal Safety Layer for Medical Summarization
Suhana Bedi, Bridget Lin, Anson Y. Zhou et al.
arXiv · 2026-06-08
CARE (Conformal Assessment for Risk Evaluation) is a post-hoc, model-agnostic safety layer that applies conformal risk control to flag hallucinations and omissions in LLM-generated medical summaries without requiring model retraining. It provides finite-sample, distribution-free guarantees through two controllers—one bounding the probability of unflagged hallucinated sentences and one bounding the expected fraction of important omissions missed—by jointly calibrating over a two-dimensional threshold space. Tested across five medical summarization tasks, CARE satisfies its target risk bound (α=0.15) with 95% confidence across 100 calibration/test resplits using roughly 100 labeled documents per domain, and surfaces up to 5× fewer sentences than alternative calibrated baselines. A preliminary clinician study of 75 document reviews found that calibrated flags improved omission detection by 28.6 percentage points on average, demonstrating a tunable mechanism for balancing residual risk against clinician review burden.
- Quality assurance
- Workforce
Research
From Statute to Control Flow: Span-Grounded Deontic Trees for Defeasible Scope Parsing
Jian Chen, Siyuan Li, Chucheng Wan et al.
arXiv · 2026-06-08
This paper introduces NormBench, a benchmark of 2,290 legal and policy provisions in Chinese, English, and cross-lingual settings, designed to measure whether AI models can correctly parse the nested exception structures in statutes and regulations. The authors identify a failure mode called Silent Scope Omission (SSO), where a model applies a general rule but silently drops nested exceptions, producing outputs that appear compliant but break on edge cases. They propose Span-Grounded Deontic Trees (SG-DT), a structured intermediate representation that anchors every logical branch to source text spans and requires explicit exclusion guards, enabling deterministic compilation and audit. Evaluations of frontier LLMs reveal two recurring problems—Recursion Decay (performance drops as exception depth increases) and an Auditability Trap (models retrieve relevant spans but fail to assemble correct control flow)—with SG-DT improving whole-tree fidelity and exception recovery especially on SSO-prone cases.
- AI policy
- Quality assurance
Research
Video - Artificial Intelligence Systems in Accounting and Auditing: A Bibliometric and Exploratory Analysis
Ioana Florina Coita, Laura Filip, Marius Vlad Pop et al.
Zenodo (CERN European Organization for Nuclear Research) · 2026-06-08
This study analyzes 729 peer-reviewed articles on AI in accounting and auditing using bibliometric methods and evaluates ten AI-based solutions against a structured framework covering AI subset, functional category, entity size, audit relevance, and architectural transparency. Results show the field converges on supervised and deep-learning approaches, and that AI tools cluster into three functional categories—process automation, analytics/business intelligence, and predictive/audit-oriented systems—with adoption patterns differing by company size: SMEs benefit most from process automation and OCR, while large entities gain more from full-population analytics and ensemble-based anomaly detection. The paper also addresses trustworthiness limits of probabilistic AI systems and discusses implications for audit assurance and emerging regulatory frameworks such as the EU AI Act and ISO/IEC 42001. These findings matter for enterprise adoption decisions, quality assurance in financial workflows, and the development of AI governance and certification standards.
- Enterprise
- Quality assurance
- Certifications
- AI policy
Research
Artificial Intelligence and Export Performance in Small and Micro-Enterprises: The Roles of Internal Capability and External Tools
Mengyang Gu, Chuyue Jin
Sustainability · 2026-06-08
This study examines how internal AI capability and external AI tool utilization jointly affect export performance in 475 small and micro-enterprises at Yiwu International Trade City. Both types of AI resources independently improve export performance, but their interaction produces diminishing marginal returns, meaning that investing heavily in both simultaneously reduces each resource's incremental contribution under resource-constrained conditions. The findings offer practical guidance for small firms on how to allocate AI-related investments efficiently to maximize international competitiveness.
- Enterprise
- Workforce
Research
The Jagged Global Economy: Frontier AI Unevenly Exposes National Economies
Arul Murugan, Tomás Aguirre, Abhishek Nagaraj et al.
arXiv (Cornell University) · 2026-06-08
This paper introduces a national AI exposure metric applied across 141 countries, combining occupation-level AI exposure scores with international employment data. It finds that high-income countries face substantially greater direct labor-market exposure to frontier AI than low-income countries, and that women are more exposed than men in 91% of countries due to their concentration in white-collar and sales roles. The study also uncovers an indirect exposure mechanism via remittance dependencies—countries like Tajikistan face above-average effective exposure because their GDP relies heavily on remittances from highly-exposed economies. The findings demonstrate that policy responses calibrated to U.S. or European labor markets will not generalize globally, underscoring the need for country-specific workforce and policy strategies.
- Workforce
- AI policy
Research
HuntGPT: Integrating Machine Learning-Based Anomaly Detection and Explainable AI with Large Language Models (LLMs)
Tarek Ali, Panos Kostakos, Saeid Sheikhi
Telecom · 2026-06-08
HuntGPT is a specialized intrusion detection dashboard that integrates a Random Forest classifier trained on the KDD99 dataset with Explainable AI frameworks (SHAP and LIME) and a GPT-3.5 Turbo conversational agent to make network anomaly detection more understandable and actionable. The system aims to address long-standing barriers to ML adoption in cybersecurity—namely false positives, opaque model outputs, and limited XAI acceptance—by delivering threat explanations in plain language through an interactive interface. The prototype was evaluated for technical accuracy using Certified Information Security Manager (CISM) Practice Exams and assessed for response readability across six metrics, with results suggesting that LLM-backed conversational agents integrated with XAI can generate explainable, actionable outputs for intrusion detection. This matters because it offers a path to increasing analyst trust and operational efficiency in threat hunting workflows.
- Workforce
- Enterprise
- Quality assurance
- Certifications