News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
Causal Stories from Sensor Traces: Auditing Epistemic Overreach in LLM-Generated Personal Sensing Explanations
Shanshan Zhu, Han Zhang, J. Doris Chi et al.
arXiv · 2026-05-09
This paper investigates a phenomenon the authors call 'epistemic overreach' (EO), where large language models (LLMs) generate explanations of personal sensing data that imply more than the available evidence can support. Using three longitudinal sensing datasets of college students (StudentLife, GLOBEM, and CollegeExperience), the researchers generated 14,922 explanations across three LLM families (Llama, Qwen, and GPT) and found that LLMs routinely attribute anomalous days to unsupported causes across datasets, anomaly types, and model families. Critically, providing richer behavioral context did not reliably reduce EO, and even prompts explicitly instructing models to stay within the data only partially helped. The findings argue that evidential grounding — distinguishing what is observed, inferred, and unknown — must become a first-order evaluation criterion for LLM-generated personal sensing explanations, alongside fluency and plausibility.
- Quality assurance
- AI policy
Research
Source or It Didn't Happen: A Multi-Agent Framework for Citation Hallucination Detection
Mingzhe Li, Zhiqiang Lin, Shiqing Ma
arXiv · 2026-05-09
This paper addresses the problem of LLMs fabricating plausible-looking but invalid citations in scientific writing. The authors introduce CiteTracer, a cascading multi-agent system that classifies citations into a 12-code taxonomy spanning Real, Potential, and Hallucinated categories, using structured extraction, multi-source retrieval, and specialist judging agents. Evaluated on a benchmark of 2,450 synthetic citations and 957 real-world fabricated citations from ICLR 2026 and desk-rejected submissions, CiteTracer achieves 97.1% accuracy on both sets. This system offers a practical tool for auditing AI-assisted scientific writing and improving the reliability of research literature.
- Quality assurance
Research
Hybrid‑Threat Intelligence: A Critical Review of Semantic Integration Challenges and the Role of the HIPSTer Ontological Framework
R. Andrew Paskauskas, Evaldas Bružė, Giedre Sabaliauskaite et al.
Journal of Intelligent Communication · 2026-05-09
This scoping review evaluates Open Source Intelligence (OSINT), Social Media Intelligence (SOCMINT), and Natural Language Processing (NLP) capabilities for detecting hybrid threats that span information, cyber, and physical domains. The authors identify a persistent 'semantic gap' in which defensive systems collect extensive data but lack integrated cross-domain reasoning to correlate cyber indicators with narrative manipulation campaigns. They assess the HIPSTer ontological framework as a solution targeting multilingual hybrid threats—particularly in Russian and Chinese contexts—achieving TRL-4 validation through semantic vectors and formal reasoning. The review also examines how European regulations including GDPR, the AI Act, and NIS2 shape operational architectures, concluding with a research agenda to advance European hybrid threat detection toward operational maturity.
- AI policy
- Enterprise
- Quality assurance
Research
SAFE Governance Standards Whitepaper: Governance Architecture for AI-Mediated Child Safeguarding
Nyree Ayne
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-09
This whitepaper introduces SAFE (Safeguarding Architecture for Foundational Environments), an independent governance framework specifically designed to fill the gap in verification and certification infrastructure for AI systems used in environments where children are present. The framework proposes structural separation among governance (SAFE Council), certification (SAFE Labs), and implementation (SAIL), along with a behavioural risk taxonomy (SAFE-001) focused on interaction trajectories rather than isolated content classification. The authors argue that independent certification infrastructure is a missing layer in current AI safety ecosystems and that SAFE can support regulatory interoperability, procurement benchmarking, and ongoing behavioural verification. This matters because it proposes a scalable, institutionally credible architecture for child safeguarding in AI-mediated environments that could inform policy and certification standards.
- Certifications
- AI policy
- Quality assurance
Research
Log analysis is necessary for credible evaluation of AI agents
Peter Kirgis, Sayash Kapoor, Stephan Rabanser et al.
arXiv · 2026-05-08
This paper argues that AI agent benchmarks that report only pass/fail outcomes are insufficient for credible evaluation, identifying three validity threats: score inflation/deflation from shortcuts and artifacts, failure to predict real-world utility due to scaffold limitations, and concealment of dangerous agent actions. The authors propose log analysis—systematic tracking of inputs, execution, and outputs—as a necessary complement to outcome metrics, and develop a taxonomy of evaluation threats alongside guiding principles for conducting such analysis. Applying their framework to the tau-Bench Airline benchmark, they find that pass^5 performance was under-elicited by nearly 50% and that deployment failure modes were invisible to outcome-only metrics. The work concludes with practical recommendations for benchmark creators, model developers, independent evaluators, and deployers to adopt log analysis in order to improve evaluation credibility.
- Quality assurance
- Certifications
Research
Human-LLM Dialogue Improves Diagnostic Accuracy in Emergency Care
Burcu Sayin, Ngoc Vo Hong, Ipek Baris Schlicht et al.
arXiv · 2026-05-08
MedSyn is an interactive system that lets physicians iteratively query an LLM with access to full clinical records while the physician initially sees only a patient's chief complaint. In a study with seven emergency medicine physicians across 52 MIMIC-IV cases, AI-assisted sessions improved residents' hard-case correctness from 0.589 to 0.734, with standardized any-match accuracy improving by 0.156 (p < 0.0001) and residents showing the largest F1 gain (Δ = 0.138; p < 0.0001). Dialogue analysis showed expertise-dependent querying strategies and increased cross-physician diagnostic concordance (Δ = 0.145; p < 0.0001). The findings suggest that interactive LLM support meaningfully enhances diagnostic reasoning, particularly for less experienced clinicians in high-stakes emergency settings.
- Workforce
- Quality assurance
Research
Teachers' Perceived Benefits and Risks of AI Across Fifty-Five Countries: An Audit of LLM Alignment and Steerability
Yan Tao, Olga Viberg, Deepak Varuvel Dennison et al.
arXiv · 2026-05-08
This study audits how well large language models (LLMs) reflect teachers' actual perceptions of AI's benefits and risks, using representative OECD TALIS survey data from 55 countries and territories. Eight state-of-the-art LLMs from four providers were benchmarked against this cross-national survey evidence, revealing that models compress country-level differences, overestimate both benefits and risks, and show limited improvement from identity prompting or enhanced reasoning. Because LLM-generated guidance increasingly shapes how teachers learn about and discuss AI, this misalignment poses real risks for global AI-in-education policy. The authors caution against substituting LLM outputs for direct teacher engagement in policy development, while noting some models partially capture cross-national ranking patterns useful for exploratory analysis.
- AI policy
- Workforce
Research
Measuring What Matters: Benchmarking Generative, Multimodal, and Agentic AI in Healthcare
Prasanna Desikan, Harshit Rajgarhia, Shivali Dalmia et al.
arXiv · 2026-05-08
This paper argues that current AI benchmarks in healthcare are inadequate for evaluating real-world clinical deployment, showing a systematic gap between high scores on narrow tasks (e.g., near-perfect on medical licensing exams) and much weaker performance on actual clinical workflows, including documentation (0.74–0.85), clinical decision support (0.61–0.76), and administrative tasks (0.53–0.63). The authors contend that most benchmarks test knowledge recall rather than reliability, safety, and clinical relevance under realistic conditions, creating a false sense of deployment readiness. They call for a principled framework for benchmark design to distinguish genuine model limitations from measurement failures, arguing this is essential before AI systems take on high-stakes clinical roles.
- Quality assurance
- Certifications
Research
Defense effectiveness across architectural layers: a mechanistic evaluation of persistent memory attacks on stateful LLM agents
Jun Wen Leong
arXiv · 2026-05-08
This paper evaluates six defenses across four architectural layers against persistent memory injection attacks on LLM agents, testing nine open-source models across 5,040 experimental runs. The core finding is that five of six defenses fail because they operate at the wrong architectural layer: input filters miss payloads that enter via RAG retrieval, while retrieval-level classifiers cannot distinguish injection from legitimate policy content. Only tool-gating at the memory layer (Memory Sandbox) reduces attack success rate to 0% for eight of nine models, but a reasoning-mode ablation reveals no single sandbox implementation is safe across both reasoning and non-reasoning model classes. A loaded-corpus frontier evaluation across 21 models and three providers finds that under realistic conditions, some models exfiltrate at up to 95% attack success rate, with nearly all OpenAI and Gemini models storing injected rules at 100% regardless of execution resistance—creating supply-chain risk in shared-memory deployments.
- Quality assurance
- Enterprise
Research
Can Language Models Identify Side Effects of Breast Cancer Radiation Treatments?
Natalie Seah, Danielle S. Bitterman, Daphna Spiegel et al.
arXiv · 2026-05-08
This study evaluates how well large language models (LLMs) can identify side effects of breast cancer radiation treatments, a task critical for informed consent and survivorship care. Using 21 breast cancer patient profiles and seven instruction-tuned LLMs tested across multiple prompting strategies, the researchers compared LLM outputs against a clinician-curated reference developed by more than seven breast radiation oncologists at two major academic medical centers. Key findings include that LLMs are sensitive to minor documentation changes, show trade-offs between precision and recall, and systematically under-recall rare and long-term side effects. Grounding LLM outputs in clinician-curated side effect lists substantially improved reliability, suggesting practical design guidance for safer, more informative oncology applications.
- Quality assurance
Research
A Semantic-Sampling Framework for Evaluating Calibration in Open-Ended Question Answering
Zhanliang Wang, Jiancong Xiao, Ruochen Jin et al.
arXiv · 2026-05-08
This paper introduces Sem-ECE (Semantic-Sampling Expected Calibration Error), a framework for evaluating how well large language models' confidence scores align with their actual accuracy in open-ended question answering. The framework samples model answers, groups them into semantic classes, and uses the resulting frequencies as a proxy for confidence, addressing shortcomings of logit-based, verbalized, and existing sampling-based calibration methods. The authors prove their two estimators are asymptotically unbiased and show experimentally across three QA benchmarks and five commercial LLMs that Sem-ECE outperforms verbalized confidence and existing sampling-based approaches. This matters for high-stakes deployment settings like medicine and law, where knowing whether a model's expressed confidence is trustworthy is critical to safe use.
- Quality assurance
Research
SecureForge: Finding and Preventing Vulnerabilities in LLM-Generated Code via Prompt Optimization
Houjun Liu, Lisa Einstein, John Yang et al.
arXiv · 2026-05-08
SecureForge is an automated pipeline that audits and reduces cybersecurity vulnerabilities in code generated by large language models (LLMs). The study finds that even frontier models asked to write secure code still produce verifiable vulnerabilities about 23% of the time across 250 benign coding prompts. SecureForge uses a Markovian sampling technique to build a diverse synthetic prompt corpus and then iteratively optimizes system prompts, achieving up to a 48% reduction in output vulnerabilities while also improving unit test success rates. The resulting secure system prompts transfer zero-shot to real-world coding agent deployments, offering a practical quality-assurance mechanism for AI-assisted software development.
- Quality assurance
- Enterprise
Research
Augmenting Human Evaluation with LLM Judges: How Many Human Reviews Do You Need?
Jane Paik Kim
arXiv · 2026-05-08
This paper addresses the lack of rigorous methodology for using large language models (LLMs) as evaluators alongside human raters. Rather than treating LLMs as full substitutes for human judgment, the authors reframe LLMs as auxiliary evaluators in a two-stage sampling design: LLM ratings are collected for all observations, while human ratings are gathered for a strategic subsample. They propose a doubly robust estimator from the missing data literature to combine these ratings and provide principled guidance on how many human reviews are needed to achieve a targeted level of statistical power, including allocating more human oversight to cases where LLM predictions are least reliable. This work is relevant to quality assurance and policy, as it provides formal study-design tools for validating AI evaluation benchmarks in high-stakes settings where human oversight requirements are otherwise undefined.
- Quality assurance
- AI policy
Research
Towards Apples to Apples for AI Evaluations: From Real-World Use Cases to Evaluation Scenarios
Yee-Yin Choong, Kristen Greene, Alice Qian et al.
arXiv · 2026-05-08
This paper addresses the challenge of inconsistent AI evaluation methodologies by proposing a structured, human-centered process for transforming high-level AI use cases into detailed, comparable evaluation scenarios. The authors introduce an AI Use Case Worksheet with six key elements—use case, sector, user, intended outcomes, expected impacts, and KPIs/metrics—and demonstrate it in the U.S. financial services sector, covering use cases such as cyber defense enablement, developer productivity, financial crime aggregation, SAR filing, credit memo generation, and internal call center support. A three-stage pipeline combining LLM prompting with iterative human reviews generates 107 scenarios from these use cases, supported by a validation rubric to assess scenario quality. The work aims to establish a more consistent and meaningful paradigm for human-centered AI evaluations, enabling more valid 'apples-to-apples' comparisons across AI systems.
- Quality assurance
- AI policy
Research
The Limits of AI-Driven Allocation: Optimal Screening under Aleatoric Uncertainty
Santiago Cortes-Gomez, Mateo Dulce Rubio, Carlos Patino et al.
arXiv · 2026-05-08
This paper examines the fundamental limits of machine learning-based resource allocation in policy and humanitarian settings, showing that even perfect predicted risk scores cannot eliminate misallocation due to irreducible (aleatoric) uncertainty about individual vulnerability status. The authors develop a two-stage framework that optimally combines algorithmic targeting with physical screening, finding that the best strategy screens units at the margin of algorithmic allocation while directly targeting the highest-risk individuals. Empirical analysis shows that efficiency gains from screening grow as population-level aleatoric uncertainty increases, with applications to income-based social protection programs and humanitarian demining in Colombia demonstrating real operational consequences of balancing screening costs against allocation efficiency.
- AI policy
Research
Creating and Evaluating K-12 GenAI Assessment Graders Through Context Engineering
Zewei Tian, Alex Liu, Lief Esbenshade et al.
arXiv · 2026-05-08
This paper evaluates large language model (LLM)-based graders for K-12 standardized assessments by measuring interrater agreement against human raters using Massachusetts Comprehensive Assessment System (MCAS) data across mathematics, science, and ELA subjects. Using metrics such as Quadratic Weighted Kappa (QWK) and Proportional Reduction in Mean-Squared Error (PRMSE), the study finds that models with more parameters—including Claude Sonnet 4, Haiku 4.5, GPT-5, and GPT-5 Mini—achieve substantial agreement with human raters in math and science, though performance varies in ELA. Teacher and student feedback shows strong acceptance of AI-generated narrative feedback but skepticism toward AI-assigned numerical scores, pointing to LLMs as more suitable for formative than summative evaluation. The authors conclude that hybrid models combining AI efficiency with teacher judgment can reduce workload, improve feedback quality, and support equitable assessment without displacing professional expertise.
- Quality assurance
- Workforce
Research
LLM Wardens: Mitigating Adversarial Persuasion with Third-Party Conversational Oversight
Lennart Wachowiak, Scott D. Blain, David Williams-King et al.
arXiv · 2026-05-08
This paper investigates how LLMs can be used as adversarial persuaders and proposes a 'warden' model—a secondary LLM that monitors human-AI conversations in real time and issues private, non-binding warnings when it detects manipulation. In a preregistered user study (N=120) across four decision-making scenarios, an adversarial LLM succeeded in steering user decisions 65.4% of the time, but adding a warden reduced that success rate to 30.4%, while only minimally affecting genuine interactions (8.6 percentage-point reduction). The researchers also released COAX-Bench, a simulation benchmark spanning 14 scenarios—including hiring, voting, and file access—where warden models reduced adversarial success from 34.7% to 12.3% across 16,212 simulated interactions. Critically, even warden models weaker than the adversary they oversee provided meaningful protection, suggesting a scalable approach to oversight of increasingly capable AI systems.
- AI policy
- Quality assurance
Research
What if AI systems weren't chatbots?
Sourojit Ghosh, Pranav Narayanan Venkit, Sanjana Gautam et al.
arXiv · 2026-05-08
This paper critiques the dominant chatbot paradigm in AI development, arguing it is not a neutral design choice but a sociotechnical configuration with broad structural downsides. The authors show that chatbot-based systems often fail users in complex or high-stakes contexts while projecting unwarranted confidence, and that normalizing chatbot-mediated interaction contributes to deskilling, homogenization of knowledge, and shifting expectations of expertise. At a societal level, the paper identifies labor displacement, concentration of economic power, and increased environmental costs as consequences of sustained investment in large-scale chatbot infrastructures. The authors call for alternative AI development paths emphasizing pluralistic system design, task-specific tools, and institutional safeguards to mitigate social and economic harm.
- Workforce
- AI policy
Research
Beyond Confidence: Rethinking Self-Assessments for Performance Prediction in LLMs
Sree Bhattacharyya, Samarth Khanna, Leona Chen et al.
arXiv · 2026-05-08
This paper challenges the common practice of using a single confidence score to predict when large language models (LLMs) will fail, proposing instead a multidimensional self-assessment framework inspired by cognitive appraisal theory from human psychology. The authors elicit six appraisal-based dimensions—including effort and ability—alongside traditional confidence, then evaluate their predictive power across 12 LLMs and 38 tasks in eight domains. They find that competence-related dimensions, especially effort and ability, consistently match or outperform confidence in predicting model errors, with effort producing less overoptimistic and more stable estimates across model sizes. These findings suggest that structured multidimensional self-assessment could meaningfully improve the reliability and safety of LLM deployment in real-world settings.
- Quality assurance
- Enterprise
Research
RuleSafe-VL: Evaluating Rule-Conditioned Decision Reasoning in Vision-Language Content Moderation
Zhifeng Lu, Dianyuan Wang, Yuhu Shang et al.
arXiv · 2026-05-08
RuleSafe-VL is a new benchmark designed to evaluate how well vision-language models (VLMs) reason through explicit policy rules in content moderation, rather than simply matching predefined labels. The benchmark formalizes 93 atomic rules and 92 typed rule relations derived from real platform moderation policies, producing 2,166 context-sensitive image-text cases across three high-risk policy families and four diagnostic tasks. Experiments on 10 frontier, open-source, and safety-oriented VLMs reveal that rule-relation recovery is the dominant bottleneck—the best model achieves only 64.8 Macro-F1, while some safety-oriented models fall below 7 Macro-F1—indicating that current models cannot reliably replicate the structured reasoning that sound content moderation requires. This work matters for quality assurance and policy because it exposes a critical gap between benchmark performance and genuine policy-compliant decision-making in automated moderation systems.
- Quality assurance
- AI policy
Research
SARC: A Governance-by-Architecture Framework for Agentic AI Systems
Gaston Besanson
arXiv · 2026-05-08
SARC is a runtime governance architecture for agentic AI systems that embeds compliance constraints directly into the agent execution loop rather than relying on prompts, dashboards, or after-the-fact documentation. The framework compiles constraint specifications into four enforcement sites—a Pre-Action Gate, Action-Time Monitor, Post-Action Auditor, and Escalation Router—and extends to multi-agent workflows through constraint propagation and attribution-preserving trace trees. In a reproducible synthetic evaluation across 50 seeds on a procurement task, SARC achieved zero hard-constraint violations under exact predicates and reduced soft-window overages by 89.5% compared to a policy-as-code-only baseline. This matters because it provides a principled architectural substrate for making regulatory obligations executable and auditable at runtime in regulated settings where post-hoc controls are structurally insufficient.
- AI policy
- Enterprise
Research
LLM hallucinations in the wild: Large-scale evidence from non-existent citations
Zhenyue Zhao, Yihe Wang, Toby Stuart et al.
arXiv · 2026-05-08
This large-scale audit of 111 million references across 2.5 million papers on arXiv, bioRxiv, SSRN, and PubMed Central finds a sharp rise in non-existent citations following widespread LLM adoption, with a conservative estimate of 146,932 hallucinated citations in 2025 alone. Hallucinated references are especially prevalent in fields with rapid AI uptake, in papers showing linguistic signatures of AI-assisted writing, and among small or early-career author teams. The study also finds that hallucinated citations disproportionately credit already prominent and male scholars, potentially reinforcing existing inequities in scientific recognition. Critically, preprint moderation and journal publication processes capture only a fraction of these errors, indicating that existing safeguards have been outpaced by the spread of hallucinated content.
- Quality assurance
- AI policy
Research
DRIP-R: A Benchmark for Decision-Making and Reasoning Under Real-World Policy Ambiguity in the Retail Domain
Hsuvas Borkakoty, Sebastian Pohl, Cheng Wang et al.
arXiv · 2026-05-08
DRIP-R is a benchmark designed to evaluate how LLM-based agents handle real-world retail policy ambiguities—situations where multiple valid interpretations exist and no single correct resolution can be defined. The benchmark pairs policy-ambiguous return scenarios with realistic customer personas, conversational simulation with tool-calling, and a multi-judge evaluation framework assessing policy adherence, dialogue quality, behavioral alignment, and resolution quality. Experiments reveal that frontier models fundamentally disagree when faced with identical ambiguous scenarios, demonstrating that policy ambiguity is a genuine and systematic challenge for LLM decision-making. This matters for enterprise deployments where AI agents handle consequential customer-facing tasks governed by imprecise policies.
- Enterprise
- Quality assurance
Research
Quality-Conditioned Agreement in Automated Short Answer Scoring: Mid-Range Degradation and the Impact of Task-Specific Adaptation
Abigail Victoria Gurin Schleifer, Moriah Ariely, Beata Beigman Klebanov et al.
arXiv · 2026-05-08
This study examines how well automated short answer scoring (ASAS) systems handle student responses of varying quality, comparing few-shot large language models (GPT and Claude), a fine-tuned BERT-based encoder, and a human expert on open-ended biology items. The key finding is that all AI models perform adequately on fully correct or fully incorrect responses but show substantial degradation in scoring accuracy on mid-range, partially correct responses — a pattern termed 'mid-range degradation.' The severity of this degradation is tied to the degree of task-specific adaptation: few-shot LLMs with minimal examples perform worst on mid-range responses, while fine-tuned encoder models perform best. The authors warn that this inequity in scoring accuracy disproportionately affects students with developing understanding, underscoring the need for quality-conditioned fairness evaluation in automated scoring.
- Quality assurance
- Certifications
Research
Safe, or Simply Incapable? Rethinking Safety Evaluation for Phone-Use Agents
Zhengyang Tang, Yi Zhang, Chenxin Li et al.
arXiv · 2026-05-08
This paper introduces PhoneSafety, a benchmark of 700 safety-critical decision points drawn from real phone interactions across more than 130 apps, designed to evaluate whether phone-use AI agents avoid harm due to genuine safety awareness or simple inability to act. The authors evaluate eight phone-use agents and find that stronger general phone-use ability does not reliably translate to safer choices at risky moments, and that failures to act are concentrated in more visually and operationally demanding settings—behaving as a capability signal rather than a safety signal. The work argues that a harmless outcome alone is insufficient evidence of safety, and that evaluations must distinguish between unsafe judgment and incapacity to act.
- Quality assurance
- Certifications