News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
Fin-Bias: Comprehensive Evaluation for LLM Decision-Making under human bias in Finance Domain
Xiaoyu Hu, Jinman Zhao
arXiv · 2026-05-09
Fin-Bias introduces a benchmark of 8,868 long firm-specific analyst reports to evaluate how large language models make investment decisions when exposed to uncertain financial contexts and potentially biased human opinions. The study finds that LLMs tend to 'herd' toward explicit biases present in the context, such as analyst investment ratings, rather than reasoning independently. The authors also develop a bias-detection method that encourages LLMs to think independently, with some models exceeding human performance in predicting future stock returns when guided by this approach. These findings raise important concerns about the reliability and alignment of LLMs deployed in financial decision-making settings.
- Enterprise
- Quality assurance
Research
BiAxisAudit: A Novel Framework to Evaluate LLM Bias Across Prompt Sensitivity and Response-Layer Divergence
Jialing Gan, Junhao Dong, Songze Li
arXiv · 2026-05-09
BiAxisAudit introduces a two-axis auditing framework for evaluating bias in large language models (LLMs) that addresses critical shortcomings in existing benchmarks used under governance frameworks like the EU AI Act. The framework reveals that meaning-preserving prompt format changes can shift bias endorsement by more than 0.7, and that within a single response, the discrete selection and free-text elaboration layers frequently take opposing stances—a 'cancellation trap' that hides internal inconsistency behind clean aggregate scores. Tested across eight LLMs with 80,200 coded responses each, the study finds that task format alone explains as much variance as model choice, and that 63.6% of pooled bias signals appear in only one coding layer, making selection-only and elaboration-only model rankings nearly uncorrelated (Spearman ρ=0.238, p=0.570). These findings matter for AI policy and quality assurance because they show that current single-scalar benchmarks can be gamed or mislead regulators without any change to model weights, undermining the reliability of compliance evaluations under emerging AI governance regimes.
- AI policy
- Quality assurance
Research
Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs
Sangyeon Yoon, Wonje Jeung, Yoonjun Cho et al.
arXiv · 2026-05-09
This paper demonstrates that Direct Preference Optimization (DPO) fine-tuning, offered through APIs like OpenAI's, creates a serious and hard-to-detect safety vulnerability in large language models. The authors show that using just 10 harmless preference pairs—where refusals are marked as dispreferred responses—is enough to broadly suppress safety refusal behavior and transfer that suppression to harmful prompts never seen during fine-tuning. Across four OpenAI models, the attack achieves jailbreak success rates ranging from 54.80% to 81.73% at costs as low as $0.10, and on open-weight models the effect can emerge from even a single benign preference pair. The findings matter because the attack is practically indistinguishable from legitimate fine-tuning requests aimed at reducing over-refusal, making it extremely difficult to audit or prevent through standard content inspection.
- AI policy
- Quality assurance
Research
Mental Health AI Safety Claims Must Preserve Temporal Evidence
Srimonti Dutta, Ratna Kandala
arXiv · 2026-05-09
This paper argues that current safety evaluations for mental health AI systems are fundamentally flawed because they assess isolated responses or aggregate outcomes rather than the temporal sequence of interactions. The authors introduce 'Temporal Safety Non-Identifiability,' a formal framework showing that safety properties depending on sequence, timing, or accumulation cannot be certified by protocols that discard those features. They develop SCOPE-MH, a reporting standard for mental health AI that preserves temporal evidence, and demonstrate its utility on the AnnoMI dataset of motivational interviewing conversations, uncovering failure mechanisms invisible to per-turn scoring. The work has direct implications for how mental health AI systems are evaluated and certified before deployment in safety-critical settings.
- Quality assurance
- Certifications
Research
FraudBench: A Multimodal Benchmark for Detecting AI-Generated Fraudulent Refund Evidence
Xinyu Yan, Boyang Chen, Jiaming Zhang et al.
arXiv · 2026-05-09
FraudBench is a new multimodal benchmark designed to detect AI-generated fraudulent refund evidence in e-commerce, food delivery, and travel-service contexts. The benchmark combines real user-review images with metadata and synthesizes fake-damaged evidence using six image editing and generation models. Experiments reveal that current multimodal large language models (MLLMs) frequently fail to detect fake-damaged evidence—with true positive rates far below 50% on most generator subsets—while specialized AI-image detectors perform better but remain inconsistent across generators and produce false positives on real-damaged samples. The work highlights a significant gap between generic AI image detection and the reliable, claim-conditioned verification needed to combat refund fraud.
- Enterprise
- Quality assurance
Research
Debugging the Debuggers: Failure-Anchored Structured Recovery for Software Engineering Agents
Chenyu Zhao, Shenglin Zhang, Yihang Lin et al.
arXiv · 2026-05-09
PROBE is a structured recovery framework for software engineering AI agents that converts runtime failure telemetry into grounded diagnoses and bounded recovery guidance, without requiring changes to the agent's policy or toolset. Evaluated on 257 unresolved cases spanning repository-level software repair, enterprise workflow recovery, and AIOps service mitigation, PROBE achieves 65.37% Top-1 diagnosis accuracy and a 21.79% recovery rate, outperforming the strongest baseline by 43.58 and 12.45 percentage points respectively. A Microsoft IcM prototype demonstrates that PROBE can operate as a non-intrusive side channel in real-world service-diagnosis workflows. The findings highlight a diagnosis-recovery gap: accurate diagnosis alone is insufficient unless translated into actionable, evidence-grounded guidance that a subsequent attempt can execute and verify.
- Enterprise
- Quality assurance
Research
AgentForesight: Online Auditing for Early Failure Prediction in Multi-Agent Systems
Boxuan Zhang, Jianing Zhu, Zeru Shi et al.
arXiv · 2026-05-09
AgentForesight introduces an online auditing framework for LLM-based multi-agent systems that predicts failures in real time rather than diagnosing them after a trajectory has completed. The authors curate AFTraj-2K, a dataset of agentic trajectories across Coding, Math, and Agentic domains with step-level annotations of decisive errors, and train AgentForesight-7B using a coarse-to-fine reinforcement learning approach. On AFTraj-2K and an external benchmark, AgentForesight-7B outperforms proprietary models including GPT-4.1 and DeepSeek-V4-Pro by up to +19.9% and achieves 3× lower step localization error, enabling deployment-time intervention before failures cascade through downstream agents.
- Quality assurance
- Enterprise
Research
When Can Human-AI Teams Outperform Individuals? Tight Bounds with Impossibility Guarantees
Dongxin Guo, Jikun Wu, Siu-Ming Yiu
arXiv · 2026-05-09
This paper derives theoretical bounds explaining when human-AI teams can outperform their best individual member. The authors show that complementarity is achievable if and only if the error correlation between human and AI falls below a critical threshold ρ*, and that no confidence-based aggregation rule can achieve complementarity when that threshold is exceeded. Their framework, combining signal detection theory and information theory, predicts observed team accuracy with high correlation (R = 0.94 on ImageNet-16H, R = 0.91 on CIFAR-10H), explaining why complementarity is rare in practice and offering actionable design guidance for building effective human-AI teams.
- Workforce
- Enterprise
Research
Explanation Fairness in Large Language Models: An Empirical Analysis of Disparities in How LLMs Justify Decisions Across Demographic Groups
Gautam Veldanda
arXiv · 2026-05-09
This paper introduces the Explanation Fairness Taxonomy (EFT), a framework for measuring whether large language models justify decisions with equal quality, depth, tone, and linguistic sophistication across demographic groups. In a controlled study spanning 80 prompt templates, four high-stakes domains (hiring, medical triage, credit assessment, and legal judgment), and five LLMs (GPT-4.1, Claude Sonnet, LLaMA 3.3 70B, GPT-OSS 120B, and Qwen3 32B), all eight EFT metrics showed statistically significant disparities (Cohen's d from small to large, all p_BH < 10^(-62)). Model choice strongly influenced disparity magnitude—for example, Qwen3 32B exhibited verbosity disparities 5.9x larger than LLaMA 3.3 70B—and while prompting-based mitigations reduced decision-linked explanation disparity by 78–95%, they had no significant effect on stylistic dimensions, suggesting those inequalities are encoded in pre-training. The findings have direct implications for AI regulation and auditing practice in consequential deployment settings.
- AI policy
- Quality assurance
Research
AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators
Aritra Mazumder, Shubhashis Roy Dipta, Nusrat Jahan Lia et al.
arXiv · 2026-05-09
AgentCollabBench introduces a diagnostic benchmark of 900 human-validated tasks spanning software engineering, DevOps, and data engineering to measure process-level failures in multi-agent AI systems that outcome-based evaluations miss. The benchmark isolates four behavioral risks—instruction decay, false-belief contagion, context leakage, and tracer durability—and evaluates four modern LLMs, revealing model-specific vulnerability profiles and finding that communication topology explains 7–40% of variance in multi-hop information survival. A key finding is that converging-DAG nodes create a synthesis bottleneck where agents discard constraints carried by minority branches, a structural flaw absent from linear chains. The work argues that multi-agent reliability is fundamentally a structural problem and that scaling model intelligence alone cannot substitute for careful architecture design.
- Quality assurance
- Enterprise
Research
The Challenges of Balancing AI Compliance and Technological Innovations in Critical Sectors: A Systematic Literature Review
Ayush Enkhtaivan, Chinazunwa Uwaoma
arXiv · 2026-05-09
This systematic literature review (2020–2025) examines how critical infrastructure sectors—healthcare, finance, energy, and defense—struggle to balance AI compliance with technological innovation. The study identifies three core challenges: fragmented regulations across jurisdictions, disproportionate compliance burdens on small and medium enterprises (SMEs), and misaligned governance models. To address these, the paper highlights strategies such as risk-tiered regulation, compliance by design, and explainable AI as pathways toward scalable, trustworthy AI deployment. The findings offer a conceptual mapping of governance challenges and actionable guidance for policymakers and practitioners seeking to harmonize oversight with innovation.
- AI policy
- Enterprise
Research
Causal Stories from Sensor Traces: Auditing Epistemic Overreach in LLM-Generated Personal Sensing Explanations
Shanshan Zhu, Han Zhang, J. Doris Chi et al.
arXiv · 2026-05-09
This paper investigates a phenomenon the authors call 'epistemic overreach' (EO), where large language models (LLMs) generate explanations of personal sensing data that imply more than the available evidence can support. Using three longitudinal sensing datasets of college students (StudentLife, GLOBEM, and CollegeExperience), the researchers generated 14,922 explanations across three LLM families (Llama, Qwen, and GPT) and found that LLMs routinely attribute anomalous days to unsupported causes across datasets, anomaly types, and model families. Critically, providing richer behavioral context did not reliably reduce EO, and even prompts explicitly instructing models to stay within the data only partially helped. The findings argue that evidential grounding — distinguishing what is observed, inferred, and unknown — must become a first-order evaluation criterion for LLM-generated personal sensing explanations, alongside fluency and plausibility.
- Quality assurance
- AI policy
Research
Source or It Didn't Happen: A Multi-Agent Framework for Citation Hallucination Detection
Mingzhe Li, Zhiqiang Lin, Shiqing Ma
arXiv · 2026-05-09
This paper addresses the problem of LLMs fabricating plausible-looking but invalid citations in scientific writing. The authors introduce CiteTracer, a cascading multi-agent system that classifies citations into a 12-code taxonomy spanning Real, Potential, and Hallucinated categories, using structured extraction, multi-source retrieval, and specialist judging agents. Evaluated on a benchmark of 2,450 synthetic citations and 957 real-world fabricated citations from ICLR 2026 and desk-rejected submissions, CiteTracer achieves 97.1% accuracy on both sets. This system offers a practical tool for auditing AI-assisted scientific writing and improving the reliability of research literature.
- Quality assurance
Research
Hybrid‑Threat Intelligence: A Critical Review of Semantic Integration Challenges and the Role of the HIPSTer Ontological Framework
R. Andrew Paskauskas, Evaldas Bružė, Giedre Sabaliauskaite et al.
Journal of Intelligent Communication · 2026-05-09
This scoping review evaluates Open Source Intelligence (OSINT), Social Media Intelligence (SOCMINT), and Natural Language Processing (NLP) capabilities for detecting hybrid threats that span information, cyber, and physical domains. The authors identify a persistent 'semantic gap' in which defensive systems collect extensive data but lack integrated cross-domain reasoning to correlate cyber indicators with narrative manipulation campaigns. They assess the HIPSTer ontological framework as a solution targeting multilingual hybrid threats—particularly in Russian and Chinese contexts—achieving TRL-4 validation through semantic vectors and formal reasoning. The review also examines how European regulations including GDPR, the AI Act, and NIS2 shape operational architectures, concluding with a research agenda to advance European hybrid threat detection toward operational maturity.
- AI policy
- Enterprise
- Quality assurance
Research
SAFE Governance Standards Whitepaper: Governance Architecture for AI-Mediated Child Safeguarding
Nyree Ayne
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-09
This whitepaper introduces SAFE (Safeguarding Architecture for Foundational Environments), an independent governance framework specifically designed to fill the gap in verification and certification infrastructure for AI systems used in environments where children are present. The framework proposes structural separation among governance (SAFE Council), certification (SAFE Labs), and implementation (SAIL), along with a behavioural risk taxonomy (SAFE-001) focused on interaction trajectories rather than isolated content classification. The authors argue that independent certification infrastructure is a missing layer in current AI safety ecosystems and that SAFE can support regulatory interoperability, procurement benchmarking, and ongoing behavioural verification. This matters because it proposes a scalable, institutionally credible architecture for child safeguarding in AI-mediated environments that could inform policy and certification standards.
- Certifications
- AI policy
- Quality assurance
Research
Log analysis is necessary for credible evaluation of AI agents
Peter Kirgis, Sayash Kapoor, Stephan Rabanser et al.
arXiv · 2026-05-08
This paper argues that AI agent benchmarks that report only pass/fail outcomes are insufficient for credible evaluation, identifying three validity threats: score inflation/deflation from shortcuts and artifacts, failure to predict real-world utility due to scaffold limitations, and concealment of dangerous agent actions. The authors propose log analysis—systematic tracking of inputs, execution, and outputs—as a necessary complement to outcome metrics, and develop a taxonomy of evaluation threats alongside guiding principles for conducting such analysis. Applying their framework to the tau-Bench Airline benchmark, they find that pass^5 performance was under-elicited by nearly 50% and that deployment failure modes were invisible to outcome-only metrics. The work concludes with practical recommendations for benchmark creators, model developers, independent evaluators, and deployers to adopt log analysis in order to improve evaluation credibility.
- Quality assurance
- Certifications
Research
Human-LLM Dialogue Improves Diagnostic Accuracy in Emergency Care
Burcu Sayin, Ngoc Vo Hong, Ipek Baris Schlicht et al.
arXiv · 2026-05-08
MedSyn is an interactive system that lets physicians iteratively query an LLM with access to full clinical records while the physician initially sees only a patient's chief complaint. In a study with seven emergency medicine physicians across 52 MIMIC-IV cases, AI-assisted sessions improved residents' hard-case correctness from 0.589 to 0.734, with standardized any-match accuracy improving by 0.156 (p < 0.0001) and residents showing the largest F1 gain (Δ = 0.138; p < 0.0001). Dialogue analysis showed expertise-dependent querying strategies and increased cross-physician diagnostic concordance (Δ = 0.145; p < 0.0001). The findings suggest that interactive LLM support meaningfully enhances diagnostic reasoning, particularly for less experienced clinicians in high-stakes emergency settings.
- Workforce
- Quality assurance
Research
Teachers' Perceived Benefits and Risks of AI Across Fifty-Five Countries: An Audit of LLM Alignment and Steerability
Yan Tao, Olga Viberg, Deepak Varuvel Dennison et al.
arXiv · 2026-05-08
This study audits how well large language models (LLMs) reflect teachers' actual perceptions of AI's benefits and risks, using representative OECD TALIS survey data from 55 countries and territories. Eight state-of-the-art LLMs from four providers were benchmarked against this cross-national survey evidence, revealing that models compress country-level differences, overestimate both benefits and risks, and show limited improvement from identity prompting or enhanced reasoning. Because LLM-generated guidance increasingly shapes how teachers learn about and discuss AI, this misalignment poses real risks for global AI-in-education policy. The authors caution against substituting LLM outputs for direct teacher engagement in policy development, while noting some models partially capture cross-national ranking patterns useful for exploratory analysis.
- AI policy
- Workforce
Research
Measuring What Matters: Benchmarking Generative, Multimodal, and Agentic AI in Healthcare
Prasanna Desikan, Harshit Rajgarhia, Shivali Dalmia et al.
arXiv · 2026-05-08
This paper argues that current AI benchmarks in healthcare are inadequate for evaluating real-world clinical deployment, showing a systematic gap between high scores on narrow tasks (e.g., near-perfect on medical licensing exams) and much weaker performance on actual clinical workflows, including documentation (0.74–0.85), clinical decision support (0.61–0.76), and administrative tasks (0.53–0.63). The authors contend that most benchmarks test knowledge recall rather than reliability, safety, and clinical relevance under realistic conditions, creating a false sense of deployment readiness. They call for a principled framework for benchmark design to distinguish genuine model limitations from measurement failures, arguing this is essential before AI systems take on high-stakes clinical roles.
- Quality assurance
- Certifications
Research
Defense effectiveness across architectural layers: a mechanistic evaluation of persistent memory attacks on stateful LLM agents
Jun Wen Leong
arXiv · 2026-05-08
This paper evaluates six defenses across four architectural layers against persistent memory injection attacks on LLM agents, testing nine open-source models across 5,040 experimental runs. The core finding is that five of six defenses fail because they operate at the wrong architectural layer: input filters miss payloads that enter via RAG retrieval, while retrieval-level classifiers cannot distinguish injection from legitimate policy content. Only tool-gating at the memory layer (Memory Sandbox) reduces attack success rate to 0% for eight of nine models, but a reasoning-mode ablation reveals no single sandbox implementation is safe across both reasoning and non-reasoning model classes. A loaded-corpus frontier evaluation across 21 models and three providers finds that under realistic conditions, some models exfiltrate at up to 95% attack success rate, with nearly all OpenAI and Gemini models storing injected rules at 100% regardless of execution resistance—creating supply-chain risk in shared-memory deployments.
- Quality assurance
- Enterprise
Research
Can Language Models Identify Side Effects of Breast Cancer Radiation Treatments?
Natalie Seah, Danielle S. Bitterman, Daphna Spiegel et al.
arXiv · 2026-05-08
This study evaluates how well large language models (LLMs) can identify side effects of breast cancer radiation treatments, a task critical for informed consent and survivorship care. Using 21 breast cancer patient profiles and seven instruction-tuned LLMs tested across multiple prompting strategies, the researchers compared LLM outputs against a clinician-curated reference developed by more than seven breast radiation oncologists at two major academic medical centers. Key findings include that LLMs are sensitive to minor documentation changes, show trade-offs between precision and recall, and systematically under-recall rare and long-term side effects. Grounding LLM outputs in clinician-curated side effect lists substantially improved reliability, suggesting practical design guidance for safer, more informative oncology applications.
- Quality assurance
Research
A Semantic-Sampling Framework for Evaluating Calibration in Open-Ended Question Answering
Zhanliang Wang, Jiancong Xiao, Ruochen Jin et al.
arXiv · 2026-05-08
This paper introduces Sem-ECE (Semantic-Sampling Expected Calibration Error), a framework for evaluating how well large language models' confidence scores align with their actual accuracy in open-ended question answering. The framework samples model answers, groups them into semantic classes, and uses the resulting frequencies as a proxy for confidence, addressing shortcomings of logit-based, verbalized, and existing sampling-based calibration methods. The authors prove their two estimators are asymptotically unbiased and show experimentally across three QA benchmarks and five commercial LLMs that Sem-ECE outperforms verbalized confidence and existing sampling-based approaches. This matters for high-stakes deployment settings like medicine and law, where knowing whether a model's expressed confidence is trustworthy is critical to safe use.
- Quality assurance
Research
SecureForge: Finding and Preventing Vulnerabilities in LLM-Generated Code via Prompt Optimization
Houjun Liu, Lisa Einstein, John Yang et al.
arXiv · 2026-05-08
SecureForge is an automated pipeline that audits and reduces cybersecurity vulnerabilities in code generated by large language models (LLMs). The study finds that even frontier models asked to write secure code still produce verifiable vulnerabilities about 23% of the time across 250 benign coding prompts. SecureForge uses a Markovian sampling technique to build a diverse synthetic prompt corpus and then iteratively optimizes system prompts, achieving up to a 48% reduction in output vulnerabilities while also improving unit test success rates. The resulting secure system prompts transfer zero-shot to real-world coding agent deployments, offering a practical quality-assurance mechanism for AI-assisted software development.
- Quality assurance
- Enterprise
Research
Augmenting Human Evaluation with LLM Judges: How Many Human Reviews Do You Need?
Jane Paik Kim
arXiv · 2026-05-08
This paper addresses the lack of rigorous methodology for using large language models (LLMs) as evaluators alongside human raters. Rather than treating LLMs as full substitutes for human judgment, the authors reframe LLMs as auxiliary evaluators in a two-stage sampling design: LLM ratings are collected for all observations, while human ratings are gathered for a strategic subsample. They propose a doubly robust estimator from the missing data literature to combine these ratings and provide principled guidance on how many human reviews are needed to achieve a targeted level of statistical power, including allocating more human oversight to cases where LLM predictions are least reliable. This work is relevant to quality assurance and policy, as it provides formal study-design tools for validating AI evaluation benchmarks in high-stakes settings where human oversight requirements are otherwise undefined.
- Quality assurance
- AI policy
Research
Towards Apples to Apples for AI Evaluations: From Real-World Use Cases to Evaluation Scenarios
Yee-Yin Choong, Kristen Greene, Alice Qian et al.
arXiv · 2026-05-08
This paper addresses the challenge of inconsistent AI evaluation methodologies by proposing a structured, human-centered process for transforming high-level AI use cases into detailed, comparable evaluation scenarios. The authors introduce an AI Use Case Worksheet with six key elements—use case, sector, user, intended outcomes, expected impacts, and KPIs/metrics—and demonstrate it in the U.S. financial services sector, covering use cases such as cyber defense enablement, developer productivity, financial crime aggregation, SAR filing, credit memo generation, and internal call center support. A three-stage pipeline combining LLM prompting with iterative human reviews generates 107 scenarios from these use cases, supported by a validation rubric to assess scenario quality. The work aims to establish a more consistent and meaningful paradigm for human-centered AI evaluations, enabling more valid 'apples-to-apples' comparisons across AI systems.
- Quality assurance
- AI policy