News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
How Coding Agents Fail Their Users: A Large-Scale Analysis of Developer-Agent Misalignment in 20,574 Real-World Sessions
Ningzhi Tang, Chaoran Chen, Gelei Xu et al.
arXiv · 2026-05-28
This paper presents a large-scale observational study of 20,574 real-world coding-agent sessions across 1,639 repositories, identifying and categorizing how AI coding agents fail to align with developer intent. The researchers define misalignment as episodes made visible by developer pushback, classifying them by form, cause, cost, and resolution across IDE and CLI workflows. Key findings include that 90.50% of misalignment episodes impose effort and trust costs rather than irreversible damage, yet 91.49% of visible resolutions still require explicit user correction—indicating that agents rarely self-correct. The study also finds that while overall misalignment rates decline over time, constraint violations and inaccurate self-reporting grow in share, with implications for how coding agents should be trained, evaluated, and interfaced to better match real developer workflows.
- Workforce
- Quality assurance
Research
FinGuard: Detecting Financial Regulatory Non-Compliance in LLM Interactions
Huaixia Dou, Jie Zhu, Minghao Wu et al.
arXiv · 2026-05-28
FinGuard addresses a critical gap in AI safety for financial services: existing guard models rely on general harm taxonomies and miss violations rooted in specific financial regulations. The authors build a regulation-driven pipeline that processes regulatory documents directly to induce a financial compliance risk taxonomy and generate grounded training data, instantiated on Chinese financial regulations. They release FinGuard-Bench, the first benchmark for financial regulatory compliance detection with expert-annotated labels, and train FinGuard on Qwen3-8B using supervised fine-tuning and self-play reinforcement learning, achieving results that substantially outperform larger general-purpose models such as Qwen3.5-397B-A17B and GPT-5.1. The system also preserves general safety capabilities and can adapt to unseen institution-specific policies using only policy documents, making it practically relevant for financial institutions managing regulatory risk.
- AI policy
- Quality assurance
Research
Offloading Score: Measuring AI Reliance Through Counterfactual Workflows
Vishakh Padmakumar, Lujain Ibrahim, Zora Zhiruo Wang et al.
arXiv · 2026-05-28
This paper introduces the 'offloading score,' a simulation-based metric that quantifies how much cognitive effort a user delegates to an AI tool by estimating a counterfactual workflow — what steps the user would have taken without the tool — and computing the fraction of steps the tool replaces. In a controlled study of 40 developers performing programming tasks, the offloading score detected a 43% increase in reliance under time pressure (p=0.018), while usage-based and self-reported measures failed to distinguish the conditions. The framework also identifies when high reliance may be inappropriate by pairing the score with task outcomes such as code understanding. The result is both a self-reflection instrument for individual users and a quantitative signal for AI agent designers seeking to mitigate overreliance.
- Workforce
- Quality assurance
Research
Attention Asymmetry in AI Layoff Discourse on X: A Computational Analysis of Capital vs Labour Amplification
Joy Bose
arXiv · 2026-05-28
This paper investigates whether pro-capital or pro-labour narratives about AI-driven layoffs receive more amplification on X (formerly Twitter). Across three studies using 763 tweets from 20 named public accounts, the researchers find that capital-aligned discourse (from tech executives and AI researchers) receives 4.18x the mean and 10.77x the median amplification of labour-aligned discourse, with the gap persisting at 2.69x even after normalising for follower count. The authors introduce two new metrics—the Amplification Ratio and Amplification Normalisation Index—to measure platform-level discourse inequality, and note the asymmetry did not replicate on Reddit, suggesting it may be specific to X's algorithmic architecture. These findings matter for understanding how platform design shapes public debate around AI-driven job displacement.
- Workforce
- AI policy
Research
Does Distributed Training Undermine Compute Governance?
Robi Rahman
arXiv · 2026-05-28
This paper examines whether advances in distributed AI training algorithms could allow developers to circumvent compute governance regimes that assume frontier AI training requires large, easily detectable computing clusters. The authors argue that hardware could be structured across distributed nodes to evade registration and monitoring requirements tied to centralized datacenters. The paper evaluates the technical feasibility of such evasion and recommends countermeasures including whistleblowing programs, chip tracking, forensic accounting, and threshold-based monitoring of memory and compute resources.
- AI policy
Research
Causal Label Recovery in Payment Networks
Gaurav Dhama
arXiv · 2026-05-28
This paper addresses a fundamental problem in payment-network fraud detection: the labels used to train models (chargebacks) are systematically distorted by four sequential impairments—authorization bias, unreported fraud, labeling delays, and label corruption from misuse or misclassification. The authors construct the Sequential Triply Robust (STR) estimator, which corrects for all four impairments simultaneously and achieves the semiparametric efficiency bound, meaning no estimator can have lower asymptotic variance. A key operational result is that the STR allows training on data that is days old rather than months old, decoupling model freshness from the chargeback maturity cycle, and it provably dominates naive chargeback-based training in mean squared error for any sample size. This matters for fraud-detection practitioners and payment enterprises because it provides a theoretically grounded, practically deployable path to better-calibrated models without waiting for full chargeback maturity.
- Enterprise
- Quality assurance
Research
AI Agents and Sarbanes-Oxley Internal Control: Six Structural Gaps in Financial Reporting
Alex Li
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-28
This paper uses assumption-violation mapping to identify six structural gaps in Sarbanes-Oxley (SOX) Sections 302 and 404 compliance when AI agents participate in financial reporting. The central finding is that CEO 'reasonable assurance' certifications may be weakened when material reporting processes rely on decision logic that is opaque, non-deterministic, or not auditably reconstructable—a problem the authors call epistemic opacity in the certification chain. The paper maps COSO Internal Control Framework components to AI agent behaviors, presents three potential material-weakness scenarios, and proposes agent-specific COSO extensions. Unlike prior work treating AI as a compliance tool, this paper treats AI as an autonomous participant within the financial reporting chain, with significant implications for auditors, regulators, and executives.
- Certifications
- AI policy
- Quality assurance
- Enterprise
Research
AI Agents and Sarbanes-Oxley Internal Control: Six Structural Gaps in Financial Reporting
Alex Li
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-28
This paper identifies six structural gaps in Sarbanes-Oxley (SOX) compliance that emerge when AI agents autonomously participate in financial reporting processes. Using assumption-violation mapping against SOX Sections 302 and 404, PCAOB standards, and the COSO Internal Control framework, the authors find that the CEO's 'reasonable assurance' certification may be weakened when material reporting decisions rely on opaque, non-deterministic, or non-reconstructable AI logic. The paper maps COSO components to AI agent behaviors, presents three potential material-weakness scenarios, and proposes agent-specific COSO extensions to address these gaps. The work is notable for treating AI as an autonomous participant within the financial reporting chain rather than merely a compliance tool.
- Certifications
- AI policy
- Enterprise
- Quality assurance
Research
Fairness at a Glance: Can We Audit Model Fairness Before Training Completes?
Yuanhao Liu, Qi Cao, Huawei Shen et al.
arXiv · 2026-05-28
This paper investigates whether AI model fairness can be assessed before a full training run completes, rather than only after. The authors find that while pre-training configurations alone cannot reliably predict final fairness outcomes, early-stage fairness metrics during training are strongly correlated with final fairness. Their approach enables early detection and stopping of unfair models, reducing training costs in fairness audit workflows by 73.8%. This has significant implications for making fairness auditing faster, cheaper, and more practical for AI developers.
- Quality assurance
- Certifications
- AI policy
Research
FIRM SIZE, DIGITAL SKILLS AND AI ADOPTION IN EUROPEAN ENTERPRISES: EVIDENCE FROM EUROSTAT DATA
Răzvan-George Cotescu
INTERNATIONAL JOURNAL OF MARKETING & HUMAN RESOURCE MANAGEMENT · 2026-05-28
This paper examines AI adoption gaps across small, medium, and large enterprises in 27 EU member states using 2025 Eurostat data. Employing descriptive comparison, OLS regression, and country fixed-effects models, the study finds a clear firm-size gradient: medium-sized firms adopt AI at substantially higher rates than small firms, and large enterprises adopt AI at even higher rates. Crucially, this gap persists after controlling for country-specific digital environments, indicating the divide is driven by firm-level organizational and human-resource conditions rather than national context alone. The findings matter for enterprise policy because they suggest that easier access to AI tools does not produce equal adoption, and that smaller firms face structural disadvantages in integrating AI into their operations.
- Enterprise
- Workforce
- AI policy
Research
A REVIEW OF THE CURRENT STATE OF ARTIFICIAL INTELLIGENCE ADOPTION IN SOUTH AFRICAN CONSTRUCTION PROJECT MANAGEMENT
T.L NKOSI, SHP CHIKAFALIMANI
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-28
This literature review examines the current state of AI adoption in South African construction project management, finding that uptake remains at an emerging stage and is concentrated among large firms and megaprojects. Applications identified include predictive analytics for cost and risk management, AI-enabled scheduling, BIM-integrated machine learning, and drone-based site monitoring. Key barriers to broader adoption include high implementation costs, inadequate digital infrastructure, a shortage of skilled professionals, fragmented data, and limited regulatory support, with SMEs particularly disadvantaged. The authors conclude that growing digital advancement and industry awareness present opportunities for future AI integration, but structural and policy gaps must be addressed first.
- Enterprise
- Workforce
- AI policy
Research
My New Colleague, ChatGPT? How German Science Journalists Perceive and Use (Generative) Artificial Intelligence
Lars Guenther, Jessica Kunert, Bernhard Goodwin
arXiv · 2026-05-28
This study examines how 30 German science journalists perceive and use generative AI in their work, drawing on the Technology Acceptance Model and semi-structured interviews. Journalists largely view AI as a productivity-enhancing colleague that can improve efficiency across news selection, production, and distribution, while insisting on human oversight for generative AI tools. The findings suggest that AI integration will likely not worsen the existing crisis in science journalism and may actually improve working conditions.
- Workforce
- Enterprise
Research
Anchoring AI Proof Certificates to Clinical Data Standards: The ARCH Framework for Adaptive Regulatory Compliance and Human Oversight in Clinical Trials
Jessica Stuyvenberg
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-28
This working paper introduces the ARCH Framework (Adaptive Regulatory Compliance and Human Oversight), a technical specification for embedding AI proof certificates directly within the CDISC Unified Study Definitions Model (USDM) used in clinical trials. The framework defines a three-gate verification system—covering deterministic regulatory compliance, formal structural verification via Lean4, and human oversight attestation—each producing cryptographically anchored certificate objects. It also addresses risk-based quality management aligned with ICH E6(R3), continuous learning governance under FDA's PCCP pathway, EU AI Act Article 10 dataset provenance requirements, and a bi-temporal audit trail satisfying 21 CFR Part 11. The framework aims to enable real-time compliance verification during protocol authoring and multi-jurisdictional regulatory alignment without requiring new standards or infrastructure.
- Certifications
- Quality assurance
- AI policy
- Enterprise
Research
Clinician and simulated patient perspectives on ambient AI scribes in psychiatric consultations: a qualitative study
Syed Ali Bokhari, Faisal A. Nawaz, Firdous M. Usman et al.
Frontiers in Psychiatry · 2026-05-28
This qualitative study explored how psychiatrists and simulated patients experienced ambient AI scribes during psychiatric consultations. Participants reported that AI scribes reduced documentation burden, enabled more authentic clinical presence, enhanced documentation quality through intelligent translation to psychiatric terminology, and prompted clinicians for missed assessments. However, the study also identified concerns around trust calibration, privacy, stigma, and psychiatric-specific consent considerations, with both groups expressing a preference for AI-assisted consultations under appropriate safeguards. The findings suggest AI scribes could meaningfully support patient-centered psychiatric care, but successful implementation requires transparent consent, clinician training, human oversight, and specialty-specific templates.
- Workforce
- Quality assurance
- Enterprise
Research
The Verification Crisis: Expert Perceptions of GenAI Disinformation and the Case for Reproducible Provenance
Alexander Loth, Martín Munizaga Kappes, Marc‐Oliver Pahl
arXiv · 2026-05-28
This paper presents findings from a longitudinal expert perception survey (N=21) involving AI researchers, policymakers, and disinformation specialists examining how Generative AI has shifted disinformation from manual fabrication to automated, large-scale manipulation. Experts identify large-scale text generation as posing a systemic risk of 'epistemic fragmentation' and 'synthetic consensus,' particularly in the political domain, while deepfake video presents more immediate shock value. The survey reveals skepticism toward technical detection tools, with experts instead favoring provenance standards and regulatory frameworks, and the authors argue that without standardized benchmarks and reproducibility checklists, tracking or countering synthetic media remains intractable. The paper proposes treating information integrity as infrastructure requiring rigorous data provenance and methodological reproducibility.
- AI policy
- Quality assurance
- Certifications
Research
Paper Agents, Paper Gains: An Empirical Analysis of DeFi Investment Agents
Jay Yu, Amy Zhao, Danning Sui
arXiv · 2026-05-27
This paper empirically examines DeFi investment agents—AI systems for autonomous on-chain crypto trading—which reached over USD 3 billion in combined token valuations since late 2024. Analyzing over 1,900 AI-tagged crypto projects and 11 Solana-based agent treasuries covering 925,323 token holders, the authors find that many deployments are little more than basic API integrations rather than truly autonomous systems, and that token holders collectively lost USD 191.7M while the top 1% of wallets captured 81.4% of all gains. Token valuations are weakly tied to treasury fundamentals, with market-cap-to-AUM ratios exceeding 10,000x, and median returns are negative on every platform with tokens declining 93% on average from all-time highs. The authors propose a maturity framework along three dimensions—autonomous execution, risk-adjusted profitability, and stakeholder alignment—to bridge the gap between current speculative deployments and future investment-grade AI agent systems.
- Enterprise
- AI policy
Research
Same Question, Different Source, Different Answer: Auditing Source-Dependence in Medical Multi-Source RAG
Yubo Li, Rema Padman, Ramayya Krishnan
arXiv · 2026-05-27
This paper identifies a critical failure mode in retrieval-augmented generation (RAG) systems: the same question can yield different answers depending on which source document is retrieved, a problem invisible to standard single-gold-answer evaluation. The authors introduce TransplantQA, a benchmark of real transplant patient questions answered by grounding generation in multiple institutional handbooks, along with HERO-QA, a hierarchical retrieval strategy, and a structured-output judge that scores inter-source relationships using a validated 5-label taxonomy. Their findings show that better retrieval reveals far more disagreement across institutional sources than prior estimates suggested. The framework generalizes beyond transplant medicine to legal and educational RAG, positioning source-dependence auditing as a broader responsibility for deployed multi-source NLP systems.
- Quality assurance
- AI policy
Research
The Importance of Out-of-Band Metadata for Safe Autonomous Agents: The Redpanda Agentic Data Plane
Tyler Akidau, Tyler Rockwood, Johannes Brüderl et al.
arXiv · 2026-05-27
This paper introduces the Redpanda Agentic Data Plane (ADP), an architecture designed to make autonomous AI agents safer in enterprise environments by routing security-critical metadata—such as access policies, data classifications, and audit trails—through infrastructure channels that agents cannot read, modify, or bypass. The core insight is that agents, unlike human employees, are prone to hallucination, misinterpretation, and adversarial manipulation while also capable of causing damage at machine speed, making it unsafe to let agents self-enforce governance rules. ADP enforces security constraints at every stage of an agent's lifecycle—scoping data access, constraining actions during execution, and producing tamper-proof audit logs—entirely outside the agent's control. The authors demonstrate the architecture with a multi-agent portfolio rebalancing system that enforces per-client data scoping and trade approval thresholds, illustrating practical applicability in high-stakes enterprise settings.
- Enterprise
- AI policy
Research
Hallucination Mitigation with Agentic AI, Nested Learning, and AI Sustainability via Semantic Caching
Diego Gosmar, Deborah A. Dahl
arXiv · 2026-05-27
This paper presents a multi-agent AI pipeline designed to reduce hallucinations in large language model outputs without requiring model retraining. Using a HOPE-inspired Nested Learning architecture with semantic caching and a three-stage agentic review process, the system achieves Total Hallucination Score reductions of -31.3% to -35.9% across multiple evaluation configurations on a 310-prompt benchmark. Semantic caching achieves a 47.3% hit rate, cutting LLM invocations nearly in half and lowering energy and CO2 emissions, making the approach viable at production scale. The findings suggest that memory-augmented, multi-agent designs can simultaneously improve factual reliability, operational efficiency, and auditability in real-world deployments.
- Enterprise
- Quality assurance
Research
When Models Disagree: Rethinking LLM Evaluation for Public Comment Analysis
Aisha Najera, Alvin Moon, Vedant Srinivasan et al.
arXiv · 2026-05-27
This paper examines the use of large language models (LLMs) to categorize public comments submitted to federal agencies, warning that standard accuracy-based evaluation fails to detect when different models organize the same public input in materially different ways. The authors propose an Interpretive Audit Pipeline that treats disagreement across multiple LLMs as a signal of interpretive complexity, directing human review toward the most ambiguous comments. Analyzing 1,260 comments on a USDA federal docket across four LLMs, they find that inter-model thematic divergence exceeds within-model prompt variation, and that a human annotator's revisions frequently introduced framings absent from the model ensemble's collective output. The paper argues that disagreement-based evaluation must complement accuracy metrics when LLMs are used to shape what policymakers see in public comment records.
- AI policy
- Quality assurance
Research
Measuring Real-World Prompt Injection Attacks in LLM-based Resume Screening
Mohan Zhang, Yuqi Jia, Zhen Tan et al.
arXiv · 2026-05-27
This paper presents the first large-scale measurement study of prompt injection attacks in real-world LLM-based resume screening, analyzing approximately 200,000 resumes collected by hireEZ over multiple years. The authors developed tailored detectors that achieve high precision and outperform general-purpose tools, finding that roughly 1% of resumes contain hidden prompt injections, that the prevalence has grown noticeably over the past one to two years, and that more than 90% of injected prompts avoid explicit instructions. These findings provide the first empirical evidence of widespread prompt injection in a real-world LLM application, highlighting a concrete security risk for enterprise hiring systems that rely on AI-driven resume screening.
- Enterprise
- Quality assurance
Research
BEAMS: Benchmarking and Evaluating AI for Modeling and Simulation
Sara Metcalf, William Schoenberg
arXiv · 2026-05-27
The BEAMS Initiative introduces a benchmarking framework for evaluating AI tools designed to support modeling and simulation tasks that inform real-world decision making. Using an open-source infrastructure (the sd ai project), the initiative implements automated tests across categories such as causal translation, model iteration, causal reasoning, conformance, model behavior explanation, and suggested model fixes. Results show that AI-enabled modeling tools perform better at discussion and basic qualitative tasks than at causal reasoning and quantitative error fixing, and no single large language model dominates across all engine types. The work matters because it establishes transparent, human-centered standards to guide responsible AI development in modeling and simulation rather than allowing these tools to replace human expertise.
- Quality assurance
- AI policy
Research
AIRGuard: Guarding Agent Actions with Runtime Authority Control
Suliu Qin, Haomin Zhuang, Yujun Zhou et al.
arXiv · 2026-05-27
AIRGuard is a runtime security layer for tool-using AI agents that enforces least-privilege authorization before any external action—such as file reads, API calls, or script executions—actually executes. The paper identifies a failure mode called 'authority confusion,' where attacker-controlled context can steer an agent's legitimate access rights into harmful side effects without producing any obviously forbidden output. On the AgentTrap benchmark, AIRGuard reduces attack success rates from 36.3% to 5.5% for Sonnet 4.6, while preserving 76.0% of benign utility on DTAP-150 compared to 52.0% for ARGUS and 42.0% for MELON. The results show that a dedicated runtime authority-control layer substantially outperforms prompt-only defenses, making it a meaningful advance for securing agentic AI systems.
- Quality assurance
- Enterprise
Research
Political Neutrality as Balanced Approval: A Large-Scale Human Evaluation of AI Responses
Jonathan Stray, David Zhai Yang, Steven Luo et al.
arXiv · 2026-05-27
This paper proposes a formal definition of AI political neutrality as 'balanced approval'—where an AI response maximizes and balances approval across groups with opposing viewpoints—and tests it via a large-scale human study. The authors release the PARETO dataset, comprising 7,434 participants and 208,152 evaluations of AI responses to controversial U.S. political issues, with prompts drawn from Reddit and responses from frontier models including GPT, Gemini, Claude, Llama, and Grok. Key findings show that neutral responses achieving high approval on both sides are attainable, that default outputs from GPT, Gemini, Claude, and Llama lean liberal while Grok does not, and that politically charged prompts are harder to answer neutrally than neutral ones. The benchmark and dataset provide a rigorous, empirically testable framework for measuring and improving AI political neutrality without assuming a single left-right axis.
- AI policy
Research
Hallucination Detection-Guided Preference Optimization for Clinical Summarization
Shamanth Kuthpadi Seethakantha, Dung Ngoc Thai, Vara Prasad Gudi et al.
arXiv · 2026-05-27
This paper introduces two methods—HDSR and HDSR-PL—to reduce hallucinations in AI-generated clinical note summaries. HDSR uses hallucination detectors at inference time to iteratively revise summaries toward factual accuracy, while HDSR-PL converts those refinement trajectories into preference pairs for fine-tuning language models. Experiments on real-world clinical notes from MIMIC-IV-Note v2.2 show that HDSR reduces hallucinations by 24% and HDSR-PL by 48% in Llama-3.1-8B-Instruct, while preserving fluency, coherence, and relevance. These results matter for healthcare AI because reducing unsupported or incorrect statements in clinical summaries is a prerequisite for safe deployment in high-stakes medical settings.
- Quality assurance