News & Research
The latest AI research and news with real-world stakes. Each item is sourced, dated, summarized in plain English and tagged by impact area, and checked against its source before it appears.
Kind
Impact area
5802 items
- ResearcharXiv2026-05-28WP
The New Pro Se: Generative AI and the Surge in Federal Civil Self-Representation · Or Cohen-Sasson
Analyzing roughly 2.8 million federal civil filings from FY2008–2025, this paper finds that the pro se plaintiff rate rose from 11.33% pre-GenAI to 16.94% post-GenAI, a 5.61 percentage-point increase that survives trend and covariate adjustments. Using stylometric AI-detection indicators, the authors estimate that 13.9% of post-GenAI non-form complaints show AI-consistent drafting; these filings are more citation-dense and disproportionately filed by first-time rather than repeat litigants. Critically, AI-flagged complaints show no improvement in win rates — they are more likely to be dismissed and to terminate at earlier procedural stages — highlighting a gap between legal formality and legal efficacy. The findings raise significant questions about access to justice and increased court screening burdens as generative AI lowers barriers to filing but not to prevailing.
- ResearcharXiv2026-05-28EP
Evolutionary Rule Extraction from Corporate Default Prediction Models · Desirè Fabbretti, Matteo Pasquino, Elia Pacioni et al.
This study examines default prediction for 50,718 Italian SMEs over 2015–2024, comparing traditional logistic regression with machine learning (ML) classifiers and finding that ML models significantly outperform the benchmark in Balanced Accuracy and PR-AUC. To tackle the 'black box' problem, the authors introduce DEXiRE-EVO, a novel evolutionary rule extraction framework combining multi-objective optimization with the Contextual Importance and Utility (CIU) explainability method. The extracted rules surface economically meaningful drivers of financial distress—weak internal liquidity generation, capital erosion, high leverage, and operational inefficiency—alongside macroeconomic context. The work matters because it shows how explainable AI can meet regulatory transparency requirements in credit risk modeling while preserving strong predictive performance.
- ResearcharXiv2026-05-28QP
Inform, Coach, Relate, Listen: Auditing LLM Caregiving Support Roles · Drishti Goel, Agam Goyal, Veda Duddu et al.
This paper investigates how different conversational support roles assigned to large language models (LLMs) affect their safety profiles in informal caregiving contexts, specifically for Alzheimer's Disease and Related Dementias (ADRD). The researchers operationalize four support roles—Inform, Coach, Relate, and Listen—grounded in social support theory and evaluate three LLMs (GPT-4o-mini, Llama-3.1-8B-Instruct, and MedGemma-1.5-4b-it) on 5,000 real-world queries from online ADRD communities. They find that the assigned support role systematically shapes both the prevalence and composition of interactional risks, and that more directive, information-oriented roles are perceived as more helpful and trustworthy despite exhibiting elevated risk profiles. This quality–safety tension has direct implications for how LLM-based caregiving tools should be designed, audited, and deployed responsibly.
- ResearcharXiv2026-05-28WQ
How Coding Agents Fail Their Users: A Large-Scale Analysis of Developer-Agent Misalignment in 20,574 Real-World Sessions · Ningzhi Tang, Chaoran Chen, Gelei Xu et al.
This paper presents a large-scale observational study of 20,574 real-world coding-agent sessions across 1,639 repositories, identifying and categorizing how AI coding agents fail to align with developer intent. The researchers define misalignment as episodes made visible by developer pushback, classifying them by form, cause, cost, and resolution across IDE and CLI workflows. Key findings include that 90.50% of misalignment episodes impose effort and trust costs rather than irreversible damage, yet 91.49% of visible resolutions still require explicit user correction—indicating that agents rarely self-correct. The study also finds that while overall misalignment rates decline over time, constraint violations and inaccurate self-reporting grow in share, with implications for how coding agents should be trained, evaluated, and interfaced to better match real developer workflows.
- ResearcharXiv2026-05-28QP
FinGuard: Detecting Financial Regulatory Non-Compliance in LLM Interactions · Huaixia Dou, Jie Zhu, Minghao Wu et al.
FinGuard addresses a critical gap in AI safety for financial services: existing guard models rely on general harm taxonomies and miss violations rooted in specific financial regulations. The authors build a regulation-driven pipeline that processes regulatory documents directly to induce a financial compliance risk taxonomy and generate grounded training data, instantiated on Chinese financial regulations. They release FinGuard-Bench, the first benchmark for financial regulatory compliance detection with expert-annotated labels, and train FinGuard on Qwen3-8B using supervised fine-tuning and self-play reinforcement learning, achieving results that substantially outperform larger general-purpose models such as Qwen3.5-397B-A17B and GPT-5.1. The system also preserves general safety capabilities and can adapt to unseen institution-specific policies using only policy documents, making it practically relevant for financial institutions managing regulatory risk.
- ResearcharXiv2026-05-28WQ
Offloading Score: Measuring AI Reliance Through Counterfactual Workflows · Vishakh Padmakumar, Lujain Ibrahim, Zora Zhiruo Wang et al.
This paper introduces the 'offloading score,' a simulation-based metric that quantifies how much cognitive effort a user delegates to an AI tool by estimating a counterfactual workflow — what steps the user would have taken without the tool — and computing the fraction of steps the tool replaces. In a controlled study of 40 developers performing programming tasks, the offloading score detected a 43% increase in reliance under time pressure (p=0.018), while usage-based and self-reported measures failed to distinguish the conditions. The framework also identifies when high reliance may be inappropriate by pairing the score with task outcomes such as code understanding. The result is both a self-reflection instrument for individual users and a quantitative signal for AI agent designers seeking to mitigate overreliance.
- ResearcharXiv2026-05-28WP
Attention Asymmetry in AI Layoff Discourse on X: A Computational Analysis of Capital vs Labour Amplification · Joy Bose
This paper investigates whether pro-capital or pro-labour narratives about AI-driven layoffs receive more amplification on X (formerly Twitter). Across three studies using 763 tweets from 20 named public accounts, the researchers find that capital-aligned discourse (from tech executives and AI researchers) receives 4.18x the mean and 10.77x the median amplification of labour-aligned discourse, with the gap persisting at 2.69x even after normalising for follower count. The authors introduce two new metrics—the Amplification Ratio and Amplification Normalisation Index—to measure platform-level discourse inequality, and note the asymmetry did not replicate on Reddit, suggesting it may be specific to X's algorithmic architecture. These findings matter for understanding how platform design shapes public debate around AI-driven job displacement.
- ResearcharXiv2026-05-28P
Does Distributed Training Undermine Compute Governance? · Robi Rahman
This paper examines whether advances in distributed AI training algorithms could allow developers to circumvent compute governance regimes that assume frontier AI training requires large, easily detectable computing clusters. The authors argue that hardware could be structured across distributed nodes to evade registration and monitoring requirements tied to centralized datacenters. The paper evaluates the technical feasibility of such evasion and recommends countermeasures including whistleblowing programs, chip tracking, forensic accounting, and threshold-based monitoring of memory and compute resources.
- ResearcharXiv2026-05-28EQ
Causal Label Recovery in Payment Networks · Gaurav Dhama
This paper addresses a fundamental problem in payment-network fraud detection: the labels used to train models (chargebacks) are systematically distorted by four sequential impairments—authorization bias, unreported fraud, labeling delays, and label corruption from misuse or misclassification. The authors construct the Sequential Triply Robust (STR) estimator, which corrects for all four impairments simultaneously and achieves the semiparametric efficiency bound, meaning no estimator can have lower asymptotic variance. A key operational result is that the STR allows training on data that is days old rather than months old, decoupling model freshness from the chargeback maturity cycle, and it provably dominates naive chargeback-based training in mean squared error for any sample size. This matters for fraud-detection practitioners and payment enterprises because it provides a theoretically grounded, practically deployable path to better-calibrated models without waiting for full chargeback maturity.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-05-28EQCP
AI Agents and Sarbanes-Oxley Internal Control: Six Structural Gaps in Financial Reporting · Alex Li
This paper uses assumption-violation mapping to identify six structural gaps in Sarbanes-Oxley (SOX) Sections 302 and 404 compliance when AI agents participate in financial reporting. The central finding is that CEO 'reasonable assurance' certifications may be weakened when material reporting processes rely on decision logic that is opaque, non-deterministic, or not auditably reconstructable—a problem the authors call epistemic opacity in the certification chain. The paper maps COSO Internal Control Framework components to AI agent behaviors, presents three potential material-weakness scenarios, and proposes agent-specific COSO extensions. Unlike prior work treating AI as a compliance tool, this paper treats AI as an autonomous participant within the financial reporting chain, with significant implications for auditors, regulators, and executives.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-05-28EQCP
AI Agents and Sarbanes-Oxley Internal Control: Six Structural Gaps in Financial Reporting · Alex Li
This paper identifies six structural gaps in Sarbanes-Oxley (SOX) compliance that emerge when AI agents autonomously participate in financial reporting processes. Using assumption-violation mapping against SOX Sections 302 and 404, PCAOB standards, and the COSO Internal Control framework, the authors find that the CEO's 'reasonable assurance' certification may be weakened when material reporting decisions rely on opaque, non-deterministic, or non-reconstructable AI logic. The paper maps COSO components to AI agent behaviors, presents three potential material-weakness scenarios, and proposes agent-specific COSO extensions to address these gaps. The work is notable for treating AI as an autonomous participant within the financial reporting chain rather than merely a compliance tool.
- ResearcharXiv2026-05-28QCP
Fairness at a Glance: Can We Audit Model Fairness Before Training Completes? · Yuanhao Liu, Qi Cao, Huawei Shen et al.
This paper investigates whether AI model fairness can be assessed before a full training run completes, rather than only after. The authors find that while pre-training configurations alone cannot reliably predict final fairness outcomes, early-stage fairness metrics during training are strongly correlated with final fairness. Their approach enables early detection and stopping of unfair models, reducing training costs in fairness audit workflows by 73.8%. This has significant implications for making fairness auditing faster, cheaper, and more practical for AI developers.
- ResearchINTERNATIONAL JOURNAL OF MARKETING & HUMAN RESOURCE MANAGEMENT2026-05-28WEP
FIRM SIZE, DIGITAL SKILLS AND AI ADOPTION IN EUROPEAN ENTERPRISES: EVIDENCE FROM EUROSTAT DATA · Răzvan-George Cotescu
This paper examines AI adoption gaps across small, medium, and large enterprises in 27 EU member states using 2025 Eurostat data. Employing descriptive comparison, OLS regression, and country fixed-effects models, the study finds a clear firm-size gradient: medium-sized firms adopt AI at substantially higher rates than small firms, and large enterprises adopt AI at even higher rates. Crucially, this gap persists after controlling for country-specific digital environments, indicating the divide is driven by firm-level organizational and human-resource conditions rather than national context alone. The findings matter for enterprise policy because they suggest that easier access to AI tools does not produce equal adoption, and that smaller firms face structural disadvantages in integrating AI into their operations.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-05-28WEP
A REVIEW OF THE CURRENT STATE OF ARTIFICIAL INTELLIGENCE ADOPTION IN SOUTH AFRICAN CONSTRUCTION PROJECT MANAGEMENT · T.L NKOSI, SHP CHIKAFALIMANI
This literature review examines the current state of AI adoption in South African construction project management, finding that uptake remains at an emerging stage and is concentrated among large firms and megaprojects. Applications identified include predictive analytics for cost and risk management, AI-enabled scheduling, BIM-integrated machine learning, and drone-based site monitoring. Key barriers to broader adoption include high implementation costs, inadequate digital infrastructure, a shortage of skilled professionals, fragmented data, and limited regulatory support, with SMEs particularly disadvantaged. The authors conclude that growing digital advancement and industry awareness present opportunities for future AI integration, but structural and policy gaps must be addressed first.
- ResearcharXiv2026-05-28WE
My New Colleague, ChatGPT? How German Science Journalists Perceive and Use (Generative) Artificial Intelligence · Lars Guenther, Jessica Kunert, Bernhard Goodwin
This study examines how 30 German science journalists perceive and use generative AI in their work, drawing on the Technology Acceptance Model and semi-structured interviews. Journalists largely view AI as a productivity-enhancing colleague that can improve efficiency across news selection, production, and distribution, while insisting on human oversight for generative AI tools. The findings suggest that AI integration will likely not worsen the existing crisis in science journalism and may actually improve working conditions.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-05-28EQCP
Anchoring AI Proof Certificates to Clinical Data Standards: The ARCH Framework for Adaptive Regulatory Compliance and Human Oversight in Clinical Trials · Jessica Stuyvenberg
This working paper introduces the ARCH Framework (Adaptive Regulatory Compliance and Human Oversight), a technical specification for embedding AI proof certificates directly within the CDISC Unified Study Definitions Model (USDM) used in clinical trials. The framework defines a three-gate verification system—covering deterministic regulatory compliance, formal structural verification via Lean4, and human oversight attestation—each producing cryptographically anchored certificate objects. It also addresses risk-based quality management aligned with ICH E6(R3), continuous learning governance under FDA's PCCP pathway, EU AI Act Article 10 dataset provenance requirements, and a bi-temporal audit trail satisfying 21 CFR Part 11. The framework aims to enable real-time compliance verification during protocol authoring and multi-jurisdictional regulatory alignment without requiring new standards or infrastructure.
- ResearchFrontiers in Psychiatry2026-05-28WEQ
Clinician and simulated patient perspectives on ambient AI scribes in psychiatric consultations: a qualitative study · Syed Ali Bokhari, Faisal A. Nawaz, Firdous M. Usman et al.
This qualitative study explored how psychiatrists and simulated patients experienced ambient AI scribes during psychiatric consultations. Participants reported that AI scribes reduced documentation burden, enabled more authentic clinical presence, enhanced documentation quality through intelligent translation to psychiatric terminology, and prompted clinicians for missed assessments. However, the study also identified concerns around trust calibration, privacy, stigma, and psychiatric-specific consent considerations, with both groups expressing a preference for AI-assisted consultations under appropriate safeguards. The findings suggest AI scribes could meaningfully support patient-centered psychiatric care, but successful implementation requires transparent consent, clinician training, human oversight, and specialty-specific templates.
- ResearcharXiv2026-05-28QCP
The Verification Crisis: Expert Perceptions of GenAI Disinformation and the Case for Reproducible Provenance · Alexander Loth, Martín Munizaga Kappes, Marc‐Oliver Pahl
This paper presents findings from a longitudinal expert perception survey (N=21) involving AI researchers, policymakers, and disinformation specialists examining how Generative AI has shifted disinformation from manual fabrication to automated, large-scale manipulation. Experts identify large-scale text generation as posing a systemic risk of 'epistemic fragmentation' and 'synthetic consensus,' particularly in the political domain, while deepfake video presents more immediate shock value. The survey reveals skepticism toward technical detection tools, with experts instead favoring provenance standards and regulatory frameworks, and the authors argue that without standardized benchmarks and reproducibility checklists, tracking or countering synthetic media remains intractable. The paper proposes treating information integrity as infrastructure requiring rigorous data provenance and methodological reproducibility.
- ResearcharXiv2026-05-27EP
Paper Agents, Paper Gains: An Empirical Analysis of DeFi Investment Agents · Jay Yu, Amy Zhao, Danning Sui
This paper empirically examines DeFi investment agents—AI systems for autonomous on-chain crypto trading—which reached over USD 3 billion in combined token valuations since late 2024. Analyzing over 1,900 AI-tagged crypto projects and 11 Solana-based agent treasuries covering 925,323 token holders, the authors find that many deployments are little more than basic API integrations rather than truly autonomous systems, and that token holders collectively lost USD 191.7M while the top 1% of wallets captured 81.4% of all gains. Token valuations are weakly tied to treasury fundamentals, with market-cap-to-AUM ratios exceeding 10,000x, and median returns are negative on every platform with tokens declining 93% on average from all-time highs. The authors propose a maturity framework along three dimensions—autonomous execution, risk-adjusted profitability, and stakeholder alignment—to bridge the gap between current speculative deployments and future investment-grade AI agent systems.
- ResearcharXiv2026-05-27QP
Same Question, Different Source, Different Answer: Auditing Source-Dependence in Medical Multi-Source RAG · Yubo Li, Rema Padman, Ramayya Krishnan
This paper identifies a critical failure mode in retrieval-augmented generation (RAG) systems: the same question can yield different answers depending on which source document is retrieved, a problem invisible to standard single-gold-answer evaluation. The authors introduce TransplantQA, a benchmark of real transplant patient questions answered by grounding generation in multiple institutional handbooks, along with HERO-QA, a hierarchical retrieval strategy, and a structured-output judge that scores inter-source relationships using a validated 5-label taxonomy. Their findings show that better retrieval reveals far more disagreement across institutional sources than prior estimates suggested. The framework generalizes beyond transplant medicine to legal and educational RAG, positioning source-dependence auditing as a broader responsibility for deployed multi-source NLP systems.
- ResearcharXiv2026-05-27EP
The Importance of Out-of-Band Metadata for Safe Autonomous Agents: The Redpanda Agentic Data Plane · Tyler Akidau, Tyler Rockwood, Johannes Brüderl et al.
This paper introduces the Redpanda Agentic Data Plane (ADP), an architecture designed to make autonomous AI agents safer in enterprise environments by routing security-critical metadata—such as access policies, data classifications, and audit trails—through infrastructure channels that agents cannot read, modify, or bypass. The core insight is that agents, unlike human employees, are prone to hallucination, misinterpretation, and adversarial manipulation while also capable of causing damage at machine speed, making it unsafe to let agents self-enforce governance rules. ADP enforces security constraints at every stage of an agent's lifecycle—scoping data access, constraining actions during execution, and producing tamper-proof audit logs—entirely outside the agent's control. The authors demonstrate the architecture with a multi-agent portfolio rebalancing system that enforces per-client data scoping and trade approval thresholds, illustrating practical applicability in high-stakes enterprise settings.
- ResearcharXiv2026-05-27EQ
Hallucination Mitigation with Agentic AI, Nested Learning, and AI Sustainability via Semantic Caching · Diego Gosmar, Deborah A. Dahl
This paper presents a multi-agent AI pipeline designed to reduce hallucinations in large language model outputs without requiring model retraining. Using a HOPE-inspired Nested Learning architecture with semantic caching and a three-stage agentic review process, the system achieves Total Hallucination Score reductions of -31.3% to -35.9% across multiple evaluation configurations on a 310-prompt benchmark. Semantic caching achieves a 47.3% hit rate, cutting LLM invocations nearly in half and lowering energy and CO2 emissions, making the approach viable at production scale. The findings suggest that memory-augmented, multi-agent designs can simultaneously improve factual reliability, operational efficiency, and auditability in real-world deployments.
- ResearcharXiv2026-05-27QP
When Models Disagree: Rethinking LLM Evaluation for Public Comment Analysis · Aisha Najera, Alvin Moon, Vedant Srinivasan et al.
This paper examines the use of large language models (LLMs) to categorize public comments submitted to federal agencies, warning that standard accuracy-based evaluation fails to detect when different models organize the same public input in materially different ways. The authors propose an Interpretive Audit Pipeline that treats disagreement across multiple LLMs as a signal of interpretive complexity, directing human review toward the most ambiguous comments. Analyzing 1,260 comments on a USDA federal docket across four LLMs, they find that inter-model thematic divergence exceeds within-model prompt variation, and that a human annotator's revisions frequently introduced framings absent from the model ensemble's collective output. The paper argues that disagreement-based evaluation must complement accuracy metrics when LLMs are used to shape what policymakers see in public comment records.
- ResearcharXiv2026-05-27EQ
Measuring Real-World Prompt Injection Attacks in LLM-based Resume Screening · Mohan Zhang, Yuqi Jia, Zhen Tan et al.
This paper presents the first large-scale measurement study of prompt injection attacks in real-world LLM-based resume screening, analyzing approximately 200,000 resumes collected by hireEZ over multiple years. The authors developed tailored detectors that achieve high precision and outperform general-purpose tools, finding that roughly 1% of resumes contain hidden prompt injections, that the prevalence has grown noticeably over the past one to two years, and that more than 90% of injected prompts avoid explicit instructions. These findings provide the first empirical evidence of widespread prompt injection in a real-world LLM application, highlighting a concrete security risk for enterprise hiring systems that rely on AI-driven resume screening.
- ResearcharXiv2026-05-27QP
BEAMS: Benchmarking and Evaluating AI for Modeling and Simulation · Sara Metcalf, William Schoenberg
The BEAMS Initiative introduces a benchmarking framework for evaluating AI tools designed to support modeling and simulation tasks that inform real-world decision making. Using an open-source infrastructure (the sd ai project), the initiative implements automated tests across categories such as causal translation, model iteration, causal reasoning, conformance, model behavior explanation, and suggested model fixes. Results show that AI-enabled modeling tools perform better at discussion and basic qualitative tasks than at causal reasoning and quantitative error fixing, and no single large language model dominates across all engine types. The work matters because it establishes transparent, human-centered standards to guide responsible AI development in modeling and simulation rather than allowing these tools to replace human expertise.