News & Research
The latest AI research and news with real-world stakes. Each item is sourced, dated, summarized in plain English and tagged by impact area, and checked against its source before it appears.
Kind
Impact area
5802 items
- ResearcharXiv2026-06-01EQ
Monitoring Agentic Systems Before They're Reliable · Marisa Ferrara Boston, Glen Hanson, Effi Georgala et al.
This paper addresses how to monitor agentic AI systems during early-stage production deployment, before they are fully reliable. The authors propose a methodology that evaluates these systems across three quality dimensions (quality, suitability, efficiency) and three monitoring scopes (within-run, cross-run, structural), using variance as a key signal and severity classification adapted from Failure Mode and Effects Analysis (FMEA) to prioritize human review. Evaluated on a synthetic testbed of 220 runs across 120 document bundles with controlled error injection, the study finds that structural defects dominate failure modes and mask task-level errors, that different monitoring scopes surface distinct failure types (e.g., within-run CV=0.02 vs. cross-run CV=1.25), and that deterministic triage routes 97% of findings to automated tracking while reserving the 2% showing variable behavior for human investigation. The authors propose a maturity-staging model for monitoring and note that the taxonomy and severity model are transferable to document-driven, multi-stage agentic workflows in regulated industries, making this work directly relevant to quality assurance and enterprise deployment of AI systems.
- ResearcharXiv2026-06-01WE
Beyond One-shot: AI Agents for Learning in Field Experiments · Junjie Luo, Ritu Agarwal, Gordon Gao
This paper tests whether tool-augmented agentic AI can learn from prior field-experiment data to design better interventions in subsequent experiments. In a two-stage healthcare messaging study covering over 693,000 patient visits, an autonomous AI system that extracted principles from Stage 1 data generated message variants that outperformed those co-designed by human behavioral experts with a chatbot, with the best AI-generated message reaching a 69.8% click-through rate (+6.5 percentage points over baseline). The results indicate that performance gains come from domain-specific experimental data rather than general LLM reasoning ability, and that general behavioral theories do not transfer uniformly to specific healthcare contexts. The work suggests agentic AI can transform A/B testing from one-shot evaluation into a scalable, cumulative learning system for intervention design.
- ResearcharXiv2026-06-01QP
Are Algorithm Registers Transparent? Perspectives from Germany · Iman Peljto, Xenia Heilmann, Mattia Cerrato
This paper examines algorithm registers—public databases listing AI systems used in public administration—and evaluates whether existing German initiatives actually deliver meaningful transparency. Using a conceptual proposal by Alina Lorenz (2025) as a structured audit framework, the authors extract checklists of transparency goals and apply them to the two main German transparency initiatives, MaKI and Lernende Systeme. The audit finds that several adaptations are needed for these registers to function as effective transparency instruments, and the authors propose a visualization of transparency levels along with concrete action items for improving the platforms. The paper also makes the audit checklists publicly available to support practitioners designing or evaluating similar registers.
- ResearcharXiv2026-06-01QP
POIROT: Interrogating Agents for Failure Detection in Multi-Agent Systems · Iñaki Dellibarda Varela, R. Sendra-Arranz, Pablo Romero-Sorozabal et al.
POIROT is a protocol for detecting and diagnosing failures in multi-agent Large Language Model systems (LLM-MAS) by repurposing the system's own agents as a distributed diagnostic layer rather than relying on a centralized evaluator. The paper shows that POIROT outperforms single-LLM evaluator baselines, with performance gains that scale with problem complexity (OR = 1.60, p = 0.008), agent count, and fault dimensionality, and that these gains persist under compound fault conditions. The authors also release POIROT as an open-source library and introduce BLAME, a benchmark for fault attribution in safety-critical multi-agent systems. This work is directly relevant to quality assurance and policy, as it addresses both the technical challenge of auditing AI system behavior and the legal-regulatory gap created by emerging AI regulation around safety-critical deployments.
- ResearcharXiv2026-06-01QP
Cross-modal linkage risk in clinical vision-language models · Soroosh Tayebi Arasteh, Mahshad Lotfinia, Sven Nebelung et al.
This paper identifies and quantifies a privacy risk in clinical vision-language models (VLMs): because these models learn a shared embedding space from paired chest X-rays and radiology reports, a de-identified radiograph can be re-linked to its original narrative report using cosine similarity alone. Evaluated on over 406,000 image-report pairs from MIMIC-CXR and CheXpert Plus, the authors found that the best-performing VLM retrieved the correct report at 50 times chance in a pool of 10,000 candidates, with the risk rising as models became more clinically specialized. To mitigate this without retraining, the authors applied differentially private optimization solely to the alignment projection heads, reducing Recall@1 by 61.8% at N=10,000 while preserving image classification performance (macro AUROC dropping only from 79.63% to 79.43%). The findings have direct implications for data-sharing policies and privacy safeguards around clinical AI systems that handle separated image and report archives.
- ResearcharXiv2026-06-01Q
Who Annotates in NLP? A Large-scale Assessment of Human Annotation Reporting between 2018 and 2025 · Maria Kunilovskaya, Gagan Bhatia, Lisa Sophie Albertelli et al.
This paper conducts the first large-scale audit of how NLP research papers report on human annotation practices, covering 1,603 papers and 2,667 annotation tasks published at ACL-venue conferences between 2018 and 2025. The authors introduce a unified taxonomy of annotation-reporting practices and validate an LLM-assisted extraction pipeline that achieves human-comparable agreement (Krippendorff's alpha of 0.606 vs. 0.585 for human-human agreement). They find that while papers commonly report operational details like recruitment and annotator expertise, critical validity information—such as annotator training, language proficiency, compensation, socio-demographics, and inter-annotator agreement—is frequently omitted, especially in model-evaluation studies. The work establishes a scalable auditing framework and minimum reporting recommendations to make annotation more reliable and reproducible.
- ResearcharXiv2026-06-01EQ
AgentRedBench: Dynamic Redteaming and Integration-Aware Defense for LLM Agents over SaaS Integrations · Hiskias Dingeto, William Leeney
AgentRedBench introduces a dynamic red-teaming benchmark (AGENTREDBENCH) of 215 subtle prompt-injection scenarios spanning 24 enterprise SaaS integrations (e.g., Gmail, Salesforce, Jira) and five attack types to evaluate how vulnerable LLM-based tool-use agents are to indirect prompt injection. Testing an eight-model panel from Anthropic, OpenAI, and Google, the authors find no-guard attack success rates ranging from 32% to 81%, revealing a serious and underappreciated production security threat. The paper also introduces AGENTREDGUARD, a purpose-built defense model that reduces attack success by 75–77 percentage points across three model families while maintaining a 0.0% false-positive rate on real-benign data, outperforming all tested open-source baselines. These results matter for enterprise AI deployments that rely on third-party integrations, highlighting both the inadequacy of existing guards and a path toward more robust agent security.
- Newsimportai.substack.com2026-06-01WP
Import AI 459: AI oversight is difficult; scaling laws for protein folding models; and pricing the extinction risk of AI systems
Import AI (Jack Clark) covers a new paper from economists at the University of Virginia, Anthropic, and the Bank of Canada estimating that the U.S. AI economy reached roughly $250 billion in nominal GDP in 2025 and is growing at approximately 2,600 percent per year in quality-adjusted real terms—yet remains largely invisible in conventional GDP statistics because per-unit prices for AI capability fall nearly as fast as quality-adjusted output rises. The newsletter warns that this measurement gap is especially alarming because, unlike semiconductors or the internet, AI may substitute rather than complement human labor at scale, meaning policymakers could be blindsided by a labor-tax-base shock. The authors recommend that statistical agencies develop AI satellite accounts, generate better primary data on training versus inference compute, and incorporate AI productive-capacity measurements into medium-term economic projections. Clark also covers an Australian government official's call for economists to formally price existential risk from AI, a UK AI Security Institute paper on the difficulties of automated alignment oversight, and Biohub's release of a new protein-structure prediction model aimed at cancer research.
- ResearcharXiv2026-06-01QP
Better with Experience: Self-Evolving LLM Agents for Evidence-Grounded Health Community Notes · Zihang Fu, Fanxiao Li, Jianyang Gu et al.
EvoNote is an agentic LLM framework that generates evidence-grounded Community Notes to correct health misinformation on social platforms, improving over time by storing and reusing lessons from prior correction episodes via a fine-grained memory system. Evaluated on MM-HealthCN, a 1,200-instance multimodal benchmark, EvoNote-generated notes were preferred over human-written notes in 89.6% of cases under a human-validated judge, and produced helpful corrections for 82.0% of posts lacking a crowd verdict. The system also cuts median correction time from over 13 hours to under 2 minutes. These results position self-evolving note generation as a scalable approach to health misinformation governance on social platforms.
- ResearcharXiv2026-06-01QP
Do Gender Cues Affect LLM Value Trade-offs? Evidence from a Controlled Decision Benchmark · Yangyang Liu, Dong Yu, Pengyuan Liu
This paper introduces the Realistic Value Decision Benchmark (RVDB), a controlled benchmark designed to test whether gender cues alter decision-making in large language models (LLMs) across seven models. The authors find that explicit gender cues cause systematic but bounded decision flips, with a consistent asymmetry favoring decisions proposed by male roles over female roles, while models themselves often attribute these flips to non-gender factors or claim no influence. Gender effects concentrate near ambiguous value boundaries and in higher-severity decision contexts, suggesting gender acts as a local boundary-shifting factor rather than a global override. The findings highlight that LLM gender bias can be behaviorally present yet self-concealed, motivating rigorous behavioral audits rather than relying on model self-explanation.
- ResearcharXiv2026-06-01QP
Model Multiplicity and Predictive Arbitrariness in Recidivism Risk Assessment · Ashwin Singh, Carlos Castillo
This paper investigates 'model multiplicity' in recidivism risk assessment — the phenomenon where many equally accurate machine learning models can produce different predictions for the same individual, raising fairness and arbitrariness concerns. The authors construct a dataset of thousands of inmate releases, learn interpretable models that improve accuracy and reduce error-rate disparities, and derive a tight lower bound on predictive agreement across any finite set of models. Their key empirical finding is that structural diversity among similarly accurate models does not necessarily translate into severe predictive arbitrariness in practice, and that a simple policy of assigning each individual the lowest risk score across models effectively addresses the problem. These results matter for the design and governance of AI-based decision support tools used in high-stakes criminal justice settings.
- ResearcharXiv2026-06-01QP
Aligning Data-Driven Predictors with Allocation: A Decision-Focused Approach to Survival Analysis · Itai Zilberstein, Ioannis Anagnostides, Tuomas Sandholm
This paper exposes a critical misalignment between standard survival-model metrics (such as the C-index) and the actual goal of organ allocation: even highly accurate predictors optimized for standard metrics can yield outcomes no better than random selection when used for allocation decisions. To close this gap, the authors introduce a decision-focused learning framework that optimizes Normalized Discounted Cumulative Gain (NDCG), proving that this metric translates into performance guarantees for allocation and also addressing the challenge of right-censored data. Empirically, applied to historical US heart transplant data, their bootstrapping approach improves NDCG of baseline models by 50–100%, which the authors project translates to tens of thousands of additional life years gained annually. The work has broad implications for any high-stakes automated decision-making system that chains predictive models to downstream allocation or resource-assignment policies.
- ResearcharXiv2026-06-01EQ
BADGER: Bridging Agentic and Deterministic Evaluation for Generative Enterprise Reasoning · Shannon Serrao, Soumitra Chatterjee, Dorina Strori et al.
BADGER is a unified evaluation framework developed at Merkle for assessing enterprise AI systems that combine natural language-to-SQL translation with multi-step agentic reasoning pipelines. The framework introduces a hybrid execution accuracy metric (Hybrid-EX) that uses an LLM to resolve column-aliasing and numeric-tolerance issues before applying deterministic cell-level scoring; validated on 150 human-annotated industry queries, Hybrid-EX achieves a Cohen's kappa of 0.717 and 87.3% balanced accuracy, outperforming six competing frameworks. BADGER also integrates existing agentic evaluation tools (RAGAS, G-Eval, and agent benchmark metrics) alongside a novel Excess Tool Usage metric into a single pipeline that runs within a client's governed data environment. The framework is designed as a continuous evaluation backbone for production enterprise AI systems rather than a one-time quality gate, addressing a gap left by academic benchmarks such as Spider and BIRD.
- ResearcharXiv2026-06-01QC
Automated Essay Scoring and Language Certification: Assessing Generalizability, Agreement and Validity for French · Rodrigo Wilkens, Rémi Cardon, Vincent Folny et al.
This paper addresses Automated Essay Scoring (AES) for French by proposing an enhanced version of the argument-based validation (ABV) framework that goes beyond minimalist benchmarking. The enhanced framework incorporates fairness analysis, correlations with linguistic features, prediction error evaluation, and model agreement compared with human raters. The authors compare 8 model architectures on a corpus of 27,000 exam essays (with 2 raters each) and a generalization corpus of 961 essays (with at least nine raters each), demonstrating how the ABV framework reveals both capabilities and pitfalls of AES models in high-stakes language testing contexts. The work advances both the methodology for validating AES systems and the state-of-the-art specifically for French language assessment.
- ResearcharXiv2026-06-01EQ
A Structured Benchmark for Text-Guided Anomaly Detection: When Language Stops Conditioning the Decision · Stefano Samele, Eugenio Lomurno, Teodora Jovanovic et al.
This paper introduces TGAD (Text-Guided Anomaly Detection), a structured benchmark revealing that current multimodal vision-language models used for industrial anomaly detection respond only superficially to textual instructions. The authors test three model paradigms across progressively demanding scenarios—prompt sensitivity, component-level instruction following, and a new Assembled Panel Dataset—finding that language rarely conditions decisions in meaningful ways, with performance collapsing dramatically (in one case below chance at 31.5 I-AUROC) when both defect-type and component-location knowledge are required. The findings suggest that existing benchmarks inherited from unimodal settings significantly overstate text-guided capabilities, and that reliable language-controlled industrial inspection systems do not yet exist. This has direct implications for deploying AI-based quality inspection in manufacturing, where operators need to trust that natural-language instructions actually constrain model behavior.
- ResearcharXiv2026-06-01QC
An NLP-Driven Framework for Curriculum-Labor Market Alignment: Schema-Constrained LLM Extraction, ESCO-Anchored Semantic Matching, and Multi-Dimensional Gap Quantification · Sherzod Turaev, Mary John, Mamoun Awad et al.
This paper presents a four-stage NLP framework for measuring how well university curricula match labor-market skill demands. It uses a two-model large-language-model ensemble with schema-constrained prompting to extract competency records from course syllabi, aligns them to the ESCO v1.2.1 occupational taxonomy using Sentence-BERT semantic matching, and quantifies supply-demand gaps with Cohen's kappa reliability metrics. Applied to the ABET-accredited BSc Computer Science program at UAE University, the framework identifies meaningful skill gaps—25.0% in general and transversal skills and 13.8% in algorithms and computational theory—while finding a near-zero 1.8% gap in AI and data science, providing actionable evidence for curriculum reform and accreditation quality assurance.
- ResearcharXiv2026-06-01EQ
Does Compression Preserve Uncertainty? A Unified Benchmark for Quantized and Sparse LLMs via Conformal Prediction · Yujia Tong, Yuxi Wang, Yunyang Wan et al.
This paper investigates whether common model compression techniques—quantization and pruning—preserve not just accuracy but also uncertainty quantification in large language models (LLMs). Using conformal prediction as a rigorous, distribution-free measure, the authors benchmark 12 LLMs across various compression settings and five NLP tasks, finding that compression frequently decouples accuracy from uncertainty, that larger models handle compression-induced uncertainty better than smaller ones, and that uncertainty inflation tends to occur suddenly rather than gradually. The results argue that accuracy-alone evaluations are insufficient for deployment readiness of compressed LLMs and that uncertainty-aware benchmarking should become a standard part of compression pipelines.
- ResearcharXiv2026-06-01P
Argument Collapse: LLMs Flatten Long-Form Public Debate · Yekyung Kim, Yapei Chang, Chau Minh Pham et al.
This paper investigates 'argument collapse,' the tendency of LLM-generated essays to converge on a narrow set of arguments, sub-arguments, and structural patterns compared to human writing. Analyzing 1,039 human responses from New York Times debates, 448 human responses from Boston Review forums, and 23,384 LLM-generated essays, the authors find that 65.3% of human main arguments are unique within a debate versus only 3.4% of LLM main arguments, and that 41.0% of human sub-arguments are unique compared to 9.1% from LLMs. LLMs also favor generalized, hedged sub-arguments and a fixed essay arc, while humans produce more concrete, topic-specific content. The findings raise concerns that widespread LLM use in drafting public-facing arguments could homogenize discourse and reduce the diversity of perspectives in public debate.
- ResearcharXiv2026-06-01EQ
Easier to Mislead Than to Correct: Harmful and Beneficial Revision in LLM Conformity · Jiaming Qu, Lucheng Fu, Yibo Hu
This paper investigates 'conformity' in large language models — the tendency to change a correct answer simply because simulated peers agree on a different one. Through controlled experiments across four open-weight LLMs and seven QA datasets, the authors find that peer agreement is much more effective at misleading initially correct models than at correcting initially wrong ones, and that authority labels cause models to favor endorsed answers regardless of accuracy. Notably, common reasoning interventions like chain-of-thought and reflection do not reliably reduce these harmful revisions, suggesting that multi-agent LLM systems need answer verification mechanisms rather than simple aggregation.
- ResearcharXiv2026-06-01QP
Identifying High-Confidence Social Biases in LLMs for Trustworthy Conversational Tutoring Agents · Aitor Arronte Alvarez, Naiyi Xie Fincham
This paper investigates social biases in large language models (LLMs) used as conversational tutoring agents, finding that these models struggle significantly more to detect stereotypical biases in naturalistic tutoring contexts than in standard benchmark evaluations. The researchers developed a new dataset generation method that embeds controlled bias into realistic student-AI tutor interactions, then assessed multiple LLMs' bias detection ability, confidence, and reasoning through computational and human evaluations. A key finding is that state-of-the-art LLMs are overconfident in their incorrect assessments of biased statements, and that this overconfidence directly shapes the reasoning and feedback delivered to learners—posing meaningful risks in educational settings. The study concludes with implications for mitigating biased, overconfident behavior in LLM-based tutoring systems.
- ResearcharXiv2026-06-01WP
The Main Barrier to AI Adoption in the Public Sector Is Lack of Training: How a Structured Method Accompanied Productivity Gains in Two Brazilian Government Cases · Vinicius Santana Gomes
This paper argues that the primary barrier to generative AI adoption in the Brazilian public sector is a lack of structured training rather than technological limitations. The authors developed a four-layer pedagogical methodology and applied it in two government audit and internal control units during 2024–2025. Official indicators from the Brazilian Federal District's Electronic Information System show that average document processing time fell by 18.2% in one unit and by 50% in another, while the second unit also saw an 85% increase in technical-report production, issued 286 formal recommendations, and analyzed matters valued at US$94.8 million. The findings suggest the training method is portable across agencies, compatible with data-protection requirements, and feasible under budget constraints using free AI models.
- ResearcharXiv2026-06-01EQ
Compliance-Scored Best-of-N Guardrail Orchestration for Multimodal Document Generation in Payments Dispute Defense · Nataraj Agaram Sundar, Tejas Morabia
This paper introduces a guardrail orchestration layer for high-stakes enterprise document generation—specifically payments dispute defense summaries—that combines parallel multi-candidate generation with a compliance scoring mechanism for early exit. The system integrates PII detection, content moderation, schema validation, and domain-specific rules into a unified pipeline, replacing fragmented sequential steps. In operational evaluations, the framework achieved 91% compliance within 20 seconds across 5 generation attempts, and dispute defense summaries produced using it showed statistically significant win-rate improvements of +11.0 percentage points overall (95% CI [6.6, 15.5], p < 0.001) and +7.5 percentage points for adjusted item-not-received cases (95% CI [0.2, 15.7], p = 0.045) compared to controls. The work also reports Responsible-AI evidence-quality signals from 770 generated-evidence reviews and documents reproducibility boundaries through scoring logic and operational evidence.
- ResearchAdvances in Economics Management and Political Sciences2026-06-01WP
Influence of Artificial Intelligence in the Labor Market · Yuehan Cai
This systematic literature review examines how artificial intelligence reshapes labor markets by synthesizing four core theoretical mechanisms: substitution, complementarity, new task creation, and skill mismatch. The paper finds that AI primarily drives substitution effects in the short term but generates complementary and creative effects over the long term, with significant variation across countries, industries, and regions. Skill mismatch is identified as the central contradiction in workforce transformation. The authors highlight gaps in existing research around local empirical evidence and micro-level task mechanisms, aiming to inform adaptive policy formulation.
- ResearchJournal of technology management & innovation2026-06-01WEP
Optimisation of Administrative Processes Through Artificial Intelligence: Analysis of Adoption and Trust in Peruvian Companies in The Telecommunications Sector · Arody Tesen Amancio, Julissa Diaz Otiniano, Liz Pacheco-Pumaleque
This study examines how Peruvian telecommunications companies adopt AI and what factors drive or hinder that adoption. Using structural equation modeling and employee survey data, the research finds that AI security positively and significantly affects adoption (p=0.000), and that employees' perceived ability directly influences their perceived ease of use (p=0.003). AI adoption in turn significantly impacts product innovation, process innovation, and AI-driven marketing (all p=0.000), suggesting that trust-building and workforce capacity development are critical levers for business efficiency gains in emerging markets.
- ResearchInternational Journal of Foreign Trade and International Business2026-06-01WEP
Artificial intelligence-enabled demand forecasting and supply-chain resilience among export-oriented firms in Italy: Evidence from industrial districts · Marco Ferretti, Giulia Romano, Luca Pietrangeli
This study of 398 export-oriented Italian SMEs across six major industrial districts finds that AI-enabled demand forecasting improves forecast accuracy by 15.8 percentage points and reduces stockout frequency by 8.4 percentage points among adopting firms. Supply-chain resilience fully mediates the relationship between AI adoption and export intensity, meaning AI's export benefits flow entirely through improved resilience rather than direct effects. The research highlights significant district-level variation in AI uptake and calls for targeted policy interventions including PNRR co-investment and Transizione 5.0 tax incentive redesign to support SME digital transformation.