News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
The Limits of AI-Driven Allocation: Optimal Screening under Aleatoric Uncertainty
Santiago Cortes-Gomez, Mateo Dulce Rubio, Carlos Patino et al.
arXiv · 2026-05-08
This paper examines the fundamental limits of machine learning-based resource allocation in policy and humanitarian settings, showing that even perfect predicted risk scores cannot eliminate misallocation due to irreducible (aleatoric) uncertainty about individual vulnerability status. The authors develop a two-stage framework that optimally combines algorithmic targeting with physical screening, finding that the best strategy screens units at the margin of algorithmic allocation while directly targeting the highest-risk individuals. Empirical analysis shows that efficiency gains from screening grow as population-level aleatoric uncertainty increases, with applications to income-based social protection programs and humanitarian demining in Colombia demonstrating real operational consequences of balancing screening costs against allocation efficiency.
- AI policy
Research
Creating and Evaluating K-12 GenAI Assessment Graders Through Context Engineering
Zewei Tian, Alex Liu, Lief Esbenshade et al.
arXiv · 2026-05-08
This paper evaluates large language model (LLM)-based graders for K-12 standardized assessments by measuring interrater agreement against human raters using Massachusetts Comprehensive Assessment System (MCAS) data across mathematics, science, and ELA subjects. Using metrics such as Quadratic Weighted Kappa (QWK) and Proportional Reduction in Mean-Squared Error (PRMSE), the study finds that models with more parameters—including Claude Sonnet 4, Haiku 4.5, GPT-5, and GPT-5 Mini—achieve substantial agreement with human raters in math and science, though performance varies in ELA. Teacher and student feedback shows strong acceptance of AI-generated narrative feedback but skepticism toward AI-assigned numerical scores, pointing to LLMs as more suitable for formative than summative evaluation. The authors conclude that hybrid models combining AI efficiency with teacher judgment can reduce workload, improve feedback quality, and support equitable assessment without displacing professional expertise.
- Quality assurance
- Workforce
Research
LLM Wardens: Mitigating Adversarial Persuasion with Third-Party Conversational Oversight
Lennart Wachowiak, Scott D. Blain, David Williams-King et al.
arXiv · 2026-05-08
This paper investigates how LLMs can be used as adversarial persuaders and proposes a 'warden' model—a secondary LLM that monitors human-AI conversations in real time and issues private, non-binding warnings when it detects manipulation. In a preregistered user study (N=120) across four decision-making scenarios, an adversarial LLM succeeded in steering user decisions 65.4% of the time, but adding a warden reduced that success rate to 30.4%, while only minimally affecting genuine interactions (8.6 percentage-point reduction). The researchers also released COAX-Bench, a simulation benchmark spanning 14 scenarios—including hiring, voting, and file access—where warden models reduced adversarial success from 34.7% to 12.3% across 16,212 simulated interactions. Critically, even warden models weaker than the adversary they oversee provided meaningful protection, suggesting a scalable approach to oversight of increasingly capable AI systems.
- AI policy
- Quality assurance
Research
What if AI systems weren't chatbots?
Sourojit Ghosh, Pranav Narayanan Venkit, Sanjana Gautam et al.
arXiv · 2026-05-08
This paper critiques the dominant chatbot paradigm in AI development, arguing it is not a neutral design choice but a sociotechnical configuration with broad structural downsides. The authors show that chatbot-based systems often fail users in complex or high-stakes contexts while projecting unwarranted confidence, and that normalizing chatbot-mediated interaction contributes to deskilling, homogenization of knowledge, and shifting expectations of expertise. At a societal level, the paper identifies labor displacement, concentration of economic power, and increased environmental costs as consequences of sustained investment in large-scale chatbot infrastructures. The authors call for alternative AI development paths emphasizing pluralistic system design, task-specific tools, and institutional safeguards to mitigate social and economic harm.
- Workforce
- AI policy
Research
Beyond Confidence: Rethinking Self-Assessments for Performance Prediction in LLMs
Sree Bhattacharyya, Samarth Khanna, Leona Chen et al.
arXiv · 2026-05-08
This paper challenges the common practice of using a single confidence score to predict when large language models (LLMs) will fail, proposing instead a multidimensional self-assessment framework inspired by cognitive appraisal theory from human psychology. The authors elicit six appraisal-based dimensions—including effort and ability—alongside traditional confidence, then evaluate their predictive power across 12 LLMs and 38 tasks in eight domains. They find that competence-related dimensions, especially effort and ability, consistently match or outperform confidence in predicting model errors, with effort producing less overoptimistic and more stable estimates across model sizes. These findings suggest that structured multidimensional self-assessment could meaningfully improve the reliability and safety of LLM deployment in real-world settings.
- Quality assurance
- Enterprise
Research
RuleSafe-VL: Evaluating Rule-Conditioned Decision Reasoning in Vision-Language Content Moderation
Zhifeng Lu, Dianyuan Wang, Yuhu Shang et al.
arXiv · 2026-05-08
RuleSafe-VL is a new benchmark designed to evaluate how well vision-language models (VLMs) reason through explicit policy rules in content moderation, rather than simply matching predefined labels. The benchmark formalizes 93 atomic rules and 92 typed rule relations derived from real platform moderation policies, producing 2,166 context-sensitive image-text cases across three high-risk policy families and four diagnostic tasks. Experiments on 10 frontier, open-source, and safety-oriented VLMs reveal that rule-relation recovery is the dominant bottleneck—the best model achieves only 64.8 Macro-F1, while some safety-oriented models fall below 7 Macro-F1—indicating that current models cannot reliably replicate the structured reasoning that sound content moderation requires. This work matters for quality assurance and policy because it exposes a critical gap between benchmark performance and genuine policy-compliant decision-making in automated moderation systems.
- Quality assurance
- AI policy
Research
SARC: A Governance-by-Architecture Framework for Agentic AI Systems
Gaston Besanson
arXiv · 2026-05-08
SARC is a runtime governance architecture for agentic AI systems that embeds compliance constraints directly into the agent execution loop rather than relying on prompts, dashboards, or after-the-fact documentation. The framework compiles constraint specifications into four enforcement sites—a Pre-Action Gate, Action-Time Monitor, Post-Action Auditor, and Escalation Router—and extends to multi-agent workflows through constraint propagation and attribution-preserving trace trees. In a reproducible synthetic evaluation across 50 seeds on a procurement task, SARC achieved zero hard-constraint violations under exact predicates and reduced soft-window overages by 89.5% compared to a policy-as-code-only baseline. This matters because it provides a principled architectural substrate for making regulatory obligations executable and auditable at runtime in regulated settings where post-hoc controls are structurally insufficient.
- AI policy
- Enterprise
Research
LLM hallucinations in the wild: Large-scale evidence from non-existent citations
Zhenyue Zhao, Yihe Wang, Toby Stuart et al.
arXiv · 2026-05-08
This large-scale audit of 111 million references across 2.5 million papers on arXiv, bioRxiv, SSRN, and PubMed Central finds a sharp rise in non-existent citations following widespread LLM adoption, with a conservative estimate of 146,932 hallucinated citations in 2025 alone. Hallucinated references are especially prevalent in fields with rapid AI uptake, in papers showing linguistic signatures of AI-assisted writing, and among small or early-career author teams. The study also finds that hallucinated citations disproportionately credit already prominent and male scholars, potentially reinforcing existing inequities in scientific recognition. Critically, preprint moderation and journal publication processes capture only a fraction of these errors, indicating that existing safeguards have been outpaced by the spread of hallucinated content.
- Quality assurance
- AI policy
Research
DRIP-R: A Benchmark for Decision-Making and Reasoning Under Real-World Policy Ambiguity in the Retail Domain
Hsuvas Borkakoty, Sebastian Pohl, Cheng Wang et al.
arXiv · 2026-05-08
DRIP-R is a benchmark designed to evaluate how LLM-based agents handle real-world retail policy ambiguities—situations where multiple valid interpretations exist and no single correct resolution can be defined. The benchmark pairs policy-ambiguous return scenarios with realistic customer personas, conversational simulation with tool-calling, and a multi-judge evaluation framework assessing policy adherence, dialogue quality, behavioral alignment, and resolution quality. Experiments reveal that frontier models fundamentally disagree when faced with identical ambiguous scenarios, demonstrating that policy ambiguity is a genuine and systematic challenge for LLM decision-making. This matters for enterprise deployments where AI agents handle consequential customer-facing tasks governed by imprecise policies.
- Enterprise
- Quality assurance
Research
Quality-Conditioned Agreement in Automated Short Answer Scoring: Mid-Range Degradation and the Impact of Task-Specific Adaptation
Abigail Victoria Gurin Schleifer, Moriah Ariely, Beata Beigman Klebanov et al.
arXiv · 2026-05-08
This study examines how well automated short answer scoring (ASAS) systems handle student responses of varying quality, comparing few-shot large language models (GPT and Claude), a fine-tuned BERT-based encoder, and a human expert on open-ended biology items. The key finding is that all AI models perform adequately on fully correct or fully incorrect responses but show substantial degradation in scoring accuracy on mid-range, partially correct responses — a pattern termed 'mid-range degradation.' The severity of this degradation is tied to the degree of task-specific adaptation: few-shot LLMs with minimal examples perform worst on mid-range responses, while fine-tuned encoder models perform best. The authors warn that this inequity in scoring accuracy disproportionately affects students with developing understanding, underscoring the need for quality-conditioned fairness evaluation in automated scoring.
- Quality assurance
- Certifications
Research
Safe, or Simply Incapable? Rethinking Safety Evaluation for Phone-Use Agents
Zhengyang Tang, Yi Zhang, Chenxin Li et al.
arXiv · 2026-05-08
This paper introduces PhoneSafety, a benchmark of 700 safety-critical decision points drawn from real phone interactions across more than 130 apps, designed to evaluate whether phone-use AI agents avoid harm due to genuine safety awareness or simple inability to act. The authors evaluate eight phone-use agents and find that stronger general phone-use ability does not reliably translate to safer choices at risky moments, and that failures to act are concentrated in more visually and operationally demanding settings—behaving as a capability signal rather than a safety signal. The work argues that a harmless outcome alone is insufficient evidence of safety, and that evaluations must distinguish between unsafe judgment and incapacity to act.
- Quality assurance
- Certifications
Research
The Moltbook Files: A Harmless Slopocalypse or Humanity's Last Experiment
William Brach, Federico Torrielli, Stine Lyngsø Beltoft et al.
arXiv · 2026-05-08
This paper introduces the Moltbook Files, a dataset of 232,000 posts and 2.2 million comments generated by OpenClaw AI agents on a Reddit-like platform over 12 days, released to study emergent behavior in AI agent populations. Analysis reveals that agents inadvertently post sensitive information such as API keys, passwords, and BIP39 seed phrases on a publicly indexed platform, and that fine-tuning a language model (Qwen2.5-14B-Instruct) on this data reduces truthfulness from 0.366 to 0.187 — though a size-matched Reddit dataset produces a comparable decrease. The authors conclude that Moltbook data represents more of a 'harmless slopocalypse' than a catastrophic risk, but identify tail risks including contamination of future training crawls and potential trait transfer to next-generation models. The study highlights the importance of control baselines when evaluating emergent misalignment in AI systems.
- Quality assurance
- AI policy
Research
Business Utility of Large Language Models as Exploratory Data Analysis Agents
Rafał Łabędzki, Patryk Miziuła, Hubert Rutkowski et al.
arXiv · 2026-05-08
This paper evaluates large language models (LLMs) as exploratory data analysis (EDA) agents in a business context, focusing on whether they are reliable enough for autonomous deployment. Using a supply chain simulation benchmark, the researchers tested 15 model-variant configurations across eight model families, scoring outputs against deterministic ground truth and introducing a 'Business utility' metric that combines mean performance with repeatability (variability discounting). Results show most configurations are not sufficiently reliable for autonomous EDA use even when average scores look acceptable, with GPT-5.4 at extra-high reasoning effort achieving the best profile (experiment-averaged mean score of 0.8748 and Business utility of 0.6952). The study argues that trustworthy EDA agents must be evaluated on average quality, repeatability, and condition sensitivity together, not average performance alone.
- Enterprise
- Quality assurance
Research
Topic Is Not Agenda: A Citation-Community Audit of Text Embeddings
Junseon Yoo
arXiv · 2026-05-08
This paper audits whether cosine similarity between text embeddings actually captures research-agenda relatedness in scientific literature, testing four state-of-the-art embedding models (Gemini, Qwen3-8B, Qwen3-0.6B, SPECTER2) against a citation graph of 3.58 million papers. The results show that while embeddings perform reasonably at the sub-field level (45–52% top-10 accuracy), they fail badly at the finer research-agenda level, with 8 out of every 10 retrieved papers being off-agenda across all four models and eight scientific domains. A simple citation-count reranking approach outperforms the best embedding-based retriever by roughly 9 percentage points at the agenda level, exposing a concrete and systematic failure mode in scientific retrieval-augmented generation (RAG) systems. This matters because RAG pipelines increasingly underpin enterprise knowledge tools and scientific search, and the findings reveal that embedding-based retrieval may silently surface topically similar but agenda-irrelevant content.
- Enterprise
- Quality assurance
Research
GAD in the Wild: Benchmarking Graph Anomaly Detection under Realistic Deployment Challenges
Jingjing Zhou, Shiyu Huang, Qing Qing et al.
arXiv · 2026-05-08
This paper introduces a benchmark for Graph Anomaly Detection (GAD) that evaluates models under realistic deployment conditions, including million-scale graphs, extreme anomaly scarcity (e.g., 0.1% anomaly ratios), and missing node attributes. Testing nine representative GAD models across five diverse graphs—including two industrial-scale datasets with over 3.7 million nodes—reveals that most GNN-based methods fail to scale due to memory constraints, detection performance collapses under realistic anomaly ratios often yielding zero recall, and reconstruction-based models are highly sensitive to attribute imputation strategies. The findings highlight a significant gap between academic benchmarks and real-world production environments, with implications for fraud detection and social platform governance. The benchmark is released as a diagnostic testbed to guide development of more robust and scalable GAD systems.
- Quality assurance
- Enterprise
Research
Beyond Single Ground Truth: Reference Monism as Epistemic Injustice in ASR Evaluation
Anna Seo Gyeong Choi, Maria Teleki, James Caverlee et al.
arXiv · 2026-05-08
This paper critiques how automatic speech recognition (ASR) systems are evaluated, arguing that the standard practice of using a single transcription convention as 'ground truth' — called reference monism — constitutes epistemic injustice toward speakers with atypical speech, particularly those with aphasia. Because transcription conventions (verbatim, non-verbatim, legal, etc.) encode normative assumptions about which speech features matter, different conventions produce different Word Error Rate (WER) scores for identical ASR output. Using AphasiaBank, the authors demonstrate empirically that WER varies meaningfully depending on which convention defines ground truth, and that speakers with aphasia are systematically penalized when clinically meaningful disfluencies are treated as errors. The paper proposes WER-Range — reporting performance across multiple legitimate conventions — and formalizes an Epistemic Injustice Distance (EID) metric to quantify the cost of reference monism.
- Quality assurance
- AI policy
Research
Digital twins and multimodal artificial intelligence in spine care: a scoping review of concepts, evidence, and translational barriers
Samer G. Salman, Rohan Phadke, Rahul Kumar et al.
Spine Deformity · 2026-05-08
This scoping review of 26 studies examines the state of multimodal AI, wearable monitoring, and digital twin concepts in spine care, finding that the field remains largely conceptual and unvalidated. Existing spine prediction models show only modest accuracy, imaging-based AI has weak links to patient outcomes like pain and disability, and no prospective studies demonstrate that digital twins improve clinical decision-making. The authors conclude that while personalized digital twin frameworks are promising, clinical readiness requires prospective validation, standardized data integration, and regulatory clarity. The findings are directly relevant to quality assurance and policy considerations around AI adoption in clinical settings.
- Quality assurance
- AI policy
- Certifications
Research
Does Artificial Intelligence Based Green Supply Chain Improve Sustainable Economic Performance? Evidence from Indonesian Firms
Rakhmawati Oktavianna, Benarda Mononym, Sri Nitta Crissiana Wirya Atmaja et al.
International Review of Management and Marketing · 2026-05-08
This study of Indonesian manufacturing firms finds that AI adoption significantly improves Green Supply Chain (GSC) practices, which in turn enhances sustainable economic efficiency. Using structural equation modeling (SEM/AMOS) with survey data, the research confirms that GSC mediates the relationship between AI adoption and sustainable economic outcomes. The findings suggest AI-powered supply chains can serve as a catalyst for national sustainability and decarbonization goals. Policymakers are urged to develop standardized sustainability policies, incentives, and metrics to scale these benefits.
- Enterprise
- AI policy
Research
Artificial Intelligence, Targeted Communication, and the Recruitment of Children in Armed Conflict
Arthur van Coller
South African Yearbook of International Law · 2026-05-08
This article examines how armed groups are using AI-driven digital tools—including algorithmic targeting, gamified content, social media echo chambers, and short-form video—to recruit children into armed conflict, replacing traditional coercive methods. The paper finds significant gaps in existing legal frameworks for regulating these digital recruitment strategies and assigns responsibility to technology companies alongside governments. The authors recommend legislative reform, multi-sector collaboration, and AI-based detection tools to better protect children in conflict-affected digital environments.
- AI policy
Research
Barriers and Facilitators to Patient Acceptance of Artificial Intelligence in Health Care: Systematic Review
Huiqin Shi, Jingying Huang, Yang Jin et al.
Journal of Medical Internet Research · 2026-05-08
This systematic review of 61 studies examines what helps or hinders patients from accepting AI in healthcare settings. Key barriers include perceived complexity, lack of trust in AI algorithms, reduced human interaction, privacy concerns, and cost, while facilitators include transparent data governance, explainable AI decisions, and human-centered design. By integrating two behavioral frameworks (UTAUT2 and TDF), the authors identified 25 behavior change techniques and 40 actionable intervention strategies for clinicians and administrators to improve patient-centered AI adoption. The findings offer a prioritized blueprint for healthcare organizations seeking to integrate AI tools more effectively into clinical practice.
- AI policy
- Enterprise
- Quality assurance
Research
Artificial Intelligence for Labor Market Analysis: Using BERT Models to align professional education and labor market demands
Rodrigo C. Michel, Yuri Lima, Cícero Braga et al.
arXiv · 2026-05-08
This study develops an AI framework using BERT-based Natural Language Processing models to extract required competencies from large-scale online job postings, with the goal of aligning vocational education and training (VET) curricula with current labor market demands. Fine-tuned transformer models—particularly BERTimbau and RoBERTa trained in Brazilian Portuguese—achieved F1-scores up to 0.76 and accuracy between 0.81 and 0.82. The findings demonstrate that linguistically adapted AI models can support data-driven curriculum design, helping VET institutions keep pace with labor market shifts driven by digitalization, automation, and AI. This matters for workforce development because it offers a scalable, automated method to close the gap between professional education offerings and actual employer needs.
- Workforce
- AI policy
Research
Redefining Accounting in the Digital Age: The Impact of Blockchain and Artificial Intelligence
Md. Kamrul Hassan Tuhin, Md. Taharim Ahmed, Md. Shahriya Mannan et al.
International Journal of Latest Technology in Engineering Management & Applied Science · 2026-05-08
This study examines how blockchain and artificial intelligence are transforming accounting practices, auditing, financial reporting, and accounting education. Using structured questionnaires from Chartered Accountants and professionals at leading audit firms, the research finds that integrating these technologies significantly improves audit quality, enhances transparency, and reduces operational costs. However, concerns around scalability, interoperability, and job displacement are identified as critical challenges. The authors recommend curriculum reform, professional re-skilling, and strategic adoption frameworks to support effective technology integration in the accounting profession.
- Workforce
- Enterprise
- Quality assurance
- Certifications
- AI policy
Research
Symmetric and asymmetric nexus between AI literacy, employee engagement, and work performance among academicians
Muhammad Asif Naveed, Muhammad Zaheer Asghar, Talha et al.
Discover Computing · 2026-05-08
This study of 300 Pakistani university academics finds that higher AI literacy is positively associated with better work performance (β=0.301), with employee engagement significantly mediating that relationship (β=0.25) and explaining nearly 40% of variance in work performance. Using PLS-SEM, fuzzy-set qualitative comparative analysis, and SHAP methods, the research shows that the combination of high AI literacy and high employee engagement is a sufficient condition for high work performance (consistency=0.875), with employee engagement emerging as the strongest individual predictor. The findings suggest that investing in AI education and literacy programs in academic workplaces can meaningfully boost both engagement and performance outcomes, with direct implications for institutional policy and workforce development.
- Workforce
- AI policy
Research
Controlled Agentic AI Systems: A Governance-Driven Architecture for Auditable and Reproducible Decision Pipelines
Tymoteusz Miller
Machine Learning and Knowledge Extraction · 2026-05-08
This paper presents Controlled Agentic AI Systems (CAIS), a formal architectural framework that embeds governance directly into AI decision pipelines as a deterministic operator rather than treating it as an afterthought. The framework integrates constraint specifications and a governance operator that transforms proposed actions into compliant ones, with formal audit trace semantics enabling deterministic replay of decision trajectories. Experiments in multi-agent and federated simulation environments show that embedding governance significantly reduces constraint violations, with projection-based repair outperforming approval-only strategies and achieving near-complete compliance without destabilizing system dynamics. This matters for regulated industries where auditability, reproducibility, and verifiable constraint adherence are legal or operational requirements.
- Enterprise
- Quality assurance
- Certifications
- AI policy
Research
Adaptive auditing of AI systems with anytime-valid guarantees
Siyu Zhou, Patrick Vossler, Venkatesh Sivaraman et al.
arXiv · 2026-05-07
This paper addresses the statistical challenges of adaptive auditing of generative AI systems, where evaluators dynamically choose which cases to test and when to stop based on interim results. The authors introduce a hypothesis-testing framework using Safe Anytime-Valid Inference (SAVI) and a 'testing by betting' approach that simultaneously tests two competing hypotheses: whether the AI system has no failure modes below a performance threshold, and whether the auditor has a strategy to uncover such failures. They prove that under a sufficiently powerful auditor, these hypotheses are asymptotically inverses, meaning passing a stringent audit can certify global robustness of the AI system. Empirically, the framework maintains valid type-I error control and can reach statistically rigorous conclusions with as few as 20 observations, making it practical for real-world AI evaluation.
- Quality assurance
- Certifications