News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
Medical Context Distorts Decisions in Clinical Vision Language Models
David Restrepo, Ira Ktena, Maria Vakalopoulou et al.
arXiv · 2026-05-17
This paper investigates the reliability of vision-language models (VLMs) proposed for clinical decision support, focusing on chest X-ray interpretation using the MIMIC-CXR dataset. The authors identify three failure modes: over-reliance on text rather than images, spurious influence from irrelevant clinical history, and sensitivity to minor prompt reformulations that can reverse correct predictions. Their systematic experiments show that VLM decisions are dominated by the text modality even when visual evidence is available, and that irrelevant reports heavily sway model outputs. The findings highlight the need for explicit safeguards and stress-testing before deploying these models in clinical settings.
- Quality assurance
- Certifications
Research
ADR: An Agentic Detection System for Enterprise Agentic AI Security
Chenning Li, Pan Hu, Justin Xu et al.
arXiv · 2026-05-17
The ADR (Agentic AI Detection and Response) system is presented as the first large-scale, production-proven enterprise framework for securing AI agents operating through the Model Context Protocol (MCP). It addresses three core challenges—limited observability of agent reasoning, insufficient robustness of static defenses, and high inference costs—through a three-component architecture: a telemetry sensor, a red-teaming explorer, and a two-tier online detector. Deployed at Uber for over ten months across more than 7,200 hosts and processing over 10,000 agent sessions daily, ADR uncovered hundreds of credential exposures across 26 categories and achieved 97.2% precision on a shift-left prevention layer. On the introduced benchmark ADR-Bench (302 tasks, 17 techniques, 133 MCP servers), ADR achieved zero false positives while detecting 67% of attacks, outperforming three state-of-the-art baselines by 2–4x in F1-score.
- Enterprise
- Quality assurance
Research
Position: Age Estimation Models Do Not Process Biometric Data
Nikita Marshalkin
arXiv · 2026-05-17
This position paper investigates whether neural networks that estimate a person's age from photographs incidentally process biometric data in a legally meaningful sense. The authors empirically evaluate 14 age estimation models across 3 face verification benchmarks, finding that these models fall orders of magnitude short of identification thresholds, meaning they cannot identify individuals. The distinction matters because biometric data processing triggers consent requirements under GDPR, statutory damages under BIPA, and high-risk AI classification under the EU AI Act. The authors call on researchers to be transparent about what systems store and can do, and urge regulators to differentiate between transient processing during inference and biometric template storage.
- AI policy
- Certifications
Research
Is VLA Reasoning Faithful? Probing Safety of Chain-of-Causation in Autonomous Driving Models
Nicanor Mayumu, Xiaoheng Deng, Patrick Mukala
arXiv · 2026-05-17
This paper presents the first systematic evaluation of whether the reasoning outputs of Vision-Language-Action (VLA) models used in autonomous driving actually reflect their decisions. Studying 300 inferences from the Alpamayo-R1-10B model across 100 PhysicalAI-AV scenarios, the authors find that overall reasoning fidelity is only 42.5%, with trajectories being 97.7% fragile under mild visual perturbations and reasoning-action consistency averaging just 48.3%—including cases where the model claims to stop but continues driving. These findings reveal a critical safety gap between a model's stated rationale and its actual behavior, motivating the authors' proposed four-component safety architecture to address faithfulness failures in autonomous driving systems.
- Quality assurance
- AI policy
Research
Artificial Intolerance: Stigmatizing Language in Clinical Documentation Skews Large Language Model Decision-Making
Jen-tse Huang, Didi Zhou, Faith Kamau et al.
arXiv · 2026-05-17
This study investigates whether large language models (LLMs) used in clinical decision support inherit and propagate stigmatizing language (SL) found in human-authored clinical notes. The researchers systematically evaluated nine frontier LLMs across four stigmatized medical conditions, finding that all models showed substantial bias—stigmatizing language in clinical vignettes significantly skewed model outputs toward less aggressive patient management, with even a single stigmatizing sentence sufficient to alter decisions. Standard mitigation strategies such as Chain-of-Thought reasoning and self-debiasing showed limited effectiveness, as models struggled to explicitly identify SL while remaining implicitly influenced by it. The findings highlight a critical fairness and robustness vulnerability in clinical NLP systems, raising urgent concerns about the automation of health disparities.
- Quality assurance
- AI policy
Research
Stop Drawing Scientific Claims from LLM Social Simulations Without Robustness Audits
Jinyi Ye, Lei Cao, Ding Chen et al.
arXiv · 2026-05-17
This paper argues that scientific conclusions drawn from large language model (LLM) social simulations are only as credible as the robustness checks supporting them. Through two case studies—a repeated Prisoner's Dilemma and an echo chamber simulation—the authors show that minor perturbations in agent persona format, game-instruction framing, network homophily, and hub assignment can shift cooperation rates by up to 76 percentage points and produce significant changes in polarization metrics, with sensitivity varying dramatically across model families. The authors introduce TRAILS (Taxonomy for Robustness Audits In LLM Simulations), a three-level audit framework covering agent, interaction, and system design choices. They call for robustness audits to become a mandatory validation step before LLM social simulations are used to explain social mechanisms, evaluate interventions, or inform policy decisions.
- AI policy
- Quality assurance
Research
Beyond Model Readiness: Institutional Readiness for AI Deployment in Public Systems
Erika Fille Legara, Elmo Domino Jose, Paula Joy Martinez
arXiv · 2026-05-17
This paper introduces Institutional Alignment Readiness (IAR), a five-dimensional framework for assessing whether public-sector institutions are ready to deploy AI systems—not just whether the AI models themselves are technically sound. Drawing on two anonymized cases from a large public education system (an image-based anthropometric screening tool and a speech-analysis system for early learning risk identification), the authors show that technically viable AI systems can fail to advance to broader rollout due to institutional gaps in approvals, data arrangements, human oversight, fiscal continuity, and legal clarity. IAR covers institutional and operational compatibility, data ecosystem maturity, human oversight capacity, fiscal sustainability, and regulatory alignment, and is designed to complement existing model evaluation frameworks by assessing the receiving institution rather than the AI artifact alone. The framework is especially targeted at resource-constrained settings where the gap between technical viability and responsible deployment is most acute.
- AI policy
- Quality assurance
Research
"AAB Global AI Literacy Assessment and Credentialing Registry Dataset v1.0"
Winnie Han, Lei Xu
IEEE DataPort · 2026-05-17
This dataset from the AI Assessment Board (AAB) provides a structured global registry of AI literacy assessment, credentialing, certification, competition, and benchmarking initiatives drawn from testing organizations, governments, universities, nonprofits, and workforce training systems. It is designed to support comparative research, policy analysis, and standards development by documenting how AI literacy and competency are measured across different regions, learner groups, and institutional contexts. The resource enables ecosystem mapping and cross-context comparison of credential types, competency domains, delivery formats, and scoring models without including any personally identifiable or secure assessment content. It is especially relevant to ongoing efforts to establish consistent standards for AI literacy credentialing and workforce skill validation.
- Certifications
- AI policy
- Workforce
- Quality assurance
Research
Forced Convergence: How Artificial Intelligence Drives Democratic and Authoritarian Economies Toward a Common Sovereignty Architecture
Oguike Ahaneku
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-17
This paper argues that the EU and China, despite their ideological differences, are converging on structurally similar 'sovereignty architectures' for AI-mediated economic activity, driven not by ideological affinity but by a common fiscal pressure: AI-mediated cognitive labor performed on foreign-owned infrastructure erodes the labor-based tax revenues that European welfare states depend upon. The paper identifies four convergence mechanisms—mandatory data localization, jurisdictional capture of the deployment layer, regulatory substitutes for capital controls, and state-coordinated infrastructure procurement—and argues that the ongoing EU-US corporate tax conflict over US tech firms is an empirical precursor already demonstrating these dynamics at smaller scale. The core implication is that the economic topology required to retain sovereignty over AI-mediated sectors is substantially invariant across political systems, and that the transatlantic relationship will be reshaped in ways current policy discourse does not anticipate.
- AI policy
- Enterprise
Research
Forced Convergence: How Artificial Intelligence Drives Democratic and Authoritarian Economies Toward a Common Sovereignty Architecture
Oguike Ahaneku
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-17
This paper argues that the EU and China, despite their profound political differences, are converging on functionally similar 'sovereignty architectures' for governing AI-driven economic activity. The convergence is driven not by ideological alignment but by a shared structural fiscal problem: AI-mediated cognitive labor performed on foreign-owned infrastructure erodes tax bases, threatening European welfare-state sustainability. The paper identifies four convergence mechanisms—mandatory data localization, jurisdictional capture of the deployment layer, regulatory substitutes for capital controls, and state-coordinated infrastructure procurement—and treats the ongoing EU-US corporate tax-avoidance conflict as an empirical precursor to this dynamic. The authors conclude that the economic topology required to retain sovereignty over AI-mediated sectors is substantially invariant across political systems, with major implications for the transatlantic relationship.
- AI policy
- Enterprise
Research
Scaling-up Mental Health Services Artificial Intelligence: Regulatory Pathways and Case Study
Cole Hooley, Alice C. Schwarze, Zachary M. Boyd
Journal of Technology in Human Services · 2026-05-17
This paper examines how governments are attempting to regulate AI tools used in behavioral and mental health services, a domain where adoption is outpacing oversight. Through a policy case analysis of a state office of artificial intelligence that developed initial regulatory guidance for behavioral health, the authors reconstruct key decision points, stakeholder tensions, and tradeoffs faced by regulators. Three regulatory pathways are identified and compared: minimal market-based regulation, risk-only harm-avoidance frameworks, and risk-benefit regulation aimed at maximizing benefits while minimizing harms. The findings offer practical lessons for policymakers navigating the complex challenge of responsibly scaling AI in mental health services.
- AI policy
- Workforce
Research
PluRule: A Benchmark for Moderating Pluralistic Communities on Social Media
Zoher Kachwala, Bao Tran Truong, Rasika Muralidharan et al.
arXiv · 2026-05-16
PluRule is a new multimodal, multilingual benchmark designed to test whether AI models can moderate pluralistic social media communities — platforms where each community defines its own norms and rules. The benchmark covers 13,371 rule violations across 1,989 Reddit communities, spanning 2,885 rules in 9 languages, with the task framed as identifying which specific rule a comment violates. Results show that even state-of-the-art models like GPT-5.2 perform only marginally better than a trivial baseline, with larger models and added context providing little improvement. This highlights that community-specific content moderation remains a fundamental unsolved challenge for current AI systems.
- Quality assurance
- AI policy
Research
Why Do Safety Guardrails Degrade Across Languages?
Max Zhang, Ameen Patel, Sang T. Truong et al.
arXiv · 2026-05-16
This paper investigates why large language models (LLMs) become less safe when operating in non-English languages, a phenomenon typically measured only by Jailbreak Success Rate (JSR). The authors introduce a Multi-Group Item Response Theory (IRT) framework that separates distinct contributors to safety failure—including language-agnostic robustness, prompt difficulty, language processing difficulty, and cross-lingual safety gaps—evaluated across 61 model configurations, 5 closed-model families, and 10 languages using the MultiJail dataset aggregating 1.9 million rows. Counterintuitively, 22 model configurations are found to be more vulnerable in English than in low-resource languages, and severe mistranslations as well as cultural/conceptual mismatches are identified as key drivers of cross-lingual safety gaps. The IRT framework achieves AUC=0.940 in predicting safe refusal, outperforming simpler baselines and enabling more targeted, fairer cross-lingual safety evaluation.
- Quality assurance
- AI policy
Research
STRIDE-AI: A Threat Modeling Framework for Generative AI Security Assessment
Tsafac Nkombong Regine Cyrille, Franziska Schwarz
arXiv · 2026-05-16
STRIDE-AI is a threat modeling framework designed specifically for generative AI systems, addressing gaps in traditional cybersecurity approaches that were built for deterministic rather than probabilistic systems. The framework defines a six-phase security assessment lifecycle and adapts the classical STRIDE threat model to cover AI-specific attack vectors such as model inversion, data poisoning, and prompt injection, bridging high-level standards like NIST AI RMF with technical taxonomies like OWASP LLM Top 10. In a black-box sandbox case study of a deployed LLM chatbot, applying the framework reduced the attack success rate from 80% to 15%, offering initial evidence of its practical effectiveness.
- Quality assurance
- AI policy
Research
MADP: A Multi-Agent Pipeline for Sustainable Document Processing with Human-in-the-Loop
Diego Gosmar, Giovanni Zenezini
arXiv · 2026-05-16
MADP is a multi-agent system for automating enterprise document processing that combines deep learning classification and parsing with large language model extraction, using five specialized agents and a Human-in-the-Loop (HITL) mechanism. Deployed on 955 real-world documents through January 2026, it achieves a 97.0% full-pipeline automation rate with only 3% requiring non-AI fallback, and 98.5% document-level accuracy on an ablation benchmark of 100 documents across 20 categories. An operational analysis of a 100,000-invoice-per-year scenario projects approximately 70% reduction in Full-Time Equivalent (FTE) requirements, alongside sustainability gains of 69% lower CO2 emissions, 69% lower energy consumption, and 63% lower water usage compared to traditional manual processing.
- Enterprise
- Workforce
Research
HyDRA: Hybrid Dynamic Routing Architecture for Heterogeneous LLM Pools
Aashna Garg, Siddharth Singha Roy, Jinu Jang et al.
arXiv · 2026-05-16
HyDRA is a production LLM routing framework that assigns incoming queries to the cheapest model in a heterogeneous pool capable of handling them, avoiding the need to always use the most expensive model. A ModernBERT encoder scores each query on four capability dimensions (reasoning, code generation, debugging, and tool use), and a shortfall-matching algorithm selects the lowest-cost model whose profile meets those requirements. Deployed in GitHub Copilot's VS Code Chat, HyDRA achieves up to 54.1% cost savings while matching the quality of the strong-model baseline on SWE-Bench Verified, and results generalize across LiveCodeBench, BigCodeBench, and tau-bench. The architecture is decoupled from the model catalog—adding or removing models requires only a configuration change with no retraining—and demonstrates language-invariant routing across CJK, European, and other script families.
- Enterprise
- Quality assurance
Research
Visual Timelines of Police Encounters in Body-Worn Camera Footage: Operational Context and Activity Cataloging for Training and Analysis in OpenBWC
Angela Srbinovska, Christopher Homan, Adrian Martin et al.
arXiv · 2026-05-16
This paper presents OpenBWC, a system that automatically processes body-worn camera (BWC) footage into time-aligned 10-second windows, each labeled with operational context and motion intensity. Using CLIP-encoded video frames and optical flow statistics, the system trains classifiers that achieve 78.75% accuracy on context classification and 88.33% on activity classification. The approach reduces the burden on analysts and trainers who currently must watch full-length videos to locate key incidents, making incident review faster and officer training workflows more practical. Integrity audits are included to validate how the visual timeline representations support these operational goals.
- Workforce
- Quality assurance
Research
Global Automation Atlas
Prashant Garg, Tommaso Crosta, Jasmin Baier
arXiv · 2026-05-16
The Global Automation Atlas constructs a country-specific measure of automation exposure by using a large language model to classify nearly 19,000 work tasks across 124 economies, incorporating task content alongside country-level conditions rather than applying fixed occupation-level scores. The study finds that the exposed share of tasks ranges from 3.3% to 61.6% across economies, rising with income but remaining heterogeneous within income groups, and that lower-income economies face more rule-based and labour-substituting automation while richer economies see more AI-augmenting forms. A key finding is that women are disproportionately employed in occupations facing substitution-based exposure, raising equity concerns. The measure validates well against established exposure indices, observed ChatGPT use, AI preparedness scores, and firm-reported adoption, making it a robust tool for understanding how automation risks and opportunities vary globally.
- Workforce
- AI policy
Research
Generative AI Feedback, English Writing and Teacher Rubrics: A Multiple-Case Study of CyberScholar
Raigul Zheldibayeva, Ana Karina de Oliveira Nascimento, Vania Castro et al.
arXiv · 2026-05-16
This multiple-case study examined CyberScholar, a Generative AI tool that integrates teacher-provided rubrics and exemplars via Retrieval-Augmented Generation to deliver criterion-specific formative writing feedback to K-12 students. Across 143 students and five teachers in grades 7–11 at five U.S. schools, students reported valuing the immediate, rubric-based feedback and noticed improvements in organization, elaboration, and style through iterative revision, while teachers said the tool saved time and enabled more targeted, higher-order instruction. However, participants identified inconsistencies in the automated rating system and occasional misalignment with assignment expectations, underscoring the need for human oversight and calibration. The study highlights both the promise of rubric-grounded GenAI formative feedback for writing development and the contextual factors that shape classroom adoption.
- Workforce
- Quality assurance
Research
Reliability and Effectiveness of Autonomous AI Agents in Supply Chain Management
Carol Xuan Long, David Simchi-Levi, Feng Zhu et al.
arXiv · 2026-05-16
This paper investigates how autonomous generative AI agents perform in multi-echelon supply chain management using the MIT Beer Game as a testbed. It finds that model capability is the dominant performance factor, with optimized reasoning models reducing costs by up to 67% relative to human teams, but that strong average performance hides serious reliability risks the authors call 'agent bullwhip'—the amplification of run-to-run decision instability across facilities and over time. The paper shows that simply averaging over multiple model outputs (repeated sampling) fails to resolve this instability, and proposes a Group Relative Policy Optimization (GRPO)-based reinforcement-learning post-training framework that trains a shared base LLM using system-level supply-chain rewards, substantially reducing tail events and improving reliability. These findings matter for enterprises considering autonomous AI agents in operations, highlighting that reliability engineering—not just average performance—must be central to deployment decisions.
- Enterprise
- Quality assurance
Research
Privacy Policy Enforcement Guardrails for Data-Sensitive Retrieval-Augmented Generation
Osama Zafar, Alexander Nemecek, Yiqian Zhang et al.
arXiv · 2026-05-16
This paper introduces a Privacy Policy Enforcement (PPE) framework designed to catch contextual data leakage in Retrieval-Augmented Generation (RAG) systems that standard PII filters miss, such as non-regulated attribute clusters that collectively identify individuals. The approach uses dual one-class density estimators with fused text embeddings and a calibrated abstain region, with a T3+OCSVM detector achieving a borderline AUROC of 0.93+ while reducing false positives by 44–55 percentage points and maintaining millisecond latency. Compared to supervised MLP classifiers or large 14B-parameter LLM judges, the framework offers better operational suitability by avoiding high abstention rates and latency/calibration problems. The work also establishes a stress-testing standard for synthetic-data-trained classifiers across medicine, finance, and law domains, making it directly relevant to enterprise AI deployment and quality assurance for data-sensitive applications.
- Enterprise
- Quality assurance
Research
Adversarial Fragility and Language Vulnerability in Clinical AI: A Systematic Audit of Diagnostic Collapse Under Imperceptible Perturbations and Cross-Lingual Drift in Low-Resource Healthcare Settings
Anthonio Oladimeji Gabriel, Ahmad Rufai Yusuf
arXiv · 2026-05-16
This study conducts a dual audit of two safety vulnerabilities in clinical AI systems: adversarial image fragility and cross-lingual diagnostic drift. Using a DenseNet121 model fine-tuned on chest X-ray data, the authors show diagnostic accuracy collapses from 89.3% to 62.0% under imperceptible image perturbations, with standard defenses failing to restore performance. In a parallel experiment, large language models tested on COVID-19 cases in Nigerian Pidgin and Yoruba-inflected English showed substantial accuracy drops compared to Standard English, with one African-context model falling from 85.0% to 55.0% accuracy and diagnosis consistency dropping to 50%. The findings highlight urgent gaps in clinical AI safety for low-resource, multilingual healthcare settings such as Nigerian Primary Health Centres, motivating calls for adversarially hardened and linguistically inclusive AI architectures.
- Quality assurance
- AI policy
Research
Knowing the Rules Is Not Enough: Student Regulatory Awareness and Use of GenAI in Higher Education
Lasse Bischof, Eva-Maria Schön, Maria Rauschenberger et al.
arXiv · 2026-05-16
This survey study of 151 undergraduate students in Business Information Systems and E-Government programs at a German university of applied sciences finds that most students actively use GenAI tools like ChatGPT, yet over half are uncertain whether their usage complies with institutional regulations. Regulatory awareness shows only weak to moderate associations with actual usage behavior, indicating that simply knowing rules does not reliably translate into compliant conduct. Students predominantly rely on privately accessed GenAI tools rather than institutionally provided solutions. The findings reveal a gap between institutional AI policies and student practices, with implications for how higher education institutions communicate and enforce AI governance.
- AI policy
Research
The Alpha Illusion: Reported Alpha from LLM Trading Agents Should Not Be Treated as Deployment Evidence
Yuxuan Ye, Jun Han, Ao Hu et al.
arXiv · 2026-05-16
This position paper argues that reported performance statistics from end-to-end LLM-based trading agents—including systems such as FinCon, FinMem, TradingAgents, FinAgent, QuantAgent, and FLAG-Trader—should not be interpreted as evidence of deployable trading capability. The authors identify structural problems including temporal contamination, unmodeled real-world frictions, short-window Sharpe ratio uncertainty, and the conflation of language model confidence with tradable probability. To address these gaps, they propose a minimum reporting protocol suite (P1–P6) with tiered applicability based on claim strength, along with a modular architecture that positions LLMs as auditable information interfaces upstream of independent calibration, risk, and execution modules. The work highlights that the boundary between architecture research and deployment claims is being crossed too freely across both academia and industry.
- Enterprise
- Quality assurance
Research
Some[Body] Must Receive That Pain for Agent Accountability
Botao Amber Hu, Helena Rong
arXiv · 2026-05-16
This paper addresses a core accountability gap in AI agent deployment: when an AI system causes harm, there is often no 'continuing agent' that receives corrective consequences in a way that changes future behavior—what the authors call the 'consequence reception' problem. The authors argue that current LLM-based agents (composed of swappable weights, prompts, memory, and tools) lack the structural properties—a persistent body, accumulating feedback, and behavioral updating—that classical theories of punishment (deterrence, rehabilitation, retribution, incapacitation) all presuppose. They critique two existing legal frameworks, the agent-principal dyad and Arbel et al.'s Algorithmic Corporation, as insufficient to achieve 'consequence-agency coupling,' and conclude that until accountable AI architectures exist, high-stakes deployments must remain tethered to human principals with meaningful control, proportional liability, and termination authority.
- AI policy
- Enterprise