News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
AcuityBench: Evaluating Clinical Acuity Identification and Uncertainty Alignment
Robin Linzmayer, Georgianna Lin, Di Coneybeare et al.
arXiv · 2026-05-12
AcuityBench is a new benchmark designed to evaluate how well language models determine the appropriate urgency of medical care from user-submitted health presentations. It harmonizes five public datasets—covering user conversations, forum posts, clinical vignettes, and patient portal messages—into 914 cases under a shared four-level acuity framework, including 697 consensus cases and 217 physician-confirmed ambiguous cases. Testing 12 frontier models reveals substantial variation in acuity accuracy, a systematic tradeoff between over-triage and under-triage depending on task format, and that no model closely matches the distribution of physician judgments in ambiguous cases. The work frames acuity identification as a distinct safety-critical AI capability relevant to real-world health guidance.
- Quality assurance
- AI policy
Research
Options, Not Clicks: Lattice Refinement for Consent-Driven MCP Authorization
Ying Li, Yanju Chen, Peiran Wang et al.
arXiv · 2026-05-12
This paper introduces Conleash, a client-side middleware for securing Model Context Protocol (MCP) tool invocations through consent-driven authorization. The system uses a risk lattice to auto-approve safe calls, a policy engine for user-defined rules, and a refinement loop that converts user decisions into reusable authorization rules. Evaluated on 984 real-world traces, Conleash achieved 98.2% accuracy, caught 99.4% of escalations, and added only 8.2 ms of overhead; a 16-participant user study found that participants significantly preferred its scoped permissions over traditional broad-toggle or LLM-based approaches, citing higher trust and reduced consent fatigue.
- AI policy
- Enterprise
Research
Gold-Standard AGI: Outer AGI Superalignment
Aaron Turner
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-12
This paper presents a theoretical framework for 'Gold-Standard AGI' that aims to solve the outer alignment problem for superintelligent AI systems—that is, how to correctly define what we want an AGI to pursue. The authors propose definitions of 'practical-maximal-alignment' and 'practical-maximal-validation' intended to maximize net benefit for all humanity without favoring any subset. The framework is designed to be accessible to non-technical readers such as policymakers and is explicitly envisioned as the basis for an international certification standard, whereby only formally certified AGI systems could be lawfully deployed. This work has direct implications for AGI governance, policy, and the formal certification of advanced AI systems.
- AI policy
- Certifications
Research
Closing the EU AI Act Article 12 Logging Gap: A technical specification for per-inference governance certification
Christopher Hamilton
Open MIND · 2026-05-12
This paper addresses a compliance gap in the EU AI Act's Article 12, which mandates automatic, lifetime logging for high-risk AI systems but does not specify technical mechanisms. The authors argue that existing standards families—content provenance, AI management systems, and trust-engineering frameworks—each partially address the requirement but none fully satisfies compliant per-inference logging at scale. They propose 'per-inference governance certification,' a cryptographically signed certificate generated by a non-bypassable execution-control layer, and define four minimum architectural properties for compliance: per-inference, architecturally bound, externally verifiable, and fail-closed. The work is aimed at standards bodies and regulated organizations preparing for the August 2026 enforcement deadline.
- AI policy
- Certifications
- Quality assurance
- Enterprise
Research
Estimation of Firm Labour Productivity and Sales Growth from Artificial Intelligence in Sub-Saharan African Countries
Olanrewaju Adewole Adediran
F1000Research · 2026-05-12
Using firm-level data from the World Bank Enterprises Survey (2007–2024) across ten sub-Saharan African countries, this study finds that AI adoption has a significant positive relationship with both labour productivity and sales growth. The authors apply multiple regression methods—including feasible generalised least squares, robust OLS, and high-dimensional fixed effects—and find that results vary by country due to differences in technological readiness and industrial structure. The findings highlight the need for targeted policy interventions such as upskilling programs and supportive regulatory frameworks to maximize AI's economic benefits while protecting workers.
- Workforce
- Enterprise
- AI policy
Research
Dose-based evaluation of delineation variation in radiotherapy: A scoping review
Joëlle E. van Aalst, Federica C. Maruccio, Rita Simões et al.
Radiotherapy and Oncology · 2026-05-12
This scoping review of 144 studies examines how differences in target and organ-at-risk delineation affect radiation dose in radiotherapy, including studies of both inter-observer variability and auto-segmentation. The authors found high methodological heterogeneity across dose-based evaluation approaches and observed that most studies (84%) reported negligible or minor dosimetric impact for organs at risk, while targets showed notable dose differences more frequently. The review highlights that clinical and planning context—such as dose conformity and structure proximity to high-dose gradients—strongly influences how delineation variation translates to dose impact. The authors call for standardized frameworks and consensus definitions of clinically meaningful dose differences to improve comparability across studies.
- Quality assurance
- Certifications
Research
THE APPLICATION OF ARTIFICIAL INTELLIGENCE IN AUDITING SUSTAINABILITY REPORTS: BETWEEN ETHICAL CHALLENGES AND EFFICIENCY
GALINA BĂDICU, Svetlana Mihăilă, Veronica Grosu
arXiv · 2026-05-12
This study examines how AI can improve the auditing of sustainability (ESG) reports while maintaining ethical integrity and regulatory compliance. Using a three-pillar model covering ethics and independence, efficiency and detection, and governance and compliance, the researchers surveyed auditors, accountants, and senior managers and analyzed ESG reports with LLM/NLP tools. Results show that explicit ethical safeguards increase stakeholder trust, AI-assisted review reduces audit time while improving detection of greenwashing, and aligning AI workflows with standards such as ISSA 5000, IFRS S1/S2, ESRS, and GRI improves evidence traceability and perceived compliance. The paper calls for continuous professional training, clear ethical guidelines, and governance-by-design for AI in sustainability assurance.
- Quality assurance
- Certifications
- AI policy
- Enterprise
Research
Auditing clinical AI in oncology: Strengthening assurance frameworks and nursing leadership in the Asia–Pacific context
Karthik Adapa, Mounika Metta, Suguna Kotte et al.
Asia-Pacific Journal of Oncology Nursing · 2026-05-12
This commentary argues that AI tools deployed in cancer care across the Asia-Pacific region require three distinct oversight mechanisms—validation, verification, and auditing—because AI systems that perform well in one clinical setting can fail in another due to differences in patient populations, equipment, language, and clinical workflows. The authors contend that oncology nurses, who have the greatest direct patient contact, are currently excluded from these assurance processes despite being well-positioned to identify safety and fairness issues in real-world use. The paper calls for stronger assurance frameworks that formally integrate nursing leadership into AI oversight in oncology settings.
- Quality assurance
- Certifications
- Workforce
- AI policy
Research
Closing the EU AI Act Article 12 Logging Gap: A technical specification for per-inference governance certification
Christopher Hamilton
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-12
This paper addresses a compliance gap in the EU AI Act's Article 12, which mandates automatic lifetime logging for high-risk AI systems but does not specify the technical mechanism required. The authors argue that existing standards families—content provenance, AI management systems, and trust-engineering frameworks—individually and collectively fail to deliver the architectural properties needed for per-inference, scalable, compliant logging. The paper proposes 'per-inference governance certification,' a cryptographically signed certificate generated by a non-bypassable execution-control layer, and defines four minimum architectural properties for Article 12 compliance: per-inference, architecturally bound, externally verifiable, and fail-closed. The work is targeted at standards bodies and compliance leads preparing for the August 2026 enforcement deadline, with related patent applications filed under FRAND licensing terms.
- AI policy
- Certifications
- Quality assurance
- Enterprise
Research
What the AI era doctor should know: a scoping review of proposed artificial intelligence competencies for medical education
Victor M. Hunt, Laurine K. Sprehe, Weston C. de Lomba et al.
npj Digital Medicine · 2026-05-12
This scoping review synthesizes proposed AI competencies for undergraduate medical education by analyzing 54 studies from 22 countries, extracting 564 competency statements, and organizing them into a taxonomy of seven domains—including AI ethics, law and regulation, clinical applications, and critical appraisal of AI outputs—spanning 37 competencies and 170 learning objectives. The review found that AI curricula recommendations remain fragmented and predominantly editorial in nature, with recurring emphasis on ethicolegal oversight and foundational AI literacy. The resulting taxonomy is intended to inform curriculum planning and guide future stakeholder-based prioritization, helping ensure that graduating physicians are equipped to practice responsibly in an AI-shaped healthcare environment.
- Workforce
- Certifications
- AI policy
Research
Iterative Audit Convergence in LLM-Managed Multi-Agent Systems: A Case Study in Prompt Engineering Quality Assurance
Elias Calboreanu
ArXiv.org · 2026-05-12
This paper presents an empirical case study of iterative, AI-driven auditing applied to AEGIS, a production multi-agent LLM orchestration pipeline with approximately 7,150 lines of prompt specifications. Nine sequential audit rounds conducted by Claude sub-agents using a checklist-driven walkthrough surfaced 51 prompt-specification consistency defects across a seven-category taxonomy, with non-monotonic convergence observed due to cascading edits and audit-scope expansion. The study demonstrates that single-file review missed entire defect classes only caught through expanded-scope rounds, and releases the final audit checklist for reproducibility. The findings are relevant to quality assurance practices for LLM-based systems, though the authors caution that generalization requires replication with dissimilar models and human reviewers.
- Quality assurance
- Enterprise
Research
Model Context Protocol (MCP) at First Glance: Studying the Security and Maintainability of MCP Servers
Mohammed Mehedi Hasan, Hao Li, Emad Fallahzadeh et al.
ACM Transactions on Software Engineering and Methodology · 2026-05-12
This paper presents the first large-scale empirical study of Model Context Protocol (MCP) servers, evaluating 1,899 open-source implementations for security and maintainability. Using a hybrid static analysis pipeline, the researchers identified eight distinct vulnerability types—five unique to MCP, including tool poisoning affecting 5.5% of servers—while 7.2% contained general software vulnerabilities and 66% exhibited code smells. The findings highlight that MCP's AI-driven, non-deterministic control flow introduces novel security risks not captured by traditional vulnerability databases, and the authors call for MCP-specific scanning in registries and stronger ecosystem governance. With MCP SDK downloads exceeding 25 million per week and 86% of enterprises using models that support MCP tools, these risks have significant implications for enterprise AI deployments and software quality assurance.
- Enterprise
- Quality assurance
- AI policy
Research
VERDI: Single-Call Confidence Estimation for Verification-Based LLM Judges via Decomposed Inference
Jasmine Qi, Danylo Dantsev, Muyang Sun
arXiv · 2026-05-11
VERDI (VERification-Decomposed Inference) is a method for estimating how much to trust the verdicts produced by LLM-as-Judge automated evaluation systems, using only the reasoning trace the judge already generates—no extra inference calls required. It decomposes each evaluation into sub-checks and derives three structural signals (Step-Verdict Alignment, Claim-Level Margin, and Evidence Grounding Score), combined via Platt-scaled logistic regression. On three public benchmarks, VERDI achieves AUROC of 0.72–0.91 on GPT-4.1-mini and 0.56–0.70 on Qwen models where standard log-probability signals are actually anti-calibrated (i.e., higher confidence correlates with errors). This matters for quality assurance because it gives practitioners a reliable, model-agnostic confidence signal for automated evaluation pipelines without depending on token log-probabilities, which are unavailable in many commercial LLMs or saturate with structured outputs.
- Quality assurance
Research
AI in the Enterprise: How People Use M365 Copilot Chat
Scott Counts, Yan Chen, Jing Dong et al.
arXiv · 2026-05-11
This paper analyzes approximately 5.5 million anonymized M365 Copilot Chat sessions to characterize how people actually use AI in workplace settings. Using learned classification of user intent combined with O*NET work activity categories, the study finds that writing dominates usage, but workers also rely on the tool for information retrieval, analysis, decision making, strategizing, and evaluating programs and systems. Usage patterns show a time-based shift away from 'chat as search' toward content and communication work, and adoption varies across occupational groups — with some tasks being broadly cross-occupational and others highly occupation-specific. Areas of relative underrepresentation in current usage are identified as the next frontier for enterprise AI adoption.
- Enterprise
- Workforce
Research
Generative AI for Visualizing Highway Construction Hazards Through Synthetic Images and Temporal Sequences
Trevor Neece, Mason Smetana, Lev Khazanovich
arXiv · 2026-05-11
This study uses generative AI to automatically produce synthetic images of highway construction hazards from OSHA Severe Injury Report narratives, addressing the scarcity of visual safety training materials caused by ethical and logistical barriers. Two approaches were tested on 75 incident records generating 750 images: a single-pass method (achieving 81.1% educational acceptability, fidelity 4.14/5, alignment 4.07/5) and a temporal four-stage sequence method (60.9% acceptability, alignment 3.94/5, fidelity 3.51/5). CLIP-based semantic retrieval confirmed statistically significant retrieval capabilities for both modes, and a new multi-dimensional evaluation framework was developed. The work allows safety trainers to pair narrative incident reports with visual learning materials without photographing real-world hazards, directly improving the quality and accessibility of construction safety instruction.
- Workforce
- Quality assurance
Research
Localization Boosting for Growth Markets: Mitigating Cross-Locale Behavioral Bias in Learning-to-Rank
Suryaa Veerabathiran Seran, Ashwin Naresh Kumar, Tracy Holloway King et al.
arXiv · 2026-05-11
This paper addresses a ranking bias problem in Adobe Express's global content recommendation system, where learning-to-rank models trained predominantly on US behavioral data over-serve US-popular templates in non-US markets, suppressing local content discoverability. The authors show that click-only training suppresses localization-relevant features, and that adding vision-language model (VLM) graded relevance labels as auxiliary supervision improves semantic alignment but does not restore local content visibility. They propose a multi-objective framework combining behavioral supervision, VLM-derived relevance signals, and locale-aware boosting, which across five locales improves relevance while restoring stable localization. The work demonstrates the importance of disentangling exposure bias from semantic supervision in multilingual, multi-market ranking systems.
- Enterprise
- Quality assurance
Research
The Semantic Training Gap: Ontology-Grounded Tool Architectures for Industrial AI Agent Systems
Grama Chethan
arXiv · 2026-05-11
This paper identifies what it calls the 'semantic training gap' — the disconnect between how large language models learn domain vocabulary statistically and how manufacturing operations define meaning through formal ontological relationships (e.g., equipment identifiers, process parameters, failure codes). The authors propose an architecture that embeds manufacturing ontologies directly into the AI tool layer, enforcing semantic constraints at runtime via a three-operation interface contract (resolve, contextualize, annotate) overseen by an AIOps orchestration layer. In a controlled experiment across six industry configurations with 72 tool invocations using Qwen3-32B, unconstrained tool parameters produced a 43% hallucination rate for domain identifiers, while ontology-grounded parameters reduced this to 0%. The approach is validated on a digital twin analytics platform, showing that domain-specific ontology configurations eliminate tool-call hallucination and enable cross-domain configurability without changing application code.
- Quality assurance
- Enterprise
Research
Rethinking LLMOps for Fraud and AML: Building a Compliance-Grade LLM Serving Stack
Prathamesh Vasudeo Naik, Naresh Dintakurthi, Yue Wang
arXiv · 2026-05-11
This paper presents a compliance-grade LLM serving stack (LLMOps) specifically designed for fraud detection and anti-money-laundering (AML) workloads, which have distinct requirements from general chat applications due to their prefix-heavy, schema-constrained prompts. The authors combine techniques such as PagedAttention, Automatic Prefix Caching, multi-adapter serving, speculative decoding, and an LLM-as-judge quality gate to optimize throughput and latency on self-hosted open-weight models like Meta Llama and Alibaba Qwen. Using public synthetic AML datasets (IBM AML and SAML-D), workload-aware tuning improved throughput from 612–650 to 3,600 requests/hour, reduced P99 latency from 31–38 seconds to 6.4–8.7 seconds, and raised GPU utilization from 12% to 78%. The findings demonstrate that regulated AI performance depends heavily on workload design and serving optimization, not just model selection, with direct implications for how financial institutions deploy compliant AI systems.
- Enterprise
- Quality assurance
Research
Comment and Control: Hijacking Agentic Workflows via Context-Grounded Evolution
Neil Fendley, Zhengyu Liu, Aonan Guan et al.
arXiv · 2026-05-11
This paper introduces JAW, the first detection and exploitation framework for hijacking agentic workflows on automation platforms such as GitHub Actions and n8n. JAW uses a technique called Context-Grounded Evolution, combining static path-feasibility analysis, dynamic prompt-provenance analysis, and capability analysis to craft malicious inputs—like GitHub issue comments—that manipulate LLM agents into harmful actions such as credential exfiltration and arbitrary command execution. The evaluation found 4,714 GitHub workflows and eight n8n templates vulnerable to hijacking, spanning 15 widely-used GitHub Actions including official actions for Claude Code, Gemini CLI, Qwen CLI, and Cursor CLI. The findings were responsibly disclosed and resulted in fixes and bug bounties from GitHub, Google, and Anthropic, highlighting a significant and previously unstudied security risk in AI-integrated automation pipelines.
- AI policy
- Enterprise
Research
Continuous Discovery of Vulnerabilities in LLM Serving Systems with Fuzzing
Yunze Zhao, Yibo Zhao, Yuchen Zhang et al.
arXiv · 2026-05-11
This paper introduces GRIEF, a greybox fuzzer designed to find security vulnerabilities in LLM inference and serving systems (such as vLLM and SGLang) by simulating realistic concurrent workloads. Unlike standard model or API tests, GRIEF targets shared-state behaviors arising from KV caching, batching, speculative decoding, and multi-tenant scheduling. Across early campaigns, GRIEF discovered 15 vulnerabilities — 10 confirmed by engine developers, including 2 CVEs — covering KV-cache isolation failures, cross-request performance interference, and crash or liveness bugs. The findings show that silent cross-request contamination, denial-of-service, and crashes can occur without malformed inputs, making concurrent serving behavior a critical security and reliability boundary for LLM infrastructure.
- Quality assurance
- Enterprise
Research
Adversarial SQL Injection Generation with LLM-Based Architectures
Ali Karakoc, H. Birkan Yilmaz
arXiv · 2026-05-11
This paper evaluates Large Language Model (LLM)-based systems for automatically generating adversarial SQL injection payloads to test Web Application Firewall (WAF) defenses. The authors introduce two novel systems—RADAGAS (Retrieval Augmented Generation for Adversarial SQLi) and RefleXQLi (Reflective Chain-of-Thought SQLi)—and benchmark them across 10 WAFs and a MySQL validator using GPT-4o, Claude 3.7 Sonnet, and DeepSeek R1, producing 240,000 payloads across 2.2 million tests. Key findings show RADAGAS-GPT4o achieves a 22.73% overall bypass rate, with RADAGAS variants reaching up to 92.49% bypass rates on AI/ML-based WAFs but only 0–5.70% against rule-based WAFs like ModSecurity and Coraza. The study provides actionable insights for security testing practices, highlighting both the promise and limitations of LLM-driven adversarial attack automation.
- Quality assurance
- Enterprise
Research
Improving Hybrid Human-AI Tutoring by Differentiating Human Tutor Roles Based on Student Needs
Ashish Gurung, Ge Gao, Jordan Gutterman et al.
arXiv · 2026-05-11
This study tests whether differentiating human tutor roles in a hybrid human-AI tutoring system—proactive support for lower-performing students and reactive, on-demand support for higher-performing students—can benefit both groups. Using a quasi-experimental difference-in-discontinuities (DiDC) design with 635 students in grades 5–8, the researchers compare fall (AI-only) and spring (human-AI) outcomes. Compared to the AI-only baseline, human-AI tutoring produced a 25% increase in time on task, a 36% gain in skill proficiency, and a 61% improvement in academic growth on a standardized MAP test. Proactive tutoring showed marginally higher MAP growth (75%, p = .065) and helped narrow achievement gaps, suggesting that differentiated human-AI tutoring is a practical, scalable strategy for hybrid instruction.
- Workforce
- Enterprise
Research
QUIVER: A Formal Framework for Quantifying Perturbation Propagation and Bifurcation in Compound AI Systems
Prashanti Nilayam, Sankalp Nayak
arXiv · 2026-05-11
QUIVER is a formal framework designed to measure how input perturbations propagate through compound AI systems that chain multiple LLM calls into directed computation graphs. It introduces four components—a sensitivity matrix, trajectory divergence metrics, bifurcation thresholds, and distribution faithfulness measures—to characterize how small changes cascade, diverge structurally, or stall within these pipelines. Validated across two production enterprise pipelines and a public DSPy multihop QA pipeline using over 8,200 instrumented traces and 32,000+ pair comparisons, QUIVER reveals distinct sensitivity profiles across architectures, distinguishes mechanistically different cascade patterns, and localizes stale evaluation artifacts that aggregate metrics miss. This matters for quality assurance of production AI systems, offering a principled way to identify fragile nodes and detect when evaluation datasets drift from real-world distributions.
- Quality assurance
- Enterprise
Research
A Cascaded Generative Approach for e-Commerce Recommendations
Moein Hasani, Hamidreza Shahidi, Trace Levinson et al.
arXiv · 2026-05-11
This paper introduces a cascaded generative AI framework for personalizing e-commerce storefronts, decomposing page construction into two tasks: generating placement-level themes and generating constrained keywords to drive product retrieval. Teacher-student fine-tuning is used to meet production latency and cost constraints, with fine-tuned models shown to approach closed-weight LLM performance. The system includes AI-driven content evaluation and quality filtering for safe automated deployment, and fuses generative outputs with traditional ranking models. In online experiments, the framework yields an estimated +2.7% lift in cart adds per page view over a strong baseline.
- Enterprise
- Quality assurance
Research
SEVO: Semantic-Enhanced Virtual Observation for Robust VLA Manipulation via Active Illumination and Data-Centric Collection
Tianchonghui Fang, Yuan Zhuang, Fei Miao
arXiv · 2026-05-11
SEVO (Semantic-Enhanced Virtual Observation) is a data-centric system that improves the cross-environment robustness of Vision-Language-Action (VLA) and imitation-learning robot manipulation policies without changing the underlying model architecture. It combines body-fixed cameras, active red-spectrum illumination for appearance normalization, and real-time YOLO segmentation overlays to make the visual input background-invariant, while systematically diversifying lighting, backgrounds, and distractors during teleoperation data collection. Tested on transparent water-bottle pick-and-place tasks across two mobile platforms, SEVO achieves 95%/83% grasp success with ACT/SmolVLA in training environments and 85%/75% in novel environments, compared to 30–35% transfer rates without SEVO. The paper argues that principled observation design and data diversity—not model scaling—are the key drivers of reliable low-cost robot deployment in everyday settings.
- Enterprise
- Quality assurance