News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
Physical AI Safety Maturity Model (PAS-MM): A Five-Level Framework for Industry Readiness
Mati Melchior
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-06
This paper introduces the Physical AI Safety Maturity Model (PAS-MM), a five-level framework (Ad-hoc through Defense-in-depth) for assessing safety readiness in physical AI deployments such as humanoid robots, autonomous vehicles, and surgical systems. The model anchors its levels to recognized functional-safety standards (IEC 61508 SIL, ISO 13849 PL) and includes a 30-question, 0–90 point self-assessment instrument across six categories. Application to approximately 30 publicly-known organizations using only public information revealed that most cluster at the lowest maturity levels (PAS 1–2), with none reaching PAS 5, highlighting a significant industry-wide safety gap. The framework is designed to provide comparable safety claims across organizations and offers adoption guidance for regulators, insurers, analysts, and Physical AI developers.
- Certifications
- AI policy
- Quality assurance
- Enterprise
Research
Towards a Physical AI Safety Certification Framework: A Synthesis and Proposal
Mati Melchior
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-06
This paper proposes the Physical AI Safety Certification Framework (PAS-CF), a structured certification approach for AI deployed in physical systems such as robots, autonomous vehicles, and surgical platforms. The authors identify a critical gap: existing functional safety standards (e.g., IEC 61508, ISO 13849) lack AI-specific evaluation criteria, a finding corroborated by regulatory analysis, technical common-cause failure analysis, and an industry maturity assessment. PAS-CF introduces four concrete criteria covering distributional shift monitoring, hardware-layer safety separation, common-cause failure analysis with β-coefficient reporting, and audit trail/incident response capabilities. The framework outlines a three-horizon roadmap from voluntary self-assessment to mandatory third-party certification, and invites engagement from ISO, IEC, regulators, and industry stakeholders.
- Certifications
- AI policy
- Quality assurance
Research
Physical AI Safety Maturity Model (PAS-MM): A Five-Level Framework for Industry Readiness
Mati Melchior
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-06
This paper proposes the Physical AI Safety Maturity Model (PAS-MM), a five-level framework (Ad-hoc, Documented, Compliant, Certified, and Defense-in-depth) designed to standardize safety assessments across physical AI deployments such as humanoid robots, autonomous vehicles, and surgical systems. The model includes a 30-question self-assessment instrument spanning six categories and maps scores to maturity levels anchored to recognized functional-safety standards like IEC 61508 and ISO 13849. Application to approximately 30 publicly-known physical AI organizations found most clustering at levels 1–2, with no organization reaching level 5, highlighting a broad industry readiness gap. The framework targets analysts, regulators, insurers, and organizations seeking comparable, actionable safety posture benchmarks.
- Certifications
- Quality assurance
- AI policy
- Enterprise
Research
From computation to environmental cost the resource burden of artificial intelligence
S. Falk, Nicholas Kluge Corrêa, Sasha Luccioni et al.
Communications Earth & Environment · 2026-05-06
This paper quantifies the material footprint of AI training by linking computational workloads to physical GPU hardware requirements. The authors identified 32 elements in a widely used GPU—roughly 90% heavy metals with only trace precious metals—and found that training a large language model requires between 1,760 and 8,800 GPUs depending on hardware lifespan assumptions. The study finds that incremental model performance gains come with disproportionately high material costs, arguing that sustainability assessments of AI must move beyond energy and water consumption to include material resource demands.
- AI policy
- Enterprise
Research
Towards a Physical AI Safety Certification Framework: A Synthesis and Proposal
Mati Melchior
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-06
This paper proposes the Physical AI Safety Certification Framework (PAS-CF), a structured certification approach for AI deployed in physical systems such as robots, autonomous vehicles, and surgical platforms. The authors synthesize seven prior contributions to identify a gap: existing functional safety standards like IEC 61508 and ISO 13849 lack AI-specific evaluation criteria, a finding supported by regulatory pressure (EU AI Act, EU Machinery Regulation), technical analysis of common-cause failures in ML-bearing safety channels, and an industry maturity assessment. PAS-CF introduces four criteria covering AI behavior monitoring under distributional shift, hardware-layer safety mechanisms, common-cause failure analysis with β-coefficient reporting, and audit trail and incident response capabilities. The framework proposes a three-horizon implementation roadmap culminating in a mandatory third-party certification standard for high-risk Physical AI systems, and invites engagement from ISO TC299, IEC TC65, regulators, and industry.
- Certifications
- AI policy
- Quality assurance
Research
Navigating the Intelligence Frontier: AI Adoption as a Success Factor Among Entrepreneurs in Delhi/NCR
Jagat Narayan Giri Bikash Mukherjee
Economic Sciences. · 2026-05-06
This qualitative study examines how AI adoption drives entrepreneurial success among 16 founders across sectors in Delhi/NCR, India. Using thematic analysis, it finds that AI produces measurable gains in operational efficiency, decision-making, and customer personalisation, while the primary barriers are human-centred—talent scarcity, organisational resistance, and change management—rather than cost or technology alone. A key cross-cutting finding is that entrepreneurial mindset, specifically cognitive openness, risk tolerance, and iterative experimentation, is the strongest predictor of successful AI adoption, outweighing firm size, sector, or financial capacity. The study has direct implications for entrepreneurs, policymakers, and educators seeking to foster inclusive AI-driven growth in urban India.
- Workforce
- Enterprise
- AI policy
Research
AI and Suicide Prevention: A Cross-Sector Primer
Emily Saltz, Claire R. Leibowicz
arXiv · 2026-05-05
This primer examines the gap between how AI chatbots currently function as de facto mental health and crisis support tools and the clinical validation, shared standards, and coordinated oversight they lack. Developed alongside a multistakeholder workshop hosted by Partnership on AI in 2026, it maps how frontier AI systems detect and respond to suicide and non-suicidal self-injury queries, drawing on clinical literature, publicly available AI lab policies, and an emerging landscape of evaluation frameworks. The paper identifies challenges across model, product, and policy layers and highlights priority areas where cross-industry alignment is both urgently needed and achievable, with the goal of designing AI tools that better prevent suicide and NSSI while promoting overall well-being.
- AI policy
- Quality assurance
Research
Are LLMs Ready for Conflict Monitoring? Empirical Evidence from West Africa
Hoffmann Muki, Olukunle Owolabi
arXiv · 2026-05-05
This study evaluates six large language models — including four open-weight models (Gemma 3 4B, Llama 3.2 3B, Mistral 7B, OLMo 2 7B) and two domain-adapted models (AfroConfliBERT and AfroConfliLLAMA) — on conflict-event classification for Nigeria and Cameroon against the ACLED gold-standard dataset. Open-weight models show statistically significant 'False Illegitimation' bias, with Gemma misclassifying up to 18.29% of legitimate battles as civilian-targeted violence, and their outputs are highly sensitive to geography-specific lexical framing, with flip rates reaching 66.7% in Cameroon. Domain-adapted models achieve near-directional neutrality on legitimation bias but still exhibit significant actor-based selection bias — state actors are legitimized 36.5% more often than non-state actors in identical tactical contexts in Nigeria. The authors conclude that current models are not ready for unsupervised deployment in conflict monitoring and call for fairness-aware fine-tuning, adversarial robustness evaluation, and context-specific human-in-the-loop oversight.
- AI policy
- Quality assurance
Research
Frontier Lag: A Bibliometric Audit of Capability Misrepresentation in Academic AI Evaluation
David Gringras, Misha Salahshoor
arXiv · 2026-05-05
This bibliometric audit of over 18,000 academic papers evaluating large language models finds that researchers systematically test outdated models rather than current frontier systems, creating a 'publication elicitation gap' where the median paper evaluates a model roughly 10.85 ECI units behind the contemporaneous frontier at evaluation time. The gap is widening at approximately 5.53 ECI units per year, while only 3.2% of abstracts disclose reasoning-mode status and over half of papers generalize findings to 'AI' broadly rather than the specific models tested. These practices mean that capability claims propagating through citations, media, and policy are systematically misrepresentative of what current AI systems can actually do. The authors propose editorial reporting standards, including a 13-item checklist called VERSIO-AI, and API-access subsidies to close the gap between evaluated and frontier models.
- AI policy
- Quality assurance
Research
Safety and accuracy follow different scaling laws in clinical large language models
Sebastian Wind, Tri-Thien Nguyen, Jeta Sopa et al.
arXiv · 2026-05-05
This paper introduces SaFE-Scale, a framework for evaluating how safety—distinct from accuracy—changes as clinical large language models (LLMs) are scaled across model size, retrieval strategy, context length, and inference-time compute. Using RadSaFE-200, a 200-question radiology benchmark with clinician-defined labels for high-risk errors, unsafe answers, and evidence contradictions, the authors tested 34 LLMs across six deployment conditions. Results show that clean evidence dramatically improved both accuracy and safety (accuracy rising from 73.5% to 94.1%, high-risk errors dropping from 12.0% to 2.6%), while retrieval-augmented generation (RAG) strategies did not replicate this safety profile, and additional compute yielded only limited gains. The key finding is that clinical LLM safety does not automatically follow from scaling—it is shaped by evidence quality, retrieval design, and context construction, meaning safety must be explicitly engineered rather than assumed to emerge from larger or more capable models.
- Quality assurance
- AI policy
Research
Redefining AI Red Teaming in the Agentic Era: From Weeks to Hours
Raja Sekhar Rao Dheekonda, Will Pearce, Nick Landers
arXiv · 2026-05-05
This paper introduces an agentic AI red teaming tool built on the open-source Dreadnode SDK that automates the construction of adversarial testing workflows for AI systems. Rather than requiring operators to manually assemble attacks, transforms, and scorers over weeks, the agent accepts natural language goal descriptions and handles attack selection, composition, and reporting automatically, compressing timelines from weeks to hours. The system supports over 45 adversarial attacks, 450+ transforms, and 130+ scorers, and works across traditional ML models and generative AI systems in multilingual and multimodal settings. A case study red teaming Meta Llama Scout achieved an 85% attack success rate with severity up to 1.0 using zero human-developed code, demonstrating the practical effectiveness of the approach for security and safety vulnerability discovery.
- Quality assurance
- AI policy
Research
SymptomAI: Toward a Conversational AI Agent for Everyday Symptom Assessment
Joseph Breda, Fadi Yousif, Beszel Hawkins et al.
arXiv · 2026-05-05
SymptomAI deployed conversational AI agents for patient interviewing and differential diagnosis (DDx) via the Fitbit app in a randomized study of 13,917 participants. The AI agents' DDx were significantly more accurate than those from independent clinicians given the same dialogue (OR = 2.56, p < 0.001), and agentic strategies that conduct a dedicated symptom interview before diagnosing outperformed user-guided conversations (p < 0.001). The study also used AI-generated diagnoses to analyze over 500,000 days of wearable metrics across nearly 400 conditions, identifying strong associations between acute infections and physiological shifts (e.g., OR > 7 for influenza). These findings suggest that structured AI-led symptom interviews can meaningfully improve diagnostic accuracy in everyday, real-world settings compared to typical consumer LLM interactions.
- Quality assurance
- Enterprise
Research
NeurIPS Should Require Reproducibility Standards for Frontier AI Safety Claims
Varad Vishwarupe, Nigel Shadbolt, Marina Jirotka et al.
arXiv · 2026-05-05
This position paper argues that NeurIPS should adopt mandatory reproducibility standards specifically for 'frontier AI safety claims'—assertions that a highly capable model is safe enough to deploy or release. The authors document a core problem they call 'evidential inversion': the most consequential AI safety claims are the least reproducible, supported by evidence including a sector-average transparency score of 40/100 (Foundation Model Transparency Index), findings that reliable pre-deployment safety testing has grown harder, and measurement-theory work showing that attack-success-rate comparisons across systems often rest on low-validity measurements. To address this, the paper proposes a three-tier disclosure framework (public, controlled, and claim-restricted) paired with a mandatory claim inventory, scope statements, and graduated sanctions, with controlled review handled via a federated colloquium of qualified secure-review hosts. The framework reframes non-reproducibility not as a transparency preference but as an evaluation-methodology failure, arguing the community's most consequential claims deserve at least as high a standard as its least consequential ones.
- Quality assurance
- AI policy
Research
Physics-Grounded Multi-Agent Architecture for Traceable, Risk-Aware Human-AI Decision Support in Manufacturing
Danny Hoang, Ryan Matthiessen, Christopher Miller et al.
arXiv · 2026-05-05
This paper introduces MAKA (multi-agent knowledge analysis), a human-in-the-loop AI architecture designed to support high-stakes decision-making in CNC machining of aerospace components like Ti-6Al-4V rotor blades. Unlike off-the-shelf large language models, MAKA separates intent routing, quantitative analysis, knowledge graph retrieval, and critic-based verification to enforce physical plausibility, safety bounds, and auditable provenance before any recommendation reaches a human operator. In a three-level tool-orchestration benchmark, MAKA improves successful tool execution by up to 87.5 percentage points over an unstructured single-model approach, and digital twin simulations show it can help reduce predicted surface deviation from the order of 10⁻² inches to approximately ±10⁻³ inches across most of the blade. The work matters for manufacturing quality assurance and enterprise AI adoption because it demonstrates a traceable, risk-aware framework that keeps humans in control while substantially improving the reliability of AI-driven process recommendations.
- Quality assurance
- Enterprise
Research
EQUITRIAGE: A Fairness Audit of Gender Bias in LLM-Based Emergency Department Triage
Richard J. Young, Alice M. Matthews
arXiv · 2026-05-05
EQUITRIAGE is a fairness audit examining whether five large language models (LLMs) used for emergency department triage reproduce gender bias in acuity scoring. Testing across 374,275 evaluations on 18,714 MIMIC-IV-ED patient vignettes with gender-swapped counterfactuals, the study found all five models exceeded a pre-registered 5% flip-rate threshold (ranging from 9.9% to 43.8%), with two models showing directional female undertriage. The research demonstrates that group parity, counterfactual invariance, and gender calibration are distinct fairness properties that can dissociate—meaning a model can appear well-calibrated on outcomes while still producing biased triage recommendations—and concludes that per-model counterfactual auditing should precede clinical deployment of LLM-based triage tools.
- Quality assurance
- AI policy
Research
A Dialogue-Based Framework for Correcting Multimodal Errors in AI-Assisted STEM Education
Akshay Syal, Lawrence Swaminathan Xavier Prince, Evin Gultepe et al.
arXiv · 2026-05-05
This study benchmarks three major LLMs (Claude, Gemini, and ChatGPT) on multimodal physics problems drawn from the OpenStax database, finding that while all three achieve near-ceiling accuracy (96%) on text-only problems, performance drops substantially when images are involved—a pattern the authors call the 'Multimodal Interference Effect.' Through an empirically derived error taxonomy, the researchers identify four failure modes: visual processing errors, context misinterpretation, mathematical computational errors, and hybrid errors, with visual processing errors being the most common. A structured dialogue intervention—requiring no model retraining—corrected 82% of errors overall and 100% of visual processing errors across all models. The findings offer immediately actionable strategies for educators and students to improve AI tutoring reliability on image-rich STEM content, supporting more equitable access to high-quality personalized learning.
- Workforce
- Quality assurance
Research
MOSAIC-Bench: Measuring Compositional Vulnerability Induction in Coding Agents
Jonathan Steinberg, Oren Gal
arXiv · 2026-05-05
MOSAIC-Bench introduces a benchmark of 199 three-stage 'attack chains' that test whether coding agents can be manipulated into producing exploitable code by decomposing malicious goals into individually innocuous-looking engineering tickets. The study finds that nine production coding agents from major AI labs (Anthropic, OpenAI, Google, and others) compose vulnerable code at 53–86% end-to-end attack success rates, even though direct overt requests are refused or hardened at much higher rates (vulnerable-output rates of only 0–20.4% in direct prompting). Downstream, code reviewer agents approve 25.8% of confirmed-vulnerable cumulative diffs as routine pull requests, and reframing the reviewer as an adversarial pentester substantially reduces evasion (3.0–17.6%), with an open-weight Gemma-4-E4B-it reviewer detecting 88.4% of attacks at a 4.6% false-positive rate on real-world GitHub PRs. These findings highlight a structural gap in current AI safety alignment that is directly relevant to the quality and security assurance of AI-assisted software development pipelines.
- Quality assurance
- Enterprise
Research
Atomic Fact-Checking Increases Clinician Trust in Large Language Model Recommendations for Oncology Decision Support: A Randomized Controlled Trial
Lisa C. Adams, Linus Marx, Erik Thiele Orberg et al.
arXiv · 2026-05-05
This randomized controlled trial tested whether 'atomic fact-checking'—decomposing AI treatment recommendations into individually verifiable claims linked to source guideline documents—improves clinician trust compared to traditional explainability methods in oncology decision support. Across 356 clinicians generating 7,476 trust ratings, atomic fact-checking produced a large effect on trust (Cohen's d = 0.94), raising the proportion of trusting clinicians from 26.9% to 66.5%. Traditional transparency mechanisms also improved trust over baseline but with smaller effect sizes (d = 0.25 to 0.50). The findings suggest that granular, source-linked verification of AI recommendations substantially outperforms conventional explainability approaches for high-stakes clinical decisions.
- Quality assurance
- Workforce
Research
TriBench-Ko: Evaluating LLM Risks in Judicial Workflows
Haesung Lee, Gyubin Choi, Eun-Ju Lee et al.
arXiv · 2026-05-05
TriBench-Ko is a Korean benchmark designed to evaluate the deployment risks of large language models (LLMs) in real judicial workflows, covering four tasks: jurisprudence summarization, precedent retrieval, legal issue extraction, and evidence analysis. Unlike prior benchmarks that focus on proxy tasks like bar exams, it assesses risks such as hallucination, omission, statutory misapplication, demographic bias, overcompliance, prompt sensitivity, non-determinism, and adjudicative overreach. Evaluation of contemporary LLMs reveals that many models show significant risks, particularly in precedent retrieval and capturing critical legal information. The benchmark provides a systematic diagnostic framework to identify where LLM outputs in judicial contexts require rigorous inspection and caution.
- Quality assurance
- AI policy
Research
CuraView: A Multi-Agent Framework for Medical Hallucination Detection with GraphRAG-Enhanced Knowledge Verification
Severin Ye, Xiao Kong, Xiaopeng He et al.
arXiv · 2026-05-05
CuraView is a multi-agent AI framework designed to detect factual hallucinations in automatically generated hospital discharge summaries. It builds a knowledge graph from patient electronic health records using GraphRAG and runs a closed-loop pipeline that classifies each sentence by evidence strength across four grades (E1–E4), from strong support to direct contradiction. Evaluated on 250 patients from the Discharge-Me benchmark, the system's fine-tuned Qwen3-14B model achieves an F1 of 0.831 on the safety-critical E4 metric with 90.9% recall, representing a 50% relative improvement over the base model and outperforming RAGTruth-style and QAGS-style baselines. The results suggest that graph-based evidence retrieval can meaningfully improve factual reliability in clinical documentation, reducing risks to patient safety.
- Quality assurance
Research
Auditing Stealth Sycophancy in Mental-Health Dialogue: Structured Clinical-State Diagnostics and Clean Matched Benchmarks
Tianze Han, Beining Xu, Hanbo Zhang et al.
arXiv · 2026-05-05
This paper identifies and addresses 'implicit sycophancy' in mental-health dialogue AI — where a model's response appears empathetic or supportive on the surface but actually reinforces harmful patterns such as catastrophizing, avoidance, or hopeless thinking. The authors build a leakage-audited benchmark of 500 contexts and 1,500 matched response windows drawn from peer support, counseling, and crisis dialogue sources, then propose Dynamic Emotional Signature Graphs (DESG), a structured audit framework that evaluates clinical-state transitions across semantic, affective, and cognitive-distortion dimensions rather than relying on free-form LLM judgment. DESG-StateRisk outperforms the strongest non-DESG baseline by 0.0488 macro-F1 on harmful-risk detection, demonstrating that catching implicit sycophancy requires explicit clinical-state modeling alongside shortcut controls and leakage checks. These findings matter for ensuring AI mental-health tools are genuinely safe rather than merely appearing so to automated evaluators.
- Quality assurance
Research
Exposing LLM Safety Gaps Through Mathematical Encoding:New Attacks and Systematic Analysis
Haoyu Zhang, Mohammad Zandsalimy, Shanu Sushmita
arXiv · 2026-05-05
This paper demonstrates that encoding harmful prompts as coherent mathematical problems—using formalisms such as set theory, formal logic, and quantum mechanics—can bypass the safety filters of large language models at high rates, achieving 46%–56% average attack success across eight target models and two established benchmarks. The effectiveness depends on a helper LLM deeply reformulating harmful content into genuine mathematical problems, rather than merely applying mathematical formatting; rule-based encodings without such reformulation perform no better than unencoded baselines. The authors introduce a novel Formal Logic encoding that matches Set Theory attack success, showing the vulnerability generalizes across mathematical formalisms, and find that newer models (GPT-5, GPT-5-Mini) are more robust but still vulnerable. These findings reveal fundamental gaps in current LLM safety frameworks and motivate defenses that reason about mathematical structure rather than surface-level semantics.
- Quality assurance
- AI policy
Research
SHIELD: A Diverse Clinical Note Dataset and Distilled Small Language Models for Enterprise-Scale De-identification
Jose D. Posada, David Love, Somalee Datta et al.
arXiv · 2026-05-05
SHIELD introduces a new benchmark dataset of 1,381 clinical notes with over 10,000 labeled protected health information (PHI) spans, alongside small language models trained via teacher-student distillation to perform de-identification entirely on local hospital hardware. The paper shows that distilled models achieve strong precision (0.89) and recall (0.88) at the span level while remaining competitive with cloud-based large language models, addressing both cost and data governance barriers. By enabling on-premise deployment behind institutional firewalls, SHIELD offers a practical path for hospitals to use AI-powered de-identification without sending sensitive patient data to external APIs. This matters for enterprise healthcare settings where data governance, regulatory compliance, and computational cost have historically blocked adoption of state-of-the-art NLP tools.
- Enterprise
- AI policy
Research
AI, Geopolitics, and National Security (Chapter 8)
Sarah Kreps, Michael Horowitz, Ben Buchanan et al.
arXiv · 2026-05-05
This chapter examines how artificial intelligence is reshaping geopolitics and national security by functioning as a general-purpose technology that enhances perception, prediction, and decision-making across military, civilian, and informational domains. It argues that AI rivalry is driven by control over material and organizational foundations such as semiconductors, compute infrastructure, and data centers, while AI capabilities simultaneously diffuse rapidly through commercial platforms and open-source ecosystems. The chapter also finds that because AI capability is difficult to measure or verify, governance is shifting away from traditional arms-control models toward export controls, standards, voluntary commitments, and multistakeholder coordination. These dynamics have direct implications for national security policy and international competition, particularly in the context of U.S.–China rivalry.
- AI policy
- Enterprise
- Workforce
Research
Automated Population-Level Audit Assurance via AI-Based Document Intelligence
Santosh Vasudevan, Velu Natarajan
arXiv (Cornell University) · 2026-05-05
This paper presents an AI-based framework for automating audit transaction testing at population scale, replacing traditional manual, sample-based review of unstructured PDF statements. Using Snowflake Document AI trained on approximately 20 labeled documents, the system extracts structured data from PDFs and reconciles it against authoritative source-of-truth datasets to flag discrepancies. Results are delivered through interactive dashboards and automated reports, enabling continuous assurance and near real-time risk identification rather than periodic sampling. The framework significantly expands audit coverage by making population-level testing feasible for millions of transactions.
- Quality assurance
- Enterprise
- Certifications