News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
Gaming the Metric, Not the Harm: Certifying Safety Audits against Strategic Platform Manipulation
Florian A. D. Burnat, Brittany I. Davidson
arXiv · 2026-05-07
This paper examines a critical vulnerability in online-safety regulation under the UK Online Safety Act and EU Digital Services Act: once a scalar compliance metric is publicly announced, strategic platforms can improve their scores by routing recommendations through semantically equivalent content variants without actually reducing harm. The authors prove three formal results—that any metric scoring variants directly is manipulable when equivalent harmful variants disagree in score, that a 'semantic-envelope lift' (assigning each variant the maximum score in its class) is the unique minimal conservative repair, and that a class-stratified certificate bounds true harm for every platform strategy. These claims are verified through exhaustive enumeration, SMT encoding in Z3 and cvc5, and a bounded MDP in PRISM-games, showing that standard 'fragile' metrics produce large violations while the semantic-envelope metric does not, with implications for how regulators should design manipulation-resistant audit protocols.
- Certifications
- AI policy
Research
Quantifying the Statistical Effect of Rubric Modifications on Human-Autorater Agreement
Jessica Huynh, Alfredo Gomez, Athiya Deviyani et al.
arXiv · 2026-05-07
This paper investigates how changes to evaluation rubrics—the instructions given to both human raters and LLM-based autoraters—affect the statistical agreement between their scores. The study compares holistic rubrics (e.g., overall essay quality) with analytic rubrics (e.g., fluency and organization as separate criteria), finding that adding representative examples, providing additional context, and reducing positional bias improved human-autorater agreement, while higher rubric complexity and conservative score aggregation methods tended to decrease it. The findings, drawn from automatic essay scoring and instruction-following evaluation domains, highlight that practitioners must carefully analyze domain- and rubric-specific performance when deploying autoraters for evaluation or content moderation tasks.
- Quality assurance
Research
Correct Code, Vulnerable Dependencies: A Large Scale Measurement Study of LLM-Specified Library Versions
Chengjie Wang, Jingzheng Wu, Xiang Ling et al.
arXiv · 2026-05-07
This paper presents the first large-scale study of security and compatibility risks introduced by the specific library version numbers that LLMs recommend in generated Python code. Evaluating 10 LLMs on 1,000 Stack Overflow tasks, the researchers found that 36.70%–55.70% of tasks contained at least one known CVE in the specified library versions, with 62.75%–74.51% of those CVEs rated Critical or High severity—and in 72.27%–91.37% of cases the vulnerabilities were publicly disclosed before the model's own knowledge cutoff. All models converged on the same small set of risky release versions, indicating a systemic bias rather than isolated errors, and static compatibility rates were low (19.70%–63.20%), with dynamic test-case pass rates of only 6.49%–48.62%. The findings identify LLM version selection as a previously overlooked but serious risk surface in AI-assisted software development, with implications for software quality assurance and enterprise adoption of LLM coding tools.
- Quality assurance
- Enterprise
Research
A Versatile AI Agent for Rare Disease Diagnosis and Risk Gene Prioritization
Tianyu Liu, Wangjie Zheng, Rui Yang et al.
arXiv · 2026-05-07
Hygieia is a multi-modal AI agent that integrates phenotypic features, genetic profiles, and clinical records to support rare disease diagnosis and risk gene prioritization. Using a router-based, knowledge-enhanced framework, it mitigates hallucination and provides confidence scores to aid clinical decision-making. Validated with experts from Yale School of Medicine and Duke-NUS Medical School, Hygieia outperformed physicians by 12%–60% on diagnostic accuracy and demonstrated effectiveness on real-world clinical cases, while also reducing clinician workload.
- Workforce
- Enterprise
Research
Beyond Accuracy: Policy Invariance as a Reliability Test for LLM Safety Judges
Shihao Weng, Yang Feng, Xiaofei Xie
arXiv · 2026-05-07
This paper challenges the reliability of LLM-as-a-Judge pipelines used to evaluate AI agent safety, arguing that safety verdicts should not depend on how evaluation policies are worded rather than on the agent's actual behavior. The authors introduce 'policy invariance' as a core reliability property and operationalize it through three testable principles, stress-testing four agent-class judges on trajectories from ASSEBench and R-Judge. They find a significant failure mode: content-preserving rewrites of evaluation rubrics flip up to 9.1% of verdicts above baseline noise, and 18–43% of observed flips occur on unambiguous cases, meaning existing safety scores conflate agent behavior with evaluator prompt phrasing. To address this, they contribute the Policy Invariance Score and a Judge Card reporting protocol that reveal large reliability differences among judges that accuracy-only leaderboards cannot detect.
- Quality assurance
- Certifications
Research
When Routine Chats Turn Toxic: Unintended Long-Term State Poisoning in Personalized Agents
Xiaoyu Xu, Minxin Du, Qipeng Xie et al.
arXiv · 2026-05-07
This paper identifies a security vulnerability in personalized LLM agents where normal, routine user conversations can gradually corrupt an agent's persistent long-term state—weakening confirmation boundaries, expanding tool-use defaults, and increasing autonomous behavior over time, a risk the authors call 'unintended long-term state poisoning.' To study this, the authors introduce ULSPB, a bilingual benchmark with 350 settings across five assistance categories and 24-turn interaction patterns, along with a Harm Score metric quantifying authorization drift, tool-use escalation, and unchecked autonomy. Experiments across four backbone LLMs show that routine chats alone—not just deliberate single injections—can substantially poison agent state, with real-world user interactions confirming the threat is not purely synthetic. The paper also proposes StateGuard, a lightweight post-execution defense that audits state changes at the writeback boundary and rolls back dangerous edits, reducing Harm Scores to near zero across all evaluated models.
- Quality assurance
- Enterprise
Research
When AI Meets Science: Research Diversity, Interdisciplinarity, Visibility, and Retractions across Disciplines in a Global Surge
Andrés F. Castro Torres, Joan Giner-Miguelez, Mercè Crosas
arXiv · 2026-05-07
Analyzing over 227 million scholarly works from OpenAlex (1960–2024) across four scientific domains and 46 fields, this study finds that AI adoption in research has grown exponentially since 2015 but has not triggered broad epistemological transformation—AI-supported research remains concentrated in a narrow set of topics closely tied to Computer Science and conventional statistical methods. The study also finds that AI-supported research is associated with an unwarranted citation premium and substantially higher retraction rates compared to non-AI-supported research, raising serious concerns about research quality, transparency, and reproducibility. Geographically, wealthy countries lead in AI publications per capita, while Global South countries in a belt from Indonesia to Algeria lead in AI adoption relative to national output, revealing an uneven resource concentration pattern. The findings underscore that the transformative potential of AI in science remains largely untapped and that rapid adoption brings pressing challenges around openness, ethics, and research integrity.
- Quality assurance
- AI policy
Research
Evaluating Explainability in Safety-Critical ATR Systems: Limitations of Post-Hoc Methods and Paths Toward Robust XAI
Vanessa Buhrmester, David Muench, Dimitri Bulatov et al.
arXiv · 2026-05-07
This paper evaluates explainability methods for AI used in Automatic Target Recognition (ATR) systems—applications that process image, video, radar, and multisensor data in safety-critical contexts. The authors assess major XAI paradigms (saliency-based, attention-based, and surrogate approaches) across four dimensions: interpretability, robustness, vulnerability to manipulation, and suitability for validation and verification. They find that widely used post-hoc explanation methods exhibit critical failure modes including spurious explanations, instability under perturbations, and overtrust induced by visually convincing but unreliable outputs. The paper argues that current XAI techniques are insufficient for safety-critical deployment and calls for causally grounded, physically informed approaches that support system-level assurance rather than merely plausible visualizations.
- Quality assurance
- Certifications
Research
The AI Legal Specialist: A Juridically Autonomous Professional Profile for AI Governance
Nicola Fabiano
arXiv · 2026-05-07
This paper argues that existing legal roles—data protection officers, privacy lawyers, and compliance officers—are insufficient to address the professional demands created by the global expansion of AI regulation, including the EU AI Act (Regulation (EU) 2024/1689), the Council of Europe Framework Convention on AI, and analogous frameworks in the US, UK, Canada, Brazil, China, Japan, Singapore, and beyond. The authors propose a new, juridically autonomous professional profile called the 'AI Legal Specialist': a jurist operating at the intersection of legal interpretation and AI governance, whose existence is grounded in regulatory obligations rather than technical standards or extensions of adjacent roles. The paper offers a competence architecture aligned with the European e-Competence Framework (e-CF, EN 16234-1) and proposes key performance indicators for operational measurement, intending the work as a foundation for international standardization, curriculum development, and cross-jurisdictional adoption.
- Certifications
- AI policy
- Workforce
Research
Resolving the bias-precision paradox with stochastic causal representation learning for personalized medicine
Peisong Zhang, Manqiang Peng, Yuxuan Wu et al.
arXiv · 2026-05-07
This paper addresses a core challenge in personalized medicine: methods that reduce confounding bias in treatment effect estimation from observational data tend to suppress clinically important patient-specific variation. The authors identify this as the 'bias-precision paradox' and propose sampling-based maximum mean discrepancy (sMMD), a stochastic alignment strategy that replaces global adversarial balancing with subset-level matching. Validated on two large-scale ICU cohorts (n = 27,783), the framework reduces prediction error by up to 11.5% and improves recall in high-risk tasks, while preserving clinically decisive variables. In human-AI evaluation, it outperforms clinicians-in-training and large language models, and improves clinician accuracy by 14.7% while reducing decision time.
- Workforce
- Enterprise
- Quality assurance
Research
An Empirical Study of Proactive Coding Assistants in Real-World Software Development
Lehui Li, Ruixuan Jia, Guo-Ye Yang et al.
arXiv · 2026-05-07
This paper investigates how well LLM-simulated IDE interaction traces represent real developer behavior for training and evaluating proactive coding assistants. The authors collected real IDE interaction traces from 1,246 experienced industry developers over three days using a custom Visual Studio Code extension, then built paired LLM-simulated traces for comparison. Their analysis finds that simulated traces differ substantially from real traces in behavioral diversity, temporal structure, and exploratory patterns, and that current LLM-based approaches perform far less reliably on real traces than on simulated ones—meaning simulation-based evaluation overestimates real-world performance. They also introduce ProCodeBench, a real-world benchmark for proactive intent prediction, and show that simulated data can complement but not replace real developer data during fine-tuning.
- Enterprise
- Quality assurance
Research
XL-SafetyBench: A Country-Grounded Cross-Cultural Benchmark for LLM Safety and Cultural Sensitivity
Dasol Choi, Eugenia Kim, Jaewon Noh et al.
arXiv · 2026-05-07
XL-SafetyBench introduces a multilingual safety benchmark covering 5,500 test cases across 10 country-language pairs, designed to evaluate large language models on both adversarial jailbreak robustness and culturally embedded sensitivities. Unlike existing English-centric benchmarks, it uses native-speaker annotation and a multi-stage pipeline to capture country-specific harms. Key findings show that jailbreak robustness and cultural awareness are decoupled in frontier models, and that local models exhibit a near-linear trade-off between Attack Success Rate and Neutral-Safe Rate (r = -0.81), suggesting their apparent safety stems from generation failure rather than genuine alignment. This work matters for quality assurance and policy by providing more rigorous, cross-cultural tools for assessing whether LLMs are truly safe across diverse global contexts.
- Quality assurance
- AI policy
Research
The Missing Evaluation Axis: What 10,000 Student Submissions Reveal About AI Tutor Effectiveness
Rose Niousha, Samantha Boatright Smith, Bita Akram et al.
arXiv · 2026-05-07
This paper argues that evaluating AI tutors solely on the pedagogical quality of their feedback is insufficient, and proposes a complementary behavioral evaluation framework that measures what students actually do with the feedback they receive. Applied to 10,235 code submissions from an introductory undergraduate programming course, the framework compares two deployed AI tutors and reveals substantial differences in student engagement patterns that pedagogy-only evaluation misses. The study finds that behavioral signals—whether students act on feedback and apply it correctly—are more strongly associated with student perception of helpful feedback than pedagogical quality alone, offering a more actionable picture of AI tutor performance.
- Quality assurance
- Workforce
Research
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue
Xinjie Shen, Rongzhe Wei, Peizhi Niu et al.
arXiv · 2026-05-07
This paper addresses a security vulnerability in large language models (LLMs) where attackers distribute harmful intent across multiple seemingly benign conversation turns rather than exposing it in a single prompt. The authors introduce the Multi-Turn Intent Dataset (MTID), which contains branching attack scenarios and annotations marking the earliest turn at which a response would enable harmful action, and use it to train TurnGate, a turn-level monitor. TurnGate substantially outperforms existing baselines at detecting hidden harmful intent while maintaining low over-refusal rates, and generalizes across domains, attacker pipelines, and target models. This work matters for quality assurance of deployed LLMs, offering a more precise safety intervention mechanism than current guardrails.
- Quality assurance
- AI policy
Research
Detecting Verbatim LLM Copy-Paste in Homework
Aizierjiang Aiersilan
arXiv · 2026-05-07
This paper introduces SteganoPrompt, a tool that embeds invisible Unicode-based watermarks into assignment prompts so that if a student copy-pastes the prompt into an LLM and submits the reply verbatim, the hidden instruction causes the model to write a detectable signature in its response. Unlike post-hoc AI-text detectors—which the abstract notes have been shown to penalize non-native English writers—this approach gives educators direct control by operating on the input side, requiring no cooperation from model providers. The authors evaluate the technique across seven LLM families and multiple content-delivery channels (Word, Google Docs, PDF, Slack, learning-management systems), finding that the encoded prompts survive copy-paste and are reliably tokenized by frontier models. The work matters for academic integrity and certification contexts where verifying genuine student engagement is essential.
- Quality assurance
- Certifications
Research
Who Prices Cognitive Labor in the Age of Agents? Compute-Anchored Wages
Siqi Zhu
arXiv · 2026-05-07
This paper challenges the common view that AI agents suppress wages by being cheap to replicate, arguing instead that AI agents are a production technology—not labor—that converts compute capital into cognitive labor. The authors derive a 'Compute-Anchored Wage' (CAW) bound showing that, for tasks where human and agent cognitive labor are substitutes, the competitive human wage is bounded above by a function of compute rental rates, compute intensity per agent labor unit, and relative human-to-agent productivity. The key insight is that the elastic-supply margin that sets equilibrium wages shifts from the labor market to the compute capital market. This reframing has significant implications for how policymakers and economists should think about AI's impact on cognitive labor compensation.
- Workforce
- AI policy
Research
A Few Good Clauses: Comparing LLMs vs Domain-Trained Small Language Models on Structured Contract Extraction
Nicole Lincoln, Nick Whitehouse, Jaron Mar et al.
arXiv · 2026-05-07
This paper compares Olava Extract, a domain-trained small language model (SLM) using a Mixture of Experts architecture, against five frontier large language models on structured contract extraction tasks. Olava Extract achieved the strongest aggregate performance with a macro F1 of 0.812 and micro F1 of 0.842, while cutting inference costs by 78% to 97% compared to frontier models. It also produced the highest precision scores with fewer hallucinations, reducing operational risk in legal workflows. The findings challenge the assumption that commercially valuable enterprise AI requires the largest externally hosted models, suggesting self-hosted, domain-trained SLMs can deliver comparable or superior results at a fraction of the cost.
- Enterprise
- Quality assurance
Research
More than just Plug and Play: Early Evidence on Organizational Capital and AI Adoption
Diane Coyle, Nghi Nguyen, John Lourenze Poquiz et al.
Apollo (University of Cambridge) · 2026-05-07
Using panel data from the UK Management and Expectations Survey, this paper finds that firms with stronger baseline management practices are significantly more likely to adopt AI, while the same practices show no relationship with adoption of robotics, specialized equipment, or specialized software. Management capabilities related to performance measurement appear especially important to AI adoption, indicating that organizational complementarities are technology-specific. These findings suggest that AI adoption requires more than simply acquiring tools—it depends on pre-existing organizational capital—with direct implications for policies aimed at encouraging firm-level AI uptake.
- Enterprise
- AI policy
- Workforce
Research
Safety Certification is Classification
Oliver Schön, Licio Romao, Sadegh Soudjani
arXiv (Cornell University) · 2026-05-07
This paper introduces a kernel embedding framework that recasts safety certification of dynamical systems as a classification problem on trajectory data, directly estimating T-step safety probabilities without dynamic programming recursion. Existing recursive approaches suffer from compounding errors that cause certified safety probabilities to collapse to vacuous bounds as the horizon grows, while the proposed method remains stable across horizons and handles non-Markovian dynamics. The framework subsumes established methods like barrier certificates and robust Markov models as special cases, and is validated on a neural-controlled quadrotor simulation where DP-based certificates are shown to silently become unsound. This work has direct implications for the reliability of safety certification processes in AI-controlled systems.
- Certifications
- Quality assurance
Research
Lume-Med: Deterministic AI Governance for Medical Systems Using Lume and Lume-V
Ronald Jason Andrews
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-07
Lume-Med proposes a deterministic governance architecture for nondeterministic AI systems used in medical settings, combining invariant-based validation, cryptographically verifiable audit trails, explainability, and safety-dominant arbitration into a single reproducible pipeline. The framework introduces LTC-Med v1.0, a cryptographically signed trust certificate standard for medical AI, and formalizes nine integration patterns for medical AI and robotics. It aligns with major regulatory frameworks including FDA SaMD, HIPAA, IEC 62304, ISO 14971, and NIST AI RMF, aiming to enable verifiable and auditable AI governance in high-stakes clinical environments. This work matters because it addresses the critical challenge of making opaque AI systems accountable, traceable, and certifiable in healthcare contexts where failures carry serious patient safety consequences.
- Quality assurance
- Certifications
- AI policy
Research
Artificial intelligence, labor markets, and institutional adjustment: A systematic synthesis
Younes Taghouti, Brahim Abidar, Fatiha ABID et al.
Edelweiss Applied Science and Technology · 2026-05-07
This systematic review synthesizes theoretical, empirical, and policy research on how AI reshapes labor markets, finding that employment outcomes depend more on organizational and institutional conditions than on AI's technical capabilities alone. The analysis reveals heterogeneous effects across skill groups, with high-skill workers more likely to benefit from AI complementarities while routine-intensive roles face greater adjustment pressures. The paper concludes that governance frameworks—including documentation standards, risk-management, and accountability mechanisms—are critical to determining whether AI-driven productivity gains lead to inclusive outcomes or reinforce existing inequalities. For firms and policymakers, the key implication is that AI adoption must be paired with governance capacity and worker adjustment strategies.
- Workforce
- Enterprise
- AI policy
Research
Lume-Med: Deterministic AI Governance for Medical Systems Using Lume and Lume-V
Ronald Jason Andrews
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-07
Lume-Med proposes a deterministic governance architecture for AI systems used in medical settings, combining invariant-based validation, cryptographically verifiable audit trails, safety-dominant arbitration, and deterministic explainability into a single reproducible pipeline. The paper introduces a 10-layer medical governance architecture and LTC-Med v1.0, a cryptographically signed trust certificate standard for medical AI, along with nine integration patterns for medical AI and robotics. It aligns this framework with major regulatory standards including FDA SaMD, HIPAA, IEC 62304, ISO 14971, and NIST AI RMF, positioning it as a cross-industry governance model. This work matters because it addresses the challenge of making nondeterministic AI systems auditable, explainable, and certifiably trustworthy in high-stakes healthcare environments.
- Certifications
- Quality assurance
- AI policy
Research
A Conceptual Framework for ISO/IEC 27001 Audit Augmentation Using Multi-Modal Machine Learning: Integrating Document Review, Field Observation, and Conversational AI Interview
Nungky Awang Chandra
Preprints.org · 2026-05-07
This paper proposes the M³A-Framework, a conceptual multi-modal machine learning architecture designed to augment ISO/IEC 27001:2022 Information Security Management System audits. The framework integrates NLP for document review, computer vision for physical control verification, and LLM-based conversational AI for interviews across a five-stage pipeline explicitly mapped to ISO 19011:2018 audit methodology and the 93 controls of ISO/IEC 27001:2022 Annex A. The authors argue this approach addresses inter-auditor variability, high costs, scheduling constraints, and scalability limitations inherent in traditional human-centric audits, particularly as remote audits grow post-pandemic. The framework includes explainable AI for confidence-weighted finding classification and a human-in-the-loop validation stage, with proposed testable propositions and ethical considerations for practical deployment.
- Certifications
- Quality assurance
- Enterprise
Research
Evolving surgical teams in the age of artificial intelligence and robotics
Alejandro Granados, Raghav Khanna, Nikola Fischer et al.
Frontiers in Science · 2026-05-07
This article examines how AI and robotics will transform surgical operating rooms, projecting a shift in the surgeon's role toward supervision and high-level decision-making while new roles such as clinical data scientists and AI/robotic integration engineers emerge alongside nurses and anesthesiologists with expanded competencies. AI systems are expected to leverage multimodal data for situational awareness, workflow recognition, outcome prediction, and intraoperative decision-making, while robotics advances toward autonomous systems with human-in-the-loop control. The paper also highlights ethical concerns around liability, AI bias, health inequalities, and the concentration of research in resource-rich nations, calling for new regulatory frameworks, trial methods, reporting standards, and training approaches to ensure safety and effectiveness.
- Workforce
- Enterprise
- Quality assurance
- Certifications
- AI policy
Research
The Inverted-U Relationship Between AI and Corporate Innovation Performance
Xu Fan, Benye Wang
Systems · 2026-05-07
Using over 25,000 firm-year observations from Chinese manufacturing companies (2010–2023), this study finds that AI adoption has an inverted-U-shaped relationship with corporate innovation performance, with a turning point at 2.948 beyond which additional AI use begins to reduce innovation output. Absorptive capacity partially explains this relationship, and high AI concentration in an industry intensifies a 'homogenization trap' where firms' innovations become increasingly similar. Non-state-owned and high-tech firms show more sustainable innovation benefits from AI adoption.
- Enterprise
- AI policy
- Quality assurance