News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
The External Governance Layer: A reference architecture for AI decision control in regulated enterprises
Rami Mohammed Kheir
Open MIND · 2026-05-24
This working paper introduces the AOS-1 External Governance Layer (EGL), a reference architecture designed to close the gap between documented AI governance policies and actual AI system behavior in production environments. The EGL functions as an out-of-process control plane that intercepts AI agent actions in real time, applies operating policy in milliseconds, returns verdicts (allow/deny/escalate/observe), and writes tamper-evident audit records to a separately governed decision vault. The paper includes regulatory cross-walks to eight frameworks including ISO/IEC 42001, the EU AI Act, NIST AI RMF, and others, and reports a median 78-day time-to-readiness based on engagements with regulated enterprises. It is positioned as an operational specification complementing the AOS-1 Operating Assurance Standard and AOS-1 Verified certification regime.
- Enterprise
- Quality assurance
- Certifications
- AI policy
Research
AOS-1 Verified: A continuous-verification certification for AI management systems
Rami Mohammed Kheir
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-24
This working paper argues that annual AI certification (e.g., ISO 42001, SOC 2) is structurally inadequate for modern AI systems that continuously drift in model behavior, retrieval corpora, and system prompts between audit cycles. The authors propose AOS-1 Verified, a continuous-verification certification framework built on a 47-control, versioned standard spanning 15 control families, mapped across eight major regulatory and standards frameworks including the EU AI Act and NIST AI RMF. The certification is expressed as a cryptographically signed JWT badge that can be verified offline and is automatically suspended when live operating signals deviate from defined thresholds. The paper matters for enterprise buyers, auditors, and regulators because it offers a real-time assurance mechanism that matches the operational tempo of deployed AI systems rather than a point-in-time attestation.
- Certifications
- Enterprise
- Quality assurance
- AI policy
Research
The GAIO Doctrine: Governance AI Optimization for non-AI-native enterprises
Rami Mohammed Kheir
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-24
This working paper introduces GAIO (Governance AI Optimization), a reference platform designed to help non-AI-native enterprises—such as mid-tier audit firms, internal audit functions, and second-line risk teams—integrate AI more effectively into audit and governance, risk, and compliance (GRC) work. The core argument is that current 'bolt-on' AI approaches accelerate drafting and evidence collation but leave the true bottleneck—partner review and judgement—untouched, failing to recover cycle-time value. The GAIO architecture proposes redirecting AI to compress the review cycle rather than the drafting cycle, preserving human auditor judgement as the firm's primary value. The paper also maps this approach against major regulatory frameworks including Sarbanes-Oxley, COSO, COBIT 2019, ISO/IEC 42001:2023, and the EU AI Act.
- Enterprise
- Quality assurance
- Certifications
- AI policy
Research
The External Governance Layer: A reference architecture for AI decision control in regulated enterprises
Rami Mohammed Kheir
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-24
This working paper introduces the AOS-1 External Governance Layer (EGL), a reference architecture designed to close the gap between formal AI governance policies and actual AI system behavior in regulated enterprises. The EGL is an out-of-process control plane that intercepts AI agent actions in real time, applies operating policies in milliseconds, returns one of four verdicts (allow/deny/escalate/observe), and writes tamper-evident audit records to a separately governed decision vault. The paper provides a regulatory crosswalk to eight frameworks including ISO/IEC 42001, the EU AI Act, NIST AI RMF, and others, and reports a median 78-day time-to-readiness based on engagements with regulated enterprises. It is positioned as an operational specification supporting both the AOS-1 Operating Assurance Standard and the AOS-1 Verified certification regime.
- Enterprise
- Quality assurance
- Certifications
- AI policy
Research
AOS-1 Verified: A continuous-verification certification for AI management systems
Rami Mohammed Kheir
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-24
This working paper argues that annual certification regimes like ISO 42001 and SOC 2 are structurally inadequate for AI systems because such systems drift continuously across models, retrieval corpora, prompts, and tools, often within a single audit cycle. The authors propose AOS-1 Verified, a continuous-verification certification built on a versioned 47-control standard spanning 15 control families, expressed as a signed JWT badge that can be verified offline and automatically suspended when live operating signals deviate from baseline. The framework includes a five-level maturity model, an eight-framework crosswalk covering major international regulations including the EU AI Act and NIST AI RMF, and a defined path toward EU AI Act Article 43 notified-body designation. The paper contends that continuous-verification certification better answers the diligence questions enterprise buyers actually ask, positioning AOS-1 Verified as a complement rather than substitute to existing assurance instruments.
- Certifications
- Enterprise
- Quality assurance
- AI policy
Research
Who judges the judges? Governance from metrics: a runtime framework for continuous LLM compliance monitoring
Jehanne Dussert
arXiv · 2026-05-23
This paper challenges the prevailing 'audit-time' model of AI compliance, arguing that systems deployed under the EU AI Act require continuous, runtime monitoring rather than one-off binary verdicts. The authors introduce 'governance from metrics,' a principle that derives regulatory compliance as an ongoing signal from system observability, and implement it in an open-source framework called govllm, which routes model selection based on accumulated compliance scores. The framework uses a panel of LLM evaluators specialized per regulatory criterion (EU AI Act, GDPR, ANSSI, accessibility), treating inter-judge disagreement as a regulatory uncertainty signal requiring human arbitration. Validated on 49 annotated prompt/response pairs across five criteria, the study finds agreement rates ranging from 51.5% to 69.1% across four small language models, with no single model dominating all criteria, and documents a judge-specific position bias that degrades agreement by up to 25 percentage points — empirically motivating a multi-model jury design for AI governance.
- AI policy
- Quality assurance
Research
Dual-Use AI Face Swap Apps Are Mostly Unsafe: A Systematic Safety Audit
Alaa Daffalla, Sarah Chao, Eric Zeng
arXiv · 2026-05-23
This paper audits 420 AI face swap apps for iOS and Android to assess whether they implement safety measures against generating synthetic non-consensual intimate imagery (SNCII). The authors manually tested 155 eligible apps and found that 70% have no technical safeguards preventing the creation of nude face swaps. The study also found that while no apps self-describe as nudification apps, the majority lack specific terms of service provisions prohibiting such harmful uses. The authors conclude that platforms and lawmakers must mandate safety filters in dual-use AI image editing applications to mitigate SNCII threats.
- AI policy
- Quality assurance
Research
Fundamental Limitation in Explaining AI
Atsushi Suzuki, Jing Wang
arXiv · 2026-05-23
This paper mathematically proves a fundamental 'quadrilemma' showing that no AI explanation can simultaneously satisfy four conditions: operating in a complex environment, achieving good AI performance, being interpretable, and being completely faithful to the AI's actual behavior. The result implies that in most real-world deployments, complete faithfulness of explanations must be sacrificed, and that practitioners should focus on explaining only the parts most relevant to their application. For AI governance and policy, this finding carries a direct consequence: regulatory frameworks and oversight mechanisms must be designed with the recognition that AI explanations are always, in principle, incomplete.
- AI policy
- Quality assurance
Research
The Open Source Economic Index of AI Adoption and Capability
Seamus Somerstep, Aritra Guha, Divesh Srivastava et al.
arXiv · 2026-05-23
This paper introduces an open-source economic index that measures both AI adoption and AI capability across occupations. For adoption, it uses publicly available user-LLM chat data combined with O*NET task data to replicate findings from frontier AI labs, identifying finance, computer science, and arts sectors as having the highest adoption rates. For capability, the authors build a benchmarking system grounded in O*NET occupations and model-context-protocol (MCP) servers, testing Kimi-k2.5 with an OpenAI agents SDK across 9 occupations; results show AI can execute high-level workflows correctly but frequently makes errors in granular details such as specific tool calls. The work matters because it provides a reproducible, open-source method for tracking how AI is displacing or augmenting discrete labor tasks across the workforce.
- Workforce
- AI policy
Research
The Governance Inversion Hypothesis: Why More AI Regulation May Produce Less Organisational Control
Victor Frimpong
arXiv · 2026-05-23
This paper introduces the Governance Inversion Hypothesis (GIH), arguing that expanding AI regulation can paradoxically reduce rather than improve organisations' actual operational control over AI systems. Drawing on institutional theory and governance scholarship, it identifies four mechanisms driving this inversion: authority fragmentation, symbolic governance expansion, externalisation of control, and authority paralysis. As regulatory frameworks become more layered and procedurally dense, organisations may lose coherent authority, technical visibility, and meaningful intervention power over increasingly opaque AI infrastructures. The paper extends institutional decoupling theory by framing governance inversion as a structural condition in which the central risk is not too little governance, but governance that appears robust while undermining effective control.
- AI policy
- Enterprise
Research
GlobalDentBench: A Multinational Benchmark for Evaluating LLM Clinical Reasoning in Dentistry with Expert Calibration
Junjie Zhao, Jingyi Liang, Zhenyang Cai et al.
arXiv · 2026-05-23
GlobalDentBench is the first multinational dental benchmark for evaluating large language models (LLMs) in clinical dentistry, covering 14 specialties across 88 countries with 8,978 expert-validated questions spanning multiple formats and three progressive reasoning levels. Evaluation of 12 frontier LLMs showed sharp performance drops as task complexity increased—accuracy fell from 81.34% on multiple-choice questions to just 22.34% on case-based questions, and from 74.01% at the knowledge-recall level to 35.71% at individualized reasoning. Most critically, risk analysis of real-world dental cases revealed an overall unsafe rate of 31.01% in LLM-generated clinical recommendations, with 4.51% posing risks of irreversible patient harm. These findings highlight fundamental safety and reasoning limitations in current LLMs and underscore the urgent need for rigorous validation before deploying AI in healthcare settings.
- Quality assurance
- Certifications
Research
Catching magnetic resonance imaging outliers in artificial intelligence-supported radiotherapy workflows: unsupervised detection and localization of image anomalies using deep learning
Mustafa Kadhim, Viktor Rogowski, Emilia Persson et al.
arXiv · 2026-05-23
This paper presents a fully automated, unsupervised deep learning framework for detecting and localizing anomalous MRI images in radiotherapy workflows. Trained on public pelvic and brain MRI datasets, the two-stage system compresses images into discrete tokens and models the distribution of normal data to flag out-of-distribution inputs. The framework achieved AUCs of 0.97 for pelvic MRI and 0.81 for brain MRI, with heatmaps showing strong spatial agreement with ground-truth anomaly locations. These results suggest unsupervised anomaly detection could serve as an automated quality-control layer to prevent unexpected AI behavior in clinical radiotherapy pipelines.
- Quality assurance
Research
Is Decentralized AI Governable? From Regulative Policy to Constitutive Protocol
Botao Amber Hu, Helena Rong
arXiv · 2026-05-23
This paper examines why existing AI governance frameworks fail when applied to decentralized AI (DeAI) systems. The authors analyze DeAI as a six-layer stack (model, training, compute, harness, identity, and ownership) and show that partial decentralization across these layers creates a 'governance vacuum' with two distinct failures: an accountability gap (no identifiable responsible party) and an incapacitation gap (even an identified party cannot alter the running system). Drawing on Lessig's modalities of regulation and Searle's distinction between regulative and constitutive rules, the authors argue governance must shift from policy-level normative address to protocol-level architectural constraint that shapes what actions are possible within AI systems. They identify four ethical conditions — legitimacy, contestability, transparency, and non-domination — that such protocol-based governance must satisfy to prevent unaccountable technocratic power.
- AI policy
Research
Structured Visual Evidence Decomposition for Evidence-Grounded Multimodal Screening of Obstructive Sleep Apnea-Hypopnea Syndrome
Chen Zhan, Yingchen Wei, Xiaoyu Tan et al.
arXiv · 2026-05-23
EviOSAHS is a multimodal AI framework designed to screen patients for obstructive sleep apnea-hypopnea syndrome (OSAHS) before formal sleep studies. It decomposes frontal facial images into seven structured anatomical queries—covering the neck, chin, mouth, facial fat, jaw, midface, and nose—converts visual findings into structured evidence cards, and then combines these with clinical data for a final large language model adjudication. Evaluated on a 642-subject cohort, the system achieved 88.47% accuracy, 94.86% sensitivity, and a 5.14% false-negative rate, outperforming direct multimodal prompting and simpler pipelines. The authors caution that EviOSAHS should serve as a triage assistant only, with prospective validation and external testing required before clinical deployment.
- Quality assurance
Research
Side-by-side Comparison Amplifies Dialect Bias in Language Models
Kritee Kondapally, Claire J. Smerdon, Pooja C. Patel et al.
arXiv · 2026-05-23
This paper investigates how language models exhibit bias against African-American Vernacular English (AAVE) compared to Standard American English (SAE), even without explicit dialect labels — a phenomenon called covert dialect bias. The authors find that this bias is significantly amplified when SAE and AAVE tweets are evaluated side by side (a ranking setting), compared to evaluating them in isolation, and worsens further when dialect labels are explicitly provided. While counterfactual fairness finetuning reduces bias in isolation evaluations, these gains do not consistently hold in side-by-side comparisons, suggesting current mitigation and evaluation approaches underestimate the problem. This matters because side-by-side comparison closely reflects real-world high-stakes decision-making contexts, such as candidate ranking, where deployed models may perpetuate racial stereotypes.
- AI policy
- Quality assurance
Research
A governance horizon for ethical-use constraints in open-weight AI models
Weiwei Xu, Hengzhi Ye, Haoran Ye et al.
arXiv · 2026-05-23
This paper audits over 2.1 million model repositories on Hugging Face Hub to test how well ethical-use constraints embedded in open-weight AI models propagate through downstream derivatives. The study finds that restriction evidence decays with a half-life of 1.31 derivation steps, and beyond seven downstream generations at least 80% of descendant models lack sufficient public evidence for a governance determination — a threshold the authors call the 'governance horizon.' Policy design emerges as the binding constraint: mandatory-declaration approaches that resolve orphaned lineage components outperform inheritance-only designs even at moderate enforcement, whereas inheritance-only rules cannot recover undecidability caused by missing upstream intent. The findings conclude that disclosure-based governance has a structurally shallow reach in open-weight AI supply chains, and that meaningful accountability requires provenance mechanisms that actively propagate governance signals through derivation itself.
- AI policy
Research
Cryptographic Attestation for AI Agent Governance under the EU AI Act: A Survey of Approaches and Standards
Anton Sokolov
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-23
This survey paper examines how cryptographic attestation can be used to satisfy governance and compliance requirements imposed on AI agents under the EU AI Act (Regulation (EU) 2024/1689). The authors argue that conventional GRC tooling and observability platforms produce only operator-side assertions rather than independently verifiable, cryptographically signed evidence of agent behavior, leaving a structural gap in conformity assessment under Articles 12, 14, 50, and 72 and Annex IV. The paper surveys emerging attestation approaches—including hardware-rooted (TEE-based) attestation, software-only cryptographic attestation, and identity-focused attestation—and introduces a six-axis taxonomy, with the OVERT 1.0 open standard highlighted as the first horizontal specification targeting this category. The work identifies seven open research problems and maps the landscape to relevant standards bodies including ETSI, CEN-CENELEC JTC 21, ISO/IEC JTC 1/SC 42, and IETF SCITT, making it directly relevant to AI policy compliance and certification efforts in the EU and beyond.
- AI policy
- Enterprise
- Certifications
- Quality assurance
Research
From Frontier to Shadow AI: A Simmering Threat to Assurance and Security in Critical Infrastructure
Mohan Baruwal Chhetri, Shahroz Tariq, Tooba Aamir et al.
arXiv (Cornell University) · 2026-05-23
This paper presents the first empirical study of 'shadow AI'—the unsanctioned use of frontier AI tools outside official organizational controls—in Australian critical infrastructure (CI) environments. Drawing on semi-structured interviews with senior executives and functional leaders across 27 CI organizations in the Communications, Energy, and Water and Sewerage sectors, the researchers develop a threat model identifying three primary mechanisms of security degradation: boundary bypass, unassessed capability expansion, and loss of observability via governance circumvention. The findings show that shadow AI erodes established assurance and oversight mechanisms, amplifying risks to data protection, decision reliability, and regulatory compliance in sectors responsible for essential services. The authors argue that existing security and compliance frameworks are fundamentally challenged by these unmanaged risks, necessitating tailored governance and control strategies.
- AI policy
- Quality assurance
- Certifications
- Enterprise
Research
Green AI adoption from academia to industry in China: the missing middle and technological heterogeneity in diffusion
Longbin Geng, Rasmus Lema
Sustainable Futures · 2026-05-23
This study examines how Green AI technologies diffuse from academic institutions to firms in China, using data from 5,201 patent licensing agreements between 2008 and 2023. The authors find a U-shaped 'missing middle' pattern where moderately capable firms adopt Green AI at below-average intensity, while both low- and high-capability firms adopt more. State-owned enterprises show substantially higher adoption rates, and the pattern is concentrated in ICT-for-greening-ICT sectors and driven by non-exclusive licensing contracts. The findings have implications for sustainable innovation policy, suggesting that diffusion support mechanisms may need to specifically target mid-capability firms.
- Enterprise
- AI policy
Research
ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions
Xianzhong Ding, Yangyang Yu, Changwei Liu et al.
arXiv · 2026-05-22
ContextEcho is a benchmark and harness for measuring how much a language model's persona shifts over long agentic-coding sessions — the kind that production deployments actually run, involving thousands of tool-using turns, in-session compaction, and hours of interaction. Using a 25-probe identity suite applied to three anonymized Claude Code sessions spanning 3,746–9,716 turns and tested across 23 frontier models, the study finds that persona drift is a general phenomenon across organizations rather than limited to specific model families, that in-session compaction does not reliably reset it, and that a single-shot anchor can restore the trained register. Downstream effects differ by mode: drift can help tool-using continuation but breaks formatting contracts and inflates output length in tool-free chat. The work matters for enterprise deployers and quality-assurance teams because it shows that evaluations performed on short dialogues may miss substantial drift that end users actually encounter.
- Quality assurance
- Enterprise
Research
Improving Labeling Consistency with Detailed Constitutional Definitions and AI-Driven Evaluation
Konstantin Berlin, Adam Swanda
arXiv · 2026-05-22
This paper proposes an AI-driven labeling workflow where a frontier large language model interprets detailed, per-category constitutional definitions to produce more consistent and accurate golden labels than human annotators working from the same documents. Tested on three content moderation categories—harassment, hate speech, and non-violent crime—the approach reduces cross-model inconsistency by up to 57x compared to paragraph-level definitions. The system keeps humans responsible for high-level definitional decisions while delegating individual labeling calls to the LLM, and introduces a dual-axis formulation that scores intent and content independently across full conversations. This matters for quality assurance in automated labeling pipelines, where label drift and inconsistency undermine downstream model training and content moderation systems.
- Quality assurance
- Enterprise
Research
How Well Do Models Follow Their Constitutions?
Arya Jakkli, Senthooran Rajamanoharan, Neel Nanda
arXiv · 2026-05-22
This paper proposes and applies a multi-method audit pipeline to measure how well frontier AI models actually follow their own published behavioral specifications under adversarial, multi-turn pressure. The pipeline decomposes each lab's specification into hundreds of atomic testable tenets, generates adversarial scenarios, and validates flagged transcripts against the relevant specification. Applied across seven models per specification, the study finds meaningful generational improvement: Claude's violation rate falls from 15.0% (Sonnet 4) to 2.0% (Sonnet 4.6), and GPT's from 11.7% (GPT-4o) to 3.6% (GPT-5.2 medium reasoning), with remaining failures clustering around AI-identity questioning, irreversible agentic actions, and fabricated quantitative claims. The findings matter for AI governance because they demonstrate that published specifications can be treated as auditable targets, while also revealing persistent gaps that policy and certification frameworks would need to address.
- AI policy
- Quality assurance
Research
Beyond Final Answers: Auditing Trajectory-Level Hallucinations in Multi-Agent Industrial Workflows
Harshada Badave, Santosh Borse, Andrea Gomez et al.
arXiv · 2026-05-22
This paper presents Trajel, a dataset and evaluation framework for detecting hallucinations in multi-agent LLM workflows at the trajectory level—meaning it audits the intermediate reasoning and action steps, not just final outputs. The authors introduce a five-type hallucination taxonomy (factual, referential, logical, procedural, and scope-based) and benchmark detection models across subtask, trajectory, and long-context levels using expert-annotated agent traces. Key findings show that existing benchmarks miss the most common failure modes, nearly half of hallucinated trajectories involve multiple hallucination types simultaneously, and even high-accuracy automated detectors struggle with the subtlest types. This work matters for quality assurance in AI deployments, demonstrating that trajectory-aware evaluation is necessary for safer use of autonomous agents in industrial workflows.
- Quality assurance
Research
Identifying and Mitigating Systemic Measurement Bias in Production LLM Inference Benchmarks
Ashok Chandrasekar, Jason Kramberger
arXiv · 2026-05-22
This paper identifies a fundamental measurement flaw in widely used LLM inference benchmarks: their single-process, asyncio-driven architectures create client-side queuing bottlenecks under high concurrency that artificially inflate key latency metrics like Time to First Token (TTFT) and Time Per Output Token (TPOT). Using M/G/1 queue modeling, the authors mathematically show how Python's Global Interpreter Lock (GIL) causes this systematic bias at scale. They propose a multi-process evaluation framework that eliminates this overhead and introduce a composite metric called Normalized Time Per Output Token (NTPOT) to more accurately capture end-to-end latency across sequence lengths. The work matters because accurate, reproducible benchmarking is essential for validating that LLM serving systems meet production Service Level Objectives at thousands of queries per second.
- Quality assurance
- Enterprise
Research
Human-AI Collaboration in Science at Scale: A Global Large-scale Randomized Field Experiment
Binglu Wang, Weixin Liang, Jiahui Xue et al.
arXiv · 2026-05-22
This large-scale randomized field experiment delivered LLM-generated feedback to authors of over 31,000 arXiv preprints across 150 fields and more than 45,000 researchers from 133 geographic regions. Authors who received AI feedback were significantly more likely to revise their manuscripts, representing a 12.55% relative increase over the baseline revision rate, and subsequently used LLM tools more in their future work. Effects were strongest among researchers from non-English-dominant regions, less-cited manuscripts, and teams with lower h-indexes and earlier career stages, suggesting AI feedback is most valuable where access to timely critique is otherwise scarce. The study provides causal evidence that structured AI interventions can redistribute scientific feedback more equitably across the global research system.
- Workforce
- AI policy