News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
RoboSafe KPI Thresholds Specification v1.0
Chang Xiong
Zenodo (CERN European Organization for Nuclear Research) · 2026-08-18
This specification document defines the quantitative safety thresholds and compliance framework for RoboSafe certification of social robots across three deployment levels: retail/corporate (99.9% reliability), hospitality/public (99.99%), and healthcare/elder care (99.999%). Six Key Performance Indicators are tracked—including hard-block accuracy, PHI redaction coverage, and response substitution latency—alongside 11 compliance checks and three possible certification verdicts. The framework is adapted from Waymo's reliability methodology and is designed to be fully reproducible by third parties without access to proprietary platforms. Released under CC BY 4.0, this document provides a concrete, auditable standard for governing AI-driven social robot deployments in sensitive environments.
- Certifications
- Quality assurance
Research
The Application of Artificial Intelligence and the Resilience of Manufacturing Enterprises: Mechanisms of Action and Heterogeneity Boundaries
Li Ran
Highlights in Business Economics and Management · 2026-08-18
Analyzing China A-share listed manufacturing firms from 2015 to 2024, this study finds that AI adoption significantly boosts enterprise resilience across three dimensions: resistance capacity, recovery capacity, and innovation capacity. The primary mechanism is technological innovation, with talent incentives and reduced management costs also playing supporting roles. Effects vary by ownership structure, firm size, region, and industry type. The findings offer empirical guidance for manufacturing firms seeking to use AI to strengthen supply chain security and withstand external shocks.
- Enterprise
Research
Does generative AI mean the “end of history” for pharmacovigilance automation? towards a framework for the future of human-AI systems
Leihong Wu, Joshua Xu, Oanh Dang et al.
Frontiers in Drug Safety and Regulation · 2026-08-18
This perspective paper examines whether large language models (LLMs) and generative AI enable full automation of pharmacovigilance (PV) workflows, concluding they do not. The authors introduce a 'computable PV' framework distinguishing routine tasks amenable to automation—such as completeness checks and duplicate detection—from complex tasks like causality assessment that still require expert judgment. They argue that hybrid architectures combining large models, small models, and rule-based components are necessary, and that final decisions must remain under human oversight with transparency and validation as priorities.
- Quality assurance
- AI policy
Research
Governing artificial intelligence-enabled labour surveillance: A multi-level framework for legal, organisational and collective governance in the digital workplace
Fu‐Hsuan Chen
Journal of Industrial Relations · 2026-08-18
This paper develops a multi-level governance framework for AI-enabled workplace surveillance, arguing that predictive scoring, affective inference, and workplace datafication create institutional misalignment across legal, organisational, and collective bargaining levels rather than a simple regulatory gap. Drawing on comparative analysis of the EU AI Act, GDPR, and platform labour regulation alongside practices in the US and selected Asian jurisdictions, it identifies a recurring triad of governance failure: ex post and abstract legal controls, unilateral organisational policies, and collective bargaining hampered by technical opacity. The paper proposes a 'sustainable digital labour contract' as a governance device that operationalises legal duties through participatory design and collectively negotiated protections for data access, grievance rights, and algorithmic accountability. The findings matter for how workers, firms, and regulators respond to AI systems that shift control from observable conduct to inferred states.
- Workforce
- AI policy
Research
AI safety evaluation in an underrepresented population: real-world performance of clinical decision support and frontier language models on Medicaid patient messaging triage
Sanjay Basu, Sadiq Patel, Parth Sheth et al.
BMC Medical Informatics and Decision Making · 2026-08-18
This retrospective study evaluated AI triage tools on 2,000 real-world messages from Medicaid patients across three U.S. states, finding that no single tool or combination—including frontier large language models and ensemble methods—met a pre-specified benchmark for autonomous (physician-unassisted) triage (sensitivity and specificity both ≥0.80). Two configurations did meet a sensitivity-floor target (0.80–0.95) when paired with clinician review of flagged messages, following a classical clinical-screening pattern. The study also found that real-world Medicaid patient messages had lower reading levels and more colloquialisms than the scripted scenarios used in prior AI triage research, highlighting a significant evaluation gap for underrepresented populations. The findings suggest that current AI triage tools require ongoing clinician oversight and cannot safely operate autonomously in this setting.
- Quality assurance
- AI policy
Research
Artificial Intelligence and National Security Governance in Africa: A Comparative Study of Nigeria, South Africa, Kenya, Rwanda and Egypt
Paul 'Seun Sosina, Odebiyi Olusola James
Zenodo (CERN European Organization for Nuclear Research) · 2026-08-18
This comparative qualitative study examines how Nigeria, South Africa, Kenya, Rwanda, and Egypt are adopting AI for national security functions such as intelligence gathering, threat detection, border surveillance, and cybersecurity. Findings show that adoption levels vary significantly—Egypt and Rwanda use centralized policy models, Kenya and South Africa face regulatory and implementation gaps, and Nigeria is constrained by infrastructure and capacity deficits. Common challenges across all five countries include algorithmic bias, poor data quality, privacy risks, inadequate oversight, and dependence on foreign technologies. The study recommends risk-based regulation, human rights–grounded governance frameworks, indigenous capacity development, and harmonized continental standards to guide responsible AI deployment in security contexts.
- AI policy
Research
Artificial Intelligence and National Security Governance in Africa: A Comparative Study of Nigeria, South Africa, Kenya, Rwanda and Egypt
Paul 'Seun Sosina, Odebiyi Olusola James
Zenodo (CERN European Organization for Nuclear Research) · 2026-08-18
This comparative qualitative study examines how five African nations—Nigeria, South Africa, Kenya, Rwanda, and Egypt—are deploying artificial intelligence for national security functions including intelligence gathering, surveillance, cybersecurity, and conflict early warning. Drawing on documentary analysis of national policies and African Union frameworks, the study finds that AI adoption improves threat detection and data processing but varies significantly by country, with Egypt and Rwanda showing centralized policy approaches and Nigeria constrained by infrastructure and capacity gaps. Common challenges across all five countries include algorithmic bias, poor data quality, inadequate oversight, skills shortages, and dependence on foreign technology. The study recommends risk-based regulation, human rights-grounded governance, indigenous capacity development, and harmonized regional standards to support responsible AI-enabled security.
- AI policy
Research
Have Your Say: stakeholder responses on EU digital-identity, trust-service and AI files (2026 snapshot)
Anton Sokolov
Open MIND · 2026-08-18
This dataset paper presents a corpus of 2,244 stakeholder responses submitted to 16 European Commission 'Have Your Say' consultations covering EU digital identity (eIDAS 2.0), trust services, and AI regulation, harvested in August 2026. The authors provide structured analysis tools and keyword-frequency counts showing that among 786 organisational responses, topics like portability and interoperability were engaged by 124 respondents, audit or certification by 70, and runtime attestation of automated systems by only 4–15, suggesting significant gaps in the professional debate over accountability for delegated automated action. The dataset is privacy-reduced (personal names removed) and includes reproducible analysis scripts, but the authors explicitly caution that over half the raw corpus is a Slovak citizen campaign opposing digital identity, and that only published responses are visible, so no representativeness or influence inferences are supported. The work is relevant to policy and certification communities tracking how stakeholders are engaging—or failing to engage—with the implementing detail of the AI Act, eIDAS 2.0, and related EU regulatory frameworks.
- AI policy
- Certifications
Research
Token Optimization and Context Window Management in Multi-Agent AI Workflows
Dvir Shamay
arXiv · 2026-08-17
This paper presents a practitioner engineering framework for reducing token costs and latency in multi-agent AI workflows, demonstrated through an internal production dashboard that extracts structured work items from meetings, email, and chat. Six optimization patterns—including context stratification, semantic caching, and inter-agent communication compression—cut cold-load latency from a baseline of roughly 3.5–10.5 minutes to 61–116 seconds and achieved an estimated 60–70% token reduction. A controlled context-composition study across 2,420 trials and 11 model configurations found that mixing high- and low-relevance items in the context window (a 'relevance-contrast context') improved relevance-score concordance by +0.077 over using high-relevance items alone (Cohen's d = 0.49, Holm-adjusted p < .001). The work provides repeatable patterns and evaluation methods that sit between model research and production deployment, offering measurable improvements in speed, cost, and reliability for enterprise AI workflows.
- Enterprise
- Quality assurance
Research
Benchmarking the Benchmarks: Evaluating Automated Safety Benchmarks for Small Language Models
Nyamtulla Shaik, Fengjun Li, Bo Luo
arXiv · 2026-08-17
This paper evaluates whether existing automated AI safety benchmarks, designed for large language models, can reliably assess safety and bias in Small Language Models (SLMs). The authors test five widely used benchmark suites across 26 open-source SLMs and find that ambiguous judgments dominate results, correlating with prompt complexity and model architecture—leading them to conclude that LLM-centric safety benchmarks are insufficient as standalone evidence for SLM safety assessment. They identify a capability-safety confound where model capability is mixed with apparent safety, and show that aggregate leaderboard rankings change significantly under different treatments of ambiguous outputs even when the underlying model responses remain unchanged. This matters because SLMs are increasingly deployed in resource-constrained and privacy-sensitive settings where unreliable safety evaluation creates real security and societal risks.
- Quality assurance
- Certifications
Research
The Acknowledgment Point Is the System: Durable Policy-Decision Receipts for AI Audit Evidence
Neeraj Kumar Singh Beshane
arXiv · 2026-08-17
This paper presents RuntimeGuard-AI, a research prototype that binds each AI policy decision to its exact policy source and commits a cryptographically signed, privacy-minimizing audit record at a caller-selected synchronization boundary, returning an Ed25519-signed receipt that explicitly states whether that durability boundary completed. The system groups committed records into chained, signed Merkle epochs verifiable by an external auditor, creating tamper-evident audit trails for AI governance. Benchmarks on Apple M4 Pro hardware reveal a concrete durability-latency trade-off: buffered signed evidence reaches ~27,193 requests/s at ~142 µs median latency, while fully synchronous per-record writes drop throughput to ~242 requests/s at ~16 ms latency, with 100,000-record epoch sealing taking 97 ms. The authors explicitly note the prototype does not prove model execution, prevent a compromised signer from forking history, or establish legal conformity, making the trade-offs transparent for policy and audit applications.
- AI policy
- Quality assurance
Research
Stranded credentials: how a skill-signaling market absorbed generative AI
Song Yao
arXiv · 2026-08-17
This study audits 444,698 participations on Kaggle (2010–2026) to test whether data science credentials retain their signaling value for future performance in the generative AI era. The authors find that competition medals remain predictive of subsequent leaderboard results almost entirely within the first year of being earned, and that fresh medals retained most of their signaling value through the AI transition. Much of the apparent decline in credential informativeness—82% lost for one competition format—is explained by institutional factors predating AI rather than by AI itself, and the platform's official lifetime-medal tiers discard 13–16% of available predictive information compared to a recency-weighted index built on pre-AI data alone. The paper concludes that credentials are informative, perishable, institution-bound, and interdependent, framing the maintenance of credential value under AI as a high-stakes problem of institutional design.
- Certifications
- AI policy
Research
Language Models Reproduce Human Reductionist Bias and Decision Inconsistency in Neurodevelopmental Disorders Assessment
Maciej Wodziński, Joanna Wodzińska, Kacper Dudzic et al.
arXiv · 2026-08-17
This study audits how large language models (LLMs) compare to human clinicians (18 physicians and 17 psychologists) when making support-eligibility decisions for neurodevelopmental disorders. Both LLMs and humans showed inconsistency between functional-level assessments and final support decisions, and neither group was meaningfully susceptible to anchoring or representativeness heuristics. LLMs expressed significantly higher intellectual humility than experts but interpreted 'basic life needs' reductively—prioritizing biological survival over communicative and social needs—mirroring a medical-model bias rather than a neurodiversity-affirming framework. The findings argue that evaluating AI in high-stakes clinical contexts requires scrutinizing the conceptual frameworks AI systems operationalize, not just measuring accuracy or bias resistance.
- AI policy
- Quality assurance
Research
Appearing Legitimate is Not Enough: Interrogating Synthetic Agents in Representational Processes through a Participatory Design Lens
Aditya Nayak, Aditi Vashistha, Alissa Centivany et al.
arXiv · 2026-08-17
This paper examines the growing use of LLM-based synthetic agents as substitutes for human participants in representational processes such as policy consultation, jury deliberation, and humanitarian diplomacy. The authors argue that legitimacy in democratic institutions depends not merely on the informational or consensus-generating contributions of participants, but on the intrinsic value of human participation itself, making synthetic agent substitution ethically and politically problematic. Applying a Participatory Design framework across three case studies—local policy, enterprise jury deliberation, and global diplomacy—they identify ethical, representational, and methodological risks and conclude by proposing soft and hard design boundaries for overseeing LLMs in these contexts.
- AI policy
- Enterprise
Research
Institution-Specific LLM Prompting Recovers PHI That De-identification Systems and Their Gold Standards Both Miss
Daniel Palacios, Matthew Brady Neeley, Angel Adetomike Otto et al.
arXiv · 2026-08-17
This study tests whether large language models (LLMs) with institution-specific prompting can better de-identify protected health information (PHI) in clinical notes than existing purpose-built systems, focusing on locally-determined PHI like hospital abbreviations and internal codes that standard tools miss. On 100 annotated pediatric oncology notes from Texas Children's Hospital, the best LLM achieved F1=0.918 versus 0.779 for the leading purpose-built system (Stanford TiDE), and a carefully calibrated single prompt reached recall=0.981 after re-annotation of 227 previously missed PHI spans surfaced by LLM outputs. The paper shows that naming missed institutional PHI categories in the prompt recovered 79% (48/61) of them, and that no multi-agent architecture outperformed well-calibrated single-pass prompting. The findings argue that institution-specific prompt development is the primary adaptation strategy for clinical de-identification, making LLMs a legitimate and auditable alternative to traditional systems for enabling secondary use of health records.
- Quality assurance
- AI policy
Research
What Do Compliance Detectors Read? An Audit of Activation Probes and Guard Models
Saisab Sadhu, Aadit Sengupta, Vinay Kumar Sankarapu et al.
arXiv · 2026-08-17
This paper audits AI compliance detectors—guard models and activation probes—used to check language model outputs against regulatory rules in domains like data protection, healthcare, and financial regulation. The authors demonstrate a failure they call 'rule blindness': deleting, permuting, or substituting the governing rule leaves detection accuracy unchanged, meaning detectors respond to surface features of scenarios rather than the actual rules they are supposed to enforce. They introduce a training-free activation readout called the Internal Compliance Score (ICS), which they subject to the same rigorous scrutiny and find it matches a simple bag-of-words baseline in generalization, though it remains useful for cheaply auditing deployed systems. The findings have direct implications for the reliability of AI-based regulatory compliance monitoring and the validity of audit controls built on current detector technology.
- Quality assurance
- AI policy
Research
Quipu: A Governed Bitemporal Knowledge Graph Store
Steve Brown
arXiv · 2026-08-17
Quipu is an embeddable bitemporal knowledge graph store designed to handle agent-written data with built-in governance, rather than relying on external dashboards or middleware. It enforces four inverted defaults: writes are gated by predicate evaluation before acceptance, data carries two time axes, trust is tracked per named graph under a non-widening lattice, and the governance specification and audit trail are stored as facts within the store itself. In evaluations using the Census benchmark, the gated store eliminated all 6 planted defects versus 6 of 6 surviving in an ungated store, and correctly re-derived 50 of 50 verdicts under bitemporal rules while a latest-only rule set would have misreported all 50. On the external DEMM-Bench benchmark covering 512 property-level governance questions across eight degradation conditions, Quipu answered all correctly with zero overclaim, while container-presence baselines overclaimed on up to 87.5% of cases.
- Quality assurance
- AI policy
Research
LAVA: Logic-Aware Validation and Augmentation Framework for Large-Scale Financial Document Auditing
Ruoqi Shu, Xuhui Wang, Isaac Wang et al.
arXiv · 2026-08-17
LAVA is a modular AI pipeline for auditing financial documents—such as payroll, tax compliance, and loan underwriting records—that combines multimodal large language models with symbolic and arithmetic verification across four stages: document-rule retrieval, layout-preserving extraction, metadata enrichment, and auditable verification. The framework addresses challenges posed by heterogeneous document formats, context-dependent content, and embedded business rules that existing pipelines handle unreliably. Evaluated on a large real-world benchmark with diverse financial documents and dozens of expert-curated validation rules, LAVA outperforms baselines in hallucination control and edge-case handling while maintaining efficient token usage. These results demonstrate its practicality for high-volume, time-critical financial auditing under strict enterprise constraints.
- Enterprise
- Quality assurance
Research
Unsupervised Anomaly Detection for Image Dataset Quality Assurance in Multi-Center Breast MRI
Chiara Tappermann, Steffen Renisch, Lars Ole Schwen et al.
arXiv · 2026-08-17
This paper addresses automated quality assurance (QA) for medical AI datasets, specifically targeting anomalous or corrupted images in multi-center dynamic contrast-enhanced breast MRI. The authors construct a benchmark of 17 realistic anomaly types drawn from six public datasets—covering protocol violations, processing errors, and incorrect anatomical regions—and evaluate four unsupervised anomaly and out-of-distribution detection methods. The best-performing approaches achieve AUROCs of up to 0.954, reliably catching medium-to-far out-of-distribution samples, while near-OOD cases and data from unseen institutions remain challenging. The work provides practical guidance for building scalable, automated dataset QA pipelines for high-risk medical AI, directly supporting regulatory and safety requirements in that domain.
- Quality assurance
- AI policy
Research
MIRROR: Multimodal Intelligent Radiology Reasoning and Observation Reporter
Vignesh Nagarajan, Sriram Venkatapathy
arXiv · 2026-08-17
MIRROR is a research prototype for automated radiology reporting that chains a multi-label image classifier, a Grad-CAM localizer, and a language layer that writes reports using only classifier outputs—never the raw image—so that every stated finding is auditable against the model's probability vector. Testing on ChestMNIST yields a macro AUROC of 0.729 across 14 labels, but the system emits no positive prediction for 11 of them at the default 0.5 threshold, and its Brier score of 0.045 is nearly matched by a predictor that ignores the image entirely (0.047). The paper's core finding is that aggregate metrics commonly used in radiology AI are misleading under class imbalance, flattering models that effectively do nothing, and must be reported against a naive baseline floor. This matters for quality assurance and certification of AI diagnostic tools, as it demonstrates that strong-looking headline numbers can obscure near-total failure at the decision level.
- Quality assurance
- Certifications
Research
Toward Better Assessment of LLMs' Performance in Clinical Error Detection
Yifan Zhang, Rahmatollah Beheshti
arXiv · 2026-08-17
This paper evaluates 15 large language models on clinical error detection across 4 benchmark datasets and 3 languages, finding that 13 of 15 models perform below random chance at paired discrimination even when achieving moderate F1 scores. The authors show that standard aggregate metrics like F1 can systematically rank the weakest discriminators highest, because F1 and pairwise accuracy are driven in opposite directions by the same underlying model bias. They also find language-dependent bias patterns — the same model may default to 'no error' in one language and over-flag errors in another — and introduce a procedure to score the evidence models cite, revealing that models locate relevant content but fail to produce the correct verdict on clean counterparts. For safety-critical clinical NLP applications, the authors advocate supplementing aggregate metrics with paired evaluations to better reflect true model reliability.
- Quality assurance
- Certifications
Research
Characterizing Agentic Flooding of Government Services
Chris Schmitz, Lewis Hammond, Alan Chan
arXiv · 2026-08-17
This paper introduces the concept of 'agentic flooding of government services,' where AI agents — particularly large language models — generate surges in demand that strain government services by automating tasks like benefits applications, policy inquiries, and public comment submissions. Drawing on a dataset of 84 potential flooding cases across 11 jurisdictions, the authors find that flooding is likely already occurring widely, with the highest near-term risk concentrated in financially attractive but administratively complex services. The authors develop a risk matrix to assess service exposure and map out potential government responses, warning that the fastest countermeasures — such as fees — risk undermining equitable access to public services. They recommend targeted near-term actions that governments can take to address flooding without sacrificing service equity.
- AI policy
- Workforce
Research
"If It Looks Like a User": Measuring Real-Time Moderation Effects via Social Media Simulation
Enrico Verdolotti, Gianluca Nogara, Luca Luceri et al.
arXiv · 2026-08-17
This paper develops a calibrated agent-based social media simulator—an extension of SimSoM—grounded in real-world vaccine discourse data from the COVID-19 pandemic to study content moderation effects. The simulator is validated against empirical data across temporal, distributional, and structural dimensions using CMA-ES optimization, and is shown to reproduce key statistical signatures such as activity distributions, post/reshare ratios, and temporal patterns. A key finding is that static (retroactive) moderation evaluations significantly overestimate the effectiveness of user bans compared to dynamic (real-time) moderation, because compensatory resharing by remaining users dampens the expected reduction in low-quality content. The work argues that simulation-based evaluation is necessary for accurately assessing content moderation policies, and provides a reusable empirical framework for doing so.
- AI policy
- Quality assurance
Research
A Regulatory Placebo? The Systemic Failure of Mandatory GenAI Labeling
Jingyi Chen, Chaofan Bu, Shibo Yan et al.
arXiv · 2026-08-17
This paper critically examines mandatory labeling requirements for generative AI (GenAI) content, arguing that such regulations are reactive and symbolic rather than effective. The authors analyze three theoretical frameworks used to justify labeling mandates—value dilution theory, information authenticity theory, and proactive regulation theory—and contend that all three reflect cognitive limitations among regulators regarding how modern AI technology actually works. The study finds that mandatory labeling creates implementation dilemmas, risks slowing AI development, and functions as a 'regulatory placebo' that masks deeper governance challenges. The authors advocate shifting from identity-label governance to content governance to better address the genuine legal and societal demands posed by GenAI.
- AI policy
Research
A Policy Algebra for Trust-Preserving Agentic AI Execution
Bhaskar Tripathi, Anurag Kumar, Ramendra Kumar et al.
arXiv · 2026-08-17
This paper addresses a critical gap in enterprise AI agent deployment: current agentic frameworks optimize for task capability but lack formal guarantees around authorization, data access, budget limits, and auditability. The authors propose a 'policy algebra' that defines a reliability envelope for agent execution, composing security profiles and runtime obligations through formal operations (joins, intersections, budget narrowing, approval inheritance) that are both trust-preserving and minimally restrictive. Their evaluated runtime catches 94.8% of policy-violating events while maintaining an 86.9% task-completion rate, eliminates observed policy-monotonicity violations, and raises audit completeness to 98.6%. This work is directly relevant to enterprises deploying AI agents, offering formal correctness conditions and trace evidence to ensure agents are not just capable but reliably and accountably capable.
- Enterprise
- AI policy