News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
5571 items
Research
Control Under Compression: Reliability Frontiers for Tool-Using Agents
Yinghan Hou, Zongyou Yang
arXiv · 2026-08-02
This paper examines what happens to AI agent reliability when the system-level instructions that govern tool use—called agent control contexts (ACCs)—are compressed to reduce token costs. Using a benchmark of 15,525 runs across nine ACCs, three task families, and six compression budgets, the authors find a nonlinear reliability frontier: at 75% retained context, top methods still achieve ~92–93% success near the full-context baseline of 93.8%, but between 50% and 35% retention methods diverge sharply, with the best achieving only 47% and the worst 19.9%, and protocols become fragile below 25%. The study shows that failures manifest primarily as tool-execution and action-parsing errors, that no single compressor ranks best across all contexts, and that ACC compression must be treated as a runtime-reliability problem evaluated through executable outcomes rather than just a token-reduction technique.
- Quality assurance
- Enterprise
Research
Don't Offer What Can't Be Done: Deterministic Executability Gating for LLM Skill Selection at Scale
Ortal Ashkenazi, Vitalii Kloz, Mykhailo Ulianchenko
arXiv · 2026-08-02
This paper describes a three-stage pipeline deployed in Wix's customer-care LLM agent ('Helpmate') to prevent the model from selecting skills it cannot actually execute given a user's current account state. A semantic matcher first narrows candidate skills, then a deterministic 'executability gate' removes any skill whose hard-stop conditions are already met, and finally the LLM chooses among the remaining valid options. In a production analysis of 756.6K user messages, the combined approach reduced skill-description context by 90.5% compared to exposing all skills to every message, and a counterfactual replay showed that without the gate the model would have selected a non-executable skill in 7.8% of tested conversations. The work demonstrates that deterministic pre-filtering meaningfully improves LLM agent reliability at scale by stopping the model from offering actions that cannot be completed.
- Enterprise
- Quality assurance
Research
DeBERTa-Sentinel: Toward Transparent and Trustworthy Detection of AI-Generated Text
Muhammad Yousaf Rehman, Muhammad Islam
arXiv · 2026-08-02
DeBERTa-Sentinel is a transformer-based framework for detecting AI-generated text that uses DeBERTa-v3's disentangled attention to identify subtle structural irregularities in synthetic content. Evaluated on the GLC-AIText dataset of 28,057 human and LLM-generated samples (GPT, LLaMA, and Claude), it achieves 97.53% test accuracy, 95.89% precision, 99.33% recall, and 99.53% ROC-AUC, outperforming the RoBERTa-Sentinel baseline. A key design feature is token-level explainability that exposes linguistic markers such as academic phrasing and formal transitions, enabling journalists, educators, and platform trust-and-safety teams to audit detection decisions. This matters for quality assurance and policy contexts where verifiable, auditable content-authenticity tools are needed to combat misinformation and protect academic integrity.
- Quality assurance
- AI policy
Research
Do people rely on ChatGPT more than their peers to detect deepfake news?
Yuhao Fu, Nobuyuki Hanaki
arXiv (Cornell University) · 2026-08-02
This experimental study examines how people weight advice from ChatGPT (GPT-4), human peers, and linguistic experts when identifying AI-generated fake news (deepfake news). Results show participants initially relied more on ChatGPT than on peers, though a 2025 follow-up found greater reliance on linguistic experts than ChatGPT, suggesting shifting trust in AI tools over time. Detection performance improved only when participants relied on high-quality advice, meaning AI-assisted deepfake detection is beneficial only if the AI tool itself is sufficiently accurate. The findings underscore generative AI's dual role as both a creator of disinformation and a potential mitigation tool.
- Quality assurance
- AI policy
Research
Conformance result: Shango MID against the AISVS C9 action-class scenario suite v1.0.0
Ishaan Ghosh
Open MIND · 2026-08-02
This conformance report documents how Shango MID, an AI write-governance layer, performed against the OWASP AISVS C9 action-class scenario suite (v1.0.0). Fourteen test cases were run in declaration-only mode; the original reading showed nine matches, four divergences, and one not-applicable, while a correction attributable to a defect in the suite itself raises the match count to twelve of fourteen. The report is notable for its transparency: it distinguishes verdicts produced by shipped code from those carried by adapter scaffolding, and retains the original contradictory record rather than overwriting it. The single remaining divergence—a composition-aware chain-sealing case—is acknowledged as an open research question unsatisfied even by the suite's own reference implementation.
- Certifications
- Quality assurance
Research
AI-Enabled Post-Process Surface Inspection in Laser-Welded Al–Cu Battery Interconnects: A Critical Review
Maricruz Hernández-Hernández, Adriana C. Flores‐Gallegos
Metals · 2026-08-02
This review paper examines how artificial intelligence can be applied to post-process surface inspection of laser-welded aluminum–copper battery interconnects used in electric vehicles. The authors detail a framework linking controlled image acquisition, defect taxonomy, weld-level datasets, and AI inference tasks—including classification, detection, segmentation, and anomaly detection—with accept–review–reject decision logic and functional validation. The review highlights that Al–Cu laser welding is prone to visible defects and hidden discontinuities due to thermophysical mismatch and intermetallic compound formation, making reliable automated inspection critical for battery pack quality. The proposed framework supports reproducible, risk-sensitive quality assurance while preserving expert oversight for uncertain or high-stakes decisions.
- Quality assurance
Research
Multi-Agent Readiness Scoring Methodology in Bioinformatics Domain
Blagojche Gjorgjioski, Djansel Bukovec, Ivana Vichentijevikj et al.
Future Internet · 2026-08-02
This paper introduces MARS (Multi-Agent Readiness Score), a standardized framework for evaluating how ready bioinformatics large language models are for deployment in autonomous, multi-agent clinical and regulatory environments. Applied to 43 genomic LLMs, the study found a widespread readiness gap: most models fell into 'Not Suitable' or 'Research Prototype' tiers, lacking technical interfaces, structured communication schemas, and provenance tracking. The authors identify a 'competence-readiness gap' where biological predictive ability scales without corresponding improvements in engineering or regulatory utility. MARS incorporates compliance criteria from the EU AI Act, HL7 FHIR, HL7 CDA, and MyHealth@EU, offering a reproducible audit metric to guide system architecture for precision medicine workflows.
- Certifications
- Quality assurance
- AI policy
Research
Children’s Rights in the Age of Artificial Intelligence: Navigating Digital Safeguarding, Algorithmic Bias and the Right to a Future
Ana Žnidarec Čučković
IntechOpen eBooks · 2026-08-02
This chapter examines how AI and data-driven technologies create significant rights risks for children, including algorithmic profiling, behavioral manipulation, biometric surveillance in schools, and commercial data extraction. It analyzes how existing frameworks like the UN Convention on the Rights of the Child and General Comment No. 25 are insufficient to address these harms, with particular focus on disproportionate impacts on marginalized children and Global South communities. The authors propose a rights-by-design framework and call for structural digital-economy regulation, mandatory algorithmic impact assessments, and participatory governance to protect children's rights—including a proposed 'right to a future' grounded in Articles 6, 8, and 12 of the UNCRC. The work is relevant to policy development around AI governance, child protection, and the legal architecture needed to safeguard children in AI-mediated environments.
- AI policy
Research
Adoption generativer KI in der Wissensarbeit: Die Rolle von beruflicher Autonomie, Arbeitsanforderungen, Tätigkeiten und Führungsverantwortung
Tim Komorowski, Thomas Süße
Zeitschrift für Arbeitswissenschaft · 2026-08-02
This German-language study examines which workplace conditions are associated with the adoption of generative AI among knowledge workers in corporate organization and strategy roles, using data from the BIBB/BAuA Employment Survey 2024. Logistic regression models show that higher occupational autonomy is significantly associated with greater generative AI adoption, while also exploring the roles of job demands, work activities, and managerial responsibility. The findings reveal social inequalities in AI adoption and illuminate how digital transformation unfolds along specific work-related conditions. The authors argue the results can inform both workplace design decisions and policy or regulatory measures around AI.
- Workforce
- AI policy
Research
Creative Labor Displacement Anxiety in the Age of Generative AI
Deeksha S, Komal.S S
International Journal For Multidisciplinary Research · 2026-08-02
This study surveyed 312 creative professionals across multiple sectors to examine what drives anxiety about AI-based job displacement. Using structural equation modelling, the researchers found that perceived threat from generative AI is the strongest predictor of displacement anxiety, while creative autonomy, trust in AI systems, and facilitating conditions reduce it. Social influence—peer narratives and public discourse—amplified anxiety. The findings suggest that human-centered AI deployment, ethical governance, and institutional support are needed to protect creative worker well-being and cultural value.
- Workforce
- AI policy
Research
Navigating codified, tacit and novel rules: Mapping the human-AI creativity frontier
Emmanuelle Walkowiak
Technovation · 2026-08-02
This paper maps the boundary between human and AI creativity by analyzing 593 tasks across 126 occupations in Australia's cultural and creative industries. Using GPT-4 to annotate task descriptions, the authors find that most tasks (86.3%) fall into a hybrid human-AI category, with only 2.7% facing likely AI replacement and 11.0% being AI-immune. Key mechanisms identified include a 'structured novelty effect' where AI autonomy is higher when defined cognitive rules combine with novel rule creation, and a 'tacit knowledge boundary' where tacit rules reduce AI autonomy feasibility, suggesting complementarity rather than substitution dominates creative work.
- Workforce
- Enterprise
Research
Identity Before Autonomy: A Universal Framework for Persistent AI Actor Identity, Permanent Traceability, Delegation, Quality Assurance, and Accountable AI Operation
Pierre-Edward Procyk
Zenodo (CERN European Organization for Nuclear Research) · 2026-08-02
AI-IDP is a proposed Canadian framework requiring that every AI agent have a persistent, tamper-evident identity before it may lawfully operate. The paper specifies 25 formal documents, 14 JSON schemas, and a reference implementation (AegisTrace) using cryptographic signatures and an append-only hash-chained ledger to ensure permanent traceability of AI actions and the human principals who authorized them. It maps current Canadian law (PIPEDA, Privacy Act, Treasury Board Directive on Automated Decision-Making) against proposed legislation (AIDA/Bill C-27) and the framework's own requirements, and includes a draft statute—the AI Actor Identity and Traceability Act. The work directly addresses accountability gaps across commercial, public-sector, and autonomous AI deployments, with impact analyses spanning business, HR, societal, and economic dimensions.
- AI policy
- Certifications
- Quality assurance
- Enterprise
Research
Identity Before Autonomy: A Universal Framework for Persistent AI Actor Identity, Permanent Traceability, Delegation, Quality Assurance, and Accountable AI Operation
Pierre-Edward Procyk
Zenodo (CERN European Organization for Nuclear Research) · 2026-08-02
This paper introduces AI-IDP, a proposed Canadian framework for ensuring every AI agent has a persistent, verifiable identity tied to an auditable chain of authority and human principals. The framework includes 25 formal specification documents, 14 JSON schemas, and a reference implementation (AegisTrace) featuring Ed25519 signatures, an append-only hash-chained ledger, and 113 passing tests. The core principle—'no valid AI actor identity, no lawful agent operation'—aims to close the accountability gap in Canadian AI governance by making AI actions permanently traceable across all deployment contexts. The authors also provide a Canadian legal analysis distinguishing current law (PIPEDA, Privacy Act, Treasury Board Directive) from proposed legislation (AIDA/Bill C-27) and contribute a draft statute, the AI Actor Identity and Traceability Act.
- AI policy
- Quality assurance
- Certifications
Research
Privacy after Publicness: A Publicness-Inference-Power Theory of AI Governance
Lakshminarasimhan Santhanam
International Journal For Multidisciplinary Research · 2026-08-02
This article introduces a Publicness-Inference-Power (PIP) theory of AI governance, arguing that existing privacy and AI regulatory frameworks fail to address how AI systems transform publicly available or voluntarily shared information into sensitive inferences, predictions, and consequential decisions. The author traces five linked stages—publicness, aggregation, inference, decision, and power—through which information acquires new meaning and generates institutional asymmetries that current rules cannot adequately address. Through comparative analysis of the EU GDPR, EU AI Act, India's DPDPA, California's CCPA, and international AI frameworks, the article finds that these regimes contain only fragments of an inference-oriented approach. The article proposes a shift from data-status governance to transformation-and-power governance, offering six theoretical propositions and an operational model built around inference registers, contextual-purpose boundaries, provenance, validation, contestability, and action controls.
- AI policy
Research
Artificial intelligence approach for predicting suicide-related behaviour in emergency departments
Tsholofelo Mokheleli, Tebogo Bokaba, Patrick Ndayizigamiye et al.
Scientific Reports · 2026-08-02
This study developed and evaluated five machine learning models to predict suicide-related behaviour within 30 days of emergency department presentation using routinely collected triage data. The best-performing model, LightGBM, achieved an AUROC of 0.88, a Recall of 0.79, and an F2-score of 0.41 on an independent test set, with SHAP analysis identifying prior psychiatric history, age, and physiological variables as key predictors. The authors argue the model could provide scalable, transparent decision support for clinician prioritisation and follow-up planning without replacing clinical judgment. External validation and prospective studies are noted as necessary next steps before deployment.
- Workforce
- Quality assurance
Research
From Local Effect to System Value: A Seven-Gate Method for Qualifying Safety Components in AI-Enabled Autonomous Systems Methods Manuscript v1.0 with Retrospective Demonstration
Karel Hrubec
Zenodo (CERN European Organization for Nuclear Research) · 2026-08-02
This paper proposes the Component Qualification Ladder (CQL), a seven-gate, non-compensatory protocol for determining whether a safety component in an AI-enabled autonomous system has demonstrated sufficient evidence to be integrated into a larger architecture. Rather than relying on isolated metrics like predictive accuracy or interpretability, CQL requires passing sequential gates covering measurement validity, local mechanism effect, marginal decision value, interaction safety, generalization, external validity, and qualified integration—with no earlier gate result compensating for a later failure. The central empirical requirement is a head-to-head comparison between a host system with and without the component, requiring meaningful benefit on primary outcomes and non-inferiority on all safety-critical outcomes. A retrospective demonstration using two components from a prior published study illustrates the method, with one component failing marginal-value and interaction-safety gates despite passing earlier ones, highlighting how the framework surfaces integration risks that isolated evaluations miss.
- Certifications
- Quality assurance
Research
The Current State and Future Trends of Automotive AI Safety Governance
Yongshou Ma
Zenodo (CERN European Organization for Nuclear Research) · 2026-08-02
This paper maps the global regulatory and governance landscape for automotive AI safety, surveying binding frameworks such as the EU AI Act (Regulation 2024/1689), UNECE R155/R156/R157, ISO 26262, ISO 21448 (SOTIF), and ISO/PAS 8800:2025, alongside internal AI governance practices disclosed by major OEMs including Mercedes-Benz, Tesla, BYD, NIO, and XPeng. It documents three simultaneous industry transitions—LLM-powered cabin assistants at scale, end-to-end neural driving stacks entering mass-market deployment, and cockpit-driving integration on unified SoCs—and argues these developments outpace existing deterministic safety frameworks. The paper proposes a forward-looking framework combining technical safeguards (explainability, formal verification, red-teaming), organizational safeguards (AI ethics committees, ISO/IEC 42001 management systems), and regulatory convergence across sectoral standards rather than a single prescriptive rulebook. The analysis is directly relevant to how AI safety standards are developed, how OEMs must demonstrate post-market AI assurance, and how certification and policy frameworks need to converge globally.
- Certifications
- AI policy
Research
The Current State and Future Trends of Automotive AI Safety Governance
Yongshou Ma
Zenodo (CERN European Organization for Nuclear Research) · 2026-08-02
This paper surveys the global regulatory and industry governance landscape for artificial intelligence embedded in road vehicles, covering frameworks from the EU AI Act and ISO/PAS 8800:2025 to internal OEM practices at companies such as Mercedes-Benz, Tesla, BYD, and NIO. It maps regulatory intensity across twelve jurisdictions, analyzes the 'agent-ization' of in-cabin, in-vehicle, and intelligent-driving AI, and identifies how end-to-end neural driving stacks and LLM-powered cabin assistants are outpacing existing safety standards like ISO 26262 and ISO 21448. The paper proposes a forward-looking framework combining technical safeguards (explainability, formal verification, red-teaming) with organizational safeguards (AI ethics committees, ISO/IEC 42001) and regulatory convergence. It argues that next-phase automotive AI safety will depend on the convergence of sectoral standards, internal AI management systems, and demonstrable post-market assurance rather than any single prescriptive rulebook.
- AI policy
- Certifications
- Quality assurance
Research
The Admissibility Protocol (AP-1): An Open Standard for Evaluating Numerical Admissibility in AI Systems
Marcus Rupp
Zenodo (CERN European Organization for Nuclear Research) · 2026-08-02
AP-1 is an open, versioned protocol for evaluating whether numerical outputs from AI-assisted systems are 'admissible'—meaning independently reproducible, traceable to authoritative source data, and defensible under audit or regulatory scrutiny. Rather than judging outputs by apparent correctness alone, the protocol assesses the computational process that produced them across seven dimensions including determinism, provenance, refusal integrity, and adversarial resistance. It is model- and architecture-agnostic and targets high-stakes domains such as financial services, aerospace, medicine, and engineering where numerical evidence must be demonstrably correct. Version 1.3, currently a draft for public comment, introduces explicit operand provenance requirements and a four-class grading of invocation evidence, with empirical evidence drawn exclusively from large language model deployments in financial services.
- Quality assurance
- Certifications
Research
From Explainable AI to Auditable AI: Developing an Integrated Artificial Intelligence Auditability Framework (AIAF) for Assuring Artificial Intelligence Systems
Mthokozisi Hlatshwayo
Zenodo (CERN European Organization for Nuclear Research) · 2026-08-02
This conceptual paper introduces 'AI Auditability' as a distinct governance construct and proposes the Artificial Intelligence Auditability Framework (AIAF), designed to support continuous, evidence-based independent assurance of AI systems throughout their lifecycle. Drawing on governance theory, assurance theory, internal auditing, and socio-technical systems theory, the framework integrates eight governance dimensions including transparency, traceability, data integrity, risk and compliance, and independent assurance. The paper argues that the next evolution of AI governance moves beyond explainability and responsibility toward AI systems that are inherently auditable, introducing original contributions such as the AI Auditability Assessment Matrix (AIAM) and the principle of 'Evidence by Design.' The framework offers practical guidance for internal and external auditors, regulators, certification bodies, and boards of directors, and is intended as a foundation for future international standards on AI auditability.
- Certifications
- AI policy
Research
From Local Effect to System Value: A Seven-Gate Method for Qualifying Safety Components in AI-Enabled Autonomous Systems Methods Manuscript v1.0 with Retrospective Demonstration
Karel Hrubec
Zenodo (CERN European Organization for Nuclear Research) · 2026-08-02
This paper proposes the Component Qualification Ladder (CQL), a seven-gate, non-compensatory protocol for determining whether a safety component genuinely improves an AI-enabled autonomous system rather than merely performing well in isolation. Each gate addresses a distinct claim—from measurement validity through external validity and qualified integration—and failure at any gate cannot be offset by success at others; weighted aggregate safety scores are explicitly prohibited from determining admission. The method requires comparing a versioned host system with and without the component, demanding meaningful benefit on at least one primary outcome and non-inferiority on all protected safety and operational outcomes. A retrospective demonstration using two components from a prior study found one failed marginal-value and interaction-safety gates despite passing earlier gates, illustrating how local predictive accuracy does not guarantee system-level value.
- Certifications
- Quality assurance
Research
The Admissibility Protocol (AP-1): An Open Standard for Evaluating Numerical Admissibility in AI Systems
Marcus Rupp
Zenodo (CERN European Organization for Nuclear Research) · 2026-08-02
AP-1 is a draft open standard for evaluating whether numerical outputs from AI-assisted systems are admissible for operational, regulatory, or safety-critical use — meaning they can be independently reproduced, traced to authoritative sources, and defended under audit. The protocol defines seven evaluation dimensions covering accuracy, determinism, provenance, refusal integrity, adversarial resistance, conflicting input handling, and computation invocation, and distinguishes deterministic computation from probabilistic inference within AI systems. Rather than judging a numerical result by its apparent correctness, AP-1 assesses the computational process that produced it, arguing that a correct output does not alone establish that the required deterministic computation actually executed. The standard is model- and architecture-agnostic, applicable across financial services, aerospace, medicine, and engineering, though empirical evidence to date is drawn solely from large language model deployments in financial services by the protocol's own author.
- Quality assurance
- Certifications
- AI policy
Research
When Algorithms Meet Institutions: Evaluation Ecologies in Public Sector Machine Learning
Nanna Bonde Thylstrup, Helene Ratner, Chiara Carboni
Digital Society · 2026-08-01
This article introduces 'evaluation ecologies' as a framework for understanding how machine learning systems are assessed in public institutions, drawing on case studies from Danish schools and Dutch psychiatric clinics. The authors show that ML evaluation is not a purely technical process but an ongoing, contested interplay of technical metrics, legal frameworks, ethical norms, professional discretion, and political objectives. In Denmark, legal and ethical concerns shut down a technically robust algorithm, while in the Netherlands, perceived operational utility overrode professional skepticism. The findings argue that power, expertise, and accountability structures—not performance metrics alone—determine algorithmic outcomes in the public sector.
- AI policy
- Quality assurance
Research
Active learning for interactive prompt clarification in safety-critical programmable logic controller code generation
Ketut Adnyana, Andreas Schwung
Applied Soft Computing · 2026-08-01
This paper proposes IPC+AL, a method combining Interactive Prompt Clarification with Active Learning to govern LLM-based code generation for Programmable Logic Controllers (PLCs) in safety-critical industrial settings. The system quantifies specification ambiguity using entropy from multi-validator disagreement and applies a fail-closed gating policy that blocks code emission unless predefined safety and dialect checks pass. Feasibility studies on Batch Mixing and Robot Pick-and-Place scenarios show improved gate pass rates and expert-assessed artifact quality compared to standard prompting baselines. The authors position IPC+AL as an auditable pre-commissioning governance layer, explicitly noting it requires downstream compilation, simulation, hardware-in-the-loop validation, formal verification, and expert approval before deployment.
- Quality assurance
- Certifications
Research
Scientific Collaboration in the Age of <scp>AI</scp> Geopolitics: Governing Openness Under Strategic Competition
Xinyi Guo, Jinghan Zeng
Politics & Policy · 2026-08-01
This policy essay argues that governments should adopt 'managed openness' to navigate the tension between AI-driven geopolitical competition and the scientific interdependence that AI innovation requires. The framework rejects both unrestricted openness and wholesale technological decoupling, instead proposing proportionate, transparent safeguards that distinguish among types of research risk while preserving international cooperation on AI safety and governance. The authors contend that scientific intermediaries and targeted resilience-building can reconcile legitimate security concerns with the collaborative norms essential to global AI progress. The piece offers a pragmatic governance model relevant to science policymakers balancing national security and international scientific exchange.
- AI policy