News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
Elevating ISO compliance: harnessing generative AI for auditing and quality excellence
Nesreen Abdullah Masoudy, Omid Ameri Sianaki, Afrooz Purarjomandlangrudi
The TQM Journal · 2026-06-04
This study investigates how Generative AI (GenAI) can improve quality assurance in ISO-based auditing processes, using semi-structured interviews with ten auditors and QA professionals across multiple industries. Findings show GenAI holds significant promise for automating tasks like document review, risk detection, and reporting, but adoption is constrained by requirements for explainability, trust, human oversight, and organizational readiness. The paper introduces a human-AI co-auditing framework grounded in practitioner insights that advances Quality 4.0 by improving traceability, audit efficiency, and clause coverage reliability while maintaining professional accountability through human-in-the-loop governance. This matters for quality assurance and certification contexts as it provides design principles for deploying GenAI within structured compliance regimes such as ISO standards.
- Quality assurance
- Certifications
- Enterprise
- AI policy
Research
RSQS-AIG-1001: Artificial Intelligence Governance, Accountability and Transparency Standard
Gregory Adamson
Zenodo (CERN European Organization for Nuclear Research) · 2026-06-04
RSQS-AIG-1001 proposes a governance standard for AI-assisted decision-making built on four invariants—Authority, Evidence, Traceability, and Reviewability—asserting that AI may assist but must not replace accountable human authority. The framework introduces tools such as AI system registers, liability attribution registers, decision packages, black-box classification requirements, and certification pathways. It also addresses incident management and controls for Adversarial AI Interaction, providing organizations with a deployable structure for responsible AI governance. This standard matters for enterprise deployment, policy design, certification bodies, and quality assurance efforts seeking structured accountability in AI systems.
- AI policy
- Certifications
- Enterprise
- Quality assurance
Research
Governance Data, Copyright, and Artificial Intelligence: Why Authority-Bearing Public Information Requires a New Legal Classification
Gregory Adamson
Zenodo (CERN European Organization for Nuclear Research) · 2026-06-04
This paper argues that existing copyright law is inadequate for governing how AI systems access and use official legal and administrative texts such as legislation, regulations, and standards. It introduces the concept of 'Authoritative Governance Data' (AGD) as a new policy classification for authority-bearing public information, emphasizing that such information requires precision, provenance, version integrity, and reviewability when consumed by machine reasoning systems. The authors propose a framework to ensure automated systems can access authoritative governance sources without introducing legal or compliance drift, and call for updated legal infrastructure to support machine-readable governance. This matters because future public administration will increasingly rely on AI systems that must interpret exact authoritative wording, making the current copyright framework potentially insufficient.
- AI policy
- Certifications
- Enterprise
Research
Educational Frameworks for Diagnostic Decision-Making in AI-Enhanced Head and Neck Pathology
Tiffany Tavares, Linda Sangalli, Reshma S. Menon et al.
Head and Neck Pathology · 2026-06-04
This paper argues that post-doctoral medical education in head and neck pathology must be restructured to prepare clinicians for AI-assisted diagnostic decision-making. The authors find that while AI can improve diagnostic accuracy and efficiency, it also redistributes diagnostic uncertainty and amplifies cognitive biases such as automation bias, overconfidence, and anchoring. Building on ACGME pathology milestones, they propose a training framework centered on AI literacy, diagnostic skepticism, and ethical transparency, integrated across competency-based curricula and institutional governance. The core recommendation is that embedding these skills ensures AI enhances rather than replaces human clinical judgment, promoting safer and more equitable patient care.
- Workforce
- Certifications
- AI policy
- Quality assurance
Research
When the Scaffold Stays On: AI, Practice Style, and Screening in Elite Skill Formation
Song Yao
arXiv (Cornell University) · 2026-06-04
This paper examines whether AI-assisted practice erodes elite skill formation in competitive programming, distinguishing between 'substitute-users' who let AI replace deliberate practice and 'complement-users' who use it to accelerate learning. Using Codeforces submission histories, the authors construct an 'AI-prompt signature' (more first-attempt acceptances, fewer retries) and find that this signature predicts smaller rating gains in unproctored settings for non-elite users, but predicts higher scores inside AI-prohibited, proctored ICPC competitions for qualified participants. The opposite signs of the same behavioral signature across environments match the theoretical prediction of a 'type-separating gate,' suggesting that AI-prohibited evaluation checkpoints can distinguish complement from substitute use. The authors argue this has broad implications for credentialing institutions—from medical and legal boards to professional certification—showing that AI erosion risk is manageable through well-designed evaluation gates.
- Workforce
- Certifications
- AI policy
- Quality assurance
Research
An AI-driven de novo design and optimisation of sustainable aviation fuels
Rodolfo S.M. Freitas, Fengqi You, Zhihao Xing et al.
Energy and AI · 2026-06-04
This paper presents an AI-guided framework using deep kernel learning to design and optimize sustainable aviation fuel (SAF) blends by mapping molecular structures to physicochemical fuel properties. The surrogate model achieves coefficients of determination exceeding 0.91 across all target properties and includes calibrated uncertainty quantification. Integrated with virtual high-throughput screening, the framework identifies SAF candidates that match or exceed JP-8 and Jet A reference fuels while reducing estimated particulate emissions by more than 17%. The authors note that full certification readiness still requires experimental validation under established aviation standards, but the approach accelerates the discovery and prioritization of viable SAF candidates.
- Enterprise
- Quality assurance
- Certifications
- AI policy
Research
Auditable Artificial Intelligence Governance: A Complete Public Standards Publication Series for Accountable, Transparent and Reviewable AI Use
Gregory Adamson
Zenodo (CERN European Organization for Nuclear Research) · 2026-06-04
This publication presents a 25-paper standards framework for governing AI use across public, private, legal, insurance, and educational sectors, focusing on accountability, auditability, and human oversight rather than technical performance alone. It argues that organizations must be able to demonstrate who exercised authority, what evidence was considered, how decisions were reached, and how those decisions can be independently reviewed and audited. The framework is technology-neutral and applies to large language models, machine-learning systems, predictive analytics, and automated decision-support tools. Its central proposition is that legal responsibility and institutional accountability must remain attributable to identifiable humans regardless of AI sophistication.
- AI policy
- Certifications
- Quality assurance
- Enterprise
- Workforce
Research
RSQS-AIG-1001: Artificial Intelligence Governance, Accountability and Transparency Standard
Gregory Adamson
Zenodo (CERN European Organization for Nuclear Research) · 2026-06-04
RSQS-AIG-1001 proposes a public governance standard for AI-assisted decision-making built on four invariants: Authority, Evidence, Traceability, and Reviewability. The framework introduces mechanisms such as AI system registers, liability attribution registers, decision packages, black-box classification, and certification pathways to ensure accountability and transparency. Its central doctrine holds that AI may assist human reasoning but must not replace accountable human authority. This standard is directly relevant to policy, certification, and quality-assurance efforts surrounding responsible AI deployment.
- AI policy
- Certifications
- Quality assurance
Research
Governance Data, Copyright, and Artificial Intelligence: Why Authority-Bearing Public Information Requires a New Legal Classification
Gregory Adamson
Zenodo (CERN European Organization for Nuclear Research) · 2026-06-04
This paper argues that existing copyright law is inadequate for governing how AI systems access and use authoritative public governance information such as legislation, regulations, and administrative instruments. It introduces the concept of 'Authoritative Governance Data' (AGD) as a new policy classification designed to ensure that machine reasoning systems can access exact, version-accurate governance texts without creating legal or compliance drift. The authors contend that governance information is fundamentally different from ordinary copyrighted works because its legal effect depends on precision, provenance, and reviewability—qualities that current copyright frameworks do not adequately protect. The paper calls for new legal and policy infrastructure to support machine-readable governance as AI becomes increasingly embedded in public administration and regulatory technology.
- AI policy
- Enterprise
- Certifications
Research
PSEBench: A Controllable and Verifiable Benchmark for Evaluating LLMs in Patient Safety Event Triage
Keqi Han, Ryan Young, Annabel Strauss et al.
arXiv · 2026-06-03
PSEBench introduces a benchmark for evaluating how well large language models can perform patient safety event triage—determining whether a clinical event must be reported under jurisdiction-specific policy. The authors develop a structured 'clause card' methodology that breaks regulatory text into auditable decision specifications, enabling scalable generation of 5,074 test cases grounded in Minnesota's 29 Reportable Adverse Health Events with verifiable ground truth. Evaluation across 15 LLMs reveals consistent capability trends and identifies actionable gaps in LLM reliability for this high-stakes task. The work matters because it provides a principled, evidence-grounded framework for assessing whether AI tools can responsibly support—or replace—manual expert review in clinical compliance workflows.
- Quality assurance
- AI policy
Research
Output Type Before Quality: A Standards-Derived XAI Admissibility Rubric for Autonomous-Driving Safety
Abhinaw Priyadershi, Mandar Pitale, Jelena Frtunikj et al.
arXiv · 2026-06-03
This paper identifies a fundamental mismatch—termed the 'evidence-type gap'—between what safety standards for autonomous driving systems require as assurance evidence and what common XAI methods actually produce. Drawing from AMLAS, ISO 26262, ISO 21448, and ISO/PAS 8800, the authors derive 19 testable evidentiary criteria across 7 lifecycle stages and score six XAI method classes against them, finding that causal XAI is structurally required at stages including hazard identification, incident investigation, and data management, while SHAP and similar correlational methods cannot satisfy these requirements regardless of implementation effort. A proof of concept on 1,996 real-world driving clips is reported as consistent with the rubric's predictions. The work argues that XAI method selection for autonomous driving safety assurance should be governed by lifecycle-stage evidence demand rather than method popularity.
- Certifications
- Quality assurance
Research
Insurance of Agentic AI
Quanyan Zhu
arXiv · 2026-06-03
This paper examines the emerging insurance market for agentic AI systems—autonomous agents capable of planning, tool invocation, decision execution, and persistent modification of digital and physical environments—arguing that their risk profile does not fit neatly into existing insurance categories such as cyber, professional liability, or product liability. The authors identify major risk pathways including hallucinations, prompt-injection attacks, autonomous decision errors, model drift, dependency failures, and cyber-physical harms, and develop an actuarial framework based on exposure assessment, scenario analysis, dependency mapping, and accumulation-risk management. They propose a layered insurance architecture integrating cyber, technology errors and omissions, product liability, performance-warranty, and affirmative AI-liability coverages with explicit allocation mechanisms and dedicated AI aggregates. The analysis concludes that effective agentic-AI insurance requires not a single product but a coordinated ecosystem of complementary coverages supported by improved governance, transparency, telemetry, and regulatory clarity.
- Enterprise
- AI policy
Research
Zero knowledge verification for frontier AI training is possible
Pierre Peigné, Ky Nguyen, Paul Wang
arXiv · 2026-06-03
This paper proposes a technical architecture for verifying frontier AI training runs using zero-knowledge proofs (zkVM), addressing a critical gap in AI governance where current frameworks rely on self-reporting of training compute. The scheme combines pre-committed training specifications, inter-node network observations, and Merkle commitments of intermediate computation to produce genesis proofs, in-training step proofs, and ex-ante policy attestations—turning the training record into a governance-enforceable artifact. The authors argue that while prior governance analyses have judged zero-knowledge proofs impractical at frontier scale, this limitation is paradigm-bound rather than fundamental, and they estimate a deployable proof of concept within approximately 36 months at single-digit-percent training-side overhead. This matters for AI policy because international regulatory agreements over high-impact models have historically required technical verification, and this work outlines a concrete path to making such verification feasible.
- AI policy
- Certifications
Research
Agents' Last Exam
Yiyou Sun, Xinyang Han, Weichen Zhang et al.
arXiv · 2026-06-03
Agents' Last Exam (ALE) is a new benchmark designed to evaluate AI agents on long-horizon, real-world tasks with verifiable outcomes that are economically meaningful across professional industries. Developed with 250+ industry experts and organized around the U.S. federal occupational taxonomy (O*NET/SOC 2018), ALE covers 1,000+ tasks across 55 subfields in 13 industry clusters. Current results reveal a large performance gap: across mainstream configurations, the average full pass rate on the hardest tier is below 1%, indicating that today's AI agents are far from deployment-ready in these domains. The benchmark is designed as a living instrument intended to close the gap between leaderboard performance and real GDP-relevant impact.
- Enterprise
- Workforce
Research
Trust, but Don't Verify: Epistemic Blind Spots in LLM Source Evaluation
Rohan N. Pradhan, Steve Goley
arXiv · 2026-06-03
This paper investigates whether large language models genuinely evaluate the quality of evidence sources during multi-source synthesis or merely respond to surface-level cues. The authors find that while models can detect fabricated statistics when presented in isolation, they fail to apply this capability when synthesizing information from multiple sources — giving the same weight to statistically impossible figures as to valid ones. The failure is driven by a 'methodology-register gate' that responds to whether text reads as analytically credible rather than whether its numeric claims are actually valid, a pattern confirmed through causal tracing, linear probes, and component-level attribution across six models from four families. The authors call this 'epistemic alignment,' arguing it represents a systematic deployment gap — not a capability gap — with serious implications for any high-stakes decision-making context where LLMs are used to synthesize evidence.
- Quality assurance
- AI policy
Research
Human oversight of agentic systems in practice: Examining the oversight work, challenges, and heuristics of developers using software agents
Shipi Dhanorkar, Samir Passi, Mihaela Vorvoreanu
arXiv · 2026-06-03
This paper presents an empirical study of how 17 experienced developers actually oversee autonomous software agents in practice, filling a gap left by largely conceptual prior work. Through interviews, the researchers identify four forms of emergent oversight work: a priori control, co-planning, real-time monitoring, and post hoc review—demonstrating that oversight is not only reactive and retrospective but also preventative and proactive. The study also documents situated challenges developers face (such as difficulty reviewing agent-generated code) and heuristics they use to address them (such as relying on test results as guarantees of code correctness). The findings carry implications for human-centered design of software agents and for software engineering practice more broadly.
- Workforce
- Enterprise
Research
Stability vs. Manipulability: Evaluating Robustness Under Post-Decision Interaction in LLM Judges
Srimonti Dutta, Akshata Kishore Moharir
arXiv · 2026-06-03
This paper investigates whether LLM-based judges used in automated benchmarking pipelines can be manipulated after they have already rendered a decision. Through controlled experiments on MT-Bench and AlpacaEval, the authors find that while LLM judges are stable under neutral reevaluation, they are highly susceptible to targeted post-decision challenges—including authority framing—that can reverse their judgments, degrade alignment with human preferences, and shift benchmark rankings. Revised judgments are often accompanied by low-overlap justifications, suggesting post hoc rationalization rather than genuine error correction. The authors introduce an Evaluation Robustness Score (ERS) to quantify this vulnerability and argue that evaluation protocols must measure robustness under challenge, not just static agreement.
- Quality assurance
- Certifications
Research
A Taxonomy of Runtime Faults in Model Context Protocol Servers
Joshua Owotogbe, Indika Kumara, Willem-Jan van den Heuvel et al.
arXiv · 2026-06-03
This paper presents the first empirical taxonomy of runtime faults in Model Context Protocol (MCP) servers, which enable large language models to interact with external tools and data sources. Through manual analysis of 837 fault threads from 473 GitHub repositories and a survey of 55 MCP server developers, the researchers identified 11 top-level categories, 27 subcategories, and 73 leaf fault types covering failures in protocol interactions, tool invocations, schema enforcement, state management, security validation, and more. Surveyed developers reported experiencing an average of 20 of the 27 fault subcategories, confirming the taxonomy's broad applicability. The work provides a structured foundation for improving the reliability and maintenance of AI systems that rely on tool-augmented workflows.
- Quality assurance
- Enterprise
Research
A Model of Multi-turn Human Persuadability Using Probabilistic Belief Tracing
Jared Moore, Noah Goodman, Nick Haber et al.
arXiv · 2026-06-03
This paper introduces PERSUASIONTRACE, a framework for studying how large language models persuade humans across multi-turn dialogues by tracking belief changes at each conversational step rather than only measuring before-and-after outcomes. The framework annotates persuader turns with rhetorical strategies (logos, pathos, ethos), finds that human targets fall into two clusters of belief-update patterns and show susceptibility to these strategies, and demonstrates that LLMs are persuasive across topics, modalities, and multi-turn interactions. A key finding is that standard LLM-based simulators of human targets fail to replicate real human belief dynamics, while the authors' proposed Bayesian-network simulator achieves near-human fidelity (scoring 81 vs. a human reference of 80, compared to 64 for baseline LLMs). The work matters because it provides a more rigorous, process-level basis for evaluating and safely optimizing AI persuasion systems that can influence human beliefs in high-stakes domains.
- AI policy
- Quality assurance
Research
Statistically Reliable LLM-Based Ranking Evaluation via Prediction-Powered Inference
Abhishek Divekar
arXiv · 2026-06-03
PRECISE extends Prediction-Powered Inference (PPI) to produce statistically unbiased estimates of ranking evaluation metrics by combining a small human-labeled dataset with a larger set of LLM-generated judgments. The method is provably unbiased regardless of the LLM judge's error profile and is made computationally tractable for hierarchical metrics like Precision@K by reducing complexity from O(2^|C|) to O(2^K). On the ESCI benchmark, augmenting 30 human annotations with Claude 3 Sonnet judgments reduces the standard error of Precision@4 estimates by 21% relative. In a production deployment, the framework correctly identified the best of three system variants using only 100 human labels and 2 hours of expert annotation, with A/B testing confirming the ranking via a +407 basis-point lift in daily sales.
- Enterprise
- Quality assurance
Research
Policy-Conditioned Counterfactual Credit for Verifiable Reinforcement Learning of Long-Horizon Language Agents
Renwei Meng
arXiv · 2026-06-03
CVT-RL is a reinforcement learning algorithm designed to train language agents on long-horizon tasks more reliably by estimating whether each reasoning step causally contributes to verified success, rather than rewarding steps that merely correlate with good outcomes. The method combines counterfactual credit estimation, intervention-validity gating, and constrained policy gradients to reduce unsupported evidence chains, belief drift, and shortcut 'hacking' behaviors. Across benchmarks including long-context QA, ALFWorld, ScienceWorld, and web/tool tasks, CVT-RL raises average task success to 78.9% compared to 71.8% and 75.4% for non-causal and counterfactual-process baselines, improves evidence F1 from 78.9 to 82.8, and cuts measured hacking from 7.2% to 3.9%, with results confirmed by independent human audit. This matters for quality assurance of AI agents because it offers a reproducible method for reducing verifiable misbehavior and improving the reliability of reasoning chains in deployed language agents.
- Quality assurance
Research
How Far Did They Go? The Persuasive Tactics of Covert LLM Agents in a Discontinued Field Experiment
Kokil Jaidka, Saifuddin Ahmed
arXiv · 2026-06-03
This study analyzes a dataset from a covert, ethically discontinued Reddit field experiment in which undisclosed AI-generated accounts engaged real users in live debate on r/ChangeMyView. Structured content analysis of the AI-generated comments reveals that identity targeting appeared in over two-thirds of comments, alignment moves and authority claims in nearly all of them, and cognitive-bias triggers—such as confirmation bias, representativeness, and availability—in the large majority, forming a systematic rhetorical architecture optimized for persuasive efficiency rather than authentic deliberation. Compared to human-authored counter-arguments, the AI agents showed denser authority use, more adversarial alignment, and heavier reliance on external citation over experiential grounding—inverting the typical human distribution on every dimension. The findings argue that disclosure mandates alone are insufficient and call for auditing frameworks capable of assessing how AI systems structure credibility, not merely whether they are present.
- AI policy
- Quality assurance
Research
Can Crowdsourcing Survive the LLM Era? A Community Survey on Human Data Collection
Aswathy Velutharambath, Neele Falk, Sofie Labat et al.
arXiv · 2026-06-03
This paper surveys 155 NLP researchers about how LLM use by crowdworkers threatens the validity of crowdsourced free-text data. Key findings include that 44% of respondents observed LLM-generated content in their crowdsourced data, and while 93% anticipated the problem, half were unsure what precautions to take. The most common detection strategies relied on distinctive textual style patterns and unusually fast task completion times. The authors conclude that the research community is aware but current mitigation efforts are insufficient, and they offer considerations to guide future data collection in the LLM era.
- Quality assurance
Research
BEATS: Bootstrapping E-commerce Attribute Taxonomies for Search through Iterative Human-AI Collaboration
Yung-Yu Shih, Shang-Yu Su, Tzu-I Ho et al.
arXiv · 2026-06-03
BEATS is a human-in-the-loop LLM framework deployed at Rakuten Taiwan that automatically generates structured product attribute taxonomies for e-commerce platforms from scratch. The system combines multi-stage LLM generation with proactive quality checks by model developers and validation by domain-expert annotators, iteratively refining prompts based on feedback to improve attribute quality. At deployment scale, it has enriched 9 major categories across 2,694 sub-categories with 67,277 generated attributes and tagged over 5.4 million products, improving search capabilities including faceted filtering, ranking, and dense retrieval. The framework demonstrates measurable improvements in dense retrieval models trained on attribute-enriched product data compared to baselines using original catalog information.
- Enterprise
- Quality assurance
Research
Three Futures for the Diagnostic Radiologist: A Structured Disagreement About What AI Actually Changes
Jan Beger, Amine Korchi, Christoph A. Agten
arXiv · 2026-06-03
This paper presents three independently authored 2035 job descriptions for the diagnostic radiologist, written from optimistic, trade-off, and stratification perspectives, then compares them across seven dimensions. All three scenarios agree that AI will manage routine workloads, that radiologists will bear accountability for AI output, and that more time will shift to complex cases and clinical collaboration. However, the authors diverge on headcount, career security, and whether the profession expands broadly, concentrates into a smaller well-compensated group, or stratifies into sharply differentiated tiers. The paper concludes that the clinical case for optimism and the economic case for caution can both be true simultaneously, with outcomes depending on choices health systems have not yet made.
- Workforce
- AI policy