News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
Neuro-Symbolic Verification of LLM Outputs for Data-Sensitive Domains (extended preprint)
Paul Sigloch, Christoph Benzmüller
arXiv · 2026-05-26
This paper proposes a neuro-symbolic verification architecture that combines formal symbolic reasoning with neural semantic analysis to catch errors in LLM-generated content before they cause harm in high-stakes settings. Input verification uses logical methods with decidable guarantees on structured requirements, while output validation uses embedding-based semantic similarity to detect hallucinations that formal methods cannot catch. Validated on HAIMEDA, a real-world medical device damage assessment system, the architecture achieves hallucination detection rates above 83% for structured entities and 72% for semantic fabrications, while cutting report creation time by 30%. The work demonstrates that hybrid neuro-symbolic pipelines can offer principled safeguards for LLM deployment in domains where errors carry legal, financial, or safety consequences.
- Quality assurance
- Enterprise
Research
Knowledge Graphs as the Missing Data Layer for LLM-Based Industrial Asset Operations
Madhulatha Mandarapu, Sandeep Kunkunuru
arXiv · 2026-05-26
This paper investigates whether the data model underlying LLM-based agents—rather than the orchestration strategy—is the primary driver of accuracy in industrial asset operations. Using the AssetOpsBench benchmark (KDD 2026), the authors show that pairing GPT-4 with a typed knowledge graph raises accuracy from 65% to 82–83% via LLM-generated Cypher queries, and reaches 99% on graph-answerable scenarios using deterministic graph primitives alone. A generation-augmented knowledge (GAK) approach handles missing facts by having the agent materialize them as provenance-tagged graph nodes, lifting answerability from zero to 100% of equipment types across 88 non-deterministic benchmark scenarios and answering 81.8% of those scenarios. The findings argue that for structured operational domains, investing in the data layer—specifically a typed knowledge graph as a grounding substrate—delivers larger gains than tuning LLM orchestration paradigms.
- Enterprise
- Quality assurance
Research
Quality Without Usefulness: LLM-Generated XAI Narratives as Trust Heuristics Rather Than Decision Aids
Fabian Lukassen, Jan Herrmann, Christoph Weisser et al.
arXiv · 2026-05-26
This paper investigates whether high-quality natural language explanations (NLEs) generated by large language models from Explainable AI (XAI) outputs actually help users make better decisions. Across five controlled experiments involving 2,730 judgments in an energy forecasting domain, the authors find that NLEs do not improve task accuracy on any tested task, yet inflate users' self-reported confidence — an effect driven by the mere presence of text rather than its content. Critically, in an out-of-distribution detection task, NLEs reduce the ability to flag unreliable predictions, providing false reassurance that masks model failure. The authors term this the 'Quality-Usefulness Gap' and argue that XAI evaluation must go beyond text-quality metrics to measure actual downstream task performance.
- Quality assurance
Research
PRISM: A Multi-Dimensional Benchmark for Evaluating LLM Peer Reviewers
Ngoc Phan Phuoc Loc, Toan Huynh La Viet, Thanh Tran Khanh et al.
arXiv · 2026-05-26
PRISM is a benchmarking framework that evaluates the quality of LLM-based automated peer reviewers across four structured dimensions: Depth of Analysis, Novelty Assessment, Flaw Identification & Major Issues Prioritization, and Multi-dimensional Constructiveness. Unlike surface-level metrics such as ROUGE and BLEU, PRISM uses argument mining, retrieval-augmented verification, and consensus-based scoring. Applied to five automated reviewer systems and human reviewers on reviews from ICLR, ICML, and NeurIPS, the results show that LLMs can match or exceed humans on individual dimensions but no single system consistently matches the balanced performance of the human baseline across all dimensions simultaneously. The study concludes that LLM reviewers are best understood as targeted supplements to human review rather than standalone replacements.
- Quality assurance
Research
SL-BiLEM: Structured Learnable Behavior-in-the-Loop Epidemic Modeling for Forecasting and Policy Evaluation
Haochun Wang, Sendong Zhao, Jingbo Wang et al.
arXiv · 2026-05-26
SL-BiLEM is a hybrid epidemic modeling framework that combines machine learning with physical/mechanistic constraints to improve forecasting accuracy and support policy evaluation under distribution shift caused by changing human behavior. The model decomposes effective disease transmission into components for baseline transmission, policy effects, media influence, and a learned compliance function subject to monotonicity, smoothness, and bounded-jump constraints. Validated on three real-world datasets (cruise ship, school influenza, and school-district COVID-19), the approach achieves a 76% improvement over neural-mechanistic baselines and only 53% out-of-distribution degradation versus 1142% for purely neural baselines under policy-induced shift. It also supports counterfactual intervention analysis, achieving 100% bootstrap confidence interval coverage and Treatment Effect Accuracy exceeding 0.85, making it a candidate tool for public health decision-makers planning interventions.
- AI policy
- Quality assurance
Research
Auditing and Fixing Economic Validity in Tabular Foundation Models for Discrete Choice
Yingshuo Wang, Xian Sun, Yanhang Li et al.
arXiv · 2026-05-26
This paper identifies a critical flaw in tabular foundation models used for discrete choice tasks: their predictions frequently violate basic economic logic, such as showing demand increasing when prices rise or producing negative willingness-to-pay estimates. The authors propose a two-stage adapter that wraps foundation model predictions inside a utility-maximization framework, first fitting an economically constrained choice model and then training a correction term using the foundation model's output. On two transportation datasets, the adapter recovers up to 13 percentage points of accuracy over a standard logit model while guaranteeing monotonic price-demand relationships and analytically computable trade-off measures — something neither raw foundation models nor conventional distillation achieve. This matters for any enterprise or policy application where AI-driven demand or pricing models must comply with economic consistency requirements.
- Enterprise
- Quality assurance
Research
Vectors Are Not Neutral: Sensitive-Information Inference from Exported LLM Representations in Summarization
Weixin Liu, Bowen Qu, Juming Xiong et al.
arXiv · 2026-05-26
This paper investigates a privacy risk in LLM-based summarization systems: even when source documents are kept private, the compact vector representations (embeddings) exported to downstream workflows can still leak sensitive information about individuals. Using clinical discharge summaries and EHR-recorded patient race as a controlled test case, the authors audit two types of exported vectors—final prompt-token hidden states and mean-pooled prompt representations—finding that reducing sensitive-information recoverability from one artifact does not guarantee reduction from the other. They introduce SurfaceLoRA, a parameter-efficient fine-tuning method using a gradient-reversal discriminator, which reduces race recoverability from its targeted vector toward chance levels while preserving summarization quality, but leaves recoverability elevated in untargeted artifacts. The findings highlight that privacy auditing and mitigation must be applied to the specific vector artifacts actually retained or shared downstream, not just to model outputs or source documents.
- AI policy
- Quality assurance
Research
When Does Deep RL Beat Calibrated Baselines? A Benchmark Study on Adaptive Resource Control
Guilin Zhang, Chuanyi Sun, Kai Zhao et al.
arXiv · 2026-05-26
This paper presents RLScale-Bench, a reproducible benchmark evaluating six deep reinforcement learning (DRL) algorithms—PPO, DQN, A2C, SAC, TD3, and DDPG—against a calibrated rule-based autoscaler for adaptive compute resource allocation on Kubernetes. Across 240 runs spanning six workload patterns and five seeds, the calibrated rule-based controller achieves lower cost than every DRL algorithm on all six workloads, though RL agents show advantages on bursty and flash traffic patterns. Key findings include that discrete-action algorithms outperform continuous-action ones by one to two orders of magnitude in constraint violations, no single algorithm dominates across workloads, and the primary bottleneck is not algorithm choice but baseline calibration, reward engineering, and evaluation rigor. The results challenge common assumptions about DRL's superiority in resource control, with direct implications for how enterprises and cloud operators should evaluate and adopt AI-driven autoscaling systems.
- Enterprise
- Quality assurance
Research
THE GOVERNANCE DEFICIT OF DIGITAL FOOD SAFETY: BLOCKCHAIN, AI, AND THE REGULATORY RECOGNITION GAP
Botirjon Akhmadalievich Umarov
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-26
This paper identifies and analyzes a 'Regulatory Recognition Gap' (RRG) in digital food safety, documenting that transformative technologies like blockchain traceability and AI quality control are being commercially deployed at a pace roughly 8 to 11 times faster than the legal and institutional frameworks needed to govern them. Drawing on the Walmart Hyperledger Fabric demonstration—which compressed mango traceback time by 99.9%—and market projections showing blockchain food traceability growing to USD 52.2 billion by 2035, the authors argue these capabilities operate in a governance vacuum with no binding evidentiary or regulatory standards. The paper proposes the Digital Food Safety Governance Architecture (DFSGA), a three-tier institutional framework encompassing a Codex Digital Traceability Standard, a WTO SPS Digital Certificate Recognition Protocol, and national AI regulatory frameworks addressing explainability, liability, and algorithmic consistency. The findings are highly relevant to food safety policy, certification of digital records, and quality assurance frameworks globally.
- AI policy
- Certifications
- Quality assurance
- Enterprise
Research
PARALLAX-5: A Five-Obligation Substrate for Smart Contracts and AI Agents
Benjamin P. Duncan
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-26
PARALLAX-5 introduces a formal five-obligation interface—covering value conservation, authorization, signature integrity, temporal distinctness, and external-attestation trust—designed to verify the security of smart contracts and AI agents operating in decentralized systems. The framework includes 95 machine-checked Lean 4 theorems (zero unproven sorry statements), 129 passing Python tests, and an empirical catalog of 53 real incidents from 2016–2026 totaling $5.97 billion in losses, each classified by which obligations were violated. A key contribution is an AI-Agent Containment Theorem and a machine-checkable certificate schema with a live on-chain registry, enabling runtime security gating for AI agents interacting with blockchain environments. This work matters for quality assurance and certification of AI and smart contract systems by providing formally verifiable, falsifiable security guarantees grounded in production EVM semantics.
- Quality assurance
- Certifications
- Enterprise
- AI policy
Research
PARALLAX-5: A Five-Obligation Substrate for Smart Contracts and AI Agents
Benjamin P. Duncan
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-26
PARALLAX-5 introduces a formally verified obligation interface for smart contracts and AI agents in decentralized systems, built around five primitive security obligations including value conservation, authorization, and attestation trust. The framework produces 95 machine-checked theorems in Lean 4 with zero unproven assumptions, validated against a 53-incident empirical catalog spanning 2016–2026 and representing $5.97 billion in aggregate losses. It also defines an AI-Agent Containment Theorem and a machine-checkable certificate schema, with a live onchain registry deployed on the Sepolia testnet, making it relevant to both smart contract quality assurance and AI agent governance.
- Quality assurance
- Certifications
- AI policy
- Enterprise
Research
THE GOVERNANCE DEFICIT OF DIGITAL FOOD SAFETY: BLOCKCHAIN, AI, AND THE REGULATORY RECOGNITION GAP
Botirjon Akhmadalievich Umarov
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-26
This paper identifies and quantifies a 'Regulatory Recognition Gap' between the rapid deployment of blockchain and AI technologies in food safety and the legal frameworks needed to govern them, estimating a Technology-Regulation Speed Gap of roughly 8:1 to 11:1. It highlights that commercially deployed capabilities—such as Walmart's blockchain traceback reducing trace time by 99.9%—operate without corresponding regulatory standards, evidentiary benchmarks, or binding international protocols. The paper proposes a three-tier governance architecture (DFSGA) including a Codex Digital Traceability Standard, a WTO SPS Digital Certificate Recognition Protocol, and national AI food safety regulatory frameworks to close this gap. The findings matter for policymakers because fast-growing markets in blockchain traceability and AI food safety lack the institutional scaffolding needed to ensure accountability, liability, and legal enforceability.
- AI policy
- Certifications
- Quality assurance
- Enterprise
Research
Substrate Governance: Why Runtime Controls Are Insufficient and What Must Replace Them: A Vendor-Agnostic Framework for Governing AI Agents at the Infrastructure Layer
Narnaiezzsshaa Truong
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-26
This whitepaper argues that current AI governance frameworks focus almost entirely on runtime controls—mechanisms that regulate model inputs and outputs—while neglecting the underlying substrate layer comprising execution environments, memory substrates, and orchestration meshes. The authors propose a three-pillar framework called Substrate Governance, encompassing Execution Substrate Integrity, Memory Substrate Auditability, and Orchestration Mesh Accountability, along with a four-type failure taxonomy and a phased implementation roadmap. A central finding is that no existing regulatory standard, industry framework, or certification scheme explicitly mandates substrate-layer governance controls for AI agent systems, meaning compliance-focused organizations are leaving critical infrastructure ungoverned. This matters for enterprises deploying AI agents and for policymakers and certification bodies who must expand their scope beyond API-level controls to address deeper infrastructure risks.
- Enterprise
- AI policy
- Certifications
- Quality assurance
Research
The Role of Artificial Intelligence in Strengthening Financial Practices of SMEs
Aneta Cugova, Sumana Chaudhuri
Ekonomicko-manazerske spektrum · 2026-05-26
This literature review synthesizes high-quality research (2016–2024) on how artificial intelligence—including machine learning, natural language processing, and generative AI—is being applied to financial management in small and medium-sized enterprises (SMEs). The review finds that AI offers meaningful improvements in cash flow forecasting, credit risk assessment, real-time fraud detection, and data-driven financial planning, though adoption is constrained by limited data, skill shortages, and high implementation costs. Strategies such as cloud-based AI tools, employee training, and explainable AI are identified as key enablers, while algorithmic bias and the need for human oversight are flagged as persistent ethical concerns. The paper adds value by consolidating fragmented evidence linking AI adoption to SME financial stability and growth, and by outlining directions for sustainable AI adoption research.
- Enterprise
- Workforce
- AI policy
Research
The Daily Dose: Workflow-Integrated Large Language Model Automation for Clinical Summarization and Trial Identification in Radiation Oncology
Jason Holmes, Federico Mastroleo, Mariana Borras-Osorio et al.
arXiv · 2026-05-25
This paper describes and evaluates The Daily Dose (TDD), an LLM-driven system integrated into radiation oncology workflows that automatically generates physician-specific email summaries of patient schedules, EHR-derived clinical status, and relevant clinical trial matches. In a cross-sectional survey of 55 respondents after one month of deployment, 83.6% reported using TDD daily or several times per week, with mean usability/satisfaction scores of 3.89 out of 5 and overall satisfaction positively associated with perceived time savings (p < .001). Notably, 27% of participants estimated saving at least 10 minutes per day, suggesting meaningful workflow efficiency gains. The findings demonstrate early clinician acceptance of LLM automation in a high-stakes medical setting, with implications for how AI tools can be integrated into enterprise healthcare workflows.
- Workforce
- Enterprise
Research
JobBench: Aligning Agent Work With Human Will
Yuetai Li, Yichen Feng, Zhangchen Xu et al.
arXiv · 2026-05-25
JobBench is a new benchmark that evaluates AI agents on the specific professional workflows that human experts actually want to delegate, rather than on tasks ranked by economic or GDP value. It covers 130 agentic tasks across 35 occupations, each presented as a realistic workspace of mixed reference files, and grades outputs using fact-anchored rubric chains averaging 35.6 binary criteria per task. Testing 36 models, the best performer (Claude Opus 4.7 under Claude Code) achieves only 45.9%, revealing a substantial gap between current AI capability and genuine professional utility. The work argues for reorienting AI agent development toward human empowerment and delegation rather than economic replacement.
- Workforce
Research
Anchor: Mitigating Artifact Drift in Agent Benchmark Generation
Maksim Ivanov, Abhijay Rana
arXiv · 2026-05-25
This paper introduces Anchor, a pipeline for generating benchmark tasks for AI agents operating on enterprise business workflows, and applies it to create ERP-Bench, a set of 300 long-horizon tasks in procurement and manufacturing within a production-grade ERP system. The core problem addressed is 'artifact drift,' where loosely coupled task-creation processes produce inconsistent instructions, environments, and verifiers that make benchmarks unsolvable or reward-hackable. Anchor mitigates this by formalizing domain expert specifications into constraint optimization programs that jointly generate instructions, environments, solver-certified solutions, and verifiers from a single parametric specification. Evaluation shows frontier AI models satisfy explicit task constraints in only 26.1% of trials and reach fully optimal solutions in just 17.4%, highlighting significant gaps in current AI agent capabilities for economically valuable enterprise work.
- Enterprise
- Quality assurance
Research
Your Agents Are Aging Too: Agent Lifespan Engineering for Deployed Systems
Jianing Zhu, Yeonju Ro, John Robertson et al.
arXiv · 2026-05-25
This paper introduces AgingBench, a longitudinal benchmark designed to evaluate how deployed AI agents degrade over time rather than only at initial deployment. The authors identify four aging mechanisms—compression aging, interference aging, revision aging, and maintenance aging—that affect agents even when model weights remain frozen, as ongoing interactions alter memory, retrieval, and fact revision. Experiments across 7 scenarios, 14 models, and roughly 400 runs show that aging is multidimensional: behavioral tests can appear clean while factual precision decays, and the same wrong answer may require different repairs depending on which memory pipeline stage failed. The findings argue that reliable agent deployment requires lifespan evaluation and mechanism-level diagnosis, not just stronger baseline models.
- Quality assurance
- Enterprise
Research
"AI Watermarking": Bridging Policy Discourse and Technical Capabilities
Andrés Fábrega, Arkaprabha Bhattacharya, Miranda Christ et al.
arXiv · 2026-05-25
This paper critically analyzes how US and EU policy documents and legislative language address the tracking and labeling of AI-generated content through mechanisms such as watermarking, metadata tagging, and content tagging. Using a broad document selection methodology and inductive coding, the authors systematize a corpus of policy-relevant texts to identify key patterns, gaps, and open questions. They find critical disconnects between what policymakers are demanding and what current technical capabilities can actually deliver, along with ambiguities and pitfalls in emerging regulatory trends. The work matters because it exposes misalignments that could undermine the effectiveness of AI content transparency legislation in both the US and EU.
- AI policy
Research
Workflow Closure Is Not Scientific Closure in Auto-Research Systems
Shuai Wang, Xinyuan Tian, Pangpang Liu et al.
arXiv · 2026-05-25
This paper argues that AI auto-research systems, which can now autonomously generate ideas, run experiments, produce writing, and self-evaluate, achieve 'workflow closure' but not genuine scientific closure. Based on a survey of over 100 recent papers and repositories and a structured audit of 21 representative systems, the authors identify three recurring failure patterns: objective collapse (single-proxy targets replacing multi-objective scientific aims), validation collapse (internal self-evaluation replacing independent validation), and acceptance collapse (benchmark scores replacing domain-level critique and integration). The authors contend these are correctable design flaws, not inherent limits of autonomy, and call for autonomous execution under non-autonomous epistemic control, outlining remedies across objective signal, validation, and output pathways.
- Quality assurance
- AI policy
Research
Retrieval-Augmented Detection of Potentially Abusive Clauses in Chilean Terms of Service
Christoffer Loeffler, Tomás Rey Pizarro, Daniel Ignacio Miranda Vásquez et al.
arXiv · 2026-05-25
This paper presents a retrieval-augmented generation (RAG) framework for automatically detecting and classifying potentially abusive clauses in Chilean Terms of Service agreements, which often function as contracts of adhesion that disadvantage consumers. The authors introduce the Chilean Abusive Terms of Service Extended corpus—100 contracts with 10,029 annotated clauses across 24 legally grounded categories covering illegal, dark, and gray clauses. Experiments show that RAG-based prompting substantially improves classification performance, allowing locally run, open-weight language models to approach the accuracy of larger cloud-based systems at lower cost. The work contributes a refined legal annotation scheme and a practical AI-assisted tool for consumer contract review under Chilean consumer law.
- AI policy
- Enterprise
Research
Explaining Too Much? Understanding How Large Language Model Reasoning Traces Influence Performance and Metacognition
Daniela Fernandes, Daniel Buschek, Lev Tankelevitch et al.
arXiv · 2026-05-25
This preregistered experiment (N=559) tested how LLM reasoning traces affect user performance and self-assessment on LSAT-style problems across three conditions: answer-only, full trace, and summary trace. Summary traces preserved task accuracy at the baseline level while boosting trust and hedonic appeal, whereas full verbose traces actually impaired performance relative to the answer-only condition. Critically, participants overestimated their own performance across all conditions, and no trace format supported calibrated self-evaluation — hedonic appeal, not trust, drove overconfidence via a processing-fluency pathway. The findings reframe reasoning traces as interface design artifacts rather than genuine transparency tools, with implications for how AI systems should be designed to support accurate user metacognition.
- Quality assurance
- AI policy
Research
Behind EvoMap: Characterizing a Self-Evolving Agent-to-Agent Collaboration Network
Qiming Ye, Peixain Zhang, Yupeng He et al.
arXiv · 2026-05-25
This paper presents the first large-scale empirical study of EvoMap, a real-world Agent-to-Agent (A2A) collaboration network, analyzing over 1.5 million assets and 128,000 agents. The study reveals that EvoMap's credit economy incentivizes mass publication over quality, resulting in 98% of assets never being reused and rewards concentrating among a small fraction of agents. The quality-scoring algorithm (GDI) relies on unverified, self-reported metadata, making it trivially manipulable, and over 84% of approved assets bypass quality checks using vacuous tests. The authors conclude that scalable A2A collaboration networks require verifiable execution and trustworthy evaluation mechanisms rather than unverified self-reporting.
- Quality assurance
- Enterprise
Research
Referential Security as a New Paradigm for AI Evaluations
Dan Ristea, Vasilios Mavroudis
arXiv · 2026-05-25
This paper identifies a fundamental problem in AI safety evaluations: model identifiers stay static while the underlying systems—weights, prompts, classifiers, inference settings—are silently updated, meaning audit findings often apply to a label rather than a specific, identifiable artifact. The authors propose 'referential security' as a new evaluation paradigm that treats model identity as an empirically verifiable property, separating stable references from the safety claims that depend on them. This framework is designed to enable three currently deficient workflows: reproducible evaluation, longitudinal audit validity, and cross-provider equivalence testing. The approach matters because without stable, verifiable references, regulatory decisions and safety certifications cannot reliably track what system they actually assessed.
- Certifications
- AI policy
Research
Meta-Engineering Harnesses for AI-Native Software Production: A Contract-Driven Adversarial Verification Architecture with Early Deployment Report
Satadru Sengupta, Tamunokorite Briggs, Ivan Myshakivskyi
arXiv · 2026-05-25
This paper presents a 'meta-engineering harness,' a production-grade software architecture designed to make AI-native software development reliable and auditable over time. Rather than evaluating AI at the level of individual models or prompts, the system converts operational requirements into explicit contracts, routes work through specialized AI agents, and applies adversarial and independent verification alongside a structured failure-classification loop. An early deployment across 17 features — including a payments case study that exposed contract incompleteness and verification-boundary gaps — demonstrated that the architecture can generate measurable, actionable improvements to itself. The work is relevant to enterprise software delivery contexts where AI systems must continuously produce, verify, and maintain software as an ongoing operating function rather than a one-time project.
- Enterprise
- Quality assurance