News & Research
The latest AI research and news with real-world stakes. Each item is sourced, dated, summarized in plain English and tagged by impact area, and checked against its source before it appears.
Kind
Impact area
5835 items
- ResearcharXiv2026-05-26QP
HARP: Measuring Harm Amplification in Multi-Agent LLM Systems · Md Hafizur Rahman, Zafaryab Haider, Tanzim Mahfuz et al.
HARP (Harm Amplification through Role Perturbation) is a trace-first evaluation methodology for measuring how localized attacks in multi-agent LLM systems propagate into broader system-level harm. The framework tracks paired clean and perturbed executions across a seven-agent finance-oriented system, recording specialist outputs, tool calls, memory reads/writes, guard events, and latency to compute a harm amplification ratio (H_global/H_local). Key findings show that single-specialist compromise produces the strongest amplification, shared-context corruption yields the highest attack success, and temporal persistence produces the largest malicious impact, while the trace-consistency defense IntegrityGuard achieves the lowest attack success and global harm but with utility and cost trade-offs. The work argues that secure multi-agent evaluation must measure not only attack bypass rates but also how orchestration spreads harm beyond the original attack point.
- ResearcharXiv2026-05-26WEP
Queue & AI: When Faster Tasks Slow Down the Workflow · Silvia Bartolucci, Pierpaolo Vivo
This paper argues that standard productivity metrics for AI tools—such as average task completion time—can be misleading in workflow settings where tasks queue up for scarce human attention. The authors formalize a 'variance wedge' concept using a queueing model, showing that AI's speed gains on individual tasks can mask system-level slowdowns when AI errors escape review and return as costly rework. Analytically, they find that under congestion, reviewers rationally reduce scrutiny of AI outputs precisely when oversight matters most, and that AI stabilizes an overloaded workflow only when both the fraction of AI-handled tasks exceeds a critical threshold and human attention for review plus expected rework is lower than for manual completion. The findings suggest AI deployment should be assessed by its effects on congestion, rework rates, and robustness of human oversight—not just average task speed.
- ResearcharXiv2026-05-26WE
An investigation of AI integration in sound designer workflows and experiences · Nelly Garcia, Joshua Reiss
This mixed-methods study surveyed 76 professional sound designers and interviewed 20 industry practitioners to examine how AI tools are being integrated into audio production workflows. Findings reveal that current AI tools perform adequately in fast-consumption media contexts but fall short of the narrative sophistication required for high-end sound design in films and immersive experiences. Practitioners prefer assistive, task-specific AI applications—particularly for audio restoration and library management—over fully generative end-to-end systems. The paper offers recommendations for developers to build more informed AI tools better aligned with professional creative needs.
- ResearcharXiv2026-05-26EP
Faults and Pitfalls in Implementing the Right to be Forgotten · Chen Sun, Nikolas Guggenberger, Supreeth Shastri
This paper examines the practical challenges of implementing the Right to be Forgotten (RTBF) under GDPR, noting that regulators issued 205 RTBF violations in the first five years of GDPR—roughly one failure every nine days on average. The authors identify computing uncertainties and risks that make RTBF difficult to enforce, and propose a two-phase approach designed to bridge the gap between legal requirements and computing practice. They demonstrate that their technique could have avoided 80% of RTBF violations in GDPR's sixth year, identify six long-standing computing and data management practices that act as anti-patterns for RTBF, and validate their approach by integrating RTBF capability into Elasticsearch, a popular open-source search engine. The work matters for policy and enterprise contexts because it provides concrete, measurable guidance for organizations struggling to comply with one of GDPR's most prominent data rights obligations.
- ResearcharXiv2026-05-26QP
Grounding Text Embeddings in Stakeholder Associations · Jonathan Rystrøm, Sofie Burgos-Thorsen, Zihao Fu et al.
This paper introduces the 'Stakeholder Grounding Exercise,' a method for evaluating whether neural text embeddings align with the semantic distinctions that domain experts actually care about. In a case study on Danish policy issues, neural embeddings were found to be substantially less reliable than human experts by 19–26 percentage points, and this misalignment directly degraded downstream clustering quality (Spearman ρ=0.9 between exercise ranking and cluster quality). A replication study on US Federal AI use cases confirmed a similar gap (16 percentage points), showing the finding holds across languages, domains, and expert communities. The method offers a practical tool for validating whether embedding models are fit for purpose before being used in high-stakes analytical workflows.
- ResearcharXiv2026-05-26QP
Detecting Is Not Resolving: The Monitoring Control Gap in Retrieval Augmented LLMs · Zhe Yu, Wenpeng Xing, Chen Ye et al.
This paper investigates a critical safety flaw in retrieval-augmented large language models (LLMs): the 'monitoring-control gap,' where models can detect contradictory or dangerous evidence in retrieved documents but still fail to act safely on that awareness. Using a multi-turn document accumulation protocol across four model families (1.5B–32B parameters) and over 50,000 turn-level evaluations, the authors show that single-turn safety benchmarks systematically overestimate real-world RAG safety, and that a model's ability to acknowledge epistemic conflict is uncorrelated with whether it resolves that conflict safely. Mechanistic analysis via hidden-state probing and attention analysis suggests danger-relevant information is internally represented and attended to, yet fails to constrain final output behavior — pointing to action selection as the core failure point. These findings matter for any high-stakes deployment of RAG systems, where evidence quality directly determines whether AI-driven recommendations are safe.
- ResearcharXiv2026-05-26QC
Semantic Robustness Probing via Inpainting: An Interactive Tool for Safety-Critical Object Detection · Nico Steckhan, Krutarth Prajapati, Weija Shao et al.
SemProbe is an interactive tool that tests object detectors in safety-critical settings by using diffusion-based inpainting to generate semantically meaningful image variations rather than simple pixel-level corruptions. Users upload deployment images, define masks manually or automatically, and select domain-relevant factors to probe how detection models respond to controlled scene changes. The system automatically runs model inference on each generated variant, displays annotated before/after comparisons with performance deltas, and logs all probes as structured artifacts to support traceable safety evaluation workflows. The authors demonstrate the tool on hand detection for dimension saws, targeting factors derived from insurance-oriented test criteria.
- ResearcharXiv2026-05-26QP
When Medical Safety Alignment Fails: A Benchmark for Evaluating LLMs on High-Risk Medical Queries · Yige Li, Jun Sun, Wei Zhao et al.
This paper introduces MedHarm, a benchmark of 1,100 medically grounded queries across 10 safety-critical categories (including toxicology, pharmacology, covert poisoning, anesthesia, and fetal harm) designed to test whether large language models handle high-risk medical prompts safely. Evaluating 15 LLMs and 4 guardrail models, the authors find a substantial gap between apparent alignment and actual medical safety: aligned models can still produce unsafe or actionable responses, medical fine-tuning can amplify harmful specificity, and external guardrails introduce brittle blocking while weakening safe helpfulness. The study concludes that medical safety cannot be inferred from general alignment or medical capability alone, underscoring the need for domain-specific stress testing before deploying LLMs in safety-critical clinical contexts.
- ResearcharXiv2026-05-26QP
Prompt Injection Detection is Regime-Dependent: A Deployment-Aware Evaluation with Interpretable Structural Signals · Akindoyin Akinrele, Shreyank N Gowda
This paper evaluates prompt injection detection—a key security threat for large language models—across a wide range of realistic deployment conditions, including out-of-distribution settings and thresholded deployment metrics. The authors compare lexical, semantic, structural, and transformer-based detectors, and introduce interpretable structural signals capturing hierarchy overrides, system prompt spoofing, role redefinition, and evasion patterns. Results show that detection performance is highly regime-dependent and sensitive to threshold selection, with no single model dominating across all settings; transformer-based models perform best overall, while structural signals provide consistent but modest gains in harder scenarios. The findings highlight a gap between ranking performance and real-world deployment effectiveness, underscoring the need to evaluate defences under realistic operational constraints.
- ResearcharXiv2026-05-26Q
JuICE: A Benchmark for Evaluating LLM-Judge in Identifying Cultural Errors · Jiho Jin, Junho Myung, Juhyun Oh et al.
JuICE introduces a multilingual benchmark of 7,470 span-level annotations covering cultural and linguistic errors in long-form LLM responses across four countries (the United States, South Korea, Indonesia, and Bangladesh). The paper finds that even the strongest LLM-judge reaches only an F1 of 0.52 on erroneous span detection and consistently misses 'thick' cultural errors that local residents readily identify. This reveals a critical gap in current LLM evaluation frameworks, which treat culture as a flat set of facts rather than accounting for the depth and situatedness of cultural meaning. The findings have direct implications for how LLM outputs are assessed for quality and appropriateness in diverse cultural contexts.
- ResearcharXiv2026-05-26QP
KZ-SafetyPrompts: A Kazakh Safety Evaluation Prompt Dataset for Large Language Models · Wajdi Zaghouani, Shimaa Amer Ibrahim, Aruzhan Muratbek et al.
KZ-SafetyPrompts introduces a dataset of 5,717 Kazakh-language prompts designed to evaluate the safety behavior of large language models (LLMs) across eleven risk categories, including self-harm, violence, child exploitation, and radicalization. The prompts are written natively in Kazakh (Cyrillic) and include English translations for cross-lingual analysis. Baseline testing with GPT-4o reveals an overall refusal rate of only 28.2%, ranging from 5.5% to 53.8% across categories, demonstrating that Kazakh prompts expose safety gaps not captured by English-only evaluations. This work highlights the need for multilingual safety benchmarks and provides a structured resource—with documented writing protocols, labeling procedures, and quality-control steps—to support broader LLM safety assessment pipelines.
- ResearcharXiv2026-05-26EQ
Neuro-Symbolic Verification of LLM Outputs for Data-Sensitive Domains (extended preprint) · Paul Sigloch, Christoph Benzmüller
This paper proposes a neuro-symbolic verification architecture that combines formal symbolic reasoning with neural semantic analysis to catch errors in LLM-generated content before they cause harm in high-stakes settings. Input verification uses logical methods with decidable guarantees on structured requirements, while output validation uses embedding-based semantic similarity to detect hallucinations that formal methods cannot catch. Validated on HAIMEDA, a real-world medical device damage assessment system, the architecture achieves hallucination detection rates above 83% for structured entities and 72% for semantic fabrications, while cutting report creation time by 30%. The work demonstrates that hybrid neuro-symbolic pipelines can offer principled safeguards for LLM deployment in domains where errors carry legal, financial, or safety consequences.
- ResearcharXiv2026-05-26EQ
Knowledge Graphs as the Missing Data Layer for LLM-Based Industrial Asset Operations · Madhulatha Mandarapu, Sandeep Kunkunuru
This paper investigates whether the data model underlying LLM-based agents—rather than the orchestration strategy—is the primary driver of accuracy in industrial asset operations. Using the AssetOpsBench benchmark (KDD 2026), the authors show that pairing GPT-4 with a typed knowledge graph raises accuracy from 65% to 82–83% via LLM-generated Cypher queries, and reaches 99% on graph-answerable scenarios using deterministic graph primitives alone. A generation-augmented knowledge (GAK) approach handles missing facts by having the agent materialize them as provenance-tagged graph nodes, lifting answerability from zero to 100% of equipment types across 88 non-deterministic benchmark scenarios and answering 81.8% of those scenarios. The findings argue that for structured operational domains, investing in the data layer—specifically a typed knowledge graph as a grounding substrate—delivers larger gains than tuning LLM orchestration paradigms.
- ResearcharXiv2026-05-26Q
Quality Without Usefulness: LLM-Generated XAI Narratives as Trust Heuristics Rather Than Decision Aids · Fabian Lukassen, Jan Herrmann, Christoph Weisser et al.
This paper investigates whether high-quality natural language explanations (NLEs) generated by large language models from Explainable AI (XAI) outputs actually help users make better decisions. Across five controlled experiments involving 2,730 judgments in an energy forecasting domain, the authors find that NLEs do not improve task accuracy on any tested task, yet inflate users' self-reported confidence — an effect driven by the mere presence of text rather than its content. Critically, in an out-of-distribution detection task, NLEs reduce the ability to flag unreliable predictions, providing false reassurance that masks model failure. The authors term this the 'Quality-Usefulness Gap' and argue that XAI evaluation must go beyond text-quality metrics to measure actual downstream task performance.
- ResearcharXiv2026-05-26Q
PRISM: A Multi-Dimensional Benchmark for Evaluating LLM Peer Reviewers · Ngoc Phan Phuoc Loc, Toan Huynh La Viet, Thanh Tran Khanh et al.
PRISM is a benchmarking framework that evaluates the quality of LLM-based automated peer reviewers across four structured dimensions: Depth of Analysis, Novelty Assessment, Flaw Identification & Major Issues Prioritization, and Multi-dimensional Constructiveness. Unlike surface-level metrics such as ROUGE and BLEU, PRISM uses argument mining, retrieval-augmented verification, and consensus-based scoring. Applied to five automated reviewer systems and human reviewers on reviews from ICLR, ICML, and NeurIPS, the results show that LLMs can match or exceed humans on individual dimensions but no single system consistently matches the balanced performance of the human baseline across all dimensions simultaneously. The study concludes that LLM reviewers are best understood as targeted supplements to human review rather than standalone replacements.
- ResearcharXiv2026-05-26QP
SL-BiLEM: Structured Learnable Behavior-in-the-Loop Epidemic Modeling for Forecasting and Policy Evaluation · Haochun Wang, Sendong Zhao, Jingbo Wang et al.
SL-BiLEM is a hybrid epidemic modeling framework that combines machine learning with physical/mechanistic constraints to improve forecasting accuracy and support policy evaluation under distribution shift caused by changing human behavior. The model decomposes effective disease transmission into components for baseline transmission, policy effects, media influence, and a learned compliance function subject to monotonicity, smoothness, and bounded-jump constraints. Validated on three real-world datasets (cruise ship, school influenza, and school-district COVID-19), the approach achieves a 76% improvement over neural-mechanistic baselines and only 53% out-of-distribution degradation versus 1142% for purely neural baselines under policy-induced shift. It also supports counterfactual intervention analysis, achieving 100% bootstrap confidence interval coverage and Treatment Effect Accuracy exceeding 0.85, making it a candidate tool for public health decision-makers planning interventions.
- ResearcharXiv2026-05-26EQ
Auditing and Fixing Economic Validity in Tabular Foundation Models for Discrete Choice · Yingshuo Wang, Xian Sun, Yanhang Li et al.
This paper identifies a critical flaw in tabular foundation models used for discrete choice tasks: their predictions frequently violate basic economic logic, such as showing demand increasing when prices rise or producing negative willingness-to-pay estimates. The authors propose a two-stage adapter that wraps foundation model predictions inside a utility-maximization framework, first fitting an economically constrained choice model and then training a correction term using the foundation model's output. On two transportation datasets, the adapter recovers up to 13 percentage points of accuracy over a standard logit model while guaranteeing monotonic price-demand relationships and analytically computable trade-off measures — something neither raw foundation models nor conventional distillation achieve. This matters for any enterprise or policy application where AI-driven demand or pricing models must comply with economic consistency requirements.
- ResearcharXiv2026-05-26QP
Vectors Are Not Neutral: Sensitive-Information Inference from Exported LLM Representations in Summarization · Weixin Liu, Bowen Qu, Juming Xiong et al.
This paper investigates a privacy risk in LLM-based summarization systems: even when source documents are kept private, the compact vector representations (embeddings) exported to downstream workflows can still leak sensitive information about individuals. Using clinical discharge summaries and EHR-recorded patient race as a controlled test case, the authors audit two types of exported vectors—final prompt-token hidden states and mean-pooled prompt representations—finding that reducing sensitive-information recoverability from one artifact does not guarantee reduction from the other. They introduce SurfaceLoRA, a parameter-efficient fine-tuning method using a gradient-reversal discriminator, which reduces race recoverability from its targeted vector toward chance levels while preserving summarization quality, but leaves recoverability elevated in untargeted artifacts. The findings highlight that privacy auditing and mitigation must be applied to the specific vector artifacts actually retained or shared downstream, not just to model outputs or source documents.
- ResearcharXiv2026-05-26EQ
When Does Deep RL Beat Calibrated Baselines? A Benchmark Study on Adaptive Resource Control · Guilin Zhang, Chuanyi Sun, Kai Zhao et al.
This paper presents RLScale-Bench, a reproducible benchmark evaluating six deep reinforcement learning (DRL) algorithms—PPO, DQN, A2C, SAC, TD3, and DDPG—against a calibrated rule-based autoscaler for adaptive compute resource allocation on Kubernetes. Across 240 runs spanning six workload patterns and five seeds, the calibrated rule-based controller achieves lower cost than every DRL algorithm on all six workloads, though RL agents show advantages on bursty and flash traffic patterns. Key findings include that discrete-action algorithms outperform continuous-action ones by one to two orders of magnitude in constraint violations, no single algorithm dominates across workloads, and the primary bottleneck is not algorithm choice but baseline calibration, reward engineering, and evaluation rigor. The results challenge common assumptions about DRL's superiority in resource control, with direct implications for how enterprises and cloud operators should evaluate and adopt AI-driven autoscaling systems.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-05-26EQCP
THE GOVERNANCE DEFICIT OF DIGITAL FOOD SAFETY: BLOCKCHAIN, AI, AND THE REGULATORY RECOGNITION GAP · Botirjon Akhmadalievich Umarov
This paper identifies and analyzes a 'Regulatory Recognition Gap' (RRG) in digital food safety, documenting that transformative technologies like blockchain traceability and AI quality control are being commercially deployed at a pace roughly 8 to 11 times faster than the legal and institutional frameworks needed to govern them. Drawing on the Walmart Hyperledger Fabric demonstration—which compressed mango traceback time by 99.9%—and market projections showing blockchain food traceability growing to USD 52.2 billion by 2035, the authors argue these capabilities operate in a governance vacuum with no binding evidentiary or regulatory standards. The paper proposes the Digital Food Safety Governance Architecture (DFSGA), a three-tier institutional framework encompassing a Codex Digital Traceability Standard, a WTO SPS Digital Certificate Recognition Protocol, and national AI regulatory frameworks addressing explainability, liability, and algorithmic consistency. The findings are highly relevant to food safety policy, certification of digital records, and quality assurance frameworks globally.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-05-26EQCP
PARALLAX-5: A Five-Obligation Substrate for Smart Contracts and AI Agents · Benjamin P. Duncan
PARALLAX-5 introduces a formal five-obligation interface—covering value conservation, authorization, signature integrity, temporal distinctness, and external-attestation trust—designed to verify the security of smart contracts and AI agents operating in decentralized systems. The framework includes 95 machine-checked Lean 4 theorems (zero unproven sorry statements), 129 passing Python tests, and an empirical catalog of 53 real incidents from 2016–2026 totaling $5.97 billion in losses, each classified by which obligations were violated. A key contribution is an AI-Agent Containment Theorem and a machine-checkable certificate schema with a live on-chain registry, enabling runtime security gating for AI agents interacting with blockchain environments. This work matters for quality assurance and certification of AI and smart contract systems by providing formally verifiable, falsifiable security guarantees grounded in production EVM semantics.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-05-26EQCP
PARALLAX-5: A Five-Obligation Substrate for Smart Contracts and AI Agents · Benjamin P. Duncan
PARALLAX-5 introduces a formally verified obligation interface for smart contracts and AI agents in decentralized systems, built around five primitive security obligations including value conservation, authorization, and attestation trust. The framework produces 95 machine-checked theorems in Lean 4 with zero unproven assumptions, validated against a 53-incident empirical catalog spanning 2016–2026 and representing $5.97 billion in aggregate losses. It also defines an AI-Agent Containment Theorem and a machine-checkable certificate schema, with a live onchain registry deployed on the Sepolia testnet, making it relevant to both smart contract quality assurance and AI agent governance.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-05-26EQCP
THE GOVERNANCE DEFICIT OF DIGITAL FOOD SAFETY: BLOCKCHAIN, AI, AND THE REGULATORY RECOGNITION GAP · Botirjon Akhmadalievich Umarov
This paper identifies and quantifies a 'Regulatory Recognition Gap' between the rapid deployment of blockchain and AI technologies in food safety and the legal frameworks needed to govern them, estimating a Technology-Regulation Speed Gap of roughly 8:1 to 11:1. It highlights that commercially deployed capabilities—such as Walmart's blockchain traceback reducing trace time by 99.9%—operate without corresponding regulatory standards, evidentiary benchmarks, or binding international protocols. The paper proposes a three-tier governance architecture (DFSGA) including a Codex Digital Traceability Standard, a WTO SPS Digital Certificate Recognition Protocol, and national AI food safety regulatory frameworks to close this gap. The findings matter for policymakers because fast-growing markets in blockchain traceability and AI food safety lack the institutional scaffolding needed to ensure accountability, liability, and legal enforceability.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-05-26EQCP
Substrate Governance: Why Runtime Controls Are Insufficient and What Must Replace Them: A Vendor-Agnostic Framework for Governing AI Agents at the Infrastructure Layer · Narnaiezzsshaa Truong
This whitepaper argues that current AI governance frameworks focus almost entirely on runtime controls—mechanisms that regulate model inputs and outputs—while neglecting the underlying substrate layer comprising execution environments, memory substrates, and orchestration meshes. The authors propose a three-pillar framework called Substrate Governance, encompassing Execution Substrate Integrity, Memory Substrate Auditability, and Orchestration Mesh Accountability, along with a four-type failure taxonomy and a phased implementation roadmap. A central finding is that no existing regulatory standard, industry framework, or certification scheme explicitly mandates substrate-layer governance controls for AI agent systems, meaning compliance-focused organizations are leaving critical infrastructure ungoverned. This matters for enterprises deploying AI agents and for policymakers and certification bodies who must expand their scope beyond API-level controls to address deeper infrastructure risks.
- ResearchEkonomicko-manazerske spektrum2026-05-26WEP
The Role of Artificial Intelligence in Strengthening Financial Practices of SMEs · Aneta Cugova, Sumana Chaudhuri
This literature review synthesizes high-quality research (2016–2024) on how artificial intelligence—including machine learning, natural language processing, and generative AI—is being applied to financial management in small and medium-sized enterprises (SMEs). The review finds that AI offers meaningful improvements in cash flow forecasting, credit risk assessment, real-time fraud detection, and data-driven financial planning, though adoption is constrained by limited data, skill shortages, and high implementation costs. Strategies such as cloud-based AI tools, employee training, and explainable AI are identified as key enablers, while algorithmic bias and the need for human oversight are flagged as persistent ethical concerns. The paper adds value by consolidating fragmented evidence linking AI adoption to SME financial stability and growth, and by outlining directions for sustainable AI adoption research.