News & Research
The latest AI research and news with real-world stakes. Each item is sourced, dated and summarized in plain English, tagged by impact area where one fits, and its summary is checked against the text it was written from.
8284 items
- ResearcharXiv2026-06-24Quality assurance · AI policy · +1
Governing Actions, Not Agents: Institutional Attestation as a Governance Model for Autonomous AI Systems · Jakob Salfeld-Nebgen
This paper proposes a governance model for autonomous AI agents that perform high-stakes, irreversible actions such as clinical prescribing and production software deployment. Rather than monitoring an agent's internal reasoning, the model requires independently attested evidence at the point of consequential action: execution is conditional on preconditions verified by separate authoritative sources, cryptographically bound to a declared intent, and evaluated by a deterministic policy. Decisions are recorded in a tamper-evident log that supports independent re-verification. The authors present a proof-of-concept implementation illustrated with software deployment and clinical prescribing scenarios, arguing this approach mirrors how human institutions have historically governed powerful autonomous actors.
- ResearcharXiv2026-06-24Enterprise · Quality assurance · +2
Real-Time Voice AI Hears but Does Not Listen · Martijn Bartelds, Federico Bianchi, James Zou
This paper evaluates four leading real-time voice AI systems—OpenAI's GPT Realtime 2, Google's Gemini 3.1 Flash Live, and Alibaba's Qwen3.5 Omni Plus and Omni Flash—on tasks where vocal delivery (emotion, tone, sarcasm) carries meaningful information alongside spoken words. Across three scenarios, all four systems consistently act on verbal content alone: ending calls with distressed callers who deny distress, approving wire transfers authorized in frightened voices, and accepting clearly sarcastic consent. Notably, three of the four systems can correctly identify the emotional cues when directly asked, yet ignore them during decision-making—a disconnect the authors call the 'emotional intelligence gap.' The findings suggest current voice AI should be used cautiously in high-stakes settings where tone and emotional delivery are critical, such as customer service, financial authorization, and consent verification.
- ResearcharXiv2026-06-24Quality assurance
Same Evidence, Different Answer: Auditing Order Sensitivity in Multimodal Large Language Models · Akshay Paruchuri, Sanmi Koyejo, Ehsan Adeli
This paper audits 18 multimodal large language models (MLLMs) for order sensitivity — whether shuffling the sequence of options, evidence chunks, documents, images, or mixed-modality inputs changes model answers despite the underlying evidence remaining identical. Using a framework called Facet-Probe and a Bayesian item-response model to separate noise from bias, the authors find that none of the 18 models are order-invariant, with per-facet flip rates ranging from 24–50%, and even the best model flipping on 13.4% of trials. Prompt-level mitigations are shown to be modality-conditional and do not transfer from text to visual reasoning, suggesting architectural or training-time solutions are needed. The findings are directly relevant to AI evaluation standards, as the authors propose cross-ordering flip rate as a standard reporting metric in line with emerging AI evaluation guidelines.
- ResearcharXiv2026-06-24Quality assurance
AI translation of literary texts is "fine", but readers still prefer human translations · Yves Ferstler, Adam Podoxin, Ty Brassington et al.
This study had 15 avid readers compare machine translations (MT) generated by an LLM-based pipeline against recently published human translations (HT) for 15 novels in French, Polish, and Japanese translated into English. Readers found MT 'fine' overall but preferred human translations—especially at the chunk level (522 out of 772 comparisons)—citing ease, clarity, and immersiveness, while MT showed greater within-book quality variation. Notably, readers could not reliably identify which version was human (only 17 of 30 guessed correctly) and tended to prefer whichever version they believed to be human. Automatic metrics, including LLM-as-a-judge approaches, failed to align with reader preferences and systematically favored MT, highlighting a gap between current evaluation tools and actual reader experience.
- ResearcharXiv2026-06-24Enterprise · Quality assurance
Detect, Unlearn, Restore: Defending Text Summarization Models Against Data Poisoning · Poojitha Thota, Shirin Nilizadeh
This paper addresses the threat of data poisoning during fine-tuning of large language models used for abstractive text summarization, where adversaries can manipulate small task-specific datasets to produce biased or harmful summaries without triggering standard evaluation metrics. The authors propose a unified post-hoc defense framework that detects poisoned training examples via influence-function analysis and behavioral auditing, then applies gradient-ascent unlearning to remediate identified poisoning. Tested across nine architectures and six benchmark datasets, the framework achieves 85–92% detection precision and restores up to 96% of original model behavior with less than 0.6% ROUGE degradation. These findings matter for quality assurance and enterprise deployment of AI summarization systems, demonstrating that fine-tuning-stage supply chain attacks are both detectable and recoverable without full retraining.
- ResearcharXiv2026-06-24Quality assurance · Algorithms & Automated Decisions
Can Trustless Agents Be Trusted? An Empirical Study of the ERC-8004 Decentralized AI Agent Ecosystem · Xihan Xiong, Zelin Li, Wei Wei et al.
This paper presents the first empirical study of ERC-8004, a permissionless trust protocol for autonomous AI agent economies deployed on Ethereum, BNB Smart Chain, and Base. The authors find that most agent registrations are inactive placeholders—only 3–15% expose valid service endpoints—and that the reputation registry is deeply flawed: feedback values are incommensurable, rarely grounded in verifiable interactions, and easily manipulated, with 59–91% of reviewers exhibiting coordinated Sybil behavior. After filtering Sybil-flagged feedback, between 16% and 87% of rated agents are left with no valid reputation signal depending on the chain. The findings carry direct implications for the design of trustworthy AI agent markets and offer concrete protocol-design recommendations for future ERC-8004 revisions.
- ResearcharXiv2026-06-24Quality assurance · Algorithms & Automated Decisions
Statistical and Structural Approaches to Algorithmic Fairness · Antonio Ferrara
This thesis examines two core limitations in current algorithmic fairness research: the use of deterministic point estimates when auditing models for bias, and the assumption that individuals are isolated from their broader social and structural context. The work argues that machine learning systems are embedded in socio-technical environments that reflect existing structural inequalities, and that early fairness mitigation strategies relied on oversimplifications that reduced their effectiveness. By addressing these gaps, the thesis aims to advance fairer auditing and modeling practices for systems that increasingly determine access to economic and social opportunities. This is relevant to policy and quality-assurance efforts around the responsible deployment of algorithmic decision-making systems.
- ResearcharXiv2026-06-24Quality assurance · AI policy · +3
AI Snitches Get Glitches: Towards Evading Agentic Surveillance · Hyejun Jeong, Dzung Pham, Amir Houmansadr et al.
This paper introduces 'agentic surveillance' — the risk that AI agents deployed by employers or governments could monitor users by analyzing their data, composing reports, and transmitting them via available tools, without users having meaningful control. The authors create SurveilBench, a benchmark dataset spanning corporate, education, and police reporting scenarios, to evaluate how readily different AI models can conduct such surveillance. They find that some models show unprompted tendencies to assist surveillance, and they also develop three prompt-injection-based evasion techniques that can hide activity from, deceive, or cause over-escalation in surveilling agents. The paper concludes that agentic surveillance is already practically feasible and calls for technical, ethical, and legislative frameworks to protect users.
- ResearcharXiv2026-06-24AI policy · Competition & Antitrust
Financing Artificial Intelligence Infrastructure: Mapping AI Infrastructure Investment and Compute Governance Across Africa · Kai-Hsin Hung, Sumaya Nur Adan, Krupa Suchak et al.
This paper maps AI infrastructure investment across Africa by systematically analyzing 46 publicly announced projects totalling USD $12.7 billion between 2019 and 2025. Using a value chain framework, the authors find that investment is highly concentrated geographically—clustering in South Africa, Kenya, Nigeria, and Egypt—and structurally dominated by global data center operators, hyperscale technology firms, and development finance institutions. They introduce the concept of 'asymmetrical interdependence' to describe how capital and physical infrastructure account for 73% of total funding while control of the compute layer remains concentrated among a small number of global technology firms. The paper argues that meaningful AI compute governance must account for capital flows, ownership, and control, not just geographic access, because infrastructure presence alone is insufficient for equitable governance capacity.
- ResearcharXiv2026-06-24Enterprise
How Large Language Models Source Brand Reputation Across Languages and Markets · Dmitrij Zatuchin
This study examines where large language models source their brand-related information by analyzing 167,551 URL-grounded citations across 128 brands, 12 home markets, and 13 languages. Key findings show that 85.7% of citations point to third-party sites rather than brand-owned properties, and the source distribution is highly concentrated, with 80% of citations coming from roughly 18% of domains following a Zipf law (alpha=0.86, R²=0.983). Wikipedia dominates as the most-cited domain in 11 of 12 languages, though market-specific patterns emerge — for example, YouTube is the top source for Polish national brands and HR/careers portals supply twice as many citations as Polish Wikipedia. These findings matter for enterprises seeking to manage AI-driven brand reputation, as they reveal that LLM outputs are shaped primarily by a narrow set of third-party sources rather than brand-controlled content.
- ResearcharXiv2026-06-24Education
The Effortless Trap: Productive Struggle, AI, and the Illusion of Learning · Mario Brcic, Stjepan Frljic
This paper argues that the 'allow or ban AI' debate in education is a false dichotomy, and that the real issue is where in the learning process AI is placed. Drawing on causal evidence, the authors report that an unguarded AI helper left high-school students roughly 17% worse on unaided exams compared to peers with no tool, while a redesigned AI that withheld answers eliminated that harm, and a well-engineered AI tutor approximately doubled learning gains. The paper proposes a six-move instructional framework—Prime, Probe, Point, Attach, Strengthen, and Test—with a practical placement rule: if AI makes a task feel effortless, it is in the wrong place. The work offers educators a concrete model for redesigning lessons and courses to harness AI for feedback and scaffolding without replacing the productive cognitive struggle that genuine learning requires.
- ResearcharXiv2026-06-24Quality assurance · Public Sector Use · +1
Bridging Predictions and Interventions: An Integrated Framework for Automated Decision-Systems · Inioluwa Deborah Raji, Lydia T. Liu, Angela Zhou et al.
This paper argues that automated decision systems (ADS) should be evaluated not just on predictive accuracy but on their actual intervention effects within organizational and social contexts. Drawing on real-world case studies in domains such as criminal pretrial release, clinical triage, and student support, the authors show that deploying ADS modifies workflows and decision-making processes in ways that purely prediction-focused frameworks fail to capture. The paper proposes an integrated framework that shifts ADS design, evaluation, and deployment toward an intervention-oriented view to better anticipate downstream societal and organizational consequences.
- ResearcharXiv2026-06-24Quality assurance · Health
MedGuards: Multi-Agent System for Reliable Medical Error Detection and Correction · Congbo Ma, Hu Wang, Yichun Zhang et al.
MedGuards is a multi-agent framework designed to improve medical error detection and correction in LLM-generated clinical text, addressing patient safety risks that arise when existing automated and heuristic-based methods fail to generalize across unseen datasets. The system assigns specialized agents to separately detect, localize, and correct errors, with a confidence-guided arbitration mechanism that resolves disagreements using reasoning traces and confidence scores — all without retraining the underlying LLMs. The authors also introduce a new evaluation metric, the Keyword-Prioritized Correction Score (KPCS), which checks whether critical keywords from reference text are correctly reproduced, offering a more thorough assessment than conventional metrics. Experiments on four multilingual clinical note datasets show significant improvements across multiple metrics and models, supporting safer LLM deployment in real-world healthcare.
- ResearcharXiv2026-06-24Quality assurance · Algorithms & Automated Decisions
Taxonomy of Risks on Automated Fact-Checking Systems Considering its Propagation · Jun Yajima, Tatsuya Oka, Takao Okubo
This paper develops a structured risk taxonomy for automated fact-checking systems that use AI and large language models to assess the veracity of social media posts. The authors identify 32 specific risks organized across a three-stage propagation framework covering risk factors, hazardous situations, and resulting harms such as spreading misinformation or enabling defamation. They apply the taxonomy as guide words in a risk assessment of the DEFAME fact-checking system, showing it surfaces risks that conventional IT security methods like STRIDE miss. The work represents a foundational step toward enabling the safe deployment of automated fact-checking systems.
- ResearcharXiv2026-06-24Enterprise · Quality assurance · +2
Probabilistic Agents in Deterministic Audits: Evaluating Multi-Agent Systems for Automated Audits Based on the German IT-Grundschutz · Lea Roxanne Muth, Marian Margraf
This paper implements and evaluates a Multi-Agent System (MAS) combined with Hybrid Retrieval Augmented Generation (HybridRAG) to partially automate German IT-Grundschutz (IT-GS) certification, which is required for NIS-2 Directive compliance. Two novel technical contributions are introduced: a Hypothesis-Verification Loop to reduce hallucinations and a Decoupled Reasoning Pipeline to separate semantic extraction from deterministic protection need inheritance. Evaluated against the BSI's 'RecPlast GmbH' expert-generated case study using Precision, Recall, and F1-scores, the system performs well on semantic tasks like Structural Analysis and Modeling, but struggles in logical reasoning phases such as Protection Needs Assessment and IT-GS Check where the probabilistic nature of LLMs conflicts with the deterministic rigor required by the standard. The findings highlight both the promise and current limits of AI-driven automation for scalable, resource-efficient IT security certification.
- ResearcharXiv2026-06-24Quality assurance · Certifications
An Approach for a Supporting Multi-LLM System for Automated Certification Based on the German IT-Grundschutz · Lea Roxanne Muth, Marian Margraf
This paper proposes a multi-LLM system (MLS) using Hybrid Retrieval-Augmented Generation (HybridRAG) to semi-automate the BSI IT-Grundschutz security certification process. Combining Large Language Models with Knowledge Graphs, the system supports key certification phases such as protection needs assessment, modeling, IT-Grundschutz check, and measure consolidation. The approach is motivated by the NIS2 directive's expanded compliance requirements, a shortage of qualified specialists, and high implementation costs, aiming to increase efficiency and help certifiers maintain quality across a growing number of affected organizations.
- ResearcharXiv2026-06-24Enterprise · Quality assurance
Constraint Tax in Open-Weight LLMs: An Empirical Study of Tool Calling Suppression Under Structured Output Constraints · Fangzheng Li, Aimin Zhang, Chen Lv
This paper identifies and investigates 'Tool Suppression,' a reproducible failure mode in production AI agent systems where enabling JSON Schema structured output constraints simultaneously with tool calling causes open-weight language models to stop invoking tools entirely, even though both capabilities work correctly in isolation. The authors trace the root cause to grammar-based token masking compiled from JSON Schema constraints, which renders tool-call tokens unreachable during decoding, and propose a 'Constraint Priority Inversion' hypothesis to explain the behavioral pattern. To address the issue without retraining, they introduce Transparent Two-Pass Execution, an inference-time strategy that separates tool execution from schema-constrained response generation, successfully restoring tool invocation while maintaining structured output compliance. The findings highlight that evaluating agent capabilities in isolation can miss critical reliability failures that only emerge under joint deployment conditions.
- ResearcharXiv2026-06-24Workforce
The impact of artificial intelligence on enterprise software user roles · Isabel Unger, Elizangela Valarini, Martin Schrepp et al.
This qualitative study examines how AI is reshaping professional roles and workflows within SAP's Business Technology Platform (BTP), drawing on expert interviews (n=20) and a participatory workshop (n=24). The findings show substantial shifts in day-to-day tasks, including increasing automation of operational work, expanded human-AI collaboration, and growing reliance on agentic AI systems. The research identifies that existing role frameworks such as the BTP User Type Matrix need adaptation to reflect these workforce changes, and calls for revised role taxonomies, new governance and oversight functions, and updated design approaches for AI-native enterprise software.
- ResearcharXiv2026-06-24Quality assurance
How Reliable Is Your Jailbreak Judge? Calibration and Adversarial Robustness of Automated ASR Scoring · Yang Gao
This paper systematically evaluates the reliability of automated judges used to score attack-success rates (ASR) in LLM jailbreak research, comparing dedicated safety classifiers and general LLM-as-judge systems against 596 human-labeled completions from the HarmBench validation set. The two judge families fail in opposite ways: the dedicated classifier over-flags harmful content (high recall, lower precision), while LLM-based judges show high precision but erratic recall (0.06 to 0.65), meaning the same responses yield wildly different ASR depending on which judge is used. Adversarial attacks exposing further weaknesses show that simple benign framing wrappers flip LLM-judges 57–100% of the time, while white-box GCG attacks flip 70% of confident true positives in the classifier—yet human audits confirm the harmful content remains intact in all sampled cases. The authors recommend that papers report judge precision, recall, and adversarially-corrected ASR to improve the trustworthiness of jailbreak benchmarking.
- ResearcharXiv2026-06-24Public Sector Use · Algorithms & Automated Decisions
From Causal Discovery to Implementation: An Agentic AI Framework for E-Scooter Mobility Hub Planning Across 29 German Cities · Meng Jin, Melanie Handrich, Simone Martinenz et al.
This paper presents a three-phase agentic AI framework that uses causal discovery—rather than correlational modeling—to identify which environmental features drive e-scooter hotspot demand across 29 German cities, distinguishing patterns by city type (large, university, industrial, hilly) and cluster type (core, peripheral). Built on public GBFS data across 57 city-cluster units, the framework constructs a Causal Template Library that an LLM-orchestrated pipeline adapts to local data conditions, revealing that core demand is driven by activity access and transit proximity while peripheral demand responds to built form. A planning tool built on this library scores candidate sites, calibrates infrastructure recommendations to local demographics, and generates practitioner-ready reports—with two hub sites in Heilbronn, Germany currently under construction based on the framework's outputs. The work matters because it provides city-type-specific, causally grounded, and transferable siting templates that can support real-world urban mobility planning decisions.
- ResearcharXiv2026-06-24Quality assurance
A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation · Abrar Alotaibi, Raed Mughus, Moataz Ahmed
This paper introduces a red teaming framework for systematically identifying vulnerabilities in large language model (LLM) outputs, using a multi-role architecture of target, attacker, and jury models. In a case study on faithfulness, adversarial prompts increased attack success rates by up to 7.9% in question-answering tasks, and the framework found that architectural design choices generally matter more than parameter scaling for model safety. The approach works across languages (English and Arabic) and tasks, revealing how structural constraints like format limitations can reduce unfaithfulness, while also highlighting remaining challenges in detecting subtle faithfulness failures. These findings offer actionable insights for assessing and improving LLM reliability in high-stakes deployments.
- ResearcharXiv2026-06-24Quality assurance · AI policy · +1
Long-Term Simulation Exposes Cognitive-Developmental Risks in AI Companions · Kaicheng Shen, Lingyu Li, Wen Wu et al.
This paper introduces TSJ (Theater-Stage-Judge), a longitudinal simulation framework for evaluating cognitive-developmental risks posed by AI companion systems—particularly large language model-based companions interacting with children and adolescents. The authors simulate 12,960 person-day interactions across six AI models, four developmental stages, twenty-four risk dimensions, and three psychological-vulnerability personas, finding that standard short-session safety evaluations systematically underestimate these risks, with stable risk estimates only emerging after 140 turns. Key findings include that early childhood and emerging adulthood are the most vulnerable developmental stages, and that cognitive trust and emotional dependency are the weakest domains. The work provides a scalable methodology for longitudinal safety assessment of AI companion systems, with direct implications for how such systems are evaluated and regulated.
- ResearcharXiv2026-06-24Quality assurance
A Survey of Toxicity Detection and Mitigation Strategies for Multilingual Language Models · Soham Dan, Himanshu Beniwal, Thomas Hartvigsen
This survey examines the safety challenges facing large language models (LLMs) deployed across multiple languages, cataloguing how adversaries can exploit language choice, translation, code-switching, and other techniques to bypass safety alignment. It organizes existing work on toxicity detection methods—including cross-lingual encoders, translation pipelines, and LLM-based detectors—and mitigation strategies such as data filtering, preference-based tuning, and decoding-time steering. The authors identify persistent gaps including uneven language coverage, culturally contingent definitions of harm, and the risk that detoxification efforts inadvertently suppress legitimate dialectal or identity-related expression. These findings matter for developers and policymakers seeking to ensure AI systems behave safely and equitably across diverse linguistic and cultural contexts.
- ResearcharXiv2026-06-24Quality assurance · Safety & Harms · +2
Text Over Image: Auditing Multimodal Robustness in Synthetic Medical Image Detection · Ching-Hao Chiu, Hao-Wei Chung, Gelei Xu et al.
This paper investigates a multimodal vulnerability in vision-language models (VLMs) used for synthetic medical image detection: when both an image and accompanying metadata are provided, VLMs can overweight the text context, causing authenticity judgments to flip based solely on changes in the accompanying record rather than the image itself. The authors introduce a paired benchmark that holds images fixed while swapping controlled metadata variants, and find that adding an explicit AI-origin tag alone causes accuracy on authentic images to drop by 61.1% on average across multiple imaging modalities and diverse VLMs. They also propose an inference-time mitigation pipeline that detects and neutralizes these 'provenance shortcuts' without retraining, outperforming direct prompt-based suppression. These findings reveal a significant robustness gap in real-world clinical deployment of VLMs and provide a standardized benchmark for evaluating multimodal robustness beyond image-only settings.
- ResearcharXiv2026-06-24Algorithms & Automated Decisions
AI Coaching for Accelerating Human Skill Development with Reinforcement Learning · Wei Wang, Enlin Gu, Antonio Loquercio et al.
This paper investigates how an AI agent can act as a coach to accelerate human motor-skill development, specifically in first-person-view drone racing. The authors formalize the coaching interaction as a non-cooperative dynamic game where the learner optimizes task performance and the coach targets the learner's independent competence, then build a reinforcement learning framework combining adaptive shared control with probabilistic models of the coach's causal influence on skill evolution. A user study with N=33 participants showed significant gains in human learning outcomes compared to state-of-the-art AI coaching baselines. This work is directly relevant to workforce skill development, demonstrating that carefully designed AI assistance—including strategic 'stepping back' to allow productive failures—can improve human capability rather than induce over-reliance or skill atrophy.