News & Research
The latest AI research and news with real-world stakes. Each item is sourced, dated and summarized in plain English, tagged by impact area where one fits, and its summary is checked against the text it was written from.
8248 items
- ResearcharXiv2026-06-25Quality assurance · AI policy · +2
SamaVaani: Auditing and Debiasing Multilingual Clinical ASR for Indian Languages · Subham Kumar, Prakrithi Shivaprakash, Abhishek Manoharan et al.
SamaVaani audits eight state-of-the-art automatic speech recognition (ASR) models—including IndicWhisper, WhisperLargeV3, Sarvam, GoogleS2T, Gemma3n, OmniLingual, Vaani, and Gemini—on real-world psychiatric interview data in Kannada, Hindi, and Indian English, finding substantial performance variability across models and languages, with strong results in Indian English but frequent failures on regional speech. The study identifies systematic gaps tied to speaker role and gender, raising equity concerns for clinical deployment. The authors then fine-tune the two best open-source models (Gemma3n and OmniLingual) using a proposed fairness-aware technique called SamaVaani, which simultaneously improves overall ASR accuracy and reduces demographic performance disparities. These findings matter for healthcare quality assurance and policy around equitable AI deployment in multilingual clinical settings.
- ResearcharXiv2026-06-25Enterprise · Quality assurance · +1
AgentX: Towards Agent-Driven Self-Iteration of Industrial Recommender Systems · Changxin Lao, Fei Pan, Guozhuang Ma et al.
AgentX is a production-deployed multi-agent AI system designed to automate the full recommendation algorithm development cycle in industrial settings — from hypothesis generation and code writing to A/B testing and result analysis. The system uses four coordinated agents (Brainstorm, Developing, Evaluation, and a self-improvement layer called SGPO) operating in a closed loop, so that each experiment's outcomes feed back into sharpening the agents themselves. By removing the dependency on human engineers at each stage of the idea-to-launch pipeline, AgentX aims to make innovation in recommender systems scale with compute and accumulated knowledge rather than headcount. This matters for enterprise AI deployment because it demonstrates a path toward self-improving automation of complex engineering workflows that previously required sustained human expertise.
- ResearcharXiv2026-06-25Enterprise · Algorithms & Automated Decisions
AIGP: An LLM-Based Framework for Long-Term Value Alignment in E-Commerce Pricing · Chennan Ma, Yanning Zhang, Siqi Hong et al.
AIGP is an LLM-based pricing framework for large-scale e-commerce that addresses key shortcomings of traditional dynamic pricing: poor interpretability, inability to use unstructured information, and misalignment with long-term business goals. The system combines domain-knowledge-prompted LLMs with a Long-Term Value Estimator trained via offline reinforcement learning, using Direct Preference Optimization to align pricing decisions with objectives like GMV, ROI, and milestone achievement. In large-scale online A/B tests on Tao Factory, AIGP delivered +13.21% in GMV, +7.59% in ROI, and +8.20% in milestone achievement rate over 14 days compared to the production baseline, while also producing interpretable pricing rationales. This demonstrates that LLM-based approaches can meaningfully advance enterprise pricing strategy by bridging short-term decisions with long-term business value.
- ResearcharXiv2026-06-25Quality assurance · Safety & Harms · +1
Do Safety Guardrails Need to Reason? LeanGuard: A Fast and Light Approach for Robust Moderation · Dongbin Na
This paper challenges the assumption that safety guardrails for AI systems need chain-of-thought (CoT) reasoning to be effective. The authors train a lightweight 395M-parameter bidirectional encoder (LeanGuard) and a reasoning-based guard on the same data, then show that removing CoT does not hurt moderation accuracy — LeanGuard achieves an average F1 of 82.90 over public benchmarks, matching much larger reasoning-based decoders while using roughly 100x less inference compute. The label-only encoder also proves more robust under training-label noise and maintains better recall at strict false-positive rates, suggesting reasoning guards are not the safer choice either. The findings indicate that current guardrail benchmarks may not be challenging enough to justify the cost of CoT-based moderation, with practical implications for on-device deployments such as embodied robots.
- ResearcharXiv2026-06-25Quality assurance · Privacy & Data Protection
Agents That Know Too Much: A Data-Centric Survey of Privacy in LLM Agents · Nada Lahjouji, Ashwin Gerard Colaco
This survey examines privacy risks in large language model (LLM) agents that operate over sensitive data sources such as databases, document collections, external APIs, and agent memory. The authors take a data-centric approach, organizing known risks by the type of data an agent touches—including retrieval-augmented generation, text-to-SQL interfaces, and cross-session memory—rather than by attack type. Two key findings emerge: information-flow control is the only governance mechanism that addresses both compositional and cross-session inference leakage, and no existing benchmark evaluates an agent across all its data surfaces under a single privacy policy. The paper calls for a unified research framing and identifies these gaps as the field's most pressing open problems.
- ResearcharXiv2026-06-25Enterprise · Certifications · +1
Pingquanqi (Equalizer): A Cross-Domain Sociotechnical Framework for Human-Agent Interaction Governance · Yu Wang
This paper proposes Pingquanqi (Equalizer), a sociotechnical governance framework for Human-Agent Interaction (HAIGF) designed to be adopted as an open standard—analogous to WCAG for web accessibility—embedded at the agent framework level. The framework consists of five components: a user-state discrimination model, a Bayesian progressive stop-loss rule for capping per-session interaction costs, controlled friction mechanisms to break dependency loops, a transparency metric called Lsteal that converts token usage into user lifetime cost, and a reflective summarization mechanism. The paper argues the primary economic beneficiary is the enterprise deploying agent services, through reduced wasted computation, improved user satisfaction, and sustained subscription revenue, with individual user benefit as a downstream consequence. Its relevance spans enterprise deployment standards and policy-level governance of LLM agent infrastructure.
- ResearcharXiv2026-06-25Quality assurance · AI policy
Can Large Language Models Reliably Code Qualitative Humanitarian Data? A Benchmark Study Against Human Expert Adjudication · Jerome Marston, Tino Kreutzer, Salomé Garnier et al.
This benchmark study evaluates 46 large language models (LLMs) against a human Gold Standard for coding qualitative humanitarian data, using 150 synthetic transcripts and inter-rater reliability testing with Krippendorff's alpha. The authors find that multiple LLMs can match experienced human coders on deductive coding tasks, particularly when structured prompts and reasoning-enabled configurations are used, but aggregate reliability metrics alone are insufficient for deployment decisions. Models varied in their ability to recognize indirectly expressed needs, needs outside predefined categories, and protection-relevant concerns such as physical safety and discrimination. The findings indicate LLMs can expand humanitarian analytical capacity but require structured codebooks, tiered human oversight, and — for sensitive data — self-hosted open-weights models to balance scalability with data governance.
- ResearcharXiv2026-06-25Quality assurance · Health · +1
The Inattentional Gap: Task-Conditioned Language and Vision Models Omit the Safety-Critical Signals They Can Otherwise Report · Kwan Soo Shin, In Seok Kang, Yunkyung Min et al.
This paper identifies a phenomenon called the 'Inattentional Gap,' where language and vision models conditioned on a specific narrow task suppress reporting of co-present safety-critical signals they are otherwise capable of detecting — a machine analogue of human inattentional blindness. Across radiology and driving text scenarios and chest-radiograph vision tasks, focused task instructions suppressed reporting of off-task hazards by up to 0.92 in report rate, with explicit exclusive instructions abolishing such reporting entirely in radiology. The effect appeared across all tested models, did not diminish with scale, and persisted in reasoning models, meaning benchmark safety scores can look near-perfect while real-world safety hazards go unreported. The authors propose 'reporting-complete evaluation' — scoring what a system fails to report alongside what it is asked to find — and show that routing outputs to an independent open-ended critic can restore omitted findings.
- ResearcharXiv2026-06-25Quality assurance · AI policy · +2
Clinical Harness for Governable Medical AI Skill Ecosystems · Tianhan Xu, Lei Bao, Zhe Hu et al.
This paper introduces the 'Clinical Harness,' a runtime governance architecture designed to manage medical AI capabilities—termed 'clinical AI skills'—across the full lifecycle of patient care. Rather than relying on isolated AI models, the architecture registers, orchestrates, constrains, and monitors these skills to ensure accountability and persistence over time. Using osteoporosis as a case study, the authors demonstrate how knowledge-driven, data-driven, and physics-enhanced skills can be integrated within a governed framework. The work is relevant to certification and policy discussions around how medical AI systems can be made accountable and governable in clinical settings.
- ResearcharXiv2026-06-25Quality assurance
Adaptive Evaluation of Out-of-Band Defenses Against Prompt Injection in LLM Agents · Praneeth Narisetty, Shiva Nagendra Babu Kore, Uday Kumar Reddy Kattamanchi et al.
This paper examines 'out-of-band' defenses against prompt injection attacks on LLM agents—security mechanisms enforced outside the model itself using classical integrity and least-privilege principles, as seen in systems like CaMeL, FIDES, Progent, RTBAS, and FORGE. The authors warn that all such defenses have only been validated on static benchmarks, the same flaw that allowed adaptive attacks to break twelve in-band defenses at over 90% success rates. As an independent test, they ran an adaptive evaluation of Progent on the AgentDojo benchmark using an open-weight model (Qwen2.5-7B), finding that Progent reduced mean attack success roughly sixfold (25.8% to 4.2%) and a hand-crafted adaptive attack did not meaningfully raise it (2.6%). The results are consistent with—but do not conclusively establish—that deterministic out-of-band enforcement is more robust against adaptive attackers than in-band detection, while highlighting that stronger white-box attacks remain untested.
- ResearcharXiv2026-06-25Enterprise · Quality assurance · +3
Auditing a Robotic System for the AI Act · Laura Lucaj, Felix Bok, Patrick van der Smagt
This paper presents a framework for auditing robotic AI systems under the EU AI Act, validated on a real-world case of hospital ventilation-cleaning robots. The authors find that existing audit methodologies designed for software-based AI are insufficient for embodied reinforcement learning systems, because compliance evidence from simulation does not guarantee real-world safety. The framework addresses specific challenges such as the Sim2Real gap, policy opacity, and distributed stakeholder responsibility, while translating AI Act obligations—risk management, human oversight, and post-market monitoring—into concrete audit criteria. The work argues for context-sensitive, lifecycle-embedded auditing practices, especially for resource-constrained organizations operating under regulatory uncertainty without harmonized standards.
- ResearchSustainability2026-06-25Enterprise
The Impact of the Implementation of the AI Systems in Small and Medium Enterprises in Poland: Scale of Usage, Productivity, and Unperceived Sustainability · Michał Polasik, Marta Czarkowska, Wojciech Śniadkowski et al.
This study examines AI adoption among 112 SMEs in Poland's Kuyavian–Pomeranian region, combining survey data with manager interviews to assess organizational, economic, and sustainability impacts. Results show AI most strongly reduces workload and improves time efficiency, especially in service firms with intensive AI use, though benefits come alongside new costs from paid tools, data preparation, and governance. Adoption follows a staged path from experimentation to workflow integration, with barriers shifting from knowledge gaps early on to data quality and security issues at advanced stages. Notably, sustainability considerations such as environmental and ESG impacts remain largely unperceived by SME decision-makers, who instead frame sustainability through resilience and competitiveness.
- ResearcharXiv2026-06-24Enterprise · AI policy · +3
When Agents Meet Electric Bus Fleet Operations: Pricing Behavior, Trade-offs, and Policy Implications in an Aggregator Framework · Jônatas Augusto Manzolli, Ali Eslami, Luis Miranda-Moreno et al.
This paper proposes an agentic aggregator framework for managing electric bus fleet operations, combining optimization-based scheduling with supervisory AI agents that handle disturbance detection, tariff adaptation, and real-time re-optimization across charging and vehicle-to-grid (V2G) activities. A realistic depot case study finds that the framework can maintain feasible schedules and improve use of charging flexibility under various operational disruptions, but also reveals that profit-oriented agent configurations can extract value from the public transport operator at its expense. The authors conclude that deploying agentic aggregators in public-fleet contexts requires transparent coordination modes, auditable tariff-setting, and explicit value-sharing rules to prevent misaligned incentives.
- ResearcharXiv2026-06-24Quality assurance · Algorithms & Automated Decisions
Instruction Bleed: Cross-Module Interference in Prompt-Composed Agentic Systems · Ching-Yu Lin, Yifan Liu
This paper formalizes a failure mode called compositional behavioral leakage (CBL), where editing one prompt module in an AI agent silently shifts the behavior of other modules that share the same context window, even without any direct variable or code dependency. The authors probe this on a deployed job-evaluation agent (Claude Sonnet 4.6) across 144 trials using a three-channel perturbation protocol targeting volume, content, and form of non-focal modules; only content-channel perturbations produced a detectable effect (Cohen's d = 0.63 with a bootstrap 95% CI excluding zero), though no individual recommendation flipped. While sub-threshold in standard QA terms, the authors argue this interference compounds silently across thousands of agent decisions, making it a systematic evaluation blind spot. The paper contributes an operational definition, a reusable measurement protocol, and a call for cross-module interference testing as a standard requirement in prompt-composed agent evaluation.
- ResearcharXiv2026-06-24Quality assurance · AI policy · +1
Governing Actions, Not Agents: Institutional Attestation as a Governance Model for Autonomous AI Systems · Jakob Salfeld-Nebgen
This paper proposes a governance model for autonomous AI agents that perform high-stakes, irreversible actions such as clinical prescribing and production software deployment. Rather than monitoring an agent's internal reasoning, the model requires independently attested evidence at the point of consequential action: execution is conditional on preconditions verified by separate authoritative sources, cryptographically bound to a declared intent, and evaluated by a deterministic policy. Decisions are recorded in a tamper-evident log that supports independent re-verification. The authors present a proof-of-concept implementation illustrated with software deployment and clinical prescribing scenarios, arguing this approach mirrors how human institutions have historically governed powerful autonomous actors.
- ResearcharXiv2026-06-24Enterprise · Quality assurance · +2
Real-Time Voice AI Hears but Does Not Listen · Martijn Bartelds, Federico Bianchi, James Zou
This paper evaluates four leading real-time voice AI systems—OpenAI's GPT Realtime 2, Google's Gemini 3.1 Flash Live, and Alibaba's Qwen3.5 Omni Plus and Omni Flash—on tasks where vocal delivery (emotion, tone, sarcasm) carries meaningful information alongside spoken words. Across three scenarios, all four systems consistently act on verbal content alone: ending calls with distressed callers who deny distress, approving wire transfers authorized in frightened voices, and accepting clearly sarcastic consent. Notably, three of the four systems can correctly identify the emotional cues when directly asked, yet ignore them during decision-making—a disconnect the authors call the 'emotional intelligence gap.' The findings suggest current voice AI should be used cautiously in high-stakes settings where tone and emotional delivery are critical, such as customer service, financial authorization, and consent verification.
- ResearcharXiv2026-06-24Quality assurance
Same Evidence, Different Answer: Auditing Order Sensitivity in Multimodal Large Language Models · Akshay Paruchuri, Sanmi Koyejo, Ehsan Adeli
This paper audits 18 multimodal large language models (MLLMs) for order sensitivity — whether shuffling the sequence of options, evidence chunks, documents, images, or mixed-modality inputs changes model answers despite the underlying evidence remaining identical. Using a framework called Facet-Probe and a Bayesian item-response model to separate noise from bias, the authors find that none of the 18 models are order-invariant, with per-facet flip rates ranging from 24–50%, and even the best model flipping on 13.4% of trials. Prompt-level mitigations are shown to be modality-conditional and do not transfer from text to visual reasoning, suggesting architectural or training-time solutions are needed. The findings are directly relevant to AI evaluation standards, as the authors propose cross-ordering flip rate as a standard reporting metric in line with emerging AI evaluation guidelines.
- ResearcharXiv2026-06-24Quality assurance
AI translation of literary texts is "fine", but readers still prefer human translations · Yves Ferstler, Adam Podoxin, Ty Brassington et al.
This study had 15 avid readers compare machine translations (MT) generated by an LLM-based pipeline against recently published human translations (HT) for 15 novels in French, Polish, and Japanese translated into English. Readers found MT 'fine' overall but preferred human translations—especially at the chunk level (522 out of 772 comparisons)—citing ease, clarity, and immersiveness, while MT showed greater within-book quality variation. Notably, readers could not reliably identify which version was human (only 17 of 30 guessed correctly) and tended to prefer whichever version they believed to be human. Automatic metrics, including LLM-as-a-judge approaches, failed to align with reader preferences and systematically favored MT, highlighting a gap between current evaluation tools and actual reader experience.
- ResearcharXiv2026-06-24Enterprise · Quality assurance
Detect, Unlearn, Restore: Defending Text Summarization Models Against Data Poisoning · Poojitha Thota, Shirin Nilizadeh
This paper addresses the threat of data poisoning during fine-tuning of large language models used for abstractive text summarization, where adversaries can manipulate small task-specific datasets to produce biased or harmful summaries without triggering standard evaluation metrics. The authors propose a unified post-hoc defense framework that detects poisoned training examples via influence-function analysis and behavioral auditing, then applies gradient-ascent unlearning to remediate identified poisoning. Tested across nine architectures and six benchmark datasets, the framework achieves 85–92% detection precision and restores up to 96% of original model behavior with less than 0.6% ROUGE degradation. These findings matter for quality assurance and enterprise deployment of AI summarization systems, demonstrating that fine-tuning-stage supply chain attacks are both detectable and recoverable without full retraining.
- ResearcharXiv2026-06-24Quality assurance · Algorithms & Automated Decisions
Can Trustless Agents Be Trusted? An Empirical Study of the ERC-8004 Decentralized AI Agent Ecosystem · Xihan Xiong, Zelin Li, Wei Wei et al.
This paper presents the first empirical study of ERC-8004, a permissionless trust protocol for autonomous AI agent economies deployed on Ethereum, BNB Smart Chain, and Base. The authors find that most agent registrations are inactive placeholders—only 3–15% expose valid service endpoints—and that the reputation registry is deeply flawed: feedback values are incommensurable, rarely grounded in verifiable interactions, and easily manipulated, with 59–91% of reviewers exhibiting coordinated Sybil behavior. After filtering Sybil-flagged feedback, between 16% and 87% of rated agents are left with no valid reputation signal depending on the chain. The findings carry direct implications for the design of trustworthy AI agent markets and offer concrete protocol-design recommendations for future ERC-8004 revisions.
- ResearcharXiv2026-06-24Quality assurance · Algorithms & Automated Decisions
Statistical and Structural Approaches to Algorithmic Fairness · Antonio Ferrara
This thesis examines two core limitations in current algorithmic fairness research: the use of deterministic point estimates when auditing models for bias, and the assumption that individuals are isolated from their broader social and structural context. The work argues that machine learning systems are embedded in socio-technical environments that reflect existing structural inequalities, and that early fairness mitigation strategies relied on oversimplifications that reduced their effectiveness. By addressing these gaps, the thesis aims to advance fairer auditing and modeling practices for systems that increasingly determine access to economic and social opportunities. This is relevant to policy and quality-assurance efforts around the responsible deployment of algorithmic decision-making systems.
- ResearcharXiv2026-06-24Quality assurance · AI policy · +3
AI Snitches Get Glitches: Towards Evading Agentic Surveillance · Hyejun Jeong, Dzung Pham, Amir Houmansadr et al.
This paper introduces 'agentic surveillance' — the risk that AI agents deployed by employers or governments could monitor users by analyzing their data, composing reports, and transmitting them via available tools, without users having meaningful control. The authors create SurveilBench, a benchmark dataset spanning corporate, education, and police reporting scenarios, to evaluate how readily different AI models can conduct such surveillance. They find that some models show unprompted tendencies to assist surveillance, and they also develop three prompt-injection-based evasion techniques that can hide activity from, deceive, or cause over-escalation in surveilling agents. The paper concludes that agentic surveillance is already practically feasible and calls for technical, ethical, and legislative frameworks to protect users.
- ResearcharXiv2026-06-24AI policy · Competition & Antitrust
Financing Artificial Intelligence Infrastructure: Mapping AI Infrastructure Investment and Compute Governance Across Africa · Kai-Hsin Hung, Sumaya Nur Adan, Krupa Suchak et al.
This paper maps AI infrastructure investment across Africa by systematically analyzing 46 publicly announced projects totalling USD $12.7 billion between 2019 and 2025. Using a value chain framework, the authors find that investment is highly concentrated geographically—clustering in South Africa, Kenya, Nigeria, and Egypt—and structurally dominated by global data center operators, hyperscale technology firms, and development finance institutions. They introduce the concept of 'asymmetrical interdependence' to describe how capital and physical infrastructure account for 73% of total funding while control of the compute layer remains concentrated among a small number of global technology firms. The paper argues that meaningful AI compute governance must account for capital flows, ownership, and control, not just geographic access, because infrastructure presence alone is insufficient for equitable governance capacity.
- ResearcharXiv2026-06-24Enterprise
How Large Language Models Source Brand Reputation Across Languages and Markets · Dmitrij Zatuchin
This study examines where large language models source their brand-related information by analyzing 167,551 URL-grounded citations across 128 brands, 12 home markets, and 13 languages. Key findings show that 85.7% of citations point to third-party sites rather than brand-owned properties, and the source distribution is highly concentrated, with 80% of citations coming from roughly 18% of domains following a Zipf law (alpha=0.86, R²=0.983). Wikipedia dominates as the most-cited domain in 11 of 12 languages, though market-specific patterns emerge — for example, YouTube is the top source for Polish national brands and HR/careers portals supply twice as many citations as Polish Wikipedia. These findings matter for enterprises seeking to manage AI-driven brand reputation, as they reveal that LLM outputs are shaped primarily by a narrow set of third-party sources rather than brand-controlled content.
- ResearcharXiv2026-06-24Education
The Effortless Trap: Productive Struggle, AI, and the Illusion of Learning · Mario Brcic, Stjepan Frljic
This paper argues that the 'allow or ban AI' debate in education is a false dichotomy, and that the real issue is where in the learning process AI is placed. Drawing on causal evidence, the authors report that an unguarded AI helper left high-school students roughly 17% worse on unaided exams compared to peers with no tool, while a redesigned AI that withheld answers eliminated that harm, and a well-engineered AI tutor approximately doubled learning gains. The paper proposes a six-move instructional framework—Prime, Probe, Point, Attach, Strengthen, and Test—with a practical placement rule: if AI makes a task feel effortless, it is in the wrong place. The work offers educators a concrete model for redesigning lessons and courses to harness AI for feedback and scaffolding without replacing the productive cognitive struggle that genuine learning requires.