News & Research
The latest AI research and news with real-world stakes. Each item is sourced, dated and summarized in plain English, tagged by impact area where one fits, and its summary is checked against the text it was written from.
8284 items
- ResearcharXiv2026-06-25Enterprise
Cluster, Route, Escalate: Cascaded Framework for Cost-Aware LLM Serving · Yasmin Moslem, Magdalena Kacmajor, Vasudevan Nedumpozhimana et al.
This paper proposes a two-stage cascaded framework for deploying large language models (LLMs) cost-efficiently in production. In Stage 1, incoming queries are clustered and routed to the most cost-effective model for each cluster, controlled by an interpretable hyperparameter tuned offline. Stage 2 adds a quality-estimation layer that escalates low-confidence outputs to a stronger, more expensive model only when needed. On test datasets, the system retains 97–99% of the strongest model's accuracy while reducing Time Per Output Token (TPOT), adapting to changes in the model pool using only task-correctness labels.
- ResearcharXiv2026-06-25Enterprise · Quality assurance · +3
LLM-Based Examination of Eligibility Criteria from Securities Prospectuses at the German Central Bank · Serhii Hamotskyi, Akash Kumar Gautam, Christian Hänig
This paper presents a case study applying Large Language Models (LLMs) to automate the examination of securities eligibility criteria from prospectuses at the German Central Bank. The system decomposes the task into extraction, normalization, and interpretation stages, replacing traditional Named Entity Recognition methods that struggled with OCR noise, bilingual content, and rigid annotation requirements. Results show the LLM-based pipeline achieves up to 91% precision in document-level eligibility decisions with a conservative profile that minimizes false acceptance. This work demonstrates how generative AI can reduce the manual burden of regulatory compliance verification in central banking operations.
- ResearcharXiv2026-06-25Quality assurance · Health
AI Healthcare Chatbots as Information Infrastructure: A Large-Scale Study of User-Reported Breakdowns · Muhammad Hassan, Ramazan Yener, Ece Gumusel et al.
This study analyzes over 15,000 user reviews from 59 AI healthcare chatbot apps to identify recurring failures users experience in everyday health information seeking and self-management. Topic modeling and interpretive analysis reveal three main breakdown categories: access barriers and service unreliability, user experience and interaction quality, and billing and customer support issues, with privacy and security concerns linked to the most negative experiences. By framing these chatbots as information infrastructures, the research highlights how failures in access, usability, and trust have real consequences for users, and offers actionable insights for designers, policymakers, and information professionals seeking to improve digital health systems.
- ResearcharXiv2026-06-25Workforce · Enterprise · +3
Prompt Injection in Automated Résumé Screening with Large Language Models: Single and Multi-Injection Settings · Preet Baxi, Jiannan Xu, Jane Yi Jiang et al.
This paper investigates prompt injection attacks in LLM-based résumé screening, where candidates embed subtle self-promotional text in their résumés to manipulate algorithmic rankings without adding real qualifications. Controlled experiments show that such injections reliably improve rankings when candidate quality is similar and few applicants inject, but effectiveness collapses as manipulation becomes widespread. In heterogeneous candidate pools, prompt injection is less effective on average but can occasionally let lower-quality candidates outrank stronger ones, raising fairness concerns. The findings suggest LLM-based hiring systems are most vulnerable when manipulation is rare and quality differences among applicants are small.
- ResearcharXiv2026-06-25AI policy · Privacy & Data Protection · +1
From Celebrities to Anyone: Characterizing AI Nudification Content, Technology, and Community Dynamics on 4chan · Chi Cui, Yixin Wu, Yang Zhang
This large-scale empirical study identifies 24,105 synthetic non-consensual sexually explicit AI-generated images and videos ('SNEACI') shared on 4chan, revealing that non-celebrity individuals now account for 55.8% of targets—up from just 4.7% in prior studies—indicating AI nudification has expanded well beyond public figures to harm people in users' personal social circles. Open-source tools dominate production, with the Stable Diffusion family responsible for 42.7% of images and Wan for 66.5% of videos, while shared fine-tuned models and accessible tutorials lower barriers to entry. A small cohort of prolific producers drives the ecosystem, with the most active individual generating 780 items, shaping community engagement, target demographics, and technical knowledge diffusion. The authors argue these findings underscore urgent needs for platform governance interventions, technical safeguards, and protections for affected individuals.
- ResearcharXiv2026-06-25Quality assurance
Inherited Circuits, Learned Semantics: How Fine-Tuning Creates Evasion Vulnerabilities Invisible to Standard Evaluation · Ryan Fetterman
This paper demonstrates that large language models fine-tuned for security classification (specifically PowerShell command detection) can pass standard held-out evaluations while becoming more vulnerable to evasion attacks introduced by the fine-tuning process itself. The authors study Foundation-Sec-8B-Instruct and its base model, finding that fine-tuning concentrates and semantically specializes an inherited late-attention classification circuit from Llama rather than building a new one, creating brittle token-level indicator rules. A three-tier evasion benchmark shows the fine-tuned model fails on behavior-preserving transformations—such as alias substitution, string construction, and case mutation—that the base model handles correctly. The authors propose a pre-deployment monitoring method using a linear probe and indicator-token sign test to identify vulnerable command families, cautioning that task-specific fine-tuning can improve accuracy metrics while silently expanding the real-world evasion surface.
- ResearcharXiv2026-06-25Quality assurance · AI policy · +1
Towards Explainable Adjudicative Variance: Quantifying Judicial Discretion via Gated Multi-Task Learning · Stanisław Sójka, Felix Steffek, Matthias Grabmair
This paper proposes a Judge-Aware Gated Multi-Task Learning architecture to predict legal outcomes in UK Employment Tribunal decisions by explicitly separating objective case facts from judge-specific discretion. Evaluated on 13,937 tribunal decisions, the approach outperforms supervised fine-tuning of a much larger Gemma-4 26B model while using an order of magnitude fewer trainable parameters, with the largest gains on the most ambiguous and rarest outcome classes. The architecture also offers interpretability through learned judge embeddings and calibration profiles that identify when adjudicative context—rather than case merit—drives predictions. These findings have implications for understanding and auditing judicial discretion in legal systems.
- ResearcharXiv2026-06-25Quality assurance · Algorithms & Automated Decisions
Adaptive Utility driven Resource Orchestration for Resilient AI (AURORA-AI) · Rahul Umesh Mhapsekar, Ilias Cherkaoui, Lizy Abraham et al.
AURORA-AI is a closed-loop resource orchestration framework that dynamically redistributes computational budget across a population of AI models to jointly optimize predictive performance, fairness, cost, latency, robustness, and interpretability under non-stationary conditions. It combines Hamilton-Jacobi-Bellman feedback control, Lyapunov-based stability monitoring, and a fairness-aware composite utility into a single adaptive policy. In a simulation stress-testing demographic bias shocks, concept drift, and black-swan disruptions, AURORA-AI recovered immediately from the black-swan event versus 88 time steps for a static baseline and 22 for PPO, lifted the alpha-quantile and super-quantile by 29% and 25% respectively, and simultaneously reduced demographic parity gaps. The results suggest that fairness-aware adaptive orchestration grounded in stability theory is a viable path toward resilient, human-centric AI deployment at enterprise scale.
- ResearcharXiv2026-06-25Quality assurance · Health
Auditing Framing-Sensitive Behavioral Instability in Large Language Models for Mental Health Interactions · Abla Bedoui, Ashley L. Greene, Mohammed Cherkaoui
This paper investigates how large language models (LLMs) used in mental health support applications respond differently to semantically similar concerns depending on how they are contextually framed. Using controlled matched prompts across multiple instruction-tuned model families, the researchers find that framing systematically alters interpretive response tendencies, with layer-wise probing showing that framing-related information is decodable throughout transformer layers. Activation steering experiments further suggest that framing-associated internal representations can partially influence downstream behavioral outputs. The findings highlight that robustness to contextual framing is an important consideration when evaluating the consistency and trustworthiness of AI systems deployed in mental-health-oriented settings.
- ResearcharXiv2026-06-25Quality assurance · Privacy & Data Protection · +1
RedVox: Safety and Fairness Gaps in Speech Models Across Languages · Beatrice Savoldi, Sara Papi, Wafa Aissa et al.
RedVox introduces a multilingual safety and fairness benchmark for speech-capable AI models, covering English, French, Italian, Spanish, and German using real human voices. The study surveys state-of-the-art model releases and finds that only 8% document any multilingual safety analysis. Evaluating eight models with RedVox, the researchers find that safety vulnerabilities persist even under non-adversarial conditions, worsen in non-English languages, and are amplified when inputs are spoken rather than text-based. The paper also highlights unique privacy and sociotechnical challenges in collecting naturalistic speech data from human participants.
- ResearcharXiv2026-06-25Quality assurance
A Deterministic Control Plane for LLM Coding Agents · Padmaraj Madatha
This paper examines how LLM coding agent configuration files (rules files, agent definitions, IDE-specific markdown) are managed across 10,008 public GitHub repositories. The study finds these configurations propagate as undeclared shared components, with 10.1% of tracked paths being SHA-256 exact duplicates across independent repositories and 75.5% of clone pairs crossing organisational boundaries; configurations are rarely revised and almost never declare permission boundaries (<1% vs 33% for CI/CD workflows). To address these gaps, the authors propose Rel(AI)Build, a deterministic control plane that treats agent definitions as a managed supply chain with content addressing, audit logs, tiered permissions, and prompt drift detection. The work highlights significant quality assurance and policy risks in how AI coding agent configurations are currently governed and distributed.
- ResearcharXiv2026-06-25Quality assurance · AI policy · +2
SamaVaani: Auditing and Debiasing Multilingual Clinical ASR for Indian Languages · Subham Kumar, Prakrithi Shivaprakash, Abhishek Manoharan et al.
SamaVaani audits eight state-of-the-art automatic speech recognition (ASR) models—including IndicWhisper, WhisperLargeV3, Sarvam, GoogleS2T, Gemma3n, OmniLingual, Vaani, and Gemini—on real-world psychiatric interview data in Kannada, Hindi, and Indian English, finding substantial performance variability across models and languages, with strong results in Indian English but frequent failures on regional speech. The study identifies systematic gaps tied to speaker role and gender, raising equity concerns for clinical deployment. The authors then fine-tune the two best open-source models (Gemma3n and OmniLingual) using a proposed fairness-aware technique called SamaVaani, which simultaneously improves overall ASR accuracy and reduces demographic performance disparities. These findings matter for healthcare quality assurance and policy around equitable AI deployment in multilingual clinical settings.
- ResearcharXiv2026-06-25Enterprise · Quality assurance · +1
AgentX: Towards Agent-Driven Self-Iteration of Industrial Recommender Systems · Changxin Lao, Fei Pan, Guozhuang Ma et al.
AgentX is a production-deployed multi-agent AI system designed to automate the full recommendation algorithm development cycle in industrial settings — from hypothesis generation and code writing to A/B testing and result analysis. The system uses four coordinated agents (Brainstorm, Developing, Evaluation, and a self-improvement layer called SGPO) operating in a closed loop, so that each experiment's outcomes feed back into sharpening the agents themselves. By removing the dependency on human engineers at each stage of the idea-to-launch pipeline, AgentX aims to make innovation in recommender systems scale with compute and accumulated knowledge rather than headcount. This matters for enterprise AI deployment because it demonstrates a path toward self-improving automation of complex engineering workflows that previously required sustained human expertise.
- ResearcharXiv2026-06-25Enterprise · Algorithms & Automated Decisions
AIGP: An LLM-Based Framework for Long-Term Value Alignment in E-Commerce Pricing · Chennan Ma, Yanning Zhang, Siqi Hong et al.
AIGP is an LLM-based pricing framework for large-scale e-commerce that addresses key shortcomings of traditional dynamic pricing: poor interpretability, inability to use unstructured information, and misalignment with long-term business goals. The system combines domain-knowledge-prompted LLMs with a Long-Term Value Estimator trained via offline reinforcement learning, using Direct Preference Optimization to align pricing decisions with objectives like GMV, ROI, and milestone achievement. In large-scale online A/B tests on Tao Factory, AIGP delivered +13.21% in GMV, +7.59% in ROI, and +8.20% in milestone achievement rate over 14 days compared to the production baseline, while also producing interpretable pricing rationales. This demonstrates that LLM-based approaches can meaningfully advance enterprise pricing strategy by bridging short-term decisions with long-term business value.
- ResearcharXiv2026-06-25Quality assurance · Safety & Harms · +1
Do Safety Guardrails Need to Reason? LeanGuard: A Fast and Light Approach for Robust Moderation · Dongbin Na
This paper challenges the assumption that safety guardrails for AI systems need chain-of-thought (CoT) reasoning to be effective. The authors train a lightweight 395M-parameter bidirectional encoder (LeanGuard) and a reasoning-based guard on the same data, then show that removing CoT does not hurt moderation accuracy — LeanGuard achieves an average F1 of 82.90 over public benchmarks, matching much larger reasoning-based decoders while using roughly 100x less inference compute. The label-only encoder also proves more robust under training-label noise and maintains better recall at strict false-positive rates, suggesting reasoning guards are not the safer choice either. The findings indicate that current guardrail benchmarks may not be challenging enough to justify the cost of CoT-based moderation, with practical implications for on-device deployments such as embodied robots.
- ResearcharXiv2026-06-25Quality assurance · Privacy & Data Protection
Agents That Know Too Much: A Data-Centric Survey of Privacy in LLM Agents · Nada Lahjouji, Ashwin Gerard Colaco
This survey examines privacy risks in large language model (LLM) agents that operate over sensitive data sources such as databases, document collections, external APIs, and agent memory. The authors take a data-centric approach, organizing known risks by the type of data an agent touches—including retrieval-augmented generation, text-to-SQL interfaces, and cross-session memory—rather than by attack type. Two key findings emerge: information-flow control is the only governance mechanism that addresses both compositional and cross-session inference leakage, and no existing benchmark evaluates an agent across all its data surfaces under a single privacy policy. The paper calls for a unified research framing and identifies these gaps as the field's most pressing open problems.
- ResearcharXiv2026-06-25Enterprise · Certifications · +1
Pingquanqi (Equalizer): A Cross-Domain Sociotechnical Framework for Human-Agent Interaction Governance · Yu Wang
This paper proposes Pingquanqi (Equalizer), a sociotechnical governance framework for Human-Agent Interaction (HAIGF) designed to be adopted as an open standard—analogous to WCAG for web accessibility—embedded at the agent framework level. The framework consists of five components: a user-state discrimination model, a Bayesian progressive stop-loss rule for capping per-session interaction costs, controlled friction mechanisms to break dependency loops, a transparency metric called Lsteal that converts token usage into user lifetime cost, and a reflective summarization mechanism. The paper argues the primary economic beneficiary is the enterprise deploying agent services, through reduced wasted computation, improved user satisfaction, and sustained subscription revenue, with individual user benefit as a downstream consequence. Its relevance spans enterprise deployment standards and policy-level governance of LLM agent infrastructure.
- ResearcharXiv2026-06-25Quality assurance · AI policy
Can Large Language Models Reliably Code Qualitative Humanitarian Data? A Benchmark Study Against Human Expert Adjudication · Jerome Marston, Tino Kreutzer, Salomé Garnier et al.
This benchmark study evaluates 46 large language models (LLMs) against a human Gold Standard for coding qualitative humanitarian data, using 150 synthetic transcripts and inter-rater reliability testing with Krippendorff's alpha. The authors find that multiple LLMs can match experienced human coders on deductive coding tasks, particularly when structured prompts and reasoning-enabled configurations are used, but aggregate reliability metrics alone are insufficient for deployment decisions. Models varied in their ability to recognize indirectly expressed needs, needs outside predefined categories, and protection-relevant concerns such as physical safety and discrimination. The findings indicate LLMs can expand humanitarian analytical capacity but require structured codebooks, tiered human oversight, and — for sensitive data — self-hosted open-weights models to balance scalability with data governance.
- ResearcharXiv2026-06-25Quality assurance · Health · +1
The Inattentional Gap: Task-Conditioned Language and Vision Models Omit the Safety-Critical Signals They Can Otherwise Report · Kwan Soo Shin, In Seok Kang, Yunkyung Min et al.
This paper identifies a phenomenon called the 'Inattentional Gap,' where language and vision models conditioned on a specific narrow task suppress reporting of co-present safety-critical signals they are otherwise capable of detecting — a machine analogue of human inattentional blindness. Across radiology and driving text scenarios and chest-radiograph vision tasks, focused task instructions suppressed reporting of off-task hazards by up to 0.92 in report rate, with explicit exclusive instructions abolishing such reporting entirely in radiology. The effect appeared across all tested models, did not diminish with scale, and persisted in reasoning models, meaning benchmark safety scores can look near-perfect while real-world safety hazards go unreported. The authors propose 'reporting-complete evaluation' — scoring what a system fails to report alongside what it is asked to find — and show that routing outputs to an independent open-ended critic can restore omitted findings.
- ResearcharXiv2026-06-25Quality assurance · AI policy · +2
Clinical Harness for Governable Medical AI Skill Ecosystems · Tianhan Xu, Lei Bao, Zhe Hu et al.
This paper introduces the 'Clinical Harness,' a runtime governance architecture designed to manage medical AI capabilities—termed 'clinical AI skills'—across the full lifecycle of patient care. Rather than relying on isolated AI models, the architecture registers, orchestrates, constrains, and monitors these skills to ensure accountability and persistence over time. Using osteoporosis as a case study, the authors demonstrate how knowledge-driven, data-driven, and physics-enhanced skills can be integrated within a governed framework. The work is relevant to certification and policy discussions around how medical AI systems can be made accountable and governable in clinical settings.
- ResearcharXiv2026-06-25Quality assurance
Adaptive Evaluation of Out-of-Band Defenses Against Prompt Injection in LLM Agents · Praneeth Narisetty, Shiva Nagendra Babu Kore, Uday Kumar Reddy Kattamanchi et al.
This paper examines 'out-of-band' defenses against prompt injection attacks on LLM agents—security mechanisms enforced outside the model itself using classical integrity and least-privilege principles, as seen in systems like CaMeL, FIDES, Progent, RTBAS, and FORGE. The authors warn that all such defenses have only been validated on static benchmarks, the same flaw that allowed adaptive attacks to break twelve in-band defenses at over 90% success rates. As an independent test, they ran an adaptive evaluation of Progent on the AgentDojo benchmark using an open-weight model (Qwen2.5-7B), finding that Progent reduced mean attack success roughly sixfold (25.8% to 4.2%) and a hand-crafted adaptive attack did not meaningfully raise it (2.6%). The results are consistent with—but do not conclusively establish—that deterministic out-of-band enforcement is more robust against adaptive attackers than in-band detection, while highlighting that stronger white-box attacks remain untested.
- ResearcharXiv2026-06-25Enterprise · Quality assurance · +3
Auditing a Robotic System for the AI Act · Laura Lucaj, Felix Bok, Patrick van der Smagt
This paper presents a framework for auditing robotic AI systems under the EU AI Act, validated on a real-world case of hospital ventilation-cleaning robots. The authors find that existing audit methodologies designed for software-based AI are insufficient for embodied reinforcement learning systems, because compliance evidence from simulation does not guarantee real-world safety. The framework addresses specific challenges such as the Sim2Real gap, policy opacity, and distributed stakeholder responsibility, while translating AI Act obligations—risk management, human oversight, and post-market monitoring—into concrete audit criteria. The work argues for context-sensitive, lifecycle-embedded auditing practices, especially for resource-constrained organizations operating under regulatory uncertainty without harmonized standards.
- ResearchSustainability2026-06-25Enterprise
The Impact of the Implementation of the AI Systems in Small and Medium Enterprises in Poland: Scale of Usage, Productivity, and Unperceived Sustainability · Michał Polasik, Marta Czarkowska, Wojciech Śniadkowski et al.
This study examines AI adoption among 112 SMEs in Poland's Kuyavian–Pomeranian region, combining survey data with manager interviews to assess organizational, economic, and sustainability impacts. Results show AI most strongly reduces workload and improves time efficiency, especially in service firms with intensive AI use, though benefits come alongside new costs from paid tools, data preparation, and governance. Adoption follows a staged path from experimentation to workflow integration, with barriers shifting from knowledge gaps early on to data quality and security issues at advanced stages. Notably, sustainability considerations such as environmental and ESG impacts remain largely unperceived by SME decision-makers, who instead frame sustainability through resilience and competitiveness.
- ResearcharXiv2026-06-24Enterprise · AI policy · +3
When Agents Meet Electric Bus Fleet Operations: Pricing Behavior, Trade-offs, and Policy Implications in an Aggregator Framework · Jônatas Augusto Manzolli, Ali Eslami, Luis Miranda-Moreno et al.
This paper proposes an agentic aggregator framework for managing electric bus fleet operations, combining optimization-based scheduling with supervisory AI agents that handle disturbance detection, tariff adaptation, and real-time re-optimization across charging and vehicle-to-grid (V2G) activities. A realistic depot case study finds that the framework can maintain feasible schedules and improve use of charging flexibility under various operational disruptions, but also reveals that profit-oriented agent configurations can extract value from the public transport operator at its expense. The authors conclude that deploying agentic aggregators in public-fleet contexts requires transparent coordination modes, auditable tariff-setting, and explicit value-sharing rules to prevent misaligned incentives.
- ResearcharXiv2026-06-24Quality assurance · Algorithms & Automated Decisions
Instruction Bleed: Cross-Module Interference in Prompt-Composed Agentic Systems · Ching-Yu Lin, Yifan Liu
This paper formalizes a failure mode called compositional behavioral leakage (CBL), where editing one prompt module in an AI agent silently shifts the behavior of other modules that share the same context window, even without any direct variable or code dependency. The authors probe this on a deployed job-evaluation agent (Claude Sonnet 4.6) across 144 trials using a three-channel perturbation protocol targeting volume, content, and form of non-focal modules; only content-channel perturbations produced a detectable effect (Cohen's d = 0.63 with a bootstrap 95% CI excluding zero), though no individual recommendation flipped. While sub-threshold in standard QA terms, the authors argue this interference compounds silently across thousands of agent decisions, making it a systematic evaluation blind spot. The paper contributes an operational definition, a reusable measurement protocol, and a call for cross-module interference testing as a standard requirement in prompt-composed agent evaluation.