News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
The New Social Image: How AI Competency and AI Proactivity Influence Self- and Peer-Perceptions in the Workplace
Kuntal Ghosh, Marc Hassenzahl, Shadan Sadeghian
arXiv · 2026-05-29
This vignette study (n=50) examines how AI competency and AI proactivity in human-AI workplace teams shape workers' self-perceptions and peer perceptions of ownership, job meaningfulness, and satisfaction. Results show that AI with low competency or low proactivity generally improved feelings of ownership, meaningfulness, satisfaction, and positive affect, while highly competent or proactive AI had the opposite effect. Notably, these outcomes differed depending on whether workers rated themselves or were rated by coworkers—for example, low AI proactivity raised job satisfaction from a self-perception standpoint but not from a peer-perception standpoint. The findings suggest that designing workplace AI purely around performance metrics may be insufficient, as highly capable AI can undermine job identity, social image, and team dynamics.
- Workforce
Research
dashi: A Python library for Dataset Shift Characterization to Support Trustworthy AI Development and Deployment
David Fernández-Narro, Pablo Ferri, Ángel Sánchez-García et al.
arXiv · 2026-05-29
This paper introduces dashi, an open-source Python library designed to detect and characterize dataset shifts—changes between training and test data distributions that can degrade AI model performance. The library offers both unsupervised methods (using information geometry and non-parametric statistical manifolds) and supervised methods (quantifying model performance degradation) across temporal and multi-source data batches. The authors demonstrate dashi on three health AI case studies involving gestational diabetes mellitus, COVID-19, and emergency medical dispatch, showing how the tool supports robust and safe machine learning pipelines through interactive visual analytics and variability metrics. This matters because uncontrolled dataset shifts pose real risks to patient safety and data quality in health AI, and accessible tooling for shift analysis has been lacking.
- Quality assurance
Research
A Visually Impaired Assistance Benchmark for VLM-as-a-Judge Evaluation
Yi Zhao, Siqi Wang, Zhe Hu et al.
arXiv · 2026-05-29
This paper introduces VIABLE, the first benchmark for evaluating whether Vision-Language Models (VLMs) can reliably act as judges for AI-based Visually Impaired Assistance (VIA) tasks. VIABLE contains over 300K judgment samples across three scenarios and uses an Effectiveness–Impartiality–Stability framework with a 12-mode failure taxonomy to assess seven VLM judges at different scales. The study finds that all existing judges are largely unreliable: even the strongest judge, GPT-5.4, achieves only 52.6% single-failure diagnostic accuracy and exhibits a 94.2% self-preference rate, while open-source models are strongly biased and adversarially fragile. To address these shortcomings, the authors propose VIA-Judge-Agent, a model-agnostic inference-time framework that adds visual evidence extraction and taxonomy-guided reasoning, improving both diagnostic accuracy and the quality of VIA responses as preferred by blind and low-vision (BLV) users.
- Quality assurance
Research
Neither Replacement nor Panacea: Comparing LLM-Based Conversational and Graphical Decision Support in Industrial Tasks
Roberto Figliè, Simone Caputo, Alan Serrano et al.
arXiv · 2026-05-29
This study compares an LLM-based conversational agent (CA) delivered through a conversational user interface (CUI) against a traditional dashboard for decision support in manufacturing settings, using a mixed factorial experiment with 134 industrial decision-makers across tasks of increasing complexity. Results showed the CUI reduced perceived mental workload and supported faster task completion for simpler tasks, but these advantages diminished as task complexity grew. Neither interface consistently outperformed the other on decision accuracy, and the CUI was not preferred as a sole basis for subsequent decisions. The findings suggest conversational AI offers conditional rather than universal benefits for industrial decision support, with complex decisions continuing to benefit from persistent, inspectable visual representations.
- Enterprise
- Workforce
Research
RealityTest: How People Probe AI Identity and Whether Models Disclose It
Anna Gausen, Sarenne Wallbridge, Bessie O'Dell et al.
arXiv · 2026-05-29
RealityTest is a large-scale multimodal and multilingual benchmark for evaluating whether AI systems disclose their identity when asked by users. Drawing on 3,152 identity-probing queries from approximately 750 participants across 49 countries and five languages, the study finds that only 31% of people ask about AI identity directly in ambiguous scenarios, and that human-generated questions are far more diverse than machine-generated ones. Testing 17 text and 6 speech models reveals substantial variation in disclosure behavior, but a single suppression instruction drops disclosure rates below 30% even in the best-performing models. The findings highlight that question phrasing and conversational context matter more than model choice, and that evaluations relying on narrow or synthetic query sets risk misrepresenting real-world AI behavior—a concern directly relevant to ongoing regulatory efforts around AI transparency.
- AI policy
- Quality assurance
Research
LLM-FACETS: A Privacy-Preserving Framework for Evaluating LLM Transparency and Accountability
Tom Lucas, Alessio Buscemi, Alfredo Capozucca et al.
arXiv · 2026-05-29
LLM-FACETS is an open-source framework designed to make auditing of Large Language Models accessible to non-technical practitioners such as domain experts and compliance officers, without requiring programming expertise or transmitting data to external services. The system structures evaluation around three stakeholder profiles aligned with the EU AI Act and NIST AI Risk Management Framework, and operationalizes transparency through token-level uncertainty visualization, multi-judge consensus, and RAG Triad hallucination metrics. Deterministic metrics run entirely on a self-hosted server with no outbound data transmission, while any LLM-judge metrics that contact external APIs do so explicitly with users retaining credential control. The authors validate the framework by cross-checking 18 metric implementations against canonical reference libraries, supporting reproducibility and decoupling AI accountability from the teams building the assessed systems.
- Quality assurance
- AI policy
- Certifications
Research
From Prompt Injection to Persistent Control: Defending Agentic Harness Against Trojan Backdoors
Jiejun Tan, Zhicheng Dou, Xinyu Yang et al.
arXiv · 2026-05-29
This paper identifies a new security threat in which LLM agents operating in local workspaces (reading/writing files, calling tools, reusing state across sessions) can be compromised through multi-step 'trojan' attacks: malicious instructions are injected into files or tool outputs, stored, and executed later, with no single step appearing harmful in isolation. The authors introduce ClawTrojan, a benchmark for evaluating these attacks, demonstrating a 95.5% attack success rate against GPT-5.4 in a simulated workspace, compared to near-zero success for existing single-turn prompt-injection attacks. They also propose DASGuard, a defense that scans sensitive local files for control-like text, traces its origin, and removes untrusted content at runtime. The work highlights critical security gaps in current agentic AI deployments and offers a concrete mitigation framework relevant to enterprise and policy considerations around safe AI agent use.
- Enterprise
- AI policy
Research
Traceable by Design: An LLM Pipeline and Dashboard for EU Regulatory Consultation Analysis
Thales Bertaglia, Haoyang Gui, Catalina Goanta et al.
arXiv · 2026-05-29
This paper presents an end-to-end LLM-based pipeline and interactive dashboard designed to help analysts process large volumes of regulatory consultation submissions. Applied to 4,322 submissions from the European Commission's Digital Fairness Act public call for evidence, the system extracted 15,368 topic annotations each grounded in verbatim quotes from source documents. The design prioritizes traceability and transparency, surfacing stakeholder concerns—such as Age Verification, Payment Processor Censorship, and Digital Ownership—that fixed-taxonomy approaches would have missed. By making consultation analysis scalable and auditable, the tool has direct implications for how policymakers engage with public input in regulatory processes.
- AI policy
Research
Cognitive Fatigue in Autoregressive Transformers: Formalization and Measurement
Riju Marwah, Ritvik Garimella, Vishal Pallagani et al.
arXiv · 2026-05-29
This paper formalizes 'cognitive fatigue' in autoregressive language models—a measurable degradation pattern during long-form generation characterized by reduced attention to the original prompt, representational drift, and entropy miscalibration. The authors introduce the Fatigue Index (FI), a lightweight, model-agnostic diagnostic that aggregates these signals under explicit axioms, achieving AUROC of 0.95 for predicting task degradation and Spearman rho of 0.94 for repetition across nine models ranging from 1B to 13B parameters. Key findings include non-monotonic scaling behavior where instruction-tuned models below 3B collapse faster than base models, with the trend reversing at 7B, and that fatigue onset accelerates under longer contexts, middle-positioned evidence, and reduced numerical precision. This matters for enterprise and quality-assurance contexts because FI enables real-time runtime monitoring of LLM reliability in production systems without requiring model modifications.
- Enterprise
- Quality assurance
Research
VisualLeakBench: Reproducible Action-Boundary Propagation Failures in Vision-Language Agents
Youting Wang, Yuan Tang, Yitian Qian et al.
arXiv · 2026-05-29
VisualLeakBench is a 500-image benchmark designed to measure a specific failure mode in vision-language agents called 'action-boundary propagation,' where sensitive or unsafe text visible in images (such as PII or unsafe content) is copied into downstream tool arguments like search queries or external handoffs. Testing four production VLM systems on a 100-image stratified subset, the study finds that sensitive text is propagated into tool arguments in 78.8% of PII cases and 85.5% of rendered unsafe-text cases at baseline. Even under a defensive system prompt, rendered unsafe-text propagation remains high at 52.6%, and PII suppression is largely achieved by disabling tool use entirely rather than preserving agent utility. These results highlight a systematic and reproducible security risk in agentic AI systems that handle real-world visual content before taking actions.
- Quality assurance
- AI policy
Research
PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say
Mingxuan Zhang, Jiahui Han, Dadi Guo et al.
arXiv · 2026-05-29
PrivacyPeek introduces a benchmark of 1,182 cases designed to audit the acquisition stage of LLM-based agents—examining what sensitive data an agent collects via tool calls, not just what it discloses in its responses. The benchmark covers 7 acquisition behaviours and 16 application domains, and experiments across 10 agents from 4 model families show that unnecessary acquisition of sensitive information is widespread. The study also finds that prompt-level defences mitigate only a small fraction of this leakage, and that higher task-completion capability correlates with greater acquisition-stage privacy risk. These findings highlight a critical gap in current privacy auditing practices for AI agents and underscore the urgency of addressing over-acquisition before it leads to outright data leaks.
- Quality assurance
- AI policy
Research
Hide-and-Seek in Trajectories: Discovering Failure Signals for VLA Runtime Monitoring
Seongheon Park, Wendi Li, Changdae Oh et al.
arXiv · 2026-05-29
Hide-and-Seek is a framework for detecting execution failures in Vision-Language-Action (VLA) robot models during runtime, without requiring expensive action resampling, external models, or step-level annotations. It formulates failure detection as a coarsely supervised learning problem, using inter-trajectory and intra-trajectory contrastive objectives to localize failure-indicative actions from trajectory-level labels alone. Evaluated on LIBERO, VLABench, and a real-world robotic platform across OpenVLA, π0, and π0.5 policies, the method achieves state-of-the-art multi-task failure detection with a practical accuracy–timeliness trade-off under conformal prediction, generalizing to both seen and unseen tasks. This matters for quality assurance of deployed robotic systems, where timely and reliable failure detection is essential for safe real-world operation.
- Quality assurance
Research
Anchoring LLM Gender Bias to Human Baselines: A Cross-Lingual Audit
Jiwoo Choi, Seonwoo Ahn, Tongxin Zhang et al.
arXiv · 2026-05-29
This paper audits six large language models (LLMs)—three English-centric (Claude, GPT, Gemini) and three East Asian (DeepSeek, Syn-Pro, HyperCLOVA X)—for gender stereotyping across English, Korean, Chinese, and Japanese, using the HEXACO-100 personality inventory anchored against a cross-cultural human dataset spanning 48 countries. The study finds that LLM gender stereotyping spans a range roughly 2.5 times wider than the full cross-country variation observed in humans, and that the effect can compound across languages—one English-centric model prompted in Korean reached five times the local human baseline, even when prompts indicated the candidate had already been hired. The authors introduce a four-pattern framework (concordance, suppression, reorganization, and amplification) to characterize model behavior across 24 model-language combinations, finding that translation does not merely rescale stereotypes but restructures which attributes are stereotyped. The results suggest no single debiasing pipeline is likely to address gender bias evenly across linguistic boundaries, with important implications for the equitable deployment of LLMs across diverse populations.
- AI policy
- Workforce
Research
PReMISE: Policy Rubrics as Measurement Specifications for LLM Judges
Swastik Roy, Rajkumar Pujari, Tharindu Kumarage et al.
arXiv · 2026-05-29
PReMISE is a framework that treats LLM evaluation rubrics as formal measurement specifications and audits them across four axes: structural adequacy, reliability, preference fit, and adversarial robustness. The paper finds that no existing rubric source simultaneously satisfies all four criteria—high inter-rater agreement does not guarantee low exploitability—while vague rubrics can reward polished but factually incorrect or intent-violating responses. PReMISE's repair operations improve judge accuracy from 65.0% to 68.6% and reduce the rate at which adversarially crafted responses receive high scores from 46.4% to 36.0%. This matters because LLM judges are increasingly used to assess open-ended AI outputs, and poorly specified rubrics undermine the validity and trustworthiness of those evaluations.
- Quality assurance
Research
ReviewGuard: Aligning LLM-Assisted Peer Review with Long-Term Scientific Impact
Abdur Rasool, Xiaohui Huang, Yanqing Hu et al.
arXiv · 2026-05-29
ReviewGuard is a two-stage framework that reorients LLM-generated peer review away from mimicking human reviewer preferences and toward predicting long-term scientific impact, measured by future citations. Trained and evaluated on 20,861 AI/ML papers from OpenReview combined with Semantic Scholar citation data, it achieves a Spearman correlation of 0.776 with future citations on rejected-then-published papers, outperforming both human reviewers (0.492) and a supervised Expert model (0.681). Critically, it flags 10.2% of high-impact rejected papers versus only 1.8% for human reviewers—a 5.6x improvement—suggesting it can serve as a complementary signal for editors identifying undervalued work. The authors frame this as augmenting, not replacing, human judgment in scientific quality control.
- Quality assurance
Research
MosaicLeaks:Privacy Risks in Querying-in-the-Open for Deep Research Agents
Alexander Gurung, Spandana Gella, Alexandre Drouin et al.
arXiv · 2026-05-29
MosaicLeaks examines a privacy vulnerability in deep research AI agents that combine private local documents with external web retrieval: the agent's external queries can inadvertently expose sensitive information, and even individually innocuous queries can collectively reveal private details through the 'mosaic effect.' The authors introduce a benchmark of 1,001 multi-hop research tasks pairing private enterprise documents with a public web corpus, then evaluate how much an adversary can infer about private information by observing only the agent's external queries. They find that models across families and sizes frequently leak private information at multiple levels, and that simple zero-shot privacy prompting reduces but does not eliminate leakage. Their proposed Privacy-Aware Deep Research (PA-DR) reinforcement learning framework improves task accuracy from 48.7% to 58.7% while cutting answer and full-information leakage from 34.0% to 9.9%, demonstrating that privacy and performance can be jointly optimized.
- Enterprise
- Quality assurance
Research
Triaging Threats to Specialized Guardrails
Wenjie Jacky Mo, Xiaofei Wen, Rui Cai et al.
arXiv · 2026-05-29
This paper addresses a key limitation of current LLM safety guardrails: existing models are trained on fragmented datasets with inconsistent risk taxonomies, making it unclear whether they generalize to real-world threats. The authors introduce GuardZoo, a human-annotated benchmark of 32,460 samples spanning 15 unsafe categories, and use it to show that monolithic guardrail models suffer from 'task interference' — different threat domains need distinct decision boundaries that a single model struggles to capture. To address this, they propose RouteGuard, a router-expert framework that routes each conversation to a specialized guardrail model, improving fine-grained threat detection, out-of-domain generalization, and modular expandability to new threats. These findings have direct implications for the reliability and robustness of safety mechanisms used when deploying LLMs in production settings.
- Quality assurance
- Enterprise
Research
Depth-Dependent Indirect Prompt Injection in Tool-Calling ReAct Agents: Injection Depth, Payload Framing, and Turn-Budget Sensitivity
Mohammadreza Rashidi
arXiv · 2026-05-29
This paper investigates how the position, framing, and turn budget of adversarial payloads affect the success of indirect prompt injection attacks against ReAct-style AI agents that use tool calls. Across 460 controlled trials on GPT-4o-mini and Claude Haiku, the study finds that attack success rate (ASR) against GPT-4o-mini decays sharply with injection depth—from 60% at the first tool observation to 0% at depths 4 and 5—while Claude Haiku resists attacks at all depths. Rhetorical framing (e.g., persona-style payloads) can swing ASR by 50 percentage points at depth 1, though this result did not reach statistical significance, and turn budget had no measurable effect on ASR. The findings identify injection depth as the dominant risk variable and suggest that sanitizing only the first tool observation would capture 67% of measured injection successes, offering a practical prioritization for defenders.
- Quality assurance
- AI policy
Research
Healthcare Mechanisms from Policy-as-Code Search under Strategic Provider Response
Zihan Wang, Xiang Xu, Hongyuan Zha et al.
arXiv · 2026-05-29
This paper addresses a fundamental gap in healthcare AI evaluation: existing benchmarks fix provider behavior rather than modeling how providers strategically respond to payment rules. The authors reframe hospital mechanism design as a program synthesis problem, using a multi-agent simulator (Medi-Sim) with five strategic provider channels to score candidate rule programs at equilibrium. Their incentive sweep reproduces classical health-economics phenomena—such as up-coding, patient selection, and Goodhart-style drift—and reveals that closing one gaming channel (coding) more than doubles another (low-complexity patient selection). An LLM-guided evolutionary search then synthesizes an inspectable mixed-objective payment rule that eliminates up-coding, halves rejection rates, and retains most of the profit-oriented baseline's funds, demonstrating that equilibrium-aware mechanism design can meaningfully improve healthcare payment policy.
- AI policy
- Quality assurance
Research
Closed-Loop Quality Assurance for Production Clinical AI Documentation
Andrew Napier, Justin Wiley, Mark Heslin
medRxiv · 2026-05-29
This paper describes a closed-loop quality assurance system for clinical AI documentation deployed across 13 US hospital sites, reporting zero regressions on 42 tracked cases over 1,089 optimization iterations and a latency reduction in note generation from 19.6 s to 10.8 s after replacing an LLM-based assembly agent with a deterministic template. The authors identify a key distinction between reasoning agents (which respond to prompt optimization) and assembly agents (which do not, and instead require deterministic post-processing or full LLM removal). They also find that binary quality checks used as optimization targets produce notes scoring 90%+ that physicians reject, while physician preference ratings show near-zero inter-rater agreement (Cohen's κ = 0.028), indicating that neither automated checks nor subjective preference alone is sufficient for quality assurance. The work has direct implications for how clinical AI systems are evaluated, deployed, and iteratively improved in real hospital environments.
- Quality assurance
- Enterprise
- Workforce
Research
LLMs in the Real World: Evaluating "AI" in Emergency Contexts
Sara Court, Lara Downing, Micha Elsner
arXiv (Cornell University) · 2026-05-29
This paper presents a case study examining an LLM-based machine translation system deployed in a text-2-911 emergency service that claimed support for 55 languages. The authors identify common misconceptions about such AI technologies in high-stakes contexts and argue that researchers have a responsibility to communicate findings to the public and policymakers. The paper concludes with concrete recommendations and best practices for all stakeholders across the development and deployment pipeline, warning that overlooked 'easy' problems—not just cutting-edge challenges—pose the greatest risks in real-world deployments.
- Quality assurance
- AI policy
- Certifications
Research
The Smarter State? Artificial Intelligence and Modern State and Local Public Finance
David R. Agrawal, William F. Fox
CESifo · 2026-05-29
This paper examines how artificial intelligence reshapes subnational public finance, finding that AI shifts income from labor toward capital and reallocates tax bases toward consumption, raising questions about sales tax treatment of digital services. AI also relaxes informational and administrative constraints for governments in taxation, enforcement, budgeting, and service delivery, while strengthening scale economies. However, the authors warn that government AI use may advantage larger jurisdictions with greater data access, raising equity and transparency concerns and increasing the value of interstate cooperation. Overall, AI reinforces rather than overturns classic trade-offs in fiscal federalism.
- AI policy
- Enterprise
- Workforce
Research
Perception of the benefits of artificial intelligence in public auditing and its impact on technology acceptance: empirical evidence from European regional audit institutions
Natalia Alonso-Morales, Alejandro Sáez-Martín, Ana Maria Plata-Díaz et al.
Journal of financial reporting & accounting · 2026-05-29
This study surveys 219 auditors from European Regional Audit Institutions to examine how perceptions of AI benefits and UTAUT model factors—performance expectancy, effort expectancy, and social influence—shape intentions to adopt AI in external public auditing. Perceived benefits emerged as the primary driver of adoption intent, with performance expectancy acting only indirectly through perceived benefits, while gender moderated the relative importance of instrumental versus social factors. The findings extend the UTAUT model by positioning perceived benefits as a central mediating variable, offering practical guidance for encouraging AI uptake in public audit contexts. The research matters for public sector accountability and efficiency as AI adoption in government auditing remains limited despite its potential for task automation, big data analysis, and risk detection.
- Workforce
- Enterprise
- Quality assurance
- AI policy
Research
Data Integrity Failures in Pharmaceutical Digital Twins and Continuous Manufacturing: An Alcoa + + Framework Integrating Human Factors and Simulation Vulnerabilities
P Ramprasath, Mohan Gandhi Bonthu
Journal of Pharmaceutical Innovation · 2026-05-29
This review paper analyzes data integrity failures in pharmaceutical digital twin and continuous manufacturing systems, classifying 248 real-world incidents using the ALCOA++ framework combined with human factors analysis. The study found that 65% of failures involved non-contemporaneous data, 22% involved non-traceable simulation inputs, and 13% stemmed from cloud synchronization issues. A proposed Digital Twin Compliance Framework integrating human-centered design, GAMP 5.2 risk assessment, hybrid audit trails, and AI-based anomaly detection reduced simulated failure rates by 68%. The findings highlight urgent gaps in virtual data governance and validation standards for pharmaceutical manufacturing, with implications for regulatory compliance and quality assurance.
- Quality assurance
- Certifications
- AI policy
- Enterprise
Research
EUDAIMONIA: Evaluating Undesirable Dynamics in AI
Jun Rui Huang, Wang Bill Zhu, Ziyi Liu et al.
arXiv · 2026-05-28
This paper introduces EUDAIMONIA, a benchmark for evaluating whether large language models (LLMs) behave safely in social and companionship contexts — specifically whether they encourage harmful intimacy, unhealthy dependence, or excessive engagement. The authors develop a 'Social AI Design Code' framework, operationalized through 969 user inputs and 3,147 design-requirement violation checks drawn from real interactions. Testing 22 recent LLMs, they find that even the best-performing models (Claude-Opus-4.7 and GPT-5.5) violate 30.7% and 27.2% of checks respectively, and that extended thinking does not reduce these failure rates. The findings suggest that social alignment harms represent a persistent, structural problem not addressable through test-time reasoning improvements alone.
- Quality assurance
- AI policy