News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
Prioritising generative AI adoption challenges in construction projects: a Fuzzy Analytic Hierarchy Process approach
Alejandro Vidal Fernández, Yixue Shen
International Journal of Construction Management · 2026-05-07
This study identifies and ranks the most significant barriers to adopting generative AI in construction projects using a Fuzzy Analytic Hierarchy Process (FAHP) informed by a systematic literature review and expert surveys. The top challenges found are data privacy, security and legal liability; poor quality and availability of training data; and lack of construction-specific knowledge and language. The findings highlight the need for strong data governance, improved data quality, and domain-specific AI capabilities. Results offer practical prioritization guidance for construction managers seeking to implement generative AI effectively.
- Enterprise
- AI policy
- Workforce
Research
FinRAG-12B: A Production-Validated Recipe for Grounded Question Answering in Banking
Denys Katerenchuk, Pablo Duboue, Keelan Evanini et al.
arXiv · 2026-05-06
FinRAG-12B is a 12-billion-parameter large language model fine-tuned for grounded question answering in banking, trained on just 143 million tokens using a pipeline that combines LLM-as-a-Judge filtering, citation annotation, and curriculum learning. The model outperforms GPT-4.1 on citation grounding and achieves a calibrated refusal rate of 12% (versus an unsafe 4.3% for the base model and an over-refusing 20.2% for GPT-4.1), addressing regulatory and accuracy demands specific to financial institutions. Deployed at 40+ financial institutions, it delivers a 7.1 percentage point improvement in query resolution (p < 0.001) while running 3–5x faster at 20–50x lower cost than GPT-4.1. The work demonstrates a practical, production-validated recipe for deploying domain-specific LLMs under real-world constraints in a heavily regulated industry.
- Enterprise
- AI policy
- Quality assurance
Research
The Refusal--Compliance Tradeoff: A Large-Scale Safety Behavior Audit of Large Language Models
Alif Al Hasan, Sumon Biswas
arXiv · 2026-05-06
This paper audits 21 open-weight large language models across four safety benchmarks to examine the tradeoff between over-refusing benign prompts and complying with harmful ones. The authors find that different model families adopt fundamentally different calibration strategies — conservative models like Llama suppress unsafe outputs but over-refuse benign content, while permissive models like DeepSeek and Qwen remain more helpful but tolerate higher harmful compliance. Demographic protection is also uneven: prominent racial and religious groups are over-protected (even benign prompts are refused), while disability-targeted harmful prompts receive weaker protection. Safety behavior is found to be stable within model families across generations and scales, suggesting post-training objectives matter more than architecture, and the authors call for joint, demographically-aware, multi-judge safety evaluation.
- Quality assurance
- AI policy
Research
From Cradle to Cloud: A Life Cycle Review of AI's Environmental Footprint
Katherine Lambert, Sasha Luccioni
arXiv · 2026-05-06
This paper presents a structured literature review of scientific papers and technical reports examining AI's environmental impacts across an eight-stage life cycle framework — from hardware manufacturing and infrastructure construction through model training, inference, and end-of-life disposal. The authors find that while 'green AI' language is increasingly common, definitions of the AI life cycle remain inconsistent, with many studies narrowly focused on training and inference while neglecting data collection, infrastructure, and embodied emissions. Reporting practices rely predominantly on coarse CO2e proxies, with limited attention to water usage, materials manufacturing, and multi-impact life cycle assessment, making cross-study comparison difficult. The paper proposes improved measurement and reporting approaches to enable more comprehensive and policy-relevant assessments of AI's environmental footprint.
- AI policy
Research
LaTA: A Drop-in, FERPA-Compliant Local-LLM Autograder for Upper-Division STEM Coursework
Jesse A. Rodríguez
arXiv · 2026-05-06
LaTA is an open-source, on-premises LLM-based autograder designed for upper-division STEM courses that processes LaTeX-formatted student submissions entirely on local hardware, avoiding FERPA violations associated with sending student data to third-party APIs. Deployed in a mechanical engineering course at Oregon State University for approximately 200 students, the system graded weekly assignments at $0 marginal cost per assignment in 1–3 minutes per submission, with an instructor-confirmed grading-error rate of roughly 0.02–0.04% per rubric line item. Compared to a traditionally-graded cohort, the LaTA-graded cohort scored approximately 11% higher on the midterm and 8% higher on the final exam, and reported statistically significant gains in self-assessed confidence across all learning objectives (N=159, Δ≥+1.49 Likert points, p<10⁻²⁷). The paper demonstrates that privacy-compliant, low-cost AI autograding is practical on commodity hardware and can meaningfully improve student outcomes and TA availability.
- Quality assurance
- AI policy
Research
Partial Evidence Bench: Benchmarking Authorization-Limited Evidence in Agentic Systems
Krti Tallam
arXiv · 2026-05-06
Partial Evidence Bench is a deterministic benchmark designed to measure a specific failure mode in enterprise AI agents: producing answers that appear complete even when material evidence falls outside the caller's authorization boundary. The benchmark includes 72 tasks across three scenario families—due diligence, compliance audit, and security incident response—with ACL-partitioned corpora and structured oracles for correctness, completeness awareness, gap-report quality, and unsafe completeness behavior. Baseline results show that silent filtering is 'catastrophically unsafe' across all scenario families, while explicit fail-and-report behavior eliminates unsafe completeness without reducing the system to trivial abstention. The benchmark's key contribution is making this governance-critical agent failure measurable without human judges or contamination-prone static corpora, which has direct implications for how enterprise agentic systems are evaluated and deployed safely.
- Enterprise
- Quality assurance
Research
Making AI Drafts Count: A Quality Threshold in Audio Description Workflows
Lana Do, Shasta Ihorn, Charity M. Pitcher-Cooper et al.
arXiv · 2026-05-06
This paper investigates how the quality of AI-generated drafts affects human editing workflows for audio description (AD), which narrates visual content for blind and low-vision audiences. Using a purpose-built generation pipeline (GenAD) and editing interface (RefineAD), a within-subjects study found that high-quality AI drafts cut completion time by more than half and significantly reduced cognitive load compared to authoring from scratch, while simple unguided AI drafts offered only modest benefits. The findings point to a minimum quality threshold for AI assistance to be effective, and qualitative results suggest this threshold rises with visual complexity. The authors propose this as a broader design principle: effective AI assistance must clear a content-appropriate quality bar, not merely be present.
- Workforce
- Quality assurance
Research
How Far Are VLMs from Privacy Awareness in the Physical World? An Empirical Study
Junran Wang, Xinjie Shen, Zehao Jin et al.
arXiv · 2026-05-06
This paper introduces ImmersedPrivacy, an audio-visual benchmark built on a Unity-based simulator that evaluates how well Vision-Language Models (VLMs) handle privacy-sensitive situations when deployed as embodied agents in physical environments like homes and hospitals. The framework tests 12 state-of-the-art models across three tiers: identifying sensitive items in cluttered scenes, adapting to shifting social contexts, and resolving conflicts between explicit commands and inferred privacy constraints. Results show consistent failures across all models—performance decays monotonically as scene complexity increases, no model exceeds 65% accuracy when social context shifts, and even the best model (gemini-3.1-pro) correctly balances task completion with privacy preservation in only 51% of conflict cases. These findings highlight that current VLMs suffer from perceptual fragility and cannot reliably apply privacy knowledge to guide behavior in real-world physical settings, raising important concerns for the safe deployment of AI-powered embodied assistants.
- AI policy
- Quality assurance
Research
Understanding Annotator Safety Policy with Interpretability
Alex Oesterling, Donghao Ren, Yannick Assogba et al.
arXiv · 2026-05-06
This paper introduces Annotator Policy Models (APMs), interpretable models that learn annotators' internal safety policies from labeling behavior alone—without requiring annotators to explain their reasoning. APMs achieve over 80% accuracy in modeling annotator safety policy, faithfully predict responses to counterfactual edits, and recover known policy differences in controlled settings. Applied to both LLM and human annotators, APMs can surface policy ambiguity (revealing how annotators interpret safety instructions differently) and value pluralism (uncovering systematic safety priority differences across demographic groups). These capabilities support more targeted, transparent, and inclusive AI safety policy design by helping distinguish operational failures, ambiguous policy wording, and genuine value differences as sources of annotation disagreement.
- AI policy
- Quality assurance
Research
Securing the Agent: Vendor-Neutral, Multitenant Enterprise Retrieval and Tool Use
Francisco Javier Arceo, Varsha Prasad Narsing
arXiv · 2026-05-06
This paper identifies a critical security gap in enterprise Retrieval-Augmented Generation (RAG) and agentic AI systems: retrieval mechanisms rank documents by relevance rather than authorization, meaning a query from one tenant can inadvertently surface another tenant's confidential data. The authors formalize this problem and additional vulnerabilities—such as tool-mediated disclosure, context accumulation across turns, and client-side orchestration bypass—that emerge when relevance is conflated with authorization. To address these issues, they introduce a layered isolation architecture combining policy-aware ingestion, retrieval-time gating, and server-side agentic orchestration, implemented in an open-source framework called OGX. Empirical evaluation shows that Attribute-Based Access Control (ABAC) gating eliminates cross-tenant data leakage while introducing negligible overhead, making the approach practical for shared enterprise infrastructure.
- Enterprise
- Quality assurance
Research
Automatically Finding and Validating Unexpected Side-Effects of Interventions on Language Models
Quintin Pope, Ajay Hayagreeve Balaji, Jacques Thibodeau et al.
arXiv · 2026-05-06
This paper presents an automated pipeline that compares a base language model against an intervened version to detect and describe behavioral changes caused by modifications such as reasoning distillation, knowledge editing, and unlearning. The system generates human-readable, statistically validated natural-language hypotheses about how the two models differ, including both intended and unexpected side-effects. Evaluated first in synthetic settings where known changes were injected, and then on real-world interventions, the pipeline reliably surfaces behavioral shifts without hallucinating differences when none exist. This provides a practical tool for post-hoc auditing of language model interventions, which is directly relevant to quality assurance and oversight of AI systems.
- Quality assurance
Research
Misaligned by Reward: Socially Undesirable Preferences in LLMs
Gayane Ghazaryan, Esra Dönmez
arXiv · 2026-05-06
This paper evaluates reward models—used to align large language models with human preferences—across four socially consequential domains: bias, safety, morality, and ethical reasoning. Testing five publicly available reward models and two instruction-tuned models used as reward proxies, the authors find that none performs best across all domains, and all fall well short of strong social intelligence, frequently preferring socially undesirable responses and producing systematically biased output distributions. The study also identifies a key trade-off: stronger bias avoidance can reduce sensitivity to context, meaning models may sacrifice contextual faithfulness to avoid biased outcomes. The findings demonstrate that standard reward benchmarks are insufficient for assessing social alignment and call for evaluations that directly measure the social preferences encoded in reward models.
- AI policy
- Quality assurance
Research
The Insurability Frontier of AI Risk: Mapping Threats to Affirmative Coverage, Silent Exposures, and Exclusions
Alex Leung, Rex Zhang, Ervin Ling et al.
arXiv · 2026-05-06
This paper analyzes how commercial insurance markets are responding to AI-related risks by mapping 55 AI threat classes against 26 insurance products, endorsements, and exclusion regimes using public carrier materials and OWASP/MITRE threat catalogs. The authors identify a four-tier 'insurability frontier' comprising affirmatively insured perils, silent-AI exposures under legacy policies (cyber, E&O, D&O, EPLI, crime, and media), actively excluded perils, and perils outside conventional private insurance structures. Key findings include early differentiation among carriers by risk emphasis (e.g., model performance, hallucination liability, IP concerns, autonomous system liability, and deepfake response), the persistence of silent-AI exposure in legacy lines where AI is an instrumentality rather than a legal cause of loss, and the identification of foundation model concentration as a genuinely novel insurability frontier because upstream model failure can correlate losses across many policyholders simultaneously. The paper matters for enterprise risk managers and policymakers who must understand where AI-related losses are covered, where coverage gaps silently exist, and where systemic risk may exceed the capacity of conventional private insurance structures.
- Enterprise
- AI policy
Research
DecodingTrust-Agent Platform (DTap): A Controllable and Interactive Red-Teaming Platform for AI Agents
Zhaorun Chen, Xun Liu, Haibo Tong et al.
arXiv · 2026-05-06
DTap is a red-teaming platform for AI agents that spans 14 real-world domains and over 50 simulation environments replicating systems like Google Workspace, PayPal, and Slack. The authors also introduce DTap-Red, an autonomous red-teaming agent that explores diverse injection vectors—including prompt, tool, skill, and environment-based attacks—to discover effective attack strategies. Using these tools, they build DTap-Bench, a large-scale dataset of attack instances paired with verifiable judges, and use it to evaluate popular AI agents, revealing systematic vulnerability patterns such as agents being manipulated into leaking API keys, deleting user data, or initiating unauthorized transactions. The work highlights serious security gaps in deployed AI agents and provides a reproducible framework for large-scale risk assessment.
- Quality assurance
- AI policy
Research
Beyond Seeing Is Believing: On Crowdsourced Detection of Audiovisual Deepfakes
Michael Soprano, Andrea Cioci, Stefano Mizzaro
arXiv · 2026-05-06
This paper investigates how well crowd workers can detect audiovisual deepfakes, using two matched crowdsourcing studies on the Prolific platform with videos from the AV-Deepfake1M and Trusted Media Challenge datasets. Across 960 judgments on 96 videos, workers rarely misclassified authentic content as fake but frequently missed real manipulations, and agreement was limited. Aggregating multiple judgments improved authenticity screening but could not recover manipulations that most workers consistently overlooked; identifying manipulation type—especially combined audio-video cases—was substantially harder than simply flagging a video as manipulated. The findings suggest crowdsourcing can serve as a scalable but imperfect screening tool for deepfake detection, while accurately attributing which modality was manipulated remains an unsolved challenge.
- Quality assurance
- AI policy
Research
AgentTrust: Runtime Safety Evaluation and Interception for AI Agent Tool Use
Chenglin Yang
arXiv · 2026-05-06
AgentTrust is a runtime safety layer that intercepts AI agent tool calls—such as file operations, shell commands, and database queries—before they execute, returning a structured verdict to allow, warn, block, or review each action. The system combines shell deobfuscation, safer-alternative suggestions (SafeFix), multi-step attack-chain detection (RiskChain), and an LLM-as-Judge for ambiguous cases. On an internal 300-scenario benchmark, the production ruleset achieves 95.0% verdict accuracy at low-millisecond latency, and on a separate 630-scenario adversarial benchmark it reaches 96.7% verdict accuracy including roughly 93% on shell-obfuscated payloads. By addressing gaps left by post-hoc benchmarks, static guardrails, and sandboxes, AgentTrust offers enterprises a practical mechanism to prevent irreversible harms from autonomous AI agents operating in real-world environments.
- Enterprise
- Quality assurance
Research
Paraphrase-Induced Output-Mode Collapse: When LLMs Break Character Under Semantically Equivalent Inputs
Aofan Liu, Jingxiang Meng
arXiv · 2026-05-06
This paper investigates whether large language models maintain their requested output format when a prompt is reworded but its meaning stays the same. The authors introduce PARACONSIST, a 900-prompt benchmark built from 150 base queries with five paraphrase variants each, and find that roughly 78% of closed-form variant responses drift away from the expected answer space into conversational prose — even at temperature zero. A new Semantic Consistency Score tracks robustness across answer consistency, semantic similarity, and length stability, revealing that task structure is a stronger predictor of this 'output-mode collapse' than model identity. The findings argue that response-format preservation should be treated as a core reliability target in LLM evaluation pipelines, not an afterthought.
- Quality assurance
Research
Ask Before You Diagnose: Safe-Psych, a Sequential Evaluation Benchmark for LLMs in Psychiatry
Oriana Presacan, Andreea Grama, Larisa Irimină et al.
arXiv · 2026-05-06
Safe-Psych is a new benchmark of over 1,000 real-world psychiatric clinical notes designed to evaluate how large language models handle evolving diagnostic uncertainty in psychiatry. Unlike existing medical benchmarks that assume complete information upfront, Safe-Psych simulates incremental evidence disclosure and labels each stage with psychiatrist-derived actions: DIAGNOSE, CLARIFY, or ABSTAIN. Key findings show that even strong LLMs are poorly calibrated under incomplete information, with under-abstention rates exceeding 60% for most models, and that models frequently commit to diagnoses prematurely before sufficient evidence is available, resulting in less accurate diagnoses. The work highlights a critical safety gap in current LLMs—their inability to recognize when clinical evidence is insufficient—with implications for the safe deployment of AI in healthcare decision support.
- Quality assurance
- AI policy
Research
Accountable Agents in Software Engineering: An Analysis of Terms of Service and a Research Roadmap
Christoph Treude
arXiv · 2026-05-06
This vision paper analyzes the Terms of Service (ToS) of widely used AI coding assistants and agent-enabled development tools to examine how accountability is allocated between tool providers and software developers. The authors find a consistent pattern of shifting responsibility for correctness, safety, and legal compliance onto users, along with significant variation across providers in areas like indemnification, data reuse, and acceptable use. Because existing policy frameworks are poorly aligned with increasingly autonomous, agent-mediated development workflows, the paper outlines a research roadmap covering responsibility modeling, governance artifact design, accountability tooling, and empirical studies of developer perceptions. This matters because as AI agents become integral to software engineering, the legal and ethical accountability gaps identified here have real consequences for developers and organizations using these tools.
- AI policy
- Enterprise
Research
An Evaluation of Chat Safety Moderations in Roblox
Priya Kaushik, Sonja Brown, Rakibul Hasan et al.
arXiv · 2026-05-06
This paper independently evaluates the effectiveness of Roblox's automated chat moderation system, which protects hundreds of millions of users—many of them minors—on a major online gaming platform. The researchers collected approximately 2 million chat messages across multiple games and age groups, manually labeled ~99,800 messages as safe or unsafe, and used the best-performing large language model among four tested to classify the full corpus. Their findings reveal that significant numbers of unsafe messages related to grooming, sexualizing minors, bullying, harassment, violence, self-harm, and sensitive information sharing evade current moderation, and that previously flagged users continue sending harmful content through active evasion techniques. The study highlights critical gaps in platform safety mechanisms that protect vulnerable underage users from online abuse.
- Quality assurance
- AI policy
Research
Deployment-Relevant Alignment Cannot Be Inferred from Model-Level Evaluation Alone
Varad Vishwarupe, Nigel Shadbolt, Marina Jirotka et al.
arXiv · 2026-05-06
This paper argues that current AI alignment benchmarks—which score model outputs under fixed inputs—cannot reliably predict how well AI systems will behave in real-world deployments. Through a structured audit of up to sixteen alignment benchmarks (coded with Cohen's kappa = 0.87 across eight dimensions), the authors find that user-facing verification support is absent across every benchmark examined, and process steerability is nearly absent. A blinded stress test using 180 transcripts across three frontier models and four scaffolds further reveals that scaffold efficacy is model-dependent: the same verification scaffold brought one model to ceiling performance while leaving another categorically unchanged, demonstrating that deployment-relevant gaps cannot be closed at the model level alone. The authors propose a system-level evaluation agenda including alignment profiles, fixed-scaffolding protocols, and reporting templates that make explicit the inferential distance between evaluation evidence and deployment claims.
- Quality assurance
- AI policy
Research
Worst-Case Discovery and Runtime Protection for RL-Based Network Controllers
Hongyu Hè, Minhao Jin, Maria Apostolaki
arXiv · 2026-05-06
ReGuard is a framework that automatically discovers worst-case operating conditions for reinforcement-learning-based network controllers (such as those used for congestion control and adaptive bitrate streaming) and protects against them at inference time without retraining. By formulating discovery as a bilevel regret-maximization problem, ReGuard produces certified lower bounds on worst-case performance gaps; it then converts discovered failure trajectories into lightweight logic rules that intervene only when a risky state is detected. Evaluated on three real RL controllers (Pensieve, Sage, and Park), ReGuard uncovers scenarios where performance is 43–64% worse than achievable, finds gaps 57% to 6× larger than the strongest baselines, and reduces those gaps by 79–85% while preserving nominal performance. This work matters for quality assurance of deployed AI systems, offering a practical path to identifying and mitigating hidden failure modes in production networking controllers.
- Quality assurance
Research
AI-Integrated Digital Twin Framework for Quality and Performance Management in Sustainable Buildings
Mohammed Abdulrazzaq Alaghbari, Aawag Mohsen Alawag, Yasser El-Kassrawy et al.
arXiv · 2026-05-06
This paper proposes an AI-integrated Semantic Digital Twin (SDT) framework designed to unify Building Information Modeling, IoT, and Building Management Systems for quality management in sustainable buildings. The framework uses ontology-driven semantics, hybrid physics-AI modeling, and a requirements-to-KPI traceability graph to link lifecycle intent, data provenance, analytic outputs, and operator decisions across a seven-layer architecture. An assurance loop validates data, quantifies uncertainty, records explainability artifacts, and enforces human approval before actuation, making the system audit-ready and governance-compliant. The work also defines a benchmark blueprint using BOPTEST and governance metrics such as energy use intensity, comfort violation hours, and fault detection performance, offering a practical foundation for trustworthy and scalable building operations.
- Quality assurance
- Enterprise
- Certifications
- AI policy
Research
Impact of artificial intelligence on cardiovascular workflow, engagement, and outcomes: a systematic review
Yi-En Lin, Shu-Mei Yang, Chi-Jung Huang et al.
npj Digital Medicine · 2026-05-06
This systematic review of 32 randomized controlled trials found that AI applied in cardiology workflows produced measurable benefits across efficiency, engagement, and clinical outcomes. Workflow time was significantly reduced (SMD -0.71), translating to 30–120 seconds shorter diagnostic times and 1.0–4.2 fewer hospital days. AI-driven behavioral nudging improved medication adherence (RR 1.59; NNT=12), and decision-support tools reduced all-cause mortality (RR 0.84; NNT=32). The authors propose a structured governance and implementation framework based on these findings, with implications for how health systems deploy and oversee clinical AI.
- Workforce
- Enterprise
- Quality assurance
- AI policy
Research
Enhancing Enterprise Risk Management and Internal Audit Practices by Applying Machine Learning Models
Reneta Duhova, Angel Duhov, P Georgieva et al.
Risks · 2026-05-06
This study applies machine learning models—specifically XGBoost and multinomial logistic regression—to 69,158 billing records from an SAP production environment to flag potentially risky or fraudulent transactions in support of enterprise risk management and internal audit. XGBoost achieved 99% overall accuracy and a macro F1-score of 0.965, substantially outperforming logistic regression (macro F1 = 0.863), demonstrating that ML can meaningfully reduce detection risk in high-volume transactional environments. The findings suggest that integrating ML-driven pattern recognition into continuous auditing and ERM processes can enable earlier investigation of suspicious activity, minimize fraud and error exposure, and strengthen internal control frameworks within ERP systems like SAP.
- Enterprise
- Quality assurance
- Certifications