News & Research
The latest AI research and news with real-world stakes. Each item is sourced, dated, summarized in plain English and tagged by impact area, and checked against its source before it appears.
- ResearcharXiv2026-04-27QC
How Sensitive Are Safety Benchmarks to Judge Configuration Choices? · Xinran Zhang
This paper investigates how sensitive AI safety benchmarks—specifically HarmBench—are to the configuration of LLM-based judges used to classify model outputs as harmful or safe. Using a factorial design with 12 prompt variants applied to six target models across 400 behaviors, the authors find that prompt wording alone can shift measured harmful-response rates by up to 24.2 percentage points, and even minor surface-level rewording causes swings of up to 20.1 percentage points. Model safety rankings are moderately unstable (mean Kendall tau = 0.89), and category-level sensitivity varies widely—from 39.6 percentage points for copyright to 0 for harassment. The findings reveal that judge prompt wording is a substantial and previously under-examined source of measurement variance, raising serious concerns about the reliability of current safety benchmarking practices.
- ResearcharXiv2026-04-27EQ
AgentPulse: A Continuous Multi-Signal Framework for Evaluating AI Agents in Deployment · Yuxuan Gao, Megan Wang, Yi Ling Yu
AgentPulse is a continuous evaluation framework that scores 50 AI agents across 10 workload categories using 18 real-time signals drawn from GitHub, package registries, IDE marketplaces, social platforms, and benchmark leaderboards. Unlike static benchmarks, it captures deployment-relevant factors including Benchmark Performance, Adoption Signals, Community Sentiment, and Ecosystem Health, which the paper shows are largely complementary rather than redundant. A key finding is that a sub-composite excluding GitHub-derived signals still predicts external adoption proxies such as GitHub stars and Stack Overflow question volume, while composite rankings diverge substantially from SWE-bench-only rankings for 9 of 11 agents with published scores. The work matters because it provides enterprises and developers with a richer, ongoing picture of how agents perform in real-world deployment rather than at a single benchmark snapshot.
- ResearcharXiv2026-04-27Q
Failure-Centered Runtime Evaluation for Deployed Trilingual Public-Space Agents · M. Meng
This paper introduces PSA-Eval, a failure-centered runtime evaluation framework designed for AI agents deployed in real public spaces that handle queries in three languages. Rather than judging a system solely by aggregate scores, PSA-Eval reframes the basic unit of analysis as a 'failure,' extending the standard evaluation pipeline to include failure-case identification, repair, and regression testing. A pilot study on a real trilingual digital front-desk system at an international financial institution (81 samples across 27 trilingual question groups) found that despite a high average score of 23.15/24, 14 of 27 groups showed cross-language score drift, with 5 groups drifting by at least 3 points and one reaching 9 points. These results suggest that aggregate scoring can mask structured inconsistencies in deployed AI systems, making failure-centered evaluation a more informative approach for real-world quality assurance.
- ResearcharXiv2026-04-27Q
Architecture Determines Observability of Transformers · Thomas Carmichael
This paper investigates why activation-based monitors can sometimes catch confident errors in autoregressive transformers that output-confidence monitoring misses. The authors find that the ability of internal activation probes to detect decision-quality signals beyond what the output already exposes is an architectural property fixed at training time — controlling for output confidence removes 60.3% of the raw activation-probe signal on average across 14 models. In controlled experiments using the Pythia model family, some training configurations preserve a readable internal signal through convergence while others erase it even as perplexity improves, showing that capability and observability are not inherently in tension. On downstream QA tasks, a WikiText-trained activation probe with no task-specific tuning catches roughly one in eight confident errors that output-confidence monitoring misses at a 20% flag rate, establishing what the authors call 'signal engineering' as a training-time design axis alongside loss and capability objectives.
- ResearcharXiv2026-04-27QP
An empirical evaluation of the risks of AI model updates using clinical data: stability, arbitrariness, and fairness · Ioannis Bilionis, Ricardo C. Berrios, Luis Fernandez-Luque et al.
This paper investigates risks that arise when AI/ML clinical decision-support models are retrained on new data, using severe hyperglycemia prediction in children with Type 1 Diabetes as a case study across four U.S. datasets (~11,300 weekly observations from 496 participants). The authors show that model updates can cause prediction instability (predictions 'flipping' for many cases), increased arbitrariness, and worsened fairness across demographic subpopulations. They propose a multi-dimensional continuous monitoring framework to detect these problems and argue it is essential for building trustworthy clinical AI systems.
- ResearchInternational Journal of Nursing Education2026-04-27WCP
A Review of Barriers, Facilitators, and Contextual Factors for the Integration of Artificial Intelligence in Nursing Education · Madhuri Meshram, Sathish Rajamani
This scoping review of 42 studies examines what makes it easier or harder for nursing schools to integrate AI into their curricula. Major barriers identified include inadequate technological infrastructure, interoperability challenges, gaps in teacher and student skills, ethical and privacy concerns, and unclear policies. Facilitators include stakeholder engagement, AI-focused training, supportive policy frameworks, and demonstrated benefits for clinical decision-making. The authors conclude that context-sensitive, multi-component strategies are needed to develop standardized AI curricula and guide future policy in nursing education.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-04-27QCP
Lume‑Med: Deterministic AI Governance for Medical Systems Using Lume and Lume‑V · Ronald Jason Andrews
Lume-Med proposes a deterministic governance architecture for AI systems used in high-stakes medical environments, combining invariant-based validation, cryptographically verifiable audit trails, and deterministic explainability into a single reproducible pipeline. The paper introduces a 10-layer medical governance architecture and LTC-Med v1.0, a cryptographically signed trust certificate standard for medical AI. It aligns this framework with major regulatory standards including FDA SaMD, HIPAA, IEC 62304, ISO 14971, and NIST AI RMF, and establishes nine integration patterns for medical AI and robotics. This work is relevant to certification and quality assurance efforts by offering a formal, auditable governance substrate designed to make nondeterministic AI systems compliant and traceable in clinical settings.
- ResearchProcesses2026-04-27WEQ
A Data-Driven Evaluation Framework for Quantifying the Impact of Artificial Intelligence on Industrial Process Performance · Qun Lu, Fengning Yang, Suhang Wang et al.
This study develops a data-driven framework to measure how AI adoption affects industrial process performance and firm value, combining the Feltham–Ohlson valuation model with AHP, Entropy Weight Method, and Fuzzy Comprehensive Evaluation. Using panel data from 3,515 Chinese A-share listed companies (20,076 firm-year observations) over 2014–2022, the authors construct a Process Performance Index (PI) covering resource allocation, coordination, and production dimensions. Results show that higher PI scores are positively linked to abnormal earnings and firm profitability, confirming that AI-enabled operational capability drives sustained enterprise value growth. The framework also reveals that AI investment intensifies digital technology spending, builds knowledge-based human capital, and improves data governance—offering practical guidance for evaluating intelligent transformation strategies in Industry 5.0 contexts.
- ResearcharXiv2026-04-26QP
Risk-Aware Robust Learning: Reducing Clinical Risk under Label Noise in Medical Image Classification · Maycon R. S. Pereira, Filipe R. Cordeiro
This paper examines whether state-of-the-art noise-robust training methods for medical image classification maintain clinical safety when labels are noisy. The authors systematically evaluate four methods—Coteaching, DivideMix, UNICON, and a GMM-based filtering approach—on binarized DermaMNIST and PathMNIST datasets under label noise rates of 20% and 40%, using a cost-sensitive Global Risk formulation that heavily penalizes false negatives (missed diagnoses). Their key finding is that robustness to label noise does not guarantee clinical safety, but integrating cost-sensitive optimization into noise-robust training significantly reduces clinical risk while preserving model utility. The work argues that noise-robust learning in medical imaging must be evaluated through a clinical risk lens, not just accuracy-oriented metrics.
- ResearcharXiv2026-04-26QP
Time-Series Forecasting in Safety-Critical Environments: An EU-AI-Act-Compliant Open-Source Package / Zeitreihenprognose in sicherheitskritischen Umgebungen: Ein KI-VO-konformes Open-Source-Paket · Thomas Bartz-Beielstein, Eva Bartz
spotforecast2-safe is an open-source Python library that embeds EU AI Act (Regulation (EU) 2024/1689), IEC 61508, ISA/IEC 62443, and Cyber Resilience Act compliance directly into the library itself rather than relying on external scanners or runtime layers. The package enforces four non-negotiable code-development rules—zero dead code, deterministic processing, fail-safe handling, and minimal dependencies—alongside process rules such as model cards, CI workflows, and REUSE-conformant licensing to operationalize a Compliance-by-Design approach for time-series point forecasting in safety-critical environments. A bidirectional traceability matrix maps each regulatory provision to a concrete mechanism in the code, and an end-to-end electricity generation, transmission, and consumption forecasting example demonstrates the approach. The work matters for both policy compliance and quality assurance, showing how safety-critical AI deployments can satisfy multiple regulatory frameworks by design rather than as an afterthought.
- ResearcharXiv2026-04-26WE
Learning Selective LLM Autonomy from Copilot Feedback in Enterprise Customer Support Workflows · Nikita Borovkov, Elisei Rykov, Olga Tsymboi et al.
This paper presents a deployed AI system that automates enterprise customer support workflows within a Business Process Management (BPM) platform. The system trains a next-UI-action policy from structured interaction traces and learns a critic model from copilot feedback—where operators accept or correct suggestions—to calibrate when the system should act autonomously versus defer to a human. In production, it automated 45% of sessions and reduced average handling time by 39% without degrading support quality, reaching selective automation within two weeks for a new process. The approach allows a single operator to supervise multiple concurrent sessions by only intervening when the system is uncertain.
- ResearcharXiv2026-04-26QP
Does Machine Unlearning Preserve Clinical Safety? A Risk Analysis for Medical Image Classification · Andreza M. C. Falcao, Filipe R. Cordeiro
This paper investigates whether machine unlearning methods—which selectively remove training data from deployed deep learning models—preserve clinical safety in binary medical image classification. The authors find that standard unlearning strategies (Fine-Tuning, Random Labeling, and SalUn) can reduce test utility and increase false-negative rates, thereby amplifying clinical risk. To address this, they propose SalUn-CRA (Clinical Risk-Aware), which uses entropy-based forgetting for malignant samples instead of random relabeling, and show it achieves lower or comparable clinical risk to full retraining on DermaMNIST and PathMNIST datasets under 20% and 50% data removal scenarios. The findings argue that clinically asymmetric error costs should be an integral part of unlearning validation in medical AI systems.
- ResearcharXiv2026-04-26QP
One Size Fits None: Heuristic Collapse in LLM Investment Advice · Jillian Ross, Andrew W. Lo
This paper investigates whether large language models (LLMs) provide genuinely individualized investment advice or instead exhibit 'heuristic collapse'—reducing complex, multi-factor decisions to a small number of dominant inputs. Using interpretable surrogate models to analyze LLM outputs, the authors find that investment allocation recommendations are largely driven by self-reported risk tolerance, with other legally required factors contributing minimally. Web search augmentation partially reduces but does not resolve this collapse, and the problem persists across model scales. The authors argue that deploying LLMs as advisors requires auditing input sensitivity, not just output quality—a finding with direct implications for financial advisory compliance and AI governance.
- ResearcharXiv2026-04-26QP
The Interlocutor Effect: Why LLMs Leak More Personal Data to Agents Than Humans · Faouzi El Yagoubi, Godwin Badu-Marfo, Ranwa Al Mallah
This paper identifies the 'Interlocutor Effect,' a phenomenon where Large Language Models leak significantly more Personally Identifiable Information (PII) when they believe they are communicating with another AI agent rather than a human user. Across 3,464 interactions spanning 222 sensitive scenarios, the researchers find that framing a recipient as an AI agent increases PII leakage by up to 23 percentage points. The authors propose the 'Attention Suppression Hypothesis' to explain this, suggesting that safety-aligned attention heads become inactive during agent-directed interactions, and experiments on Llama-3.1-8B-Instruct support this by showing that deactivating one safety head induces leakage while reactivating it restores privacy protections. These findings carry significant implications for the security and trustworthiness of multi-agent AI systems.
- ResearcharXiv2026-04-26QP
Benchmarking Web Agent Safety under E-commerce Deceptive Interfaces · Zijing Shi, Meng Fang, Ling Chen
This paper introduces WebDecept, a configurable plugin framework that injects realistic deceptive interface patterns—such as targeted advertisements, domain redirection, and shopping manipulation—into e-commerce web environments to evaluate the safety of autonomous web agents. Testing multiple multimodal web agents against seven deceptive patterns, the study finds that current agents are highly susceptible to these manipulations and that prompt-based constraints alone are often insufficient to prevent failures. The findings underscore significant safety gaps that must be addressed before web agents can be reliably deployed in real-world settings.
- ResearcharXiv2026-04-26QP
FAIR_XAI: Improving Multimodal Foundation Model Fairness via Explainability for Wellbeing Assessment · Sophie Chiang, Tom Brennan, Fethiye Irmak Dogan et al.
This paper investigates how Vision-Language Models (VLMs) perform on mental health and depression assessment tasks, specifically examining diagnostic accuracy and demographic fairness across laboratory (AFAR-BSFT) and naturalistic (E-DAIC) datasets. The authors find substantial variation in model performance—Phi3.5-Vision reached 80.4% accuracy on E-DAIC while Qwen2-VL achieved only 33.9%—and document meaningful bias patterns, with Qwen2-VL showing higher gender disparities and Phi-3.5-Vision exhibiting more racial bias. An Explainable AI (XAI) intervention framework yielded mixed results: fairness prompting achieved perfect equal opportunity for one model but at a severe accuracy cost, and some interventions amplified racial bias rather than reducing it. The findings highlight a persistent gap between procedural transparency and equitable outcomes, and the authors offer recommendations for future work that jointly optimises accuracy, demographic parity, and cross-domain generalisation.
- ResearcharXiv2026-04-26QP
Personality Shapes Gender Bias in Persona-Conditioned LLM Narratives Across English and Hindi: An Empirical Investigation · Tanay Kumar, Shreya Gautam, Aman Chadha et al.
This study examines how personality traits assigned to LLM personas affect gender bias in AI-generated stories, using a controlled experiment spanning English and Hindi across six state-of-the-art LLMs. The researchers generated 23,400 stories featuring working professionals in India with systematically varied persona gender, occupational role, and personality traits drawn from the HEXACO and Dark Triad frameworks. They find that Dark Triad personality traits are consistently associated with higher gender-stereotypical representations compared to socially desirable HEXACO traits, and that these associations vary across models and languages. The results indicate that gender bias in LLMs is context-dependent rather than static, meaning persona-conditioned systems deployed in education, customer service, or social platforms may reinforce gender stereotypes unevenly across generated content.
- ResearcharXiv2026-04-26QP
When AI reviews science: Can we trust the referee? · Jialiang Wang, Yuchen Liu, Hang Xu et al.
This paper investigates whether large language models (LLMs) can be trusted as peer reviewers in scientific publishing, motivated by the growing gap between submission volume and available human referees. The authors develop a taxonomy of attacks across the AI review lifecycle—covering training, desk review, deep review, rebuttal, and system-level stages—and test four adversarial probes on ICLR 2025 submissions using two LLM-based referees. Their experiments demonstrate that prestige framing, assertion strength, rebuttal sycophancy, and contextual poisoning (including prompt injection) can systematically skew AI-generated review scores, exposing concrete reliability and security failures. The findings provide an evidence-based baseline for evaluating AI peer review integrity and identifying targets for mitigation.
- ResearcharXiv2026-04-26QP
FinGround: Detecting and Grounding Financial Hallucinations via Atomic Claim Verification · Dongxin Guo, Jikun Wu, Siu Ming Yiu
FinGround is a three-stage pipeline designed to detect and correct hallucinations in financial AI systems by decomposing answers into atomic claims, verifying them using finance-specific strategies (including arithmetic re-verification against structured tables), and rewriting unsupported claims with precise citations to source documents. The abstract reports that existing hallucination detectors miss 43% of computational errors, while FinGround reduces hallucination rates by 68% over the strongest baseline under controlled retrieval conditions, and by 78% relative to GPT-4o across the full pipeline. A distilled 8B model retains 91.4% F1 at 18x lower latency, enabling low-cost deployment at $0.003 per query. The work is directly motivated by regulatory risk, citing the EU AI Act's high-risk enforcement deadline of August 2026 as a driver for reliable, grounded financial AI outputs.
- ResearcharXiv2026-04-26WEP
ComplianceNLP: Knowledge-Graph-Augmented RAG for Multi-Framework Regulatory Gap Detection · Dongxin Guo, Jikun Wu, Siu Ming Yiu
ComplianceNLP is an end-to-end NLP system designed to automate regulatory monitoring, obligation extraction, and compliance gap detection across major frameworks including SEC, MiFID II, and Basel III. The system combines a knowledge-graph-augmented retrieval-augmented generation (RAG) pipeline, multi-task legal information extraction using LEGAL-BERT, and severity-aware gap scoring against institutional policies, achieving 87.7 F1 on gap detection and outperforming GPT-4o+RAG by 3.5 F1 points. In a four-month parallel deployment at a financial institution processing 9,847 regulatory updates, the system reached 96.0% estimated recall, 90.7% precision, and a 3.1× analyst efficiency gain. The work directly addresses the challenge of tracking over 60,000 regulatory events annually and the industry's USD 300 billion in post-2008 fines, with implications for both compliance workforce productivity and enterprise regulatory risk management.
- ResearcharXiv2026-04-26EQ
AgentEval: DAG-Structured Step-Level Evaluation for Agentic Workflows with Error Propagation Tracking · Dongxin Guo, Jikun Wu, Siu Ming Yiu
AgentEval is a framework that evaluates multi-step AI agent workflows by modeling executions as directed acyclic graphs (DAGs), where each node is assessed by a calibrated LLM judge and linked to upstream dependencies for automated root cause attribution. The paper reports that DAG-based dependency modeling alone adds +22 percentage points to failure detection recall and +34 pp to root cause accuracy over flat step-level evaluation, while the full system achieves 2.17x higher failure detection recall than end-to-end evaluation (0.89 vs. 0.41) with strong expert agreement (Cohen's kappa = 0.84). In a 4-month pilot with 18 engineers, the framework detected 23 pre-release regressions via CI/CD integration and reduced median root-cause identification time from 4.2 hours to 22 minutes. These results suggest structured intermediate-step evaluation can meaningfully improve reliability and debugging efficiency for production agentic systems.
- ResearcharXiv2026-04-26EQ
RouteNLP: Closed-Loop LLM Routing with Conformal Cascading and Distillation Co-Optimization · Dongxin Guo, Jikun Wu, Siu Ming Yiu
RouteNLP is a closed-loop framework that routes NLP queries across a tiered portfolio of large and small language models to cut inference costs while meeting per-task quality constraints. It combines a difficulty-aware router, conformal-prediction-based cascading for threshold setting, and a distillation-routing co-optimization loop that clusters escalation failures and applies targeted knowledge distillation to cheaper models. In an 8-week enterprise pilot processing roughly 5,000 queries per day, the system reduced inference costs by 58%, cut p99 latency from 1,847 ms to 387 ms, and maintained 91% response acceptance; on a six-task benchmark it achieved 40–85% cost reduction while retaining 96–100% quality on structured tasks. The paper is directly relevant to enterprises facing high LLM inference costs, demonstrating that intelligent routing and distillation co-optimization can dramatically reduce spend without sacrificing output quality.
- ResearcharXiv2026-04-26EP
Do Transaction-Level and Actor-Level AML Queues Agree? An Empirical Evaluation of Granularity Effects on the Elliptic++ Graph · Ankur Malik
This paper investigates whether anti-money laundering (AML) systems that score suspicious blockchain activity at the transaction level versus the actor (address) level produce meaningfully different investigation queues when operating under fixed review budgets. Using the public Elliptic++ Bitcoin dataset (203,769 transactions; 822,942 address occurrences), the authors train independent random forest classifiers at each granularity level and compare resulting review queues via Jaccard overlap, yield metrics, and burden decomposition. They find very low agreement between the two queue types—temporal evaluation yields a mean Jaccard of only 0.374, and static evaluation just 0.087—meaning the same data and budget lead to substantially different sets of addresses being investigated depending on scoring granularity. The findings establish that granularity is a consequential design choice for AML compliance systems, with practical implications for how financial crime detection pipelines should be structured and evaluated.
- ResearcharXiv2026-04-26WQ
Your Students Don't Use LLMs Like You Wish They Did · Sebastian Kobler, Matthew Clemson, Angela Sun et al.
This paper introduces six computational metrics for evaluating whether student-AI conversations in educational settings actually achieve pedagogical goals, validated on 12,650 messages across 500 conversations from four courses. The study finds a fundamental misalignment: educators build conversational tutors to foster sustained learning dialogue, but students predominantly use them to extract answers, including copying verbatim assignment questions. Deployment context—whether tools are optional or course-integrated—is the strongest predictor of how students use AI, outweighing system design or student preference. These metrics give researchers and educators a practical way to measure pedagogical alignment beyond engagement or satisfaction proxies.
- ResearcharXiv2026-04-26Q
Agentic Adversarial Rewriting Exposes Architectural Vulnerabilities in Black-Box NLP Pipelines · Mazal Bethany, Kim-Kwang Raymond Choo, Nishant Vishwamitra et al.
This paper introduces a two-agent adversarial framework that tests the robustness of multi-component NLP pipelines under strict real-world conditions: binary-only feedback, no gradient access, and a 10-query budget. Evaluated against four misinformation detection pipelines, the framework achieves evasion rates of 19.95–40.34% on modern LLM-based systems, far exceeding the ≤3.90% rates of token-level baselines, while a legacy lexical retrieval system shows near-total vulnerability at 97.02%. The work identifies three architectural properties—evidence retrieval mechanism, retrieval-inference coupling, and baseline classification accuracy—that govern how susceptible a pipeline is to attack, and proposes a pattern-informed defense that reduces evasion rates by up to 65.18%. These findings matter for quality assurance of deployed NLP systems, revealing how architectural design choices directly determine the attack surface in high-stakes decision pipelines.