News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
Navigating the Sea of LLM Evaluation: Investigating Bias in Toxicity Benchmarks
Regina Gugg, Selina Niederländer, Andreas Stöckl et al.
arXiv · 2026-05-11
This paper investigates hidden biases in toxicity benchmarks used to evaluate the safety of large language models (LLMs). The researchers find that changing evaluation conditions—such as shifting tasks from text completion to summarization, or switching input data domains—causes significant inconsistencies in benchmark results, including increased rates of flagging content as harmful. They also identify model-specific instabilities that undermine the reliability of current safety assessments. These findings matter because organizations rely on these benchmarks to certify LLMs for customer-facing and automated moderation applications, meaning undetected evaluation biases could lead to the deployment of unsafe or poorly calibrated systems.
- Certifications
- Quality assurance
Research
The Open-Box Fallacy: Why AI Deployment Needs a Calibrated Verification Regime
Phongsakon Mark Konrad, Tim Lukas Adam, Ane Cathrine Holst Merrild et al.
arXiv · 2026-05-11
This paper argues that AI deployment authorization in high-stakes domains—healthcare, credit, employment, criminal justice—should not hinge on mechanistic interpretability of model internals, but instead on 'calibrated verification': a structured regime that is domain-scoped, independently checkable, monitored post-release, accountable, contestable, and revocable. The authors draw an analogy to how societies govern opaque human expertise through credentials, liability, and appeal rather than demanding full mechanistic explanation. They cite a 53-percentage-point gap between internal representations and output correction as evidence that interpretability does not reliably translate into behavioral control, and note that only 9.0% of FDA-approved AI/ML device documents included a prospective post-market surveillance study, highlighting current regulatory gaps. To address this, they propose 'Verification Coverage,' a six-component reportable standard intended to sit alongside capability scores in model cards, leaderboards, and regulatory disclosures.
- AI policy
- Certifications
Research
Acceptance Cards:A Four-Diagnostic Standard for Safe Fine-Tuning Defense Claims
Phongsakon Mark Konrad, Toygar Tanyel, Serkan Ayvaz
arXiv · 2026-05-11
This paper introduces 'Acceptance Cards,' a four-part evaluation protocol for assessing whether AI safety fine-tuning defenses genuinely reduce harmful outputs or merely appear to do so due to noise, capability loss, or non-transferable mechanisms. The four diagnostics check statistical reliability, semantic generalization, mechanism alignment, and cross-task transfer before crediting a defense with a valid 'gap reduction.' Applying this protocol to SafeLoRA on Gemma-2-2B-it, the authors find that no cell in a 46-cell audit achieves a full-card pass under strict criteria, and the best-performing approach still fails on fresh-subject generalization, transfer, and incurs a measurable deployment-accuracy cost. The work matters because it establishes a rigorous, reproducible evidential standard for safety claims in fine-tuning research, reducing the risk that weak or misleading defenses are adopted in practice.
- Quality assurance
- Certifications
Research
Agent-First Tool API: A Semantic Interface Paradigm for Enterprise AI Agent Systems
Kai Pan
arXiv · 2026-05-11
This paper identifies five structural mismatches between conventional CRUD-based APIs and the needs of autonomous AI agents in enterprise settings, then proposes the 'Agent-First Tool API' paradigm to address them. The paradigm includes a Six-Verb Semantic Protocol (search, resolve, preview, execute, verify, recover), a Normalized Tool Contract with confidence scores and suggested next actions, and a dual-layer governance pipeline. Validated on a production multi-tenant SaaS platform with 85 tools across 6 business domains, the approach achieves an 88% task success rate versus 64% for optimized CRUD baselines—a 37.5% improvement—while reducing required human interventions by 72.7% and improving autonomous error recovery by 5.8x. These results matter for enterprises deploying AI agents at scale, as better-designed tool interfaces can dramatically reduce operational friction and human oversight burden.
- Enterprise
- Quality assurance
Research
DuetFair: Coupling Inter- and Intra-Subgroup Robustness for Fair Medical Image Segmentation
Yiqi Tian, Sangjoon Park, Bo Zeng et al.
arXiv · 2026-05-11
DuetFair introduces a dual-axis fairness framework for medical image segmentation that addresses both inter-subgroup disparities and a newly identified problem called 'intra-group hidden failure,' where high-loss cases within a subgroup are masked by the subgroup average. The proposed method, FairDRO, combines distribution-aware mixture-of-experts with subgroup-conditioned distributionally robust optimization to simultaneously adapt across subgroups and reduce hidden failures within them. Evaluated on three medical imaging benchmarks, FairDRO achieves the best equity-scaled performance on Harvard-FairSeg, improves worst-case subgroup performance on HAM10000, and improves worst-group Dice by up to 4.1 points (7.4%) on a 3D radiotherapy cohort. This work matters for quality assurance in AI-assisted medical imaging, where uneven model performance across patient subgroups can have serious clinical consequences.
- Quality assurance
Research
Consistency as a Testable Property: Statistical Methods to Evaluate AI Agent Reliability
Harsh Raj, Niranjan Orkat, Suvrorup Mukherjee et al.
arXiv · 2026-05-11
This paper introduces a statistical framework for measuring AI agent reliability by quantifying how consistently agents behave under semantically equivalent task variations. It uses U-statistics for output-level reliability and kernel-based metrics for trajectory-level stability, validated across three agentic benchmarks. A key finding is that minor task-level variations can cause complete strategy breakdowns even when an agent possesses the required knowledge, and that trajectory-level consistency metrics offer far greater diagnostic sensitivity than traditional pass@1 rates. The framework helps identify architectural weaknesses that impede deployment in high-stakes settings.
- Quality assurance
- Certifications
Research
StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs
Pierre Le Jeune, Étienne Duchesne, Weixuan Xiao et al.
arXiv · 2026-05-11
StereoTales introduces a multilingual dataset and evaluation pipeline for studying social bias in open-ended text generation by large language models (LLMs), covering 10 languages, 79 socio-demographic attributes, and over 650,000 stories generated by 23 LLMs. The study finds that every evaluated model produces harmful stereotypes regardless of size or capability, and that these biases are broadly shared across providers rather than being isolated incidents. Prompt language significantly shapes which stereotypes emerge, with harmful associations adapting culturally and amplifying bias against locally salient protected groups. Human and LLM harmfulness judgments show broad alignment (Spearman ρ=0.62), and the authors release all code, data, and annotations to support further research.
- AI policy
- Quality assurance
Research
TourMart: A Parametric Audit Instrument for Commission Steering in LLM Travel Agents
Yao Liu
arXiv · 2026-05-11
TourMart is a proposed audit instrument designed to detect and measure commission-driven steering in LLM-based online travel agents (OTAs) such as Booking, Trip.com, and Expedia. The paper introduces two governance parameters—lambda and kappa—that quantify how much a commission-aware prompt shifts a traveler's perceived recommendation quality compared to a neutral factual template, using a paired counterfactual design. Empirical tests show that at standard deployment settings, a Qwen-14B reader model exhibits +7.69 percentage points of commission steering (McNemar p=0.003), and a Llama-3.1-8B reader shows +2.96–3.50 pp in the same direction, with results surviving family-wise correction. The instrument produces a quotable compliance output—e.g., '7.7 extra commission-steered recommendations per 100 paired traveler sessions'—addressing a measurable accountability gap that existing disclosure banners and generic LLM safety scores do not cover.
- AI policy
- Quality assurance
Research
Toward an Engineering of Science: Rebalancing Generation and Verification in the Age of AI
Jiaqi W. Ma
arXiv · 2026-05-11
This paper argues that AI systems can now generate scientific artifacts—papers, reviews, and surveys—far more cheaply than verification systems can evaluate them, creating structural 'epistemic pollution' risk. The authors frame this as an engineering problem: the traditional cost of producing a plausible paper once served as a quality filter, but AI removes that filter without reducing verification costs. To address this imbalance, they propose 'blueprints,' structured research artifacts that represent claims, evidence, assumptions, and definitions as typed graph components, designed to shift cost toward generation and make downstream verification cheaper and more distributed. A proof-of-concept prototype is presented to demonstrate the approach.
- Quality assurance
- AI policy
Research
GuardAD: Safeguarding Autonomous Driving MLLMs via Markovian Safety Logic
Tianyuan Zhang, Peng Yue, Zihao Peng et al.
arXiv · 2026-05-11
GuardAD is a model-agnostic safety layer for multimodal large language models (MLLMs) used in autonomous driving systems. It formulates driving safety as an evolving Markovian logical state, using Neuro-Symbolic Logic Formalization to track safety predicates across heterogeneous traffic participants over time, enabling detection of emerging and latent hazards beyond single-step observations. Rather than blocking unsafe actions outright, it performs Logic-Driven Action Revision to refine decisions without altering the underlying MLLM. Experiments on multiple benchmarks show GuardAD reduces accident rates by 32.07% while slightly improving task performance by 6.85%, with results further validated in closed-loop simulation and physical-world vehicle tests.
- Quality assurance
- Certifications
Research
Generative AI Fuels Solo Entrepreneurship, but Teams Still Lead at the Top
Hyunso Kim, Hyo Kang, Jaeyong Song
arXiv · 2026-05-11
Analyzing over 160,000 product launches on Product Hunt, this study finds that the public release of ChatGPT-3.5 significantly increased entrepreneurial entry, with the growth driven disproportionately by solo entrepreneurs—even in categories that historically favored teams. However, this surge largely reflects low-commitment, experimental activity: solo entrepreneurs remain underrepresented among the highest-quality outcomes, while team-based ventures have become increasingly dominant at the top tiers of platform rankings. The findings suggest that generative AI lowers barriers to entry for solo founders but does not eliminate the structural advantages that teams hold in producing top-quality entrepreneurial outcomes.
- Workforce
- Enterprise
Research
Knowledge Poisoning Attacks on Medical Multi-Modal Retrieval-Augmented Generation
Peiru Yang, Haoran Zheng, Tong Ju et al.
arXiv · 2026-05-11
This paper introduces M³Att, an attack framework that poisons the retrieval databases used in medical multimodal Retrieval-Augmented Generation (RAG) systems without requiring prior knowledge of specific user queries. The approach injects covert misinformation into text while using subtle, imperceptible perturbations to paired visual data as query-agnostic retrieval triggers, and exploits the inherent ambiguity of medical diagnosis to evade the self-correction capabilities of large language models. Experiments across five LLMs and datasets show the method consistently causes clinically plausible but incorrect outputs, highlighting a serious reliability and safety risk for AI-assisted medical diagnosis systems. The findings have direct implications for quality assurance and policy around the deployment of RAG-based medical AI.
- Quality assurance
- AI policy
Research
Social Policy of Large Language Models: How GPT, Claude, DeepSeek and Grok Allocate Social Budgets in Spain and Germany
Claudia Benavides Cantos, Eduardo C. Garrido-Merchán
arXiv · 2026-05-11
This study examines how four major large language models—Claude, GPT-4o, DeepSeek, and Grok—allocate hypothetical national social budgets across twelve public expenditure categories for Spain and Germany, comparing results against OECD reference benchmarks. Across 48 independent trials, all four models showed a systematic implicit social policy that diverges significantly from real European spending: pensions were under-allocated by roughly a factor of three, while housing and employment were over-allocated by factors of four and two respectively. Contrary to expectations, the main difference between models was not geopolitical bias but rather how concentrated or dispersed their allocations were, with only Claude showing meaningful sensitivity to national context. The findings suggest LLMs can support but should not replace expert deliberation in public budgeting decisions.
- AI policy
Research
To Redact, or not to Redact? A Local LLM Approach to Deliberative Process Privilege Classification
Maik Larooij, David Graus
arXiv · 2026-05-11
This paper investigates using small, locally deployable Large Language Models (LLMs) to automatically classify whether sentences in government documents qualify for redaction under FOIA Exemption 5's deliberative process privilege. The authors test eight prompting variants on a 9-billion-parameter model (Qwen3.5 9B) and find that combining Chain-of-Thought prompting with few-shot prompting using error-based examples outperforms prior classification models in recall and F2 score, closely matching the commercial Gemini 2.5 Flash model. Because processing uncleared documents via third-party cloud APIs is often legally or politically untenable, the local deployment approach offers a practical path for governments to automate sensitivity review without compromising data sovereignty. The work also identifies linguistic markers of deliberativeness—particularly first-person phrasing combined with opinion-expressing verbs—providing interpretable signals to support redaction decisions.
- AI policy
- Quality assurance
Research
LegalCiteBench: Evaluating Citation Reliability in Legal Language Models
Sijia Chen, Hang Yin, Shunfan Zhou
arXiv · 2026-05-11
LegalCiteBench is a new benchmark designed to evaluate how reliably large language models (LLMs) can handle legal citations in a closed-book setting, without access to external retrieval systems. Constructed from 1,000 real U.S. judicial opinions and roughly 24,000 evaluation instances, the benchmark covers five tasks including citation retrieval, error detection, and case verification. Results across 21 LLMs are stark: even the best-performing models score below 7/100 on citation retrieval and completion, and Misleading Answer Rates exceed 94% for 20 of 21 models on retrieval-heavy tasks — meaning models frequently return plausible but incorrect or fabricated case authorities. These findings are directly relevant to legal professionals relying on AI-assisted drafting and research, highlighting a serious reliability gap that neither model scale nor legal-domain pretraining resolves.
- Quality assurance
- Enterprise
Research
Scaling Vision Models Does Not Consistently Improve Localisation-Based Explanation Quality
Mateusz Cedro, Marcin Chlebus
arXiv · 2026-05-11
This study tests whether scaling up computer vision models (ResNet, DenseNet, and Vision Transformer families, totaling 11 models) leads to better post-hoc explanation quality, measured by how well saliency maps align with ground-truth segmentation masks. Using five explainable AI methods and two localisation metrics—including a newly proposed Dual-Polarity Precision—across three image datasets, the authors find that increasing architectural depth and parameter count does not consistently improve explanation quality, and smaller models often match or outperform larger ones. The paper also highlights cases where high predictive accuracy coexists with near-zero localisation precision, meaning strong performance metrics do not guarantee that a model's predictions are grounded in the correct image regions. The authors conclude that explainability must be assessed explicitly during model selection, especially for safety-sensitive applications.
- Quality assurance
- Certifications
Research
Useful for Exploration, Risky for Precision: Evaluating AI Tools in Academic Research
Anthea Dathe, Kiran Hoffmann, Aline Mangold
arXiv · 2026-05-11
This paper proposes and applies a benchmarking framework that combines human-centered and computer-centered metrics to evaluate AI-based question-answering and literature review tools used in academic research workflows. The findings show that Q&A tools provide useful overviews but are unreliable for precise information extraction, with explainable AI (xAI) features frequently failing to link highlighted source passages to generated answers, pushing validation burden back onto researchers. Literature review tools supported exploratory searches but exhibited low reproducibility, limited transparency about sources and databases, and inconsistent source quality, making them unsuitable for systematic reviews. The study concludes that while AI tools can boost efficiency in early-stage and shallow research tasks, human verification remains essential, and better explainability features are needed for safe integration into research workflows.
- Quality assurance
- Workforce
Research
Speech-based Psychological Crisis Assessment using LLMs
Terumi Chiba, Yang Luo, Ziyun Cui et al.
arXiv · 2026-05-11
This paper presents an LLM-based framework for automatically classifying psychological crisis levels in speech recordings from mental health support hotlines. The system introduces a paralinguistic injection method that adds non-verbal emotional cues from speech into text transcripts, allowing the LLM to reason over acoustic signals alongside words. A reasoning-enhanced training strategy generates diagnostic reasoning chains as an auxiliary task, acting as a regularizer to improve classification. Combined with data augmentation, the system achieves a macro F1-score of 0.802 and accuracy of 0.805 on a three-class crisis classification task, offering a potential tool to reduce variability in human operator assessments and address staffing constraints.
- Quality assurance
- Workforce
Research
Medical Incident Causal Factors and Preventive Measures Generation Using Tag-based Example Selection in Few-shot Learning
Yuna Haseyama, Tomoki Ito, Hiroki Sakaji et al.
arXiv · 2026-05-11
This study proposes a tag-based few-shot example selection method to prompt large language models (LLMs) to generate causal factors and preventive measures from medical incident reports. Using the Japanese Medical Incident Dataset (JMID) of 3,884 real-world medical accident and near-miss reports, the researchers compare random sampling, cosine similarity-based selection, and their tag-based approach with GPT-4o and LLaMA 3.3. Results show the tag-based method achieves the highest precision and most stable generation behavior, while similarity-based selection frequently triggers unintended outputs and safety filter activation. The findings suggest that human-interpretable tags can improve reliability of LLM-generated clinical insights, which matters for quality assurance in high-stakes healthcare settings.
- Quality assurance
Research
Position: Academic Conferences are Potentially Facing Denominator Gaming Caused by Fully Automated Scientific Agents
Rong Shan, Te Gao, Hang Zheng et al.
arXiv · 2026-05-11
This position paper identifies a structural vulnerability in top AI academic conferences called 'Agentic Denominator Gaming,' where a malicious actor could deploy AI agents to flood conferences with large volumes of low-quality submissions—not to get those papers accepted, but to inflate the submission pool and thereby increase the acceptance probability of a targeted set of legitimate papers under stable acceptance-rate policies. The authors analyze the practical feasibility of this threat and its downstream consequences, including reviewer burnout and degraded review quality. They argue that addressing this threat requires system-level policy and incentive reforms, not just technical detection methods.
- AI policy
- Quality assurance
Research
Fairness of Explanations in Artificial Intelligence (AI): A Unifying Framework, Axioms, and Future Direction toward Responsible AI
Gideon Popoola, John Sheppard
arXiv · 2026-05-11
This survey paper identifies a critical gap at the intersection of algorithmic fairness and explainable AI (XAI): a model may satisfy standard fairness criteria in its outputs while remaining deeply unfair in its reasoning process, which the authors term 'procedural bias.' The authors introduce a conditional invariance framework that formalizes explanation fairness as a single unifying principle, from which all existing explanation fairness metrics can be derived. They also develop a seven-dimensional taxonomy, identify three generative mechanisms of explanation inequity, and propose a six-step evaluation workflow for conducting explanation fairness audits in practice. This work matters because high-stakes decisions in criminal justice, healthcare, credit, and employment increasingly rely on ML models, and ensuring fairness only in outputs — without scrutinizing the reasoning — may leave significant inequities unaddressed.
- AI policy
- Quality assurance
Research
TRACE-TEST: An Evidence-Linked Software Quality Assurance Framework for AI-Augmented Testing in Regulated Platforms
Rajeew Vishvakarma
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-11
TRACE-TEST is a software quality assurance framework designed to make AI-augmented testing practices viable in regulated industries such as banking, healthcare, insurance, and public services. The paper combines an integrative review of empirical testing studies with a design-science artifact proposal, identifying five key research gaps and synthesizing evidence on coverage, mutation effectiveness, traceability, and human oversight. The framework links regulations, requirements, risk classification, AI-generated artifacts, reviewer actions, execution evidence, and release sign-off into a single quality-assurance chain, specifying a minimum evidence object and a governed release-assurance workflow. Its core contribution is making AI-assisted testing more reviewable, measurable, and aligned with software quality standards in high-assurance environments, rather than simply optimizing for speed or coverage alone.
- Quality assurance
- Certifications
- AI policy
- Enterprise
Research
Gold-Standard AGI: Outer AGI Superalignment
Aaron Turner
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-11
This paper proposes a foundational theory of AGI alignment focused on 'outer alignment'—defining a correct final goal for a superintelligent AGI system—and introduces the concept of 'Gold-Standard AGI' that is both maximally aligned and maximally validated. The authors argue that their definitions of practical-maximal-alignment and practical-maximal-validation could form the basis of an international standard for AGI certification, enabling formal certification by competent authorities so that only certified systems could be lawfully deployed. Written in an accessible, pedagogic style aimed at policymakers and technical readers alike, the paper frames AGI alignment as a governance challenge with civilization-scale stakes. The work has direct implications for how AI certification and regulatory frameworks might be designed at an international level.
- AI policy
- Certifications
- Quality assurance
Research
A kind of magic? Artificial intelligence-assisted recording: developing principles for an AI-informed social work curriculum
Richard; id_orcid 0000-0002-6367-4024 Ingram
Social Work Education · 2026-05-11
This paper examines the workforce and educational implications of deploying Magic Note, an AI-assisted recording tool, across a Scottish Local Authority social work department. Drawing on surveys of 152 practitioners, focus groups, and over 255 free-text responses, the study identifies five core tensions—such as administrative relief vs. reflective writing and efficiency vs. relational practice—that emerge when social workers encounter AI in practice. The authors translate these empirical tensions into five curriculum principles, arguing that social work education must proactively prepare students to engage critically and ethically with AI tools rather than simply react to their adoption. The paper's central message is that students should enter practice already equipped to act as agents in AI-mediated environments rather than passive subjects of those systems.
- Workforce
- AI policy
- Certifications
Research
HUMAN RESOURCE MANAGEMENT CHALLENGES IN THE CONTEXT OF THE DEVELOPMENT OF ARTIFICIAL INTELLIGENCE SYSTEMS. A CONTENT ANALYSIS
Alic Bîrcă, Christiana Brigitte Sandu, GENU ALEXANDRU CĂRUNTU
arXiv · 2026-05-11
This content analysis reviews 42 scientific papers indexed in Web of Science to examine how AI systems are reshaping human resource management (HRM). The study finds that AI is being applied across traditional HRM functions—including recruitment, training, talent management, and performance management—as well as less obvious areas like job design and employee engagement. The authors conclude that AI integration is driving a fundamental reconfiguration of the HR function and requiring HR specialists to update their professional competencies. A key concern identified across the literature is ensuring ethical principles, fairness, and objectivity in AI-driven HR processes.
- Workforce
- Enterprise