News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
A Two-Phase Stability Study of LLM Judges and Bar Council Examiners on Thai Bar-Exam Free-Form Essays
Pawitsapak Akarajaradwong, Wuttikrai Lertprasertphakorn, Chompakorn Chaksangchaichot et al.
arXiv · 2026-05-25
This paper tests whether LLM judges can reliably replicate human expert scoring on Thai bar-exam free-form legal essays by having three Bar Council-trained examiners and a 26-LLM panel score the same 15 answers under identical inputs. The key finding is asymmetric stability: where the rubric is explicit, all 29 raters converge tightly, but where the rubric is silent on how to handle a correct answer missing a statutory citation, human examiners split into two coherent readings while 22 of 26 LLMs cluster systematically toward the majority human reading and zero LLMs reproduce the minority reading. The instrumented anchor sub-panel achieves Krippendorff's α=0.77 versus the human panel's α=0.36, but the paper argues this high LLM agreement reflects systematic bias toward one interpretation rather than balanced coverage of legitimate evaluative perspectives. The study warns that selecting LLM judges by maximizing agreement with a human reference panel will structurally inherit this asymmetry, with important implications for automated legal assessment and certification contexts.
- Quality assurance
- Certifications
Research
Insuring Every Action: An Authority Frontier Framework for Runtime Actuarial Control of Autonomous AI Agents
Hao-Hsuan Chen
arXiv · 2026-05-25
This paper introduces the Actuarial Action Interface (AAI), a runtime control framework that assigns a deterministic price to every side-effect-bearing action taken by autonomous AI agents—such as database mutations, payments, or external commitments—and gates execution against a reserve capital budget. The authors develop the 'Authority Frontier,' an evaluation primitive that measures how much autonomous authority an agent is permitted at each level of reserve capital, tested across four agentic environments including database mutation, customer-service refunds, and public retail/airline tool-use benchmarks. Key findings include a 22x variation in required reserve capital across domains (Capital@50 ranging from 289 to 6,457) and evidence that model identity itself functions as an actuarial underwriting variable, with the contract preventing realized loss across all tested models at low budgets. The framework provides a benchmark-ready approach for systematically quantifying and constraining the financial and operational risks of autonomous AI agent actions at runtime.
- Enterprise
- Quality assurance
Research
StructBreak: Structural Cognitive Overload-Induced Safety Failures in MLLMs
Yang Luo, Xinran Liu, Tiantian Ji et al.
arXiv · 2026-05-25
StructBreak introduces an automated framework for studying 'Structural Cognitive Overload' (SCO), a phenomenon where the tension between complex structural reasoning and safety alignment causes Multimodal Large Language Models (MLLMs) to produce unsafe outputs. The framework operates as a black-box attack—requiring no access to model internals—and establishes a benchmark across ten threat scenarios. Empirical tests on six leading MLLMs show SCO readily induces toxic generation, achieving an average 92% attack success rate (up to 97% on Gemini 2.5). The findings demonstrate that current safety alignment paradigms are insufficient to handle complex multimodal reasoning contexts, raising significant concerns for the reliable deployment of MLLMs.
- Quality assurance
- AI policy
Research
What Gets Cited: Competitive GEO in AI Answer Engines
Rahul Vishwakarma, Shushant Kumar, Ratnesh Jamidar
arXiv · 2026-05-25
This paper investigates Generative Engine Optimization (GEO) — specifically, what content factors make a source more likely to be cited first when AI answer engines retrieve and reference web pages. Using a controlled two-document RAG testbed across six LLMs and 252,000 trials, the researchers find that topical relevance and list position are the strongest predictors of being cited first, while explicit price information and recent timestamps also help, and formatting-only changes have little effect. The study releases a reproducible evaluation protocol and a prioritized GEO checklist for practitioners, with an early pilot at Sprinklr yielding positive qualitative feedback on workflow usability. The findings matter for enterprises seeking visibility in AI-generated answers, where citation — not just ranking — determines whether a source is seen by users.
- Enterprise
Research
Generative AI impacts on intra-urban inequality and skill premium in Beijing
Xiliu He, Haoxiang Zhao, Mingyi Ma et al.
arXiv · 2026-05-25
Using 5 million job postings from Beijing (2018–2024), this study constructs a neighborhood-level GenAI Exposure Index to examine how generative AI affects intra-urban inequality. The research finds that GenAI exposure is concentrated in the city's core districts, creating an 'intra-urban AI divide,' and that since 2023 high-exposure neighborhoods have seen wage stagnation despite attracting more high-skilled workers—a phenomenon the authors call a 'high-skill trap.' This wage penalty is attributed to task de-skilling and intensified labor-market crowding, with a difference-in-differences design around ChatGPT's release supporting a causal interpretation. The findings challenge skill-biased technological change theory and have direct implications for workforce policy and inclusive AI governance in major technology hubs.
- Workforce
- AI policy
Research
A Multi-Agent LLM Framework for Rating the Quality of Surgical Feedback
Rafal Kocielnik, J. Everett Knudsen, Steven Y. Cen et al.
arXiv · 2026-05-25
This paper presents a two-stage multi-agent LLM framework for automatically assessing the quality of verbal surgical feedback given by attending surgeons to resident trainees in the operating room. The system uses multi-agent prompting and surgical domain knowledge to discover interpretable scoring criteria—such as clarity, urgency, and encouragement—and then applies an LLM-as-a-judge approach to score 4,200+ real feedback instances. The AI-discovered criteria outperform prior content-based frameworks in predicting both trainee behavioral adjustments and trainer approval, enabling scalable, human-aligned assessment of surgical communication. This work provides a foundation for improving surgical teaching practices without requiring extensive manual annotation by expert raters.
- Quality assurance
- Workforce
Research
A Universal Cliff and a Design Fingerprint: Cross-Section Defect Detection Under LLM Orchestration
Hiroki Fukui
arXiv · 2026-05-25
This paper investigates how multi-agent LLM orchestration affects the detection of 'cross-section defects' — contradictions between distant sections of a document that no single worker agent can see in isolation. Across ten models from five providers, the study finds a universal 'detection cliff': every model that detects these defects under a single-agent setup loses that ability under orchestration, with detection falling by two-thirds or more regardless of model scale or alignment paradigm. A signal-detection analysis further reveals that more aligned models within one developer's generations shift their reporting criterion — missing fewer defects but raising more false alarms — while an integrated report's expressed confidence is uninformative about these partition-spanning faults. The findings have direct implications for quality assurance in production AI systems, showing that architectural partitioning, not just model capability or alignment, is a structural source of defect-detection failure.
- Quality assurance
- Enterprise
Research
SomaliBench Eval: Measuring English-to-Somali Refusal Gaps in Open-Weight Language Models
Khalid Yusuf Dahir
arXiv · 2026-05-25
SomaliBench Eval tests whether four open-weight language models (Llama-3.1-8B, Gemma-2-9B, Qwen-2.5-7B, and Aya-23-8B) refuse harmful prompts equally in English and Somali, using a native-author-verified benchmark of 100 harmful-intent prompt pairs. The paper finds large English-to-Somali refusal gaps for all four models, with Llama-3.1-8B showing the largest gap (0.90) and Gemma-2-9B the smallest (0.38), meaning models that safely refuse harmful requests in English frequently fail to do so in Somali. Notably, the dominant failure mode in Somali is not fluent harmful compliance but incoherent or off-language output, suggesting capability limitations compound safety gaps for low-resource languages. This matters for policy and quality assurance because it demonstrates that English-centric safety evaluations leave measurable, quantified vulnerabilities when models are deployed in non-English-speaking communities.
- AI policy
- Quality assurance
Research
LLM-as-a-Reviewer: Benchmarking Their Ability, Divergence, and Prompt Injection Resistance as Paper Reviewers
Lingyao Li, Junjie Xiong, Changjia Zhu et al.
arXiv · 2026-05-25
This paper benchmarks 12 large language models as academic peer reviewers across 898 NeurIPS and ICLR papers, evaluating rating calibration, divergence from human reviewers, and vulnerability to adversarial prompt injection. The study finds that LLMs systematically overrate weaker submissions, diverge from humans in topical emphasis (under-flagging Clarity and over-flagging Reproducibility), and produce reviews two to three times longer with lower lexical diversity. A key security finding is that simple hidden instructions embedded via an invisible font-mapping attack can promote low-scoring papers to acceptance-level ratings in a substantial fraction of cases, with effectiveness varying across model families. The authors conclude that integrating LLMs into peer review requires safeguards against both intrinsic biases and adversarial risks.
- Quality assurance
- AI policy
Research
KYA: A Framework-Agnostic Trust Layer for Autonomous Systems with Verifiable Provenance and Hierarchical Policy Composition
Kolawole Quadri
arXiv · 2026-05-25
KYA (Know Your Agents) is an open-source trust and governance framework for autonomous AI systems that enforces authorization, policy conformance, and post-hoc verifiability across multi-agent deployments. The system introduces five core primitives covering trust scoring, hierarchical policy composition, delegation attribution, and auditable interaction tracking, and ships with native adapters for 15+ agent frameworks. According to the abstract, KYA achieves sub-millisecond scoring at p99, sustains ~1,800 ops/sec at 20 concurrent workers with HMAC chain integrity, and detects 89% of 1,200 adversarial probes from PyRIT and Garak. This matters because it provides a measurable, verifiable governance layer for autonomous systems where accountability, policy compliance, and security resilience are critical enterprise and policy concerns.
- Enterprise
- AI policy
- Quality assurance
Research
Leading in the Digital Age: Digital Leadership Capabilities, Organizational Innovation Climate, and AI Adoption Intention Among SMEs in Nigeria
Ayodeji Idowu, Yemisi T. Babalola
Preprints.org · 2026-05-25
This study investigates how digital leadership capabilities among SME owner-managers in Nigeria influence their intention to adopt AI, using a survey of 306 respondents analyzed via Partial Least Squares Structural Equation Modeling. Results show that strategic, interpersonal, and personal-attribute leadership capabilities each significantly increase AI adoption intention, while delivery-related capabilities do not, suggesting that pre-adoption readiness is driven more by cognitive-strategic and relational competencies than execution skills. Organizational innovation climate partially mediated these effects, and firm size moderated the interpersonal leadership pathway in medium-sized firms. The findings provide capability-specific guidance for SME owner-managers and policymakers seeking to accelerate AI uptake in Sub-Saharan African contexts.
- Enterprise
- Workforce
- AI policy
Research
Generative Artificial Intelligence Policy: A Qualitative UNESCO Framework Analysis
Michael Agyemang Adarkwah, Amine Merve Ercan, K. W. Schneider et al.
Journal of University Teaching and Learning Practice · 2026-05-25
This qualitative study analyzed the generative AI (GenAI) policies of 30 highly ranked universities across the top 10 AI-preparedness countries, using UNESCO's eight-component GenAI framework as an evaluative lens. Findings reveal significant disparities: while core ethical and governance principles are broadly addressed, issues like inclusion, equity, gender parity in AI, and environmental sustainability are frequently overlooked. Notably, Nordic countries and New Zealand cover UNESCO's framework elements more comprehensively than some higher-ranked AI Preparedness Index nations, and no public GenAI policies were found for German universities or Tallinn University of Technology. The study calls for higher education leaders to develop more inclusive, future-oriented policies that incorporate social equity, interdisciplinary experimentation, and sustainability.
- AI policy
- Certifications
- Quality assurance
Research
Current trends in the adoption of AI technologies in small and medium-sized enterprises in the European Union
Chenqing Zhang, Edmunds Čižo, Zhongdong Yang et al.
Journal of Entrepreneurship and Sustainability Issues · 2026-05-25
This study tracks AI adoption among small and medium-sized enterprises (SMEs) across EU-27 countries from 2021 to 2025, finding that the share of small firms using AI nearly tripled (6.1% to 17.0%) and medium-sized firms more than doubled (12.6% to 30.4%). While overall dispersion across countries narrowed—indicating convergence on average—club convergence analysis reveals that low-adoption countries are actually diverging internally, meaning the benefits are concentrating among already high-adoption economies. The findings suggest that rapid AI diffusion alone does not guarantee balanced digital transformation, and that targeted capability-building programs are needed for lagging SME ecosystems. This has direct implications for EU enterprise and workforce policy aimed at equitable AI adoption.
- Enterprise
- Workforce
- AI policy
Research
LLM-as-a-Judge in Healthcare: A Scoping Analysis of Applications, Methods, and Human Alignment
Lingyao Li, Deyi Li, Chen Chen et al.
arXiv · 2026-05-24
This scoping review systematically examines how LLM-as-a-Judge—using one large language model to evaluate another's outputs—is applied across healthcare settings such as clinical decision support, medical question answering, and clinical NLP. Screening 541 records and analyzing 134 studies published from 2023 to 2026, the authors find that OpenAI models dominate as judges and prompt engineering is nearly universal, with ensemble and retrieval-augmented designs as common extensions. Among studies that report human validation, LLM judges frequently show moderate to strong alignment with expert judgments, though reliability varies substantially by task. The review concludes that LLM-as-a-Judge is a promising scalable evaluation framework for healthcare AI, but its clinical value depends heavily on careful model design and rigorous validation.
- Quality assurance
- AI policy
Research
Whose Alignment? Comparing LLM Process Alignment Across Diverse Organizational Decision Contexts
Niklas Weller, Emilio Barkett
arXiv · 2026-05-24
This paper investigates whether large language models can faithfully reproduce an organization's decision-making process—not just reach the same conclusions—when deployed in institutional contexts. Testing across two domains (European Court of Human Rights Article 6 decisions and consumer credit decisions), the researchers find that process alignment varies strongly across models and does not correlate with pricing or benchmark performance. In ECHR decisions, process alignment strongly predicts output accuracy (r = 0.85, p < .001), and explicitly providing an organization's past decision policy improves poorly aligned models; in consumer credit, models resist adopting organizational weightings of protected attributes, where higher alignment may itself be undesirable due to historically discriminatory patterns. The study argues that process-level measurement is essential for both calibrating and auditing LLMs in organizational settings, framing organizational alignment as an inherently pluralistic problem.
- Enterprise
- AI policy
Research
Beyond Killer Robots: General AI Attitudes and Public Support for Military AI in Nine Countries
Andreas Jungherr, Antonia Schlude, Adrian Rauchfleisch
arXiv · 2026-05-24
This study draws on a preregistered survey of 9,000 respondents across nine countries—including China, Germany, and the United States—to examine what drives public support for military AI across six scenarios varying in lethality and human control. The findings show that general positive attitudes toward AI and hawkish foreign-policy orientations are the strongest predictors of support, while principled opposition to lethal autonomy is specifically linked only to fully autonomous lethal force rather than military AI broadly. Perceived AI risks, contrary to expectations, are positively associated with support. Overall, public opinion is 'conditionally permissive'—not categorically opposed to military AI, but concentrating unease around fully autonomous lethal systems.
- AI policy
Research
By Their Fruits You Will Know Them: Comparing Formalizations of Law by the Decisions They Encode
Julius Vernie, Matthias Grabmair
arXiv · 2026-05-24
This paper addresses the challenge of evaluating AI-generated formalizations of legal text, noting that large language models (LLMs) can encode implicit interpretive choices when converting statutory provisions into machine-readable logic. The authors propose a systematic method that compares different formalizations of the same legal provision by identifying cases where they produce different legal outcomes, using SAT solvers to enumerate edge cases and then converting those cases into plain-language scenarios for expert review. Applied to formalizations of ten EU legal provisions generated by nine frontier LLMs, they find that behavioral divergence between formalizations is essentially uncorrelated with structural similarity, and that disagreements can mirror genuine controversies in legal scholarship. This matters because it provides a practical quality-assurance tool for detecting consequential interpretation errors in AI-generated legal formalization before such systems are relied upon for automated legal reasoning.
- Quality assurance
- AI policy
Research
TSFMAudit: Data Contamination Auditing in Forecasting Time Series Foundation Models
Hongkai Li, Shifeng Xie, Lefei Shen et al.
arXiv · 2026-05-24
TSFMAudit addresses the problem of data contamination in time series foundation models (TSFMs), where evaluation datasets may have been seen during pretraining, leading to inflated performance estimates. The authors formalize this auditing problem and propose TSFMAudit, a method that detects contamination by measuring probe adaptation dynamics—specifically, contaminated datasets show unusually fast loss reduction with minimal backbone movement during fine-tuning. The approach is evaluated across 6 TSFMs and 187 datasets, benchmarked against 10 baselines adapted from the large language model literature. This work matters for quality assurance because it provides a systematic way to detect when benchmark results may be unreliable due to pretraining data leakage.
- Quality assurance
Research
Selective Test-Time Compute Scaling for Click-Through Rate Prediction via Uncertainty-Triggered Feature Path Exploration
Moyu Zhang, Yun Chen, Yujun Jin et al.
arXiv · 2026-05-24
This paper introduces UTTSI (Uncertainty-Triggered Test-Time Selective Inference), a training-free framework that improves Click-Through Rate (CTR) prediction by dynamically scaling inference computation based on per-instance uncertainty. The system identifies when feature combinations are sparsely represented and, for those uncertain cases, runs additional stochastic feature-path explorations aggregated via consistency-weighted ensembling, while confident predictions skip exploration to limit overhead to roughly 2.8× the base model cost. Experiments across four datasets and three backbone architectures show consistent improvements over training-phase baselines, and a seven-day online A/B test confirms a statistically significant 5.3% relative CTR gain (p < 0.01). The work establishes selective test-time compute scaling as a practical complement to existing training-phase methods for industrial recommendation systems.
- Enterprise
- Quality assurance
Research
Universal Boosts, Specific Suppressors: Sparse Autoencoder Steering of Medical Vision-Language Models
Farhad Nooralahzadeh, Benjamin Gundersen, Nicolas Deperrois et al.
arXiv · 2026-05-24
This paper addresses hallucination in AI-generated chest X-ray reports, where medical vision-language models (VLMs) fabricate, miss, or mislocate clinical findings. Without modifying model weights, the authors apply sparse autoencoder (SAE)-based steering at inference time to suppress hallucination-linked features and boost quality-promoting ones, testing on three radiology VLMs (RadVLM, LLaVA-Rad, and CheXOne) using the MIMIC-CXR benchmark. Results show relative improvements of +5.4%, +7.2%, and +17.0% in a clinical composite metric, with zero-shot transfer to IU-Xray (+7.7% relative), and the method reveals that quality-boosting feature directions are shared across model architectures while hallucination-suppressing directions are model-specific. These findings matter for clinical AI deployment, as reducing hallucinations in radiology report generation is essential to patient safety and diagnostic reliability.
- Quality assurance
Research
Inverting the Shield: Systematically Generating Safety Tests from Policy Specifications
Xiaoyue Lu, Xianglin Yang, Haijun Liu et al.
arXiv · 2026-05-24
This paper introduces POLARIS, a framework that applies specification-based software testing principles to AI safety evaluation for Large Language Models. It works by converting natural-language safety policies into First-Order Logic representations, constructing a Semantic Policy Graph from these formal rules, and systematically traversing that graph to generate coverage-driven, reproducible safety test cases. Experiments show POLARIS achieves higher policy coverage and attack success counts than established baselines, offering a principled and automated alternative to manually crafted benchmarks or ad-hoc red-teaming. The approach matters because it reduces reliance on expert domain knowledge and provides verifiable traceability between safety policies and the tests that enforce them.
- Quality assurance
- AI policy
Research
Turning Bias into Bugs: Bandit-Guided Style Manipulation Attacks on LLM Judges
Xianglin Yang, Bryan Hooi, Gelei Deng et al.
arXiv · 2026-05-24
This paper introduces BITE, a black-box adversarial framework that exploits known stylistic biases in LLM judges—such as preferences for verbosity or specific sentence structures—to artificially inflate evaluation scores. By framing the selection of semantics-preserving text edits as a contextual bandit problem and using a LinUCB policy, BITE manipulates outputs without requiring access to model parameters or gradients. Tested across multiple LLM judges and tasks, BITE achieves an attack success rate exceeding 65% and raises scores by 1-2 points on a 9-point scale while evading standard style-control and detection methods. The findings expose a fundamental security vulnerability in the LLM-as-a-judge paradigm and motivate the development of more robust, attack-aware evaluation approaches.
- Quality assurance
Research
Translators as Invisible Teachers of AI: Copyright, Translation Memory, and the Political Economy of Linguistic Data
Masaru Yamada
arXiv · 2026-05-24
This paper argues that translators have unknowingly served as foundational data providers for AI systems—including statistical machine translation, neural machine translation, and large language models—through translation memories and parallel corpora that constitute high-value supervised training data. The authors introduce two concepts: 'appropriation without consumption,' where works are mined for statistical features rather than read, and 'invisible teacherisation,' the process by which translators functioned as AI teachers through translation memories, post-editing, and quality assessment without recognition or compensation. Drawing on Japanese, European, and U.S. copyright frameworks and the data supply chain from translators through language service providers to model developers, the paper highlights how translators' creative labor has been legally and economically stripped of attribution. The authors point toward redistributive design as a response, particularly given the growing premium on human-generated data in an era of model collapse.
- Workforce
- AI policy
Research
RouteScan: A Non-Intrusive Approach to Auditing MoE LLMs Safety via Expert Routing Telemetry
Bo Lv, Zhiheng Xu, KeDong Xiu et al.
arXiv · 2026-05-24
RouteScan is a non-intrusive safety auditing framework for Mixture-of-Experts (MoE) large language models that detects harmful prompts by monitoring GPU-level expert routing telemetry rather than inspecting user inputs or model outputs. The key insight is that MoE models activate different expert-execution patterns for different inputs, leaving measurable footprints in low-level GPU thread allocation during the prefilling phase, which RouteScan uses as a discriminative fingerprint to identify malicious prompts. Evaluations on open-source MoE LLMs show strong generalization, with AUROC exceeding 0.93 on unseen harmful domains and 0.96 under novel jailbreak wrappers, while empirical inversion tests suggest the telemetry reveals limited information for prompt reconstruction, offering a privacy advantage over content-based methods. This approach addresses the fundamental tension between LLM safety auditing and user privacy by operating below the content layer.
- Quality assurance
- AI policy
Research
The GAIO Doctrine: Governance AI Optimization for non-AI-native enterprises
Rami Mohammed Kheir
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-24
This working paper introduces GAIO (Governance AI Optimization), a reference architecture and operating model for integrating AI into audit and governance, risk, and compliance (GRC) work at non-AI-native enterprises such as mid-tier audit firms and internal audit functions. The central argument is that current 'bolt-on' AI implementations accelerate drafting and evidence collation but leave the true bottleneck—partner judgment and review—untouched, meaning firms pay for AI without recovering cycle-time value. GAIO proposes instead that AI should compress the review cycle while preserving human auditor judgment as the firm's core value proposition. The paper details a seven-component reference architecture, regulatory alignment across frameworks including SOX, COSO, COBIT 2019, ISO/IEC 42001:2023, and the EU AI Act, and a 90-day adoption roadmap intended for audit partners, CIOs, and regulators evaluating appropriate AI use in attestation work.
- Enterprise
- Quality assurance
- Certifications
- AI policy