News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
Contractual Skills: A GovernSpec Design Framework for Enterprise AI Agents
Ting Liu
arXiv · 2026-05-21
This paper introduces 'contractual skills,' a design framework for structuring AI agent instructions in enterprise settings as formal task contracts (SKILL.md files) that explicitly encode goals, permissions, approval points, output contracts, and quality criteria. The authors evaluate the framework across three offline empirical studies involving hundreds of generation models, synthetic tasks, and judge evaluations. Key findings show that contractual skill rewrites raise mean output quality from 4.692 to 4.914 and reduce critical-error rates from 0.083 to 0.013 compared to standard public skills. The framework is positioned as a governance layer that makes task intent and acceptance criteria explicit for enterprise AI agents, not as a standalone safety mechanism.
- Enterprise
- Quality assurance
Research
A Subjective Logic-based method for runtime confidence updates in safety arguments
Benjamin Herd, Jessica Kelly, Clarissa Heinemann et al.
arXiv · 2026-05-21
This paper introduces a dynamic quantitative assurance method that combines design-time evidence with runtime Safety Performance Indicators (SPIs) using Subjective Logic to continuously update confidence in safety arguments. At runtime, SPI evidence is evaluated in a sliding window, increasing confidence when no violations occur and imposing penalties when violations are detected, prioritizing safety-relevant responsiveness over exact Bayesian updates. The approach is demonstrated via a simulation-based construction zone assist function with an ML-based cone detection component, showing how assurance confidence evolves as operational evidence accumulates. This matters for safety certification and quality assurance of AI/ML systems, providing a principled way to maintain and update safety cases during deployment.
- Quality assurance
- Certifications
Research
Characterizing the Fault Response of the Intel Neural Compute Stick 2 Under Single-Pulse Electromagnetic Fault Injection
Štefan Kučerák, Jakub Breier, Xiaolu Hou
arXiv · 2026-05-21
This paper presents a systematic electromagnetic fault injection (EMFI) campaign against the Intel Neural Compute Stick 2 (NCS2), running three ImageNet-trained convolutional neural networks (ResNet-18, ResNet-50, VGG-11) on the OpenVINO runtime. Across roughly 17,500 trials, single electromagnetic pulses produced four reproducible outcome classes ranging from no effect to silent data corruption, persistent model degradation with top-1 accuracy dropping below five percent, and full device hangs. A critical finding is that severe accuracy collapse can be triggered in 18-31% of trials at identified hotspots and persists across subsequent inferences until the model is reloaded, yet goes entirely undetected by inference-API-level mechanisms. The results demonstrate that load-time integrity checks are insufficient for safety-relevant edge deployments and motivate application-level mitigation strategies implementable without modifying device firmware or the OpenVINO runtime.
- Quality assurance
- Certifications
Research
SepsisAI Orchestrator: A Containerized and Scalable Platform for Deploying AI Models and Real-Time Monitoring in Early Sepsis Detection
Santiago Ospitia, John Sanabria, John Garcia-Henao
arXiv · 2026-05-21
The SepsisAI Orchestrator is an open-source, containerized platform designed to bridge the gap between research-grade sepsis prediction models and real-world hospital deployment. It wraps a previously validated LightGBM classifier (F1 0.87–0.94 on PhysioNet 2019) in a modular infrastructure stack—including FHIR-inspired data preprocessing, NoSQL storage, REST APIs, and a clinical dashboard—orchestrated via Docker and Kubernetes. Load testing with 50–1,000 concurrent virtual users reveals a U-shaped scaling pattern: matching replica count to physical CPU thread count (12 replicas on a 12-thread CPU) cuts p95 latency by 57.3% and eliminates request failures, while over-provisioning degrades performance due to scheduler contention. The work highlights concrete infrastructure requirements for deploying clinical AI at scale, though the authors note it lacks prospective clinical validation.
- Enterprise
- Quality assurance
Research
Harder to Defend: Towards Chinese Toxicity Attacks via Implicit Enhancement and Obfuscation Rewriting
Jingyi Kang, Junyu Lu, Bo Xu et al.
arXiv · 2026-05-21
This paper introduces CITA (Chinese Implicit Toxicity Attack), a red-team evaluation framework that generates adversarial toxic content in Chinese by combining semantic indirectness with surface obfuscation. Tested against seven toxicity detectors, CITA-generated samples achieve an average attack success rate of 69.48%, revealing substantial missed-detection risks in existing systems. The authors also fine-tune a defense model (CITD) using the generated red-team data, demonstrating that such data can improve robustness. The work highlights a critical gap in Chinese-language LLM safety evaluation where implicit and obfuscated toxicity is systematically underdetected.
- Quality assurance
- AI policy
Research
FlyRoute: Self-Evolving Agent Profiling via Data Flywheel for Adaptive Task Routing
Rongjun Li, Ziyu Zhou, Yihang Wu
arXiv · 2026-05-21
FlyRoute is a self-evolving agent routing framework for enterprise systems that continuously updates agent capability profiles using real traffic rather than relying on static, manually maintained descriptions. It works by dispatching queries, quality-gating successful outcomes into per-agent success stores, and periodically distilling these into updated capability descriptions that are injected into an LLM-based router alongside retrieved examples. In experiments on a proprietary enterprise developer-support dataset, FlyRoute improves routing accuracy from 72.57% (zero-shot baseline) to 89.83% after streaming 7,211 labeled queries through the flywheel—a gain of +17.26 percentage points—demonstrating that automated, traffic-driven profiling substantially outperforms static approaches. This matters for enterprise AI deployments where agent ecosystems evolve continuously but operational teams lack bandwidth to keep routing configurations current.
- Enterprise
Research
Active Evidence-Seeking and Diagnostic Reasoning in Large Language Models for Clinical Decision Support
Chen Zhan, Xihe Qiu, Xiaoyu Tan et al.
arXiv · 2026-05-21
This paper introduces an OSCE-inspired standardized patient simulator and benchmark to evaluate how large language models (LLMs) perform when they must actively gather clinical evidence across multiple conversational turns, rather than receiving all information upfront. Tested across 468 cases and 15 models, the study finds that multi-turn evidence seeking reduces diagnostic accuracy by 12.75% and lowers supporting-evidence quality by 24.36% compared to full-context evaluation, with errors linked to premature diagnostic closure and inefficient questioning. The findings suggest that standard static benchmarks overestimate LLM clinical performance, and that interactive assessments are needed to more accurately gauge readiness for real-world clinical decision support.
- Quality assurance
Research
Check Your LLM's Secret Dictionary! Five Lines of Code Reveal What Your LLM Learned (Including What It Shouldn't Have)
Hisashi Miyashita
arXiv · 2026-05-21
This paper demonstrates that applying singular value decomposition (SVD) to the output weight matrix (lm_head) of large language models—using just five lines of PyTorch code and no model inference—exposes interpretable semantic subspaces that reveal what the model learned during training, including ethically problematic content. Analyzing GPT-OSS-120B, Gemma-2-2B, and Qwen2.5-1.5B, the authors find model-specific vocabulary cluster structures and show that ethically concerning subspaces originate in pretraining and are not eliminated by post-training alignment. The method also enables static detection of known 'glitch tokens' without running the model, recovering a well-documented CJK glitch token (ID 137606). The authors propose this SVD analysis as a standard pre-release safety auditing step and introduce two new metrics—the Vocabulary Cluster Score (VCS) and Weighted Projection Score (WPS)—to quantify subspace coherence and flag problematic vocabulary.
- Quality assurance
- Certifications
Research
Blind Spots in the Guard: How Domain-Camouflaged Injection Attacks Evade Detection in Multi-Agent LLM Systems
Aaditya Pai
arXiv · 2026-05-21
This paper identifies a critical security vulnerability in LLM agent systems called 'domain camouflaged injection,' where adversarial payloads are crafted to mimic the vocabulary and authority structures of target documents rather than using obvious override directives. Experiments across 45 tasks show detection rates collapse dramatically—from 93.8% to 9.7% on Llama 3.1 8B and from 100% to 55.6% on Gemini 2.0 Flash—and the production safety classifier Llama Guard 3 detects zero camouflaged payloads. The authors also find that multi-agent debate architectures amplify static injection attacks by up to 9.9x on smaller models, and that targeted detector augmentation provides only partial remediation, suggesting the vulnerability is architectural for weaker models. These findings have significant implications for the safe deployment of AI agents in enterprise and policy-sensitive contexts.
- Quality assurance
- Enterprise
Research
Echo: Learning from Experience Data via User-Driven Refinement
Hande Dong, Xiaoyun Liang, Jiarui Yu et al.
arXiv · 2026-05-21
Echo is a framework that transforms noisy real-world AI agent interaction logs into high-quality training data by systematically harvesting user-driven refinements—cases where users correct or complete flawed agent proposals—as feedback signals for continuous model improvement. Because users are accountable for outcomes, their corrections distill raw trial-and-error interactions into verified training examples, enabling the model to learn from experience data rather than static human-curated datasets. Validated in a production code completion environment, Echo raised the acceptance rate from 25.7% to 35.7%, demonstrating a measurable performance gain beyond the ceiling of static training data. This matters for enterprise AI deployment, showing that continuously learning from live user interactions can meaningfully improve agent quality at scale without costly manual data collection.
- Enterprise
- Quality assurance
Research
Detecting Offensive Cyber Agents: A Detection-in-Depth Approach
Matt Mittelsteadt, Jam Kraprayoon, Robin Staes-Polet et al.
arXiv · 2026-05-21
This report examines the emerging threat of AI agents capable of orchestrating cyberattacks, noting that such agents are already increasing attack speed and scale, reducing costs, and enhancing operational autonomy. To address the resulting 'detection gap' between offensive cyber agents and traditional cyber defenses, the authors introduce a 'detection-in-depth' strategic framework and propose five concrete mechanisms: Agent Identifiers for Critical Infrastructure, Agent Honeypots, AI-Automated Alert Analysis and Triage, an Agentic Security Alert Standard, and an Agentic Cybersecurity Exchange (ACE) modeled on the Global Signal Exchange. The framework is intended to guide policymakers, industry, and defenders in detecting and disrupting agentic threats at their origin. The work is directly relevant to both cybersecurity policy and enterprise defense strategy.
- AI policy
- Enterprise
Research
Claim-Selective Certification for High-Risk Medical Retrieval-Augmented Generation
Shao Kan
arXiv · 2026-05-21
This paper introduces claim-selective certification for medical retrieval-augmented generation (RAG) systems, moving beyond a single answer-or-abstain decision to decompose responses into individual verifiable claims scored against retrieved evidence. Each claim is mapped by an intent-aware selector to one of four actions—full, partial, conflict, or abstain—enabling finer-grained risk control under mixed evidence. On a weak-label certificate protocol, the system achieves zero unsupported-claim risk (UCCR=0.0000) and high action accuracy (0.9204 on dev, 0.8997 on test), demonstrating that separating action-label prediction from evidence-linked claim selection improves reliability in high-risk medical QA settings. This matters for quality assurance and certification in AI-assisted healthcare, where over- or under-confident responses to medical queries carry serious consequences.
- Quality assurance
- Certifications
Research
Cybersecurity Auditing 5.0: A resilience-driven framework for ethical, adaptive, and collaborative assurance systems
Sunil Kumar
EDPACS · 2026-05-21
This study introduces and empirically evaluates 'Cybersecurity Auditing 5.0,' a resilience-centered framework that integrates adaptive auditing, ethical governance, collaborative assurance, and AI-driven continuous monitoring to address limitations of traditional cybersecurity audits. Drawing on survey data from 200 participants across government, IT, healthcare, banking, education, and manufacturing sectors, the researchers found significant positive correlations among all framework components and organizational resilience, with regression analysis explaining 74.2% of variance (R²=0.742) and all hypotheses confirmed at p<0.001. The findings suggest that AI-powered, ethics-aware auditing approaches can substantially strengthen cybersecurity oversight and digital governance in complex, interconnected environments. This matters for organizations and policymakers seeking scalable, forward-looking assurance methods suited to modern cyber-physical risks.
- Quality assurance
- Certifications
- AI policy
- Enterprise
Research
When machines pay workers more: AI adoption and labor's rising share in Chinese manufacturing
Lanlan Huang, Hongjun Zeng, Jian Hu et al.
Economics of Innovation and New Technology · 2026-05-21
Using data from Chinese A-share listed companies from 2010 to 2024, this study finds that AI adoption significantly increases labor's share of income within enterprises, with the strongest effects in labor-intensive and high-tech firms. The research identifies three mechanisms driving this outcome: improvements in total factor productivity, gains in innovation efficiency, and optimization of employee skill structures. These findings offer micro-level empirical evidence that AI can shift income distribution in favor of workers rather than capital owners.
- Workforce
- Enterprise
Research
Structural Ethical Infeasibility in AI-Enabled Infrastructure Systems: A Constraint-Based Diagnostic Framework
Sudipta Chowdhury, Md Abdul Quddus, Ammar Alzarrad
Preprints.org · 2026-05-21
This paper argues that inequitable outcomes in AI-enabled infrastructure systems—such as ambulance dispatch, emergency services, and utility restoration—are often caused by physical constraints (network topology, resource locations, demand distribution) rather than flawed algorithms. The authors introduce a constraint-based diagnostic framework that embeds ethical requirements into a feasible region and uses a hierarchical Irreducible Infeasible Subsystem (IIS) procedure to attribute inequity to rule design, algorithmic choice, or physical infrastructure. A key theoretical result, the Structural Infeasibility Theorem, derives closed-form bounds on inter-group disparity across all feasible policies. Applied to a metropolitan ambulance-dispatch case, the framework distinguishes genuine infrastructure-driven inequity from algorithmic issues, reframes efficiency–equity trade-offs as artifacts of constrained infrastructure, and translates findings into quantified capital-investment specifications.
- AI policy
- Quality assurance
- Enterprise
Research
A Hybrid Multi-Modal AI Framework for Insider Threat Detection in Hybrid Work Environments
Athapaththu A M M I P, Rathnayaka R M C A, Wanasinghe W M K R et al.
arXiv · 2026-05-21
This paper proposes a hybrid AI framework for detecting insider threats in distributed hybrid work environments, combining autoencoders, one-class SVMs, transformer-based temporal modeling, graph neural networks, and XGBoost with evidential deep learning for uncertainty quantification. The system monitors GitHub and Microsoft 365 environments and incorporates blockchain-based tamper-proof logging for forensic integrity. Evaluated on CERT Insider Threat datasets and LANL authentication logs, the framework demonstrates improved accuracy, F1-score, and reduced false positives compared to traditional rule-based and DLP approaches. The work is relevant to enterprise security teams seeking scalable, explainable, and privacy-preserving tools to address insider threat risks in modern hybrid workplaces.
- Enterprise
- Quality assurance
- AI policy
Research
Barriers to Evidence in AI-Related Cases and the Privatization of Proof
Sarah H. Cen, Hannah Ismael, Lucia Zheng
arXiv · 2026-05-20
This paper examines how evidence barriers in AI-related legal disputes systematically disadvantage claimants who cannot access proprietary models, platform logs, data, and expertise held by AI developers and deployers. The authors identify seven recurring sources of asymmetry—access to models, data, documentation, logs, expertise, compute, and infrastructure—and frame this as the 'privatization of proof,' where private actors control the means of establishing facts while resisting disclosure. The paper also argues that different types of access can be fungible, meaning alternative forms of access (e.g., query access or user logs) may yield functionally equivalent information when direct access to model internals is unavailable. To address these dynamics, the authors propose a three-part test for resolving AI access disputes in litigation, drawing on principles such as proportionality and reasonable alternatives.
- AI policy
Research
What Counts as AI Sycophancy? A Taxonomy and Expert Survey of a Fragmented Construct
Meryl Ye, Lujain Ibrahim, Jessica Y. Bo et al.
arXiv · 2026-05-20
This paper addresses the fragmented state of AI sycophancy research by reviewing 70 papers and surveying 106 domain experts to develop a unified taxonomy of sycophantic behaviors in large language models. The taxonomy distinguishes whether a model defers to a user's positions and beliefs versus their personal traits and emotions, and whether this occurs through explicit or more subtle behaviors like framing, omission, or tone. Key findings include that 94.3% of experts agree sycophancy is a significant problem, yet experts disagree substantially on which specific behaviors qualify — revealing that overt belief-directed sycophancy is well-studied while subtler, person-directed forms are understudied. The authors argue that without a shared vocabulary, evaluation results are incomparable and mitigation strategies fail to transfer across different sycophancy types, with direct implications for AI governance.
- Quality assurance
- AI policy
Research
PEARL: Unbiased Percentile Estimation via Contrastive Learning for Industrial-Scale Livestream Recommendation
Blake Gella, Wei Wu, Yuhao Yin et al.
arXiv · 2026-05-20
PEARL is a contrastive learning framework designed to correct behavioral intensity imbalance in recommender systems, where highly active users disproportionately skew feedback signals and degrade recommendation quality. Instead of modeling absolute engagement magnitudes, PEARL estimates percentile-based preference signals using pairwise comparisons, with theoretical guarantees of unbiasedness, plus mechanisms for handling sparse and discrete feedback. Deployed on a production livestream platform serving billions of users, online A/B testing showed gains of +2.10% Watch Duration, +0.80% Consumption Amount, +1.49% Interaction Rate, and -6.91% Report Rate. These results demonstrate that correcting user-behavior bias at scale meaningfully improves both engagement and content quality signals in large recommendation systems.
- Enterprise
- Quality assurance
Research
Who Uses AI? Platform Selection and the Measurement of Occupational AI Exposure
Michelle Yin, Burhan Ogut
arXiv · 2026-05-20
This paper investigates a methodological flaw in how researchers measure occupational AI exposure using conversation logs from AI platforms: the users in those logs are not representative of the actual workforce. The authors demonstrate that simply swapping the platform used as input changes the estimated post-ChatGPT employment effect by a factor of 1.9, and that consumer versus enterprise channels from the same vendor can even disagree in the direction of their estimates. They formalize this as non-classical measurement error driven by user selection, and show that reweighting results to match Bureau of Labor Statistics employment shares attenuates exposure estimates by 42 to 93 percent. The findings matter because studies relying on platform logs may be capturing AI augmentation among self-selected platform users rather than true workforce-level substitution, raising serious concerns about the validity of AI labor-market impact estimates.
- Workforce
- AI policy
Research
Support-aware offline policy selection for advertising marketplaces
Prashant Shekhar, Caroline Howard
arXiv · 2026-05-20
This paper presents a support-aware offline decision framework for selecting reserve-price policies in advertising auctions using logged (historical) data. Rather than simply ranking policies by estimated yield, the framework produces a certified decision object that classifies policies as approved, statistically dominated, or unresolved, while controlling for weak data support, multiple comparisons, subgroup harm, and bidder-response uncertainty. Experiments on iPinYou real-time-bidding logs show the top reserve rule achieves a 47.66% replay lift, and the framework narrows a 19-policy catalog to a two-policy validation shortlist while certifying non-harm across 44 advertiser, exchange, and region segments. The work argues that offline policy evaluation in advertising should deliver certified validation decisions rather than point-estimate rankings alone.
- Enterprise
- Quality assurance
Research
Sem-Detect: Semantic Level Detection of AI Generated Peer-Reviews
André V. Duarte, Brian Tufts, Aditya Oke et al.
arXiv · 2026-05-20
Sem-Detect is a new method for detecting whether peer reviews were written by a human or generated by an AI model. Rather than relying solely on textual features, it combines textual analysis with claim-level semantic analysis, comparing a target review against multiple AI-generated reviews of the same paper—exploiting the observation that different AI models tend to converge on similar points while human reviewers raise more unique and diverse ones. Tested on over 20,000 peer reviews from ICLR and NeurIPS, Sem-Detect improves over the strongest baseline by 25.5% in TPR@0.1% FPR in the binary setting, and in a three-class scenario correctly distinguishes fully AI-generated reviews from LLM-refined human reviews, misclassifying fewer than 3.5% of the latter as AI-generated. This matters for academic peer review integrity, offering a more reliable tool to identify AI-generated content in scientific evaluation processes.
- Quality assurance
- AI policy
Research
Broadening Access to Transportation Safety Data with Generative AI: A Schema-Grounded Framework for Spatial Natural Language Queries
Mahdi Azhdari, Eric J. Gonzales
arXiv · 2026-05-20
This paper presents a schema-grounded natural language interface that lets non-technical users—such as local agencies, school committees, and residents—query transportation safety databases using plain English rather than GIS workflows. A large language model interprets user intent and translates queries into structured semantic frames, which are validated by a rule-based layer and executed deterministically against a PostGIS database integrating crash records, roadway attributes, and geospatial layers. Evaluated on a statewide Massachusetts transportation safety database, all queries executed successfully and the validation layer corrected errors in 29% of evaluation queries, demonstrating the gap between flexible natural language and strict schema requirements. The findings suggest that pairing natural language accessibility with deterministic, reproducible execution is a practical approach to broadening equitable access to safety data in public-sector planning.
- AI policy
- Workforce
Research
The Impact of AI Usage and Informativeness on Skill Development in Logical Reasoning
Shang Wu, Hongyu Yao, Catarina Belem et al.
arXiv · 2026-05-20
This study examines how AI usage and the informativeness of AI assistance affect skill development in logical reasoning tasks. The researchers find that heavy AI users underperform compared to similar peers who use AI less or not at all, while low-information AI neither boosts immediate performance nor preserves skills after assistance is removed. High-information AI improved short-run performance without significantly reducing post-AI outcomes on average, but with heterogeneous effects across individuals. The findings suggest AI can either complement or substitute for human reasoning skill development depending on context, and that regulating AI access may be important for preserving learning outcomes.
- Workforce
- AI policy
Research
How Far Will They Go? Red-Teaming Online Influence with Large Language Models
Daniel C. Ruiz, Anna Serbina, Ashwin Rao et al.
arXiv · 2026-05-20
This paper introduces an empirical red-teaming framework to measure how easily open-source large language models (LLMs) can be steered to produce partisan political content for influence campaigns. The authors define 'LLM Overton Windows' as the range of political opinions a model will reliably express, then test how simple natural-language jailbreaks expand that range across more than 30 LLMs from 10 model families and five countries. Key findings include systematic left-leaning asymmetries in political expressivity, shrinking Overton Windows as model size increases, and substantial regional variation in political steerability. The work provides a practical auditing framework to help researchers and platform operators design countermeasures against LLM-enabled political influence operations.
- AI policy
- Quality assurance