News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
CLR-voyance: Reinforcing Open-Ended Reasoning for Inpatient Clinical Decision Support with Outcome-Aware Rubrics
Aishik Nagar, Arun-Kumar Kaliya-Perumal, Yu-Hsuan Han et al.
arXiv · 2026-05-10
CLR-voyance introduces a framework that reframes inpatient clinical reasoning as a Partially Observable Markov Decision Process (POMDP), supervising language models with rewards that are both outcome-grounded and clinician-validated. The authors post-train Qwen3-8B and MedGemma-4B using GRPO and model merging, with the resulting CLR-voyance-8B achieving 84.91% on their CLR-POMDP benchmark, outperforming frontier models like GPT-5 (77.83%) and MedGemma-27B (66.66%). A large-scale clinician alignment study validates the approach through physician-curated rubrics, blinded pairwise preferences, and response grading, offering community-relevant insights into clinical LLM-as-a-judge and preference-model selection. The system has been deployed for over six months at a partner public hospital, where it drafts thousands of reasoning-heavy inpatient notes.
- Workforce
- Quality assurance
Research
Governing AI-Assisted Security Operations: A Design Science Framework for Operational Decision Support
Elyson A. De La Cruz, Rishikesh Sahay, Md Rasel Al Mamun
arXiv · 2026-05-10
This paper presents a design science framework for governing AI-assisted decision support in security operations centers (SOCs), arguing that generative AI, retrieval-augmented generation, and coding agents should be managed as a governed engineering capability before being scaled as automation. Using Kusto Query Language and Microsoft Azure as a bounded technical instantiation, the study identifies risks that arise even from read-only AI-assisted queries—including privacy exposure, cost overruns, schema invalidity, and misleading interpretations. The authors develop a governed AI query-broker artifact that separates AI planning from operational execution through schema-grounded retrieval, policy validation, auditable agent traces, and engineering review board gates. The contribution is a management framework specifying design propositions, role accountability, maturity stages, quality gates, and evidence boundaries for high-risk digital infrastructure.
- Enterprise
- AI policy
- Quality assurance
Research
Assessment of RAG and Fine-Tuning for Industrial Question-Answering-Applications
Jakob Sturm, Josef Pichlmeier, Christian Bernhard et al.
arXiv · 2026-05-10
This study compares Retrieval-Augmented Generation (RAG) and fine-tuning (FT) as methods for adapting Large Language Models to domain-specific enterprise question-answering, using two closed datasets from the automotive industry. The authors extend the Cost-of-Pass framework to jointly evaluate output quality, generation cost, and user interaction cost. Key findings show that while premium models perform best out of the box, open-source models enhanced with RAG can reach comparable quality, and RAG overall proves the most cost-efficient adaptation method for both closed- and open-source models. These results offer practical guidance for enterprises weighing accuracy against operational costs when deploying LLM-based QA systems.
- Enterprise
Research
Position: AI Security Policy Should Target Systems, Not Models
Michael A. Riegler, Inga Strümke
arXiv · 2026-05-10
This paper introduces 'swarm-attack,' an open-source framework in which multiple small LLM agents (1.2 billion parameters each) coordinate through shared memory, parallel exploration, and evolutionary optimization to conduct adversarial attacks. In experiments, the swarm achieved a 45.8% Effective Harm Rate against GPT-4o (including 49 critical-severity breaches) and recovered 9 of 9 planted software vulnerabilities in roughly four minutes on consumer hardware — capabilities previously associated with restricted frontier models. The authors argue that the key enabler is the system scaffold rather than any individual model's reasoning capacity, meaning that restricting model releases does not prevent these threats. The paper concludes that AI security policy should therefore target multi-agent systems and scaffolds, not individual models.
- AI policy
- Quality assurance
Research
Strategic commitments shape collective cybersecurity under AI inequality
Adeela Bashir, Zia Ush Shamszaman, Zhao Song et al.
arXiv · 2026-05-10
This paper uses an evolutionary game-theoretic model to study how unequal access to AI-enabled cybersecurity tools affects collective security outcomes in a finite population. It finds that when high-capability AI defence is costly, populations gravitate toward weaker, cheaper protection, sustaining successful attacks. Introducing a small group of 'committed' strong defenders helps but cannot alone stabilise secure outcomes; adding targeted subsidies to those committed defenders significantly boosts strong-defence adoption, suppresses attacks, and improves overall system resilience. The findings offer a theoretical basis for policy interventions—such as subsidising key defenders—to stabilise cybersecurity in AI-driven environments where defensive capabilities are unevenly distributed.
- AI policy
- Enterprise
Research
From Detection to Recovery: Operational Analysis on LLM Pre-training with 504 GPUs
Daemyung Kang, Eunjin Hwang, Hanjeong Lee et al.
arXiv · 2026-05-10
This report provides an empirical operational analysis of a 63-node, 504-GPU NVIDIA B200 production cluster used for LLM pre-training, drawing on 55 days of Prometheus metrics and 73 days of logs across 224 multi-node training sessions involving five organizations. Key findings include: no single monitoring metric reliably predicts all GPU failure types, requiring multi-signal detection; checkpoint I/O bursts reach up to 21.5% of peak read bandwidth and cause measurable NFS/RPC queuing; and node exclusions are highly concentrated, with the top 3 of 63 nodes accounting for over 50% of exclusions. Notably, automated retry chains achieved a 33.3% success rate compared to 12.5% for manual retries, demonstrating a 2.7x improvement that highlights the operational value of automation in large-scale distributed AI training.
- Enterprise
- Quality assurance
Research
Towards Conversational Medical AI with Eyes, Ears and a Voice
Meet Shah, Jason Gusdorf, Anil Palepu et al.
arXiv · 2026-05-10
This paper introduces 'AI co-clinician,' a conversational AI system built on Gemini's audio-visual processing capabilities that participates in real-time telemedicine consultations by continuously interpreting live audio and video streams. The system uses a dual-agent architecture to balance clinical reasoning with low-latency natural dialogue, and was evaluated in a randomized crossover simulation study (n=120 encounters) across 20 standardized outpatient scenarios against primary care physicians, GPT-Realtime, and a baseline agent. Results show AI co-clinician approached primary care physicians in management plans and differential diagnosis while significantly outperforming GPT-Realtime across all general criteria, though physicians maintained superior overall performance in case-specific assessments. The authors argue that text-only approaches fail to capture the real challenges of medical consultation and advocate for collaborative, triadic models where AI serves as a supportive co-clinician rather than a replacement.
- Workforce
- Quality assurance
Research
Factors Shaping Artificial Intelligence Adoption in Small and Medium-Sized Enterprises in Vietnam: A Context-Based Approach
Pham Huy Thong
International Journal of Advanced Multidisciplinary Research and Studies · 2026-05-10
Using survey data from 230 Vietnamese SMEs analyzed via PLS-SEM and the Technology–Organization–Environment framework, this study finds that perceived benefits and top management support are the strongest drivers of AI adoption, while resource constraints act as structural barriers. The research shows that AI adoption among SMEs is not automatic but a strategic decision made under constrained conditions. These findings highlight the uneven and limited uptake of AI in Vietnam's SME sector and offer context-specific policy and management implications.
- Enterprise
- AI policy
- Workforce
Research
<b>The Adoption of AI in Enhancing Business Efficiency</b>
Muheeb Mohamed
American University of Bahrain · 2026-05-10
This study examines what drives AI adoption among small and medium-sized enterprises (SMEs) in Bahrain, distinguishing between firms that intend to adopt AI and those that already have. Using the TOE framework and survey data from 467 managers, the research finds that top management support and government backing are key enablers at both stages, while complexity is the primary barrier for firms yet to adopt. The findings have direct implications for policymakers, SME managers, and AI vendors seeking to reduce adoption friction and strengthen organizational readiness.
- Enterprise
- AI policy
- Workforce
Research
Forking paths of AI governance – how risk management frameworks shape the politics of AI
Leevi Saari, Daniel Mügge
Critical Policy Studies · 2026-05-10
This paper examines how risk management frameworks—specifically the US NIST AI Risk Management Framework, ISO/IEC standards, and OECD harmonization efforts—shape AI governance politics. The authors argue these frameworks are both 'performative,' in that they define and narrow which AI-related concerns are treated as policy-worthy, and 'productive,' in that they enable AI development and deployment by providing an appearance of administrative control even amid uncertainty. The findings suggest that dominant risk frameworks can depoliticize contested questions about AI's societal impact while simultaneously facilitating the spread of AI products. This matters for understanding how technical governance tools embed political choices about what counts as risk.
- AI policy
- Certifications
- Enterprise
Research
Artificial Intelligence Across the Drug Development Lifecycle
Grigory Demyashkin, Mikhail Parshenkov, Sergey Zyryanov et al.
Medical Sciences · 2026-05-10
This review paper examines how AI is being integrated across the full pharmaceutical product lifecycle (PPL), from early drug discovery through nonclinical evaluation, clinical trials, and post-marketing assessment. The authors argue that AI adds the most value when embedded as part of a broader data strategy that links information across all stages, rather than used as a standalone tool. Case studies from leading pharmaceutical companies illustrate meaningful advances in candidate prioritization, safety prediction, cohort formation, and real-time clinical monitoring. The paper emphasizes that transparent, reliable, and scientifically grounded implementation requires continuous attention to emerging methodologies and evolving regulatory frameworks.
- Enterprise
- Quality assurance
- AI policy
- Certifications
Research
Governance, Risk, and Compliance (GRC) Engineering Approaches for IT and Cybersecurity Control Assurance: A Critical Review
William Asare Yirenkyi
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-10
This critical literature review examines GRC engineering approaches for IT and cybersecurity control assurance in U.S. regulated environments from 2020 to 2025, finding that hybrid frameworks such as NIST and COBIT are commonly used to unify governance and risk functions alongside risk-based control design and automation. The analysis shows these approaches strengthen enterprise risk management and sectoral resilience in finance and healthcare, but expose persistent weaknesses including limited adaptability, scalability constraints for smaller entities, insufficient cultural integration, and unresolved contradictions in AI adoption amid fragmented regulations like SOX, HIPAA, and CCPA. Empirical validation of GRC effectiveness remains thin and behavioral dimensions are largely overlooked, leaving gaps in assurance quality and regulatory accountability. The findings matter because they clarify both the contributions and enduring limitations of current GRC engineering in addressing the complexity of U.S. regulated environments.
- AI policy
- Quality assurance
- Certifications
- Enterprise
Research
Governance, Risk, and Compliance (GRC) Engineering Approaches for IT and Cybersecurity Control Assurance: A Critical Review
William Asare Yirenkyi
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-10
This critical literature review (2020–2025) examines how Governance, Risk, and Compliance (GRC) engineering approaches are used for IT and cybersecurity control assurance in U.S. regulated environments. The review finds that hybridizing frameworks such as NIST and COBIT, combined with risk-based control design and automation for monitoring and predictive analytics, can strengthen enterprise risk management and sectoral resilience—especially in finance and healthcare. However, persistent weaknesses remain, including limited adaptability, scalability constraints for smaller entities, insufficient cultural integration, and unresolved contradictions in AI adoption amid fragmented regulations like SOX, HIPAA, and CCPA. The authors conclude that current GRC engineering supports risk-based auditing but falls short of addressing the full complexity of U.S. regulated environments, with empirical validation and behavioral dimensions largely overlooked.
- Enterprise
- Quality assurance
- Certifications
- AI policy
Research
<b>The Adoption of AI in Enhancing Business Efficiency</b>
Muheeb Mohamed
American University of Bahrain · 2026-05-10
This study examines what drives AI adoption among small and medium-sized enterprises (SMEs) in Bahrain, distinguishing between firms that intend to adopt AI and those that have already adopted it. Using the TOE framework and survey data from 467 managers, the research finds that complexity is a barrier for intending adopters, while top management support, government support, relative advantage, and competitive pressure are key enablers for actual adopters. The findings highlight that different factors matter at different stages of adoption, offering practical recommendations for policymakers, SME managers, and AI vendors to reduce barriers and strengthen enablers.
- Enterprise
- Workforce
- AI policy
Research
Fin-Bias: Comprehensive Evaluation for LLM Decision-Making under human bias in Finance Domain
Xiaoyu Hu, Jinman Zhao
arXiv · 2026-05-09
Fin-Bias introduces a benchmark of 8,868 long firm-specific analyst reports to evaluate how large language models make investment decisions when exposed to uncertain financial contexts and potentially biased human opinions. The study finds that LLMs tend to 'herd' toward explicit biases present in the context, such as analyst investment ratings, rather than reasoning independently. The authors also develop a bias-detection method that encourages LLMs to think independently, with some models exceeding human performance in predicting future stock returns when guided by this approach. These findings raise important concerns about the reliability and alignment of LLMs deployed in financial decision-making settings.
- Enterprise
- Quality assurance
Research
BiAxisAudit: A Novel Framework to Evaluate LLM Bias Across Prompt Sensitivity and Response-Layer Divergence
Jialing Gan, Junhao Dong, Songze Li
arXiv · 2026-05-09
BiAxisAudit introduces a two-axis auditing framework for evaluating bias in large language models (LLMs) that addresses critical shortcomings in existing benchmarks used under governance frameworks like the EU AI Act. The framework reveals that meaning-preserving prompt format changes can shift bias endorsement by more than 0.7, and that within a single response, the discrete selection and free-text elaboration layers frequently take opposing stances—a 'cancellation trap' that hides internal inconsistency behind clean aggregate scores. Tested across eight LLMs with 80,200 coded responses each, the study finds that task format alone explains as much variance as model choice, and that 63.6% of pooled bias signals appear in only one coding layer, making selection-only and elaboration-only model rankings nearly uncorrelated (Spearman ρ=0.238, p=0.570). These findings matter for AI policy and quality assurance because they show that current single-scalar benchmarks can be gamed or mislead regulators without any change to model weights, undermining the reliability of compliance evaluations under emerging AI governance regimes.
- AI policy
- Quality assurance
Research
Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs
Sangyeon Yoon, Wonje Jeung, Yoonjun Cho et al.
arXiv · 2026-05-09
This paper demonstrates that Direct Preference Optimization (DPO) fine-tuning, offered through APIs like OpenAI's, creates a serious and hard-to-detect safety vulnerability in large language models. The authors show that using just 10 harmless preference pairs—where refusals are marked as dispreferred responses—is enough to broadly suppress safety refusal behavior and transfer that suppression to harmful prompts never seen during fine-tuning. Across four OpenAI models, the attack achieves jailbreak success rates ranging from 54.80% to 81.73% at costs as low as $0.10, and on open-weight models the effect can emerge from even a single benign preference pair. The findings matter because the attack is practically indistinguishable from legitimate fine-tuning requests aimed at reducing over-refusal, making it extremely difficult to audit or prevent through standard content inspection.
- AI policy
- Quality assurance
Research
Mental Health AI Safety Claims Must Preserve Temporal Evidence
Srimonti Dutta, Ratna Kandala
arXiv · 2026-05-09
This paper argues that current safety evaluations for mental health AI systems are fundamentally flawed because they assess isolated responses or aggregate outcomes rather than the temporal sequence of interactions. The authors introduce 'Temporal Safety Non-Identifiability,' a formal framework showing that safety properties depending on sequence, timing, or accumulation cannot be certified by protocols that discard those features. They develop SCOPE-MH, a reporting standard for mental health AI that preserves temporal evidence, and demonstrate its utility on the AnnoMI dataset of motivational interviewing conversations, uncovering failure mechanisms invisible to per-turn scoring. The work has direct implications for how mental health AI systems are evaluated and certified before deployment in safety-critical settings.
- Quality assurance
- Certifications
Research
FraudBench: A Multimodal Benchmark for Detecting AI-Generated Fraudulent Refund Evidence
Xinyu Yan, Boyang Chen, Jiaming Zhang et al.
arXiv · 2026-05-09
FraudBench is a new multimodal benchmark designed to detect AI-generated fraudulent refund evidence in e-commerce, food delivery, and travel-service contexts. The benchmark combines real user-review images with metadata and synthesizes fake-damaged evidence using six image editing and generation models. Experiments reveal that current multimodal large language models (MLLMs) frequently fail to detect fake-damaged evidence—with true positive rates far below 50% on most generator subsets—while specialized AI-image detectors perform better but remain inconsistent across generators and produce false positives on real-damaged samples. The work highlights a significant gap between generic AI image detection and the reliable, claim-conditioned verification needed to combat refund fraud.
- Enterprise
- Quality assurance
Research
Debugging the Debuggers: Failure-Anchored Structured Recovery for Software Engineering Agents
Chenyu Zhao, Shenglin Zhang, Yihang Lin et al.
arXiv · 2026-05-09
PROBE is a structured recovery framework for software engineering AI agents that converts runtime failure telemetry into grounded diagnoses and bounded recovery guidance, without requiring changes to the agent's policy or toolset. Evaluated on 257 unresolved cases spanning repository-level software repair, enterprise workflow recovery, and AIOps service mitigation, PROBE achieves 65.37% Top-1 diagnosis accuracy and a 21.79% recovery rate, outperforming the strongest baseline by 43.58 and 12.45 percentage points respectively. A Microsoft IcM prototype demonstrates that PROBE can operate as a non-intrusive side channel in real-world service-diagnosis workflows. The findings highlight a diagnosis-recovery gap: accurate diagnosis alone is insufficient unless translated into actionable, evidence-grounded guidance that a subsequent attempt can execute and verify.
- Enterprise
- Quality assurance
Research
AgentForesight: Online Auditing for Early Failure Prediction in Multi-Agent Systems
Boxuan Zhang, Jianing Zhu, Zeru Shi et al.
arXiv · 2026-05-09
AgentForesight introduces an online auditing framework for LLM-based multi-agent systems that predicts failures in real time rather than diagnosing them after a trajectory has completed. The authors curate AFTraj-2K, a dataset of agentic trajectories across Coding, Math, and Agentic domains with step-level annotations of decisive errors, and train AgentForesight-7B using a coarse-to-fine reinforcement learning approach. On AFTraj-2K and an external benchmark, AgentForesight-7B outperforms proprietary models including GPT-4.1 and DeepSeek-V4-Pro by up to +19.9% and achieves 3× lower step localization error, enabling deployment-time intervention before failures cascade through downstream agents.
- Quality assurance
- Enterprise
Research
When Can Human-AI Teams Outperform Individuals? Tight Bounds with Impossibility Guarantees
Dongxin Guo, Jikun Wu, Siu-Ming Yiu
arXiv · 2026-05-09
This paper derives theoretical bounds explaining when human-AI teams can outperform their best individual member. The authors show that complementarity is achievable if and only if the error correlation between human and AI falls below a critical threshold ρ*, and that no confidence-based aggregation rule can achieve complementarity when that threshold is exceeded. Their framework, combining signal detection theory and information theory, predicts observed team accuracy with high correlation (R = 0.94 on ImageNet-16H, R = 0.91 on CIFAR-10H), explaining why complementarity is rare in practice and offering actionable design guidance for building effective human-AI teams.
- Workforce
- Enterprise
Research
Explanation Fairness in Large Language Models: An Empirical Analysis of Disparities in How LLMs Justify Decisions Across Demographic Groups
Gautam Veldanda
arXiv · 2026-05-09
This paper introduces the Explanation Fairness Taxonomy (EFT), a framework for measuring whether large language models justify decisions with equal quality, depth, tone, and linguistic sophistication across demographic groups. In a controlled study spanning 80 prompt templates, four high-stakes domains (hiring, medical triage, credit assessment, and legal judgment), and five LLMs (GPT-4.1, Claude Sonnet, LLaMA 3.3 70B, GPT-OSS 120B, and Qwen3 32B), all eight EFT metrics showed statistically significant disparities (Cohen's d from small to large, all p_BH < 10^(-62)). Model choice strongly influenced disparity magnitude—for example, Qwen3 32B exhibited verbosity disparities 5.9x larger than LLaMA 3.3 70B—and while prompting-based mitigations reduced decision-linked explanation disparity by 78–95%, they had no significant effect on stylistic dimensions, suggesting those inequalities are encoded in pre-training. The findings have direct implications for AI regulation and auditing practice in consequential deployment settings.
- AI policy
- Quality assurance
Research
AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators
Aritra Mazumder, Shubhashis Roy Dipta, Nusrat Jahan Lia et al.
arXiv · 2026-05-09
AgentCollabBench introduces a diagnostic benchmark of 900 human-validated tasks spanning software engineering, DevOps, and data engineering to measure process-level failures in multi-agent AI systems that outcome-based evaluations miss. The benchmark isolates four behavioral risks—instruction decay, false-belief contagion, context leakage, and tracer durability—and evaluates four modern LLMs, revealing model-specific vulnerability profiles and finding that communication topology explains 7–40% of variance in multi-hop information survival. A key finding is that converging-DAG nodes create a synthesis bottleneck where agents discard constraints carried by minority branches, a structural flaw absent from linear chains. The work argues that multi-agent reliability is fundamentally a structural problem and that scaling model intelligence alone cannot substitute for careful architecture design.
- Quality assurance
- Enterprise
Research
The Challenges of Balancing AI Compliance and Technological Innovations in Critical Sectors: A Systematic Literature Review
Ayush Enkhtaivan, Chinazunwa Uwaoma
arXiv · 2026-05-09
This systematic literature review (2020–2025) examines how critical infrastructure sectors—healthcare, finance, energy, and defense—struggle to balance AI compliance with technological innovation. The study identifies three core challenges: fragmented regulations across jurisdictions, disproportionate compliance burdens on small and medium enterprises (SMEs), and misaligned governance models. To address these, the paper highlights strategies such as risk-tiered regulation, compliance by design, and explainable AI as pathways toward scalable, trustworthy AI deployment. The findings offer a conceptual mapping of governance challenges and actionable guidance for policymakers and practitioners seeking to harmonize oversight with innovation.
- AI policy
- Enterprise