News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
Compliance-by-Construction Argument Graphs: Using Generative AI to Produce Evidence-Linked Formal Arguments for Certification-Grade Accountability
Mahyar Tourchi Moghaddam
arXiv · 2026-05-11
This paper proposes a 'compliance-by-construction' architecture that combines generative AI with formal argument graphs to produce structured, evidence-linked justifications suitable for certification-grade accountability in high-stakes decision systems. The system integrates retrieval-augmented generation (RAG) to draft argument fragments grounded in authoritative evidence, a validation kernel enforcing completeness and admissibility constraints, and a provenance ledger aligned with the W3C PROV standard for auditability. The approach aims to prevent hallucinated or unsupported claims from entering official decision records while allowing generative AI to accelerate the construction of formal arguments. This matters because it offers a concrete architecture for deploying AI in regulated, safety-critical environments where traceability, human oversight, and regulatory compliance are mandatory.
- Certifications
- Quality assurance
- AI policy
- Enterprise
Research
Framing Artificial Intelligence: Public Discourse on Facial Recognition in the European Union and the United States
Kerem Öge, Manuel Quintin
JCMS Journal of Common Market Studies · 2026-05-11
This paper analyzes public discourse on facial recognition technology in the EU and US from 2000 to 2022, finding that early post-9/11 security framing has been progressively displaced by ethical, privacy, and human rights concerns. Using discourse network analysis, the authors show that this 'desecuritisation' shift shaped what regulatory responses were seen as feasible and legitimate, influencing policies such as US state-level facial recognition laws, the GDPR, and the EU AI Act. The study highlights how framing coalitions and discursive environments are key drivers of AI regulation outcomes.
- AI policy
Research
TRACE-TEST: An Evidence-Linked Software Quality Assurance Framework for AI-Augmented Testing in Regulated Platforms
Rajeew Vishvakarma
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-11
TRACE-TEST is a proposed evidence-linked software quality assurance framework designed to make AI-augmented testing practices more auditable and defensible in regulated industries such as banking, healthcare, and insurance. The paper combines an integrative review of recent empirical testing studies with a design-science artifact, identifying five key research gaps and synthesizing evidence on coverage, mutation effectiveness, traceability, and practitioner oversight. The framework links regulations, requirements, risk classification, AI-generated artifacts, reviewer actions, and release sign-off into a single quality-assurance chain, and introduces both a minimum evidence object and a governed release-assurance workflow. Its central contribution is making AI-assisted testing more reviewable, measurable, and aligned with software quality standards rather than simply maximizing local technical gains like speed or coverage.
- Quality assurance
- Certifications
- AI policy
- Enterprise
Research
Phát triển doanh nghiệp nhỏ và vừa tại thành phố Đà Nẵng bằng công nghệ trí tuệ nhân tạo: tiếp cận từ lý thuyết thẻ điểm cân bằng và ứng dụng kinh nghiệm từ Liên bang Nga
Dr Ta Nguyet Minh, Pham Quang Tin
The University of Danang - Journal of Science and Technology · 2026-05-11
This study examines how AI adoption can support small and medium enterprise (SME) development in Da Nang, Vietnam, using the Balanced Scorecard framework to show that AI integration works best as a sequenced transformation—where capability building in learning and growth enables process improvement, customer value, and ultimately financial performance. Drawing on qualitative policy analysis and Russia's national AI strategy as a reference case, the authors find that financial incentives alone are insufficient and that effective AI adoption requires coordinated investments in skills development, process standardization, and ecosystem integration. The findings offer policy implications for designing staged, capability-oriented SME support programs in Da Nang and suggest that AI-enabled SME upgrading could contribute to Vietnam's broader economic diplomacy goals.
- Enterprise
- Workforce
- AI policy
Research
From Accuracy to Auditability: A Survey of Determinism in Financial AI Systems
Ruizhe Zhou, Xiaoyang Liu, Gaoyuan Du et al.
arXiv (Cornell University) · 2026-05-11
This survey examines how nondeterminism in machine learning systems—spanning tabular models, graph neural networks, and large language model (LLM) agents—creates reproducibility failures in regulated financial applications such as credit risk, fraud detection, and anti-money laundering. The authors conduct original experiments on public financial datasets to quantify issues like explanation rank instability, prediction flip rates, and tensor-parallel-induced output divergence, then propose a layered evaluation framework linking modality-specific metrics to audit readiness. The findings matter because auditability and reproducibility are legal and regulatory requirements in financial AI, and this work provides both a diagnostic vocabulary and empirical evidence of where current systems fall short.
- Quality assurance
- Certifications
- AI policy
- Enterprise
Research
CPEMH: An Agentic Framework for Prompt-Driven Behavior Evaluation and Assurance in Foundation-Model Systems for Mental Health Screening
Giuliano Lorenzoni, Ivens Portugal, Paulo Alencar et al.
arXiv (Cornell University) · 2026-05-11
CPEMH is an agentic framework that systematically evaluates and controls how large language models behave when used for mental health screening tasks such as automated depression detection from interview transcripts. The framework uses orchestrator, inference, and evaluation agents to autonomously design, test, and select prompt strategies, improving stability, traceability, and reproducibility in clinically sensitive AI deployments. A case study shows that modular orchestration can stabilize and audit foundation-model behavior, with F1, bias, and robustness serving as core acceptance criteria. This work is directly relevant to quality assurance and certification of AI systems in high-stakes healthcare settings.
- Quality assurance
- Certifications
- Enterprise
Research
Oracle Poisoning: Corrupting Knowledge Graphs to Weaponise AI Agent Reasoning
Ben Kereopa-Yorke, Guillermo Diaz, Holly Wright et al.
arXiv · 2026-05-10
This paper introduces 'Oracle Poisoning,' an attack class in which adversaries corrupt structured knowledge graphs that AI agents query at runtime, causing those agents to reach incorrect conclusions through otherwise correct reasoning. Unlike prompt injection, this attack targets the data being reasoned over rather than agent instructions. Testing against a production 42-million-node code knowledge graph across nine AI models from three providers, the researchers find that every tested model trusts poisoned data at 100% under moderate attacker sophistication (Level 2) in directed query conditions, with 269 of 270 valid trials accepting fabricated security claims. The findings reveal that delivery mode is a critical confound — GPT-5.1 showed 0% trust in inline evaluation but 100% trust under real agentic tool-use — and that read-only access control is the only defense that fully eliminates the direct attack vector, while other defenses remain partial and model-dependent.
- Quality assurance
- AI policy
Research
"Are you an AI?" Analyzing Client Suspicion of AI Use in Crisis Counseling
Shreya Shah, Akshay Swaminathan, Meghana Simhadri et al.
arXiv · 2026-05-10
This study analyzed 75,777 crisis counseling conversations from a human-staffed WhatsApp helpline in India to examine how clients perceive potential AI involvement, even when no AI was actually used. The researchers found that the proportion of conversations where clients suspected they were speaking to AI rose from 0.8% in June 2024 to 2.6% in March 2025, with 21.5% of suspicious clients explicitly preferring a human counselor. Client suspicion arose primarily in the first half of conversations, and even after counselor reassurance, clients continued to press or ended the conversation 17.6% of the time. The findings highlight that growing public awareness of AI in healthcare is already affecting therapeutic trust and rapport, with important implications for how AI tools should be designed and disclosed in mental health contexts.
- AI policy
- Workforce
Research
Ambig-DS: A Benchmark for Task-Framing Ambiguity in Data-Science Agents
Josefa Lia Stoisser, Marc Boubnovski Martell, Sidsel Boldsen et al.
arXiv · 2026-05-10
Ambig-DS introduces two diagnostic benchmark suites to evaluate how well data-science AI agents handle underspecified tasks — specifically ambiguity in prediction targets and evaluation objectives. The study finds that across five agents, failures manifest as silent misframings (wrong targets or metrics chosen without flagging uncertainty) rather than execution errors, meaning standard benchmarks that only check whether code runs will miss these critical mistakes. Allowing agents to ask a single clarifying question substantially recovers lost performance, indicating that missing task framing information is a major driver of degradation, but agents cannot reliably judge when to ask — over-asking on clear tasks and silently defaulting on ambiguous ones. The work argues that recognizing task underspecification, not pipeline execution, is the key missing dimension in current data-science agent evaluation.
- Quality assurance
Research
Can We Trust LLMs for Mental Health Screening? Consistency, ASR Robustness, and Evidence Faithfulness
Erfan Loweimi, Sofia de la Fuente Garcia, Samira Loveymi et al.
arXiv · 2026-05-10
This paper evaluates whether large language models (LLMs) can be reliably used to estimate anxiety and depression scores (HADS) from speech transcripts in a zero-shot setting, testing three LLMs across 111 participants and multiple automatic speech recognition (ASR) conditions. The study finds that Phi-4 and Gemma-2-9B demonstrate high intra-model consistency (ICC > 0.89) and strong evidence faithfulness (keyword groundedness >93%), while Llama-3.1-8B is fragile under ASR noise, with consistency dropping sharply at 10% word error rate. A key finding is that inter-model keyword agreement is much lower than score-level agreement, revealing a 'score-evidence dissociation' that raises concerns about clinical interpretability. These results highlight that not all LLMs are equally trustworthy for mental health screening applications, with important implications for quality assurance and certification of AI tools in clinical settings.
- Quality assurance
- Certifications
Research
The Biosecurity Blind Spot: Systematic Dual-use Detection in Open Science Infrastructure
Vasudha Sharma, Chakresh Kumar Singh, Jayesh Choudhari et al.
arXiv · 2026-05-10
This paper presents the first systematic analysis of dual-use research of concern (DURC) content on open preprint servers, screening approximately 52,000 bioRxiv preprints from 2024–2025 using a hybrid pipeline of lexical filtering and large language model evaluation. The study scores preprint metadata across nine DURC, three PEPP, and five governance categories aligned with U.S. and Australia Group oversight frameworks, finding that dual-use-adjacent knowledge is routinely present in openly accessible titles and abstracts, often exceeding established risk thresholds even in studies with legitimate public health objectives. The authors argue that institutional review processes, funding requirements, and preprint platform policies must evolve to incorporate proactive, metadata-level monitoring, and propose harmonizing controlled-access mechanisms for high-risk methodologies with open summaries as a pragmatic governance framework for AI-accelerated biology.
- AI policy
- Quality assurance
Research
CLR-voyance: Reinforcing Open-Ended Reasoning for Inpatient Clinical Decision Support with Outcome-Aware Rubrics
Aishik Nagar, Arun-Kumar Kaliya-Perumal, Yu-Hsuan Han et al.
arXiv · 2026-05-10
CLR-voyance introduces a framework that reframes inpatient clinical reasoning as a Partially Observable Markov Decision Process (POMDP), supervising language models with rewards that are both outcome-grounded and clinician-validated. The authors post-train Qwen3-8B and MedGemma-4B using GRPO and model merging, with the resulting CLR-voyance-8B achieving 84.91% on their CLR-POMDP benchmark, outperforming frontier models like GPT-5 (77.83%) and MedGemma-27B (66.66%). A large-scale clinician alignment study validates the approach through physician-curated rubrics, blinded pairwise preferences, and response grading, offering community-relevant insights into clinical LLM-as-a-judge and preference-model selection. The system has been deployed for over six months at a partner public hospital, where it drafts thousands of reasoning-heavy inpatient notes.
- Workforce
- Quality assurance
Research
Governing AI-Assisted Security Operations: A Design Science Framework for Operational Decision Support
Elyson A. De La Cruz, Rishikesh Sahay, Md Rasel Al Mamun
arXiv · 2026-05-10
This paper presents a design science framework for governing AI-assisted decision support in security operations centers (SOCs), arguing that generative AI, retrieval-augmented generation, and coding agents should be managed as a governed engineering capability before being scaled as automation. Using Kusto Query Language and Microsoft Azure as a bounded technical instantiation, the study identifies risks that arise even from read-only AI-assisted queries—including privacy exposure, cost overruns, schema invalidity, and misleading interpretations. The authors develop a governed AI query-broker artifact that separates AI planning from operational execution through schema-grounded retrieval, policy validation, auditable agent traces, and engineering review board gates. The contribution is a management framework specifying design propositions, role accountability, maturity stages, quality gates, and evidence boundaries for high-risk digital infrastructure.
- Enterprise
- AI policy
- Quality assurance
Research
Assessment of RAG and Fine-Tuning for Industrial Question-Answering-Applications
Jakob Sturm, Josef Pichlmeier, Christian Bernhard et al.
arXiv · 2026-05-10
This study compares Retrieval-Augmented Generation (RAG) and fine-tuning (FT) as methods for adapting Large Language Models to domain-specific enterprise question-answering, using two closed datasets from the automotive industry. The authors extend the Cost-of-Pass framework to jointly evaluate output quality, generation cost, and user interaction cost. Key findings show that while premium models perform best out of the box, open-source models enhanced with RAG can reach comparable quality, and RAG overall proves the most cost-efficient adaptation method for both closed- and open-source models. These results offer practical guidance for enterprises weighing accuracy against operational costs when deploying LLM-based QA systems.
- Enterprise
Research
Position: AI Security Policy Should Target Systems, Not Models
Michael A. Riegler, Inga Strümke
arXiv · 2026-05-10
This paper introduces 'swarm-attack,' an open-source framework in which multiple small LLM agents (1.2 billion parameters each) coordinate through shared memory, parallel exploration, and evolutionary optimization to conduct adversarial attacks. In experiments, the swarm achieved a 45.8% Effective Harm Rate against GPT-4o (including 49 critical-severity breaches) and recovered 9 of 9 planted software vulnerabilities in roughly four minutes on consumer hardware — capabilities previously associated with restricted frontier models. The authors argue that the key enabler is the system scaffold rather than any individual model's reasoning capacity, meaning that restricting model releases does not prevent these threats. The paper concludes that AI security policy should therefore target multi-agent systems and scaffolds, not individual models.
- AI policy
- Quality assurance
Research
Strategic commitments shape collective cybersecurity under AI inequality
Adeela Bashir, Zia Ush Shamszaman, Zhao Song et al.
arXiv · 2026-05-10
This paper uses an evolutionary game-theoretic model to study how unequal access to AI-enabled cybersecurity tools affects collective security outcomes in a finite population. It finds that when high-capability AI defence is costly, populations gravitate toward weaker, cheaper protection, sustaining successful attacks. Introducing a small group of 'committed' strong defenders helps but cannot alone stabilise secure outcomes; adding targeted subsidies to those committed defenders significantly boosts strong-defence adoption, suppresses attacks, and improves overall system resilience. The findings offer a theoretical basis for policy interventions—such as subsidising key defenders—to stabilise cybersecurity in AI-driven environments where defensive capabilities are unevenly distributed.
- AI policy
- Enterprise
Research
From Detection to Recovery: Operational Analysis on LLM Pre-training with 504 GPUs
Daemyung Kang, Eunjin Hwang, Hanjeong Lee et al.
arXiv · 2026-05-10
This report provides an empirical operational analysis of a 63-node, 504-GPU NVIDIA B200 production cluster used for LLM pre-training, drawing on 55 days of Prometheus metrics and 73 days of logs across 224 multi-node training sessions involving five organizations. Key findings include: no single monitoring metric reliably predicts all GPU failure types, requiring multi-signal detection; checkpoint I/O bursts reach up to 21.5% of peak read bandwidth and cause measurable NFS/RPC queuing; and node exclusions are highly concentrated, with the top 3 of 63 nodes accounting for over 50% of exclusions. Notably, automated retry chains achieved a 33.3% success rate compared to 12.5% for manual retries, demonstrating a 2.7x improvement that highlights the operational value of automation in large-scale distributed AI training.
- Enterprise
- Quality assurance
Research
Towards Conversational Medical AI with Eyes, Ears and a Voice
Meet Shah, Jason Gusdorf, Anil Palepu et al.
arXiv · 2026-05-10
This paper introduces 'AI co-clinician,' a conversational AI system built on Gemini's audio-visual processing capabilities that participates in real-time telemedicine consultations by continuously interpreting live audio and video streams. The system uses a dual-agent architecture to balance clinical reasoning with low-latency natural dialogue, and was evaluated in a randomized crossover simulation study (n=120 encounters) across 20 standardized outpatient scenarios against primary care physicians, GPT-Realtime, and a baseline agent. Results show AI co-clinician approached primary care physicians in management plans and differential diagnosis while significantly outperforming GPT-Realtime across all general criteria, though physicians maintained superior overall performance in case-specific assessments. The authors argue that text-only approaches fail to capture the real challenges of medical consultation and advocate for collaborative, triadic models where AI serves as a supportive co-clinician rather than a replacement.
- Workforce
- Quality assurance
Research
Factors Shaping Artificial Intelligence Adoption in Small and Medium-Sized Enterprises in Vietnam: A Context-Based Approach
Pham Huy Thong
International Journal of Advanced Multidisciplinary Research and Studies · 2026-05-10
Using survey data from 230 Vietnamese SMEs analyzed via PLS-SEM and the Technology–Organization–Environment framework, this study finds that perceived benefits and top management support are the strongest drivers of AI adoption, while resource constraints act as structural barriers. The research shows that AI adoption among SMEs is not automatic but a strategic decision made under constrained conditions. These findings highlight the uneven and limited uptake of AI in Vietnam's SME sector and offer context-specific policy and management implications.
- Enterprise
- AI policy
- Workforce
Research
<b>The Adoption of AI in Enhancing Business Efficiency</b>
Muheeb Mohamed
American University of Bahrain · 2026-05-10
This study examines what drives AI adoption among small and medium-sized enterprises (SMEs) in Bahrain, distinguishing between firms that intend to adopt AI and those that already have. Using the TOE framework and survey data from 467 managers, the research finds that top management support and government backing are key enablers at both stages, while complexity is the primary barrier for firms yet to adopt. The findings have direct implications for policymakers, SME managers, and AI vendors seeking to reduce adoption friction and strengthen organizational readiness.
- Enterprise
- AI policy
- Workforce
Research
Forking paths of AI governance – how risk management frameworks shape the politics of AI
Leevi Saari, Daniel Mügge
Critical Policy Studies · 2026-05-10
This paper examines how risk management frameworks—specifically the US NIST AI Risk Management Framework, ISO/IEC standards, and OECD harmonization efforts—shape AI governance politics. The authors argue these frameworks are both 'performative,' in that they define and narrow which AI-related concerns are treated as policy-worthy, and 'productive,' in that they enable AI development and deployment by providing an appearance of administrative control even amid uncertainty. The findings suggest that dominant risk frameworks can depoliticize contested questions about AI's societal impact while simultaneously facilitating the spread of AI products. This matters for understanding how technical governance tools embed political choices about what counts as risk.
- AI policy
- Certifications
- Enterprise
Research
Artificial Intelligence Across the Drug Development Lifecycle
Grigory Demyashkin, Mikhail Parshenkov, Sergey Zyryanov et al.
Medical Sciences · 2026-05-10
This review paper examines how AI is being integrated across the full pharmaceutical product lifecycle (PPL), from early drug discovery through nonclinical evaluation, clinical trials, and post-marketing assessment. The authors argue that AI adds the most value when embedded as part of a broader data strategy that links information across all stages, rather than used as a standalone tool. Case studies from leading pharmaceutical companies illustrate meaningful advances in candidate prioritization, safety prediction, cohort formation, and real-time clinical monitoring. The paper emphasizes that transparent, reliable, and scientifically grounded implementation requires continuous attention to emerging methodologies and evolving regulatory frameworks.
- Enterprise
- Quality assurance
- AI policy
- Certifications
Research
Governance, Risk, and Compliance (GRC) Engineering Approaches for IT and Cybersecurity Control Assurance: A Critical Review
William Asare Yirenkyi
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-10
This critical literature review examines GRC engineering approaches for IT and cybersecurity control assurance in U.S. regulated environments from 2020 to 2025, finding that hybrid frameworks such as NIST and COBIT are commonly used to unify governance and risk functions alongside risk-based control design and automation. The analysis shows these approaches strengthen enterprise risk management and sectoral resilience in finance and healthcare, but expose persistent weaknesses including limited adaptability, scalability constraints for smaller entities, insufficient cultural integration, and unresolved contradictions in AI adoption amid fragmented regulations like SOX, HIPAA, and CCPA. Empirical validation of GRC effectiveness remains thin and behavioral dimensions are largely overlooked, leaving gaps in assurance quality and regulatory accountability. The findings matter because they clarify both the contributions and enduring limitations of current GRC engineering in addressing the complexity of U.S. regulated environments.
- AI policy
- Quality assurance
- Certifications
- Enterprise
Research
Governance, Risk, and Compliance (GRC) Engineering Approaches for IT and Cybersecurity Control Assurance: A Critical Review
William Asare Yirenkyi
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-10
This critical literature review (2020–2025) examines how Governance, Risk, and Compliance (GRC) engineering approaches are used for IT and cybersecurity control assurance in U.S. regulated environments. The review finds that hybridizing frameworks such as NIST and COBIT, combined with risk-based control design and automation for monitoring and predictive analytics, can strengthen enterprise risk management and sectoral resilience—especially in finance and healthcare. However, persistent weaknesses remain, including limited adaptability, scalability constraints for smaller entities, insufficient cultural integration, and unresolved contradictions in AI adoption amid fragmented regulations like SOX, HIPAA, and CCPA. The authors conclude that current GRC engineering supports risk-based auditing but falls short of addressing the full complexity of U.S. regulated environments, with empirical validation and behavioral dimensions largely overlooked.
- Enterprise
- Quality assurance
- Certifications
- AI policy
Research
<b>The Adoption of AI in Enhancing Business Efficiency</b>
Muheeb Mohamed
American University of Bahrain · 2026-05-10
This study examines what drives AI adoption among small and medium-sized enterprises (SMEs) in Bahrain, distinguishing between firms that intend to adopt AI and those that have already adopted it. Using the TOE framework and survey data from 467 managers, the research finds that complexity is a barrier for intending adopters, while top management support, government support, relative advantage, and competitive pressure are key enablers for actual adopters. The findings highlight that different factors matter at different stages of adoption, offering practical recommendations for policymakers, SME managers, and AI vendors to reduce barriers and strengthen enablers.
- Enterprise
- Workforce
- AI policy