News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
5608 items
Research
Do These Violent Delights Have Violent Ends? Measuring the Post-Merge Fate of Agentic Code
Chunqiu Steven Xia, Courtney Miller
arXiv · 2026-07-10
This paper conducts a longitudinal empirical study of AI-generated ('agentic') code contributions across 182 repositories, tracking what happens after those contributions are merged into real projects. While overall maintenance rates are comparable to human contributions, agentic code requires significantly higher rates of corrective maintenance and introduces more security weaknesses and dependency vulnerabilities. The study also finds that projects with higher no-review rates see roughly a 6% increase in agentic maintenance burden per 10 percentage-point rise in that rate. The findings suggest that evaluating agentic coding tools should go beyond merge acceptance to also assess long-term security and maintainability outcomes.
- Quality assurance
- Enterprise
- AI policy
Research
What You Train Is What You Get: Gender Bias, Training Composition, and Post-Hoc Mitigation in Audio Deepfake Detection
Aishwarya R. Fursule, Vamshi Nallaguntla, Shruti Kshirsagar et al.
arXiv · 2026-07-10
This paper investigates gender bias in audio deepfake detection models trained on the ASVspoof5 dataset, finding that the gender composition of training data strongly predicts which demographic group performs worse at test time. Using a ResNet18 classifier with two feature types (LogSpectrogram and WavLM-Base+), the researchers show that WavLM-Base+ features produce gender performance gaps 3.0 to 4.3 times larger than LogSpectrogram under identical conditions, and that balanced training reduces LogSpectrogram bias but leaves WavLM bias largely intact. Critically, all six post-hoc threshold calibration methods tested — including an oracle approach with full test-set label access — failed to close the Equal Error Rate gap, confirming that score distribution disparities cannot be corrected after training. These findings indicate that achieving gender fairness in audio deepfake detection requires interventions at training time rather than post-hoc adjustments.
- Quality assurance
- AI policy
Research
ConceptSMILE: Auditing the Trustworthiness of Concept-Based Explainable AI
Mohadeseh Mollapour, Koorosh Aslansefat, Zeinab Dehghani et al.
arXiv · 2026-07-10
ConceptSMILE is a model-agnostic auditing framework that evaluates the reliability of concept-based explainable AI (XAI) explanations by perturbing input regions, measuring concept-response shifts, and fitting a surrogate model to approximate local concept behaviour. Reliability is assessed across five dimensions: attribution accuracy, surrogate fidelity, faithfulness, stability, and consistency. Applied to retinal fundus images, the framework finds that reliability varies by concept and pathway—MedSAM-derived visual concepts achieve higher spatial attribution and surrogate fidelity (R²=0.8503), while VLM-based semantic concepts show stronger vessel faithfulness and stability under certain artefact conditions. This work matters for quality assurance and certification in medical AI, providing an independent audit layer to assess whether concept-level explanations can actually be trusted before deployment.
- Quality assurance
- Certifications
- AI policy
Research
TrustX Agent Risk Classification Framework (ARC): Risk-Tiering Internally Created Agentic AI Systems
Hannah M. Liu, Rhea Saxena, Shiv Asthana
arXiv · 2026-07-10
This paper introduces the TrustX Agent Risk Classification Framework (ARC), a structured instrument designed to classify and govern internally created agentic AI systems in enterprise and public-sector contexts. The framework applies across seven types of agentic AI systems using a twelve-dimension scoring rubric, a GPA + IAT classification model, and a five-level autonomy framework derived from existing literature, producing a three-tier governance output with mapped control recommendations. A specialized extension for Coding Assistants is also included to handle nuances specific to that system type. ARC addresses a gap where general-purpose AI risk frameworks have failed to keep pace with the proliferation of agentic AI, and is intended for AI governance practitioners, risk officers, developers, and regulators.
- Enterprise
- AI policy
- Certifications
Research
Deceptive Grounding: Entity Attribution Failure in Clinical Retrieval-Augmented Generation
Cedric Caruzzo, Donggeun Yoo, Tae Soo Kim
arXiv · 2026-07-10
This paper identifies a new failure mode in clinical retrieval-augmented generation (RAG) systems called 'deceptive grounding' (DG), where a model presents real clinical evidence about one drug as if it applies to a different queried drug. The authors show this failure is invisible to standard hallucination, faithfulness, and citation checks because every claim is sourced from a real document — just the wrong entity. Across a controlled benchmark of 13 models, DG rates range from 8–87% under adversarial conditions, with medical and biomedical fine-tuned models reaching up to 86.7%, meaning domain specialization worsens rather than prevents the problem. In a deployed RAG system evaluated over 740 drug-disease pairs, the overall DG rate was 7.8%, rising to 13.6% for recently approved drugs — a serious patient safety concern that existing evaluation frameworks do not address, though entity-attribution verification detected DG at 97.0% precision and 98.7% recall.
- Quality assurance
- AI policy
- Certifications
Research
Letting the Data Speak: Extracting Keywords from Crowdsourced Collections with AI
Miguel Arana-Catania, Catherine Conisbee, Matthew Kidd
arXiv · 2026-07-10
This paper evaluates three Natural Language Processing approaches—Named Entity Recognition, Keyword Extraction, and Topic Modelling—for automating keyword assignment in crowdsourced digital collections, using the Their Finest Hour Online Archive at the University of Oxford as a case study. Testing methods ranging from traditional statistical techniques to modern generative AI, the authors find that NLP offers real potential for large-scale keyword extraction but that no single method is complete and model choice significantly shapes results. The study also highlights distinct ethical stewardship responsibilities arising when metadata is produced through engagement with living contributors, concluding that open-weight, extractive models are best suited for responsible deployment while generative AI introduces accountability risks that collection managers should carefully weigh.
- Enterprise
- Quality assurance
- AI policy
Research
Geopolitical alignment: Endorsement effects in large language models
Maxim Chupilkin
arXiv · 2026-07-10
This paper investigates whether large language models (LLMs) are implicitly influenced by geopolitical cues when evaluating international economic and security policies. Using an endorsement experiment, the researcher found that GPT-5, Claude Sonnet, and Gemini consistently rated identical policies lower when described as supported by China or Russia compared to the United States or the European Union, with DeepSeek behaving differently. Adding justification prompts left the Western/non-Western gap intact for some models, attenuated it for Gemini, and sharply activated China and Russia penalties in DeepSeek, with justifications revealing that Western endorsement was treated as a credibility signal while Chinese and Russian endorsement triggered concerns about data security, surveillance, and geopolitical risk. These findings matter for policy and enterprise applications because LLM evaluations of policy-relevant information can be systematically shaped by the identity of a foreign endorser even when policy content is held constant.
- AI policy
- Enterprise
- Quality assurance
Research
A Personalized Computational Framework for Assessing the Sufficiency of Partially Observed Data in Healthcare AI models
Qingchu Jin, Felistas Mazhude, Jamie B. Rabb et al.
arXiv · 2026-07-10
This paper introduces Feature Sufficiency Analysis (FSA), a personalized computational framework that determines whether a subset of available clinical features is sufficient for an AI model to achieve the same predictive performance as when all features are present (termed full-feature-capacity, or FFC). By estimating the distributions of missing variables conditioned on available ones, FSA provides patient-specific assessments of whether additional data collection is necessary before making a prediction. The authors demonstrate FSA in two clinical case studies—predicting postoperative prolonged ventilation after heart surgery and 10-year mortality in an outpatient cohort—and show it can also rank features by predictive sufficiency, identify hard-to-predict patient subgroups, and support cost-aware data acquisition. This work matters for healthcare AI deployment by offering a principled method to determine when incomplete clinical data is trustworthy enough to support AI-assisted decision-making.
- Quality assurance
- Enterprise
- AI policy
Research
MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation
Runhan Shi, Quan Zhou, Yuqian Xu et al.
arXiv · 2026-07-10
MedRealMM introduces a large-scale benchmark of 5,620 real-world multimodal medical consultation cases drawn from de-identified patient-doctor interactions at a Chinese internet hospital, spanning 64 clinical departments. Unlike existing benchmarks that rely on synthetic conversations or multiple-choice metrics, it uses a Multimodal Clinical Challenge Point (MCCP) framework to identify clinically demanding moments and evaluates model responses against physician-refined rubrics that reward desirable clinical behaviors and penalize unsafe or contradictory outputs. Evaluation of 19 general-purpose and medical-specialized LLMs finds that image information is critical for reliable clinical performance and that current frontier models fall below online physician response quality, particularly in avoiding safety-sensitive errors. The benchmark highlights that error avoidance remains a central bottleneck for deploying AI in clinical consultation settings.
- Quality assurance
- AI policy
- Enterprise
Research
Beyond Metadata: CAPRA for Hidden Subgroup Analysis under Missing Metadata in Medical Imaging
Yawen Li, Yan Li, Zhe Xue et al.
arXiv · 2026-07-10
CAPRA is a framework for auditing medical imaging AI models when demographic and acquisition metadata are unavailable at deployment. It predicts image-derived semantic axes, calibrates posteriors on a small labeled split, and organizes them into a subgroup interface that reveals disparity patterns and failure modes otherwise masked by strong aggregate performance. Tested across fundus, dermoscopy, and chest radiography datasets, CAPRA uncovers disparities missed by metadata-only slicing and produces subgroup partitions that align more closely with explicit failure axes than baseline approaches. The framework is reusable by downstream robust learners, making hidden subgroup auditing more interpretable and practical in real-world deployment settings.
- Quality assurance
- Certifications
- AI policy
Research
The Admissibility Threshold: A Sector-Specific Certification Standard for High-Stakes AI, with Health as the First Mandatory Domain (Version 2)
Siddiqui Jameel Ahmed
Zenodo (CERN European Organization for Nuclear Research) · 2026-07-10
This paper introduces 'Domain Admissibility,' a proposed certification framework requiring AI systems to demonstrate fitness for a specific high-stakes domain before deployment, rather than relying on uniform horizontal governance standards. The authors argue that existing medical-device regulation leaves a critical gap for general-purpose AI entering clinical settings without any formal approval pathway, and they propose four testable pillars of admissibility along with a certification threshold. The framework is illustrated through a clinical case study, with health proposed as the first domain where the threshold should be mandatory. The work aims to lay constitutional groundwork for future domain-specific AI certification institutions.
- Certifications
- AI policy
- Quality assurance
Research
AI-related disclosure intensity and financial transparency: evidence from Chinese listed companies
Yulun Gu
Cogent Business & Management · 2026-07-10
Analyzing a balanced panel of 1,028 Chinese firms listed on the Shanghai and Shenzhen Stock Exchanges from 2011 to 2023, this study finds that greater AI-related disclosure intensity in annual reports is associated with higher external financial transparency ratings. The relationship holds after controlling for firm characteristics, year effects, and alternative model specifications, and is particularly pronounced for smaller firms and non-state-owned enterprises. These findings suggest that AI-related corporate communications may serve as a meaningful signal of disclosure quality, with implications for how investors and regulators assess financial transparency in China's rapidly growing AI market.
- Enterprise
- AI policy
Research
Artificial Intelligence and Labor Productivity in Construction: A Comparative Systems Analysis Across European Economies
Claudiu George Bocean, Adriana Scrioșteanu, Sorina Gîrboveanu et al.
Systems · 2026-07-10
This study examines how AI adoption affects labor productivity in the construction sectors of EU countries using 2023–2024 data. Employing multivariate log-linear regressions and cluster analysis, the researchers find that construction output is the primary driver of labor productivity, while AI adoption shows a small, negative association with productivity, interpreted as short-term adjustment costs during early digital transformation. Cluster analysis reveals diverse country profiles in AI use, productivity, and labor intensity, suggesting that blanket digital adoption policies may be insufficient. The authors conclude that gradual, organization- and skills-focused change management is needed to realize long-term productivity gains from AI in construction.
- Workforce
- Enterprise
- AI policy
Research
The process of creating an artificial intelligence-based agent for solving management ophthalmology tasks.
A. I. Bursov, A. V. Belogurova
Manager Zdravookhranenia · 2026-07-10
This paper describes the development of an AI agent designed to automate competitive intelligence for private ophthalmology clinics, replacing manual monitoring of competitor websites and price lists. Built on Python, LangChain, retrieval-augmented generation (RAG), and locally deployed large language models (Qwen 2.5 and Llama 3.1), the agent crawls clinic websites, structures data, and generates comparative reports in response to management queries. Testing showed that typical query execution time dropped from approximately 2 hours of manual work to under 12 minutes (as low as 1.5–2 minutes with cloud models), demonstrating significant labor cost reduction. The authors argue the framework is adaptable to medical organizations beyond ophthalmology and supports more informed, timely managerial decision-making.
- Enterprise
- Workforce
Research
From AI Use to Sustainable Value Creation Through Entrepreneurial Reconfiguration and Business Model Innovation in SMEs: Evidence from an Emerging Economy
Alexander Sánchez-Rodríguez, Jesús Rodríguez-Flores, Reyner Pérez-Campdesuñer et al.
Sustainability · 2026-07-10
This study examines how AI use translates into sustainable value creation for small and medium-sized enterprises (SMEs) in Ecuador, using survey data from 385 firms across four sectors. Using PLS-SEM, the researchers found positive associations between AI use, entrepreneurial reconfiguration capability, business model innovation, and sustainable value creation, suggesting that AI adoption alone is insufficient without developing the organizational capabilities to connect AI with new business models. The findings are particularly relevant to emerging-economy contexts where AI adoption is often fragmented. The study shifts focus from mere AI adoption to AI-enabled entrepreneurial transformation as the pathway to measurable business impact.
- Enterprise
- Workforce
Research
The Admissibility Threshold: A Sector-Specific Certification Standard for High-Stakes AI, with Health as the First Mandatory Domain (Version 2)
Siddiqui Jameel Ahmed
Zenodo (CERN European Organization for Nuclear Research) · 2026-07-10
This paper introduces 'Domain Admissibility,' a sector-specific certification framework requiring AI systems to prove fitness for a particular high-stakes domain before deployment, rather than satisfying only generic horizontal governance standards. The authors argue that uniform AI governance is inadequate when error costs vary drastically across domains—most critically in healthcare, where mistakes can be irreversible and fatal. They define four testable pillars of admissibility, propose a certification threshold for consequential settings, and make the case that health should be the first domain where this threshold becomes mandatory. The work positions itself as a foundational constitutional framework on which future certification institutions could be built, addressing a gap left by existing medical-device regulation that does not cover general-purpose AI entering clinical settings.
- Certifications
- AI policy
- Quality assurance
Research
Runtime assurance for enterprise agentic AI systems: A policy-gated control model with quantitative autonomy-risk scoring
Kwan Hong Tan
World Journal of Advanced Research and Reviews · 2026-07-10
This paper presents a Runtime Assurance Architecture (RAA) for enterprise agentic AI systems that introduces a quantitative Autonomy-Risk Exposure (ARE) score to govern when AI agents can act autonomously, require sandboxing, need human approval, or must be blocked. Evaluated across 2,000 simulated enterprise agent episodes, the full RAA configuration reduced mean ARE scores by 31.5%, cut policy-conflicting actions from 9.8% to 4.6%, eliminated unsupervised high-risk pass-throughs, and improved audit evidence coverage from 0.61 to 0.89—all with a mean latency overhead of 95 ms. The findings are directly relevant to organizations deploying agentic AI in regulated or high-consequence workflows, offering a practical reference architecture and policy decision algorithm grounded in AI risk-management standards and security guidance.
- Enterprise
- Quality assurance
- Certifications
- AI policy
Research
The Patchwork Problem in LLM-Generated Code
Viraaji Mothukuri, Reza M. Parizi
arXiv · 2026-07-09
This paper identifies and formalizes what it calls the 'patchwork problem': LLM-generated code that compiles and passes tests but is globally incoherent due to structural failures such as missing configuration keys, nonexistent packages, or omitted authentication guards. The authors model structural coherence as consistency invariants over graph representations of repository artifacts and introduce an eight-category failure taxonomy, distinguishing defects unique to LLM generation from those merely amplified by it. Their hybrid verification framework combines mature static analysis tools with purpose-built detectors, and empirical evaluation shows the vast majority of these structural failures evade type checking, testing, and existing SAST tools entirely. External validation on real-world AI-generated repositories confirms these failures are widespread wherever LLMs write code with minimal human oversight, posing a significant and growing risk to software quality.
- Quality assurance
- Enterprise
- Workforce
Research
Eluna: An Agentic LLM System for Automating Warehouse Operations with Reasoning and Task Execution
Ning Liu, Kalle Kujanpää, Zhaoxuan Zhu et al.
arXiv · 2026-07-09
Eluna is a production-deployed agentic AI system designed to automate warehouse operations by encoding Standard Operating Procedures (SOPs) as directed acyclic graphs within a multi-agent framework that enforces procedural compliance and handles complex, multi-system decision logic under strict time constraints. The system uses asymmetric episodic distillation, where a strong teacher model is refined through episodic error memories and a smaller student model is fine-tuned on corrected trajectories, allowing smaller models to match or exceed larger off-the-shelf baselines without inference-time overhead. On a 13-task benchmark and two production applications, Eluna's fine-tuned models achieve 94% expert agreement on a ticket processing application, demonstrating reliable SOP execution at scale. This work is significant for enterprise warehouse automation, showing that purpose-built agentic frameworks can outperform general large language models on complex operational tasks.
- Enterprise
- Workforce
- Quality assurance
Research
Trivial Prompt Reframing Bypasses Safety Guardrails in Googleś MedGemma-4B
Avi-ad Avraam Buskila
arXiv · 2026-07-09
This paper evaluates the safety guardrails of MedGemma-4B-it, Google's open-weight medical language model, by testing whether simple, non-technical prompt reformulations can bypass restrictions the model card prohibits—such as recommending drug dosages, issuing diagnoses, or advising patients to skip emergency care. Using a benchmark of 4,500 generated responses across five guarded behaviors, six attack styles, and three judge methods, the authors find an overall Attack Success Rate of 38.0%, with reframing a question as a 'medical board exam' item raising success rates to 53.1% and appeals to claimed doctor authority reaching 43.7%. The drug-interaction guardrail proved nearly absent at 83.2% attack success, while emergency-deferral was more robust at 4.7% but still breachable via authority framing. The findings demonstrate a substantial gap between model card intent and actual robustness, motivating stronger deployment-time safety measures for open medical AI models.
- Quality assurance
- AI policy
- Certifications
Research
L2-Bench: An Evaluation Benchmark for Measuring LLM Capabilities in Second Language Education
James Edgell, Wm. Matthew Kennedy, Ben Knight et al.
arXiv · 2026-07-09
L2-Bench introduces an open-source benchmark of 1,000+ task-response pairs designed to evaluate large language models on their ability to apply second language (L2) education principles, not merely recall them. The benchmark includes a taxonomy of 12 competencies and 31 subcompetencies validated by 200+ expert practitioners, alongside a rubric-based evaluation methodology intended to generalize to other open-ended, qualitative educational domains. Among tested models, Claude Opus 4.7 performs best overall at 85.5%, though performance drops notably on harder tasks. The work equips education stakeholders with more rigorous tools for making informed decisions about adopting, using, and governing AI-powered educational systems.
- Quality assurance
- Certifications
- AI policy
- Enterprise
Research
Trusting sovereign language models as scientific instruments: evidence from Portugal's AMALIA
Manuel Pita
arXiv · 2026-07-09
This paper audits whether nationally sovereign language models can be trusted as valid scientific measurement instruments, using Portugal's publicly funded AMALIA model as a test case. The authors introduce the 'recovery gap' metric, which checks whether a model's coding performance can actually be attributed to the theoretical construct it is supposed to measure—rather than surface-level correlates—by decomposing a codebook into theory-defined clauses and recombining them. Applied to moral foundation coding in European Portuguese, AMALIA achieves competitive agreement with human coders but only about half of its coding performance on the authority foundation can be traced back to the underlying theory, while a larger multilingual model closes this gap. The findings argue that public ownership and linguistic specialization earn operational trust but not epistemic trust, and that the proposed audit method is inexpensive and portable across models, languages, and tasks.
- Quality assurance
- AI policy
- Certifications
Research
SolarChain-Eval: A Physics-Constrained Benchmark for Trustworthy Economic Agents in Decentralized Energy Markets
Shilin Ou, Yifan Xu, Luyao Zhang
arXiv · 2026-07-09
SolarChain-Eval introduces a physics-constrained benchmark for evaluating autonomous AI agents operating in decentralized energy markets, framing market governance as a Markov Decision Process assessed across dimensions such as market utility, physical safety, slippage, action smoothness, spatial fairness, and auditability. The benchmark incorporates an LLM-based Planner/Auditor layer that defines action bounds, reviews high-risk decisions, and logs all interventions with structured audit traces. Experiments comparing static, random, myopic, RL, and RL+LLM policies reveal a clear utility-safety trade-off: RL agents boost market utility but can still produce unsafe behavior, and removing physics penalties causes reward-maximizing agents to exploit invalid generation data and inflate artificial liquidity. The findings demonstrate that trustworthy agentic AI evaluation in cyber-physical settings requires both hard physical constraints and transparent intervention records, with implications for quality assurance and policy governance of autonomous economic systems.
- Quality assurance
- AI policy
- Enterprise
Research
The complexities of patient-centred conversational artificial intelligence
João Matos, Olivia Buege, Donny Cheung et al.
arXiv · 2026-07-09
This paper analyzes 2,053 real patient-chatbot conversations to show that communication patterns and emotional expression vary widely among users—far beyond what cooperative, idealized simulated patients capture. The researchers built a patient simulator modeling clinical content, emotional state, conversational strategy, and communication style, producing conversations so realistic that human graders achieved only 55% accuracy distinguishing them from real ones. Testing four large language models across 1,164 clinician-graded cases using five distinct patient personae, they found that communication style significantly alters triage outcomes. The findings warn that health chatbots designed for idealized interactions risk underperforming and amplifying health disparities when deployed with real, diverse patients.
- Quality assurance
- AI policy
- Workforce
Research
Towards Precision Therapy in Hepatocellular Carcinoma: A Clinical-Reasoning LLM for Risk Stratification and Treatment Guidance
Peng Cui, Jitao Wang, Siyan Xue et al.
arXiv · 2026-07-09
HCC-STAR is a large language model designed to support clinical decision-making in hepatocellular carcinoma (HCC) by reading electronic medical record narratives and jointly producing risk-based staging, ranked treatment recommendations, and individualized survival estimates. Trained on approximately 30,000 SEER-derived cases expanded into EMR-style narratives, the model was evaluated on a multi-center cohort of 6,668 patients across 12 hospitals in China, achieving state-of-the-art performance compared to clinical guidelines and leading models including GPT-5 and Gemini-2.5 Pro. Hypothetical overall-survival analysis showed a median survival of 51 months under HCC-STAR recommendations versus 29 and 32 months under BCLC and CNLC guidelines respectively. Blinded hepatobiliary specialists rated HCC-STAR's reasoning as trustworthy, and the model outperformed resident and attending physicians in treatment accuracy while helping clinicians make more accurate decisions faster.
- Workforce
- Enterprise
- Quality assurance