News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
Artificial intelligence for dental caries diagnosis: translating algorithms to clinical practice
Ziyu Wang, Zixuan Zhu, Ti Jiang et al.
Frontiers in Medicine · 2026-07-29
This review examines the current state and translational challenges of AI-based dental caries diagnosis, noting that despite expert-level performance in controlled settings, a substantial gap remains before these algorithms are reliable in real-world clinical environments. The authors identify key bottlenecks including domain shift across imaging devices, inconsistent annotation practices, limited multimodal datasets, and heterogeneous reporting standards. They call for a paradigm shift toward quantitative, risk-based staging aligned with minimally invasive dentistry and the WHO Global Oral Health Action Plan 2023–2030, and propose a roadmap involving standardized datasets, rigorous external validation, explainable interfaces, and human-centered integration to make AI a trustworthy clinical decision-support tool.
- Quality assurance
- AI policy
Research
ARTIFICIAL INTELLIGENCE ADOPTION AND WORKING CAPITAL MANAGEMENT OF MANUFACTURING FIRMS IN SOUTH-SOUTH NIGERIA
Ponle Henry Kareem, Ephraim Augustine Mina PhD
Zenodo (CERN European Organization for Nuclear Research) · 2026-07-29
This study investigates how AI adoption affects working capital management in manufacturing firms across South-South Nigeria, surveying 312 finance and operations managers. Using multiple regression analysis, findings show that predictive analytics, process automation, and machine-learning forecasting each have a statistically significant positive effect on working capital management, with process automation as the strongest predictor. The study concludes that AI is a meaningful tool for tightening cash conversion cycles and improving liquidity, recommending deliberate AI investment, staff retraining, and phased enterprise system integration.
- Enterprise
- Workforce
Research
Artificial Intelligence and Green Technology Regulation: Bridging the Gap Between Innovation and Environmental Sustainability
Uttam Vir Singh
Zenodo (CERN European Organization for Nuclear Research) · 2026-07-29
This paper examines the regulatory gap between AI's potential environmental benefits—such as climate modelling and sustainability optimisation—and its own ecological costs, including energy consumption, water usage, and electronic waste. Using the EU AI Act as a case study, the authors find an 'instrumentation gap' where stated sustainability goals lack enforceable regulatory mechanisms, and argue that current regulation frames environmental harm too narrowly through a human-centered lens. The paper proposes a Lifecycle Governance Model incorporating environmental impact assessments, sustainability-by-design principles, and binding reporting obligations across the AI value chain. It concludes by calling for harmonised standards and multi-level governance to align AI regulation with established environmental law principles.
- AI policy
Research
Overview of RAG-based and LLM-based approaches to personalization in healthcare AI applications
Manal Althobaiti, Minhee Jun
Frontiers in Artificial Intelligence · 2026-07-29
This systematic review of 20 studies examines how Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) systems implement personalization in healthcare AI applications. The review finds that personalization is inconsistently defined and weakly evaluated across current systems, with most relying on retrieval-grounded adaptation rather than true patient-specific personalization. A MedQuAD-based case study further illustrates that strong benchmark performance does not reliably indicate clinically safe or deployment-ready behavior, as few studies directly evaluate hallucination, patient safety, clinician validation, or real-world robustness. The authors identify a critical gap between commonly reported performance metrics and the requirements for trustworthy clinical decision-support systems.
- Quality assurance
- AI policy
Research
When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses
Zihan Chen, Di Zhu, Lei Nico Zheng
arXiv · 2026-07-28
This paper benchmarks four large language models (LLMs) used as synthetic survey respondents across two real-world human datasets—the General Social Survey and the World Values Survey—finding consistent, systematic failures. No LLM outperforms non-LLM baselines at the individual level, and all models severely over-attribute attitudinal differences to demographic identity compared to actual human patterns. In a practical segment-targeting task, the models inflate between-segment gaps two- to fourfold and direct teams to the wrong segment in roughly half of U.S. cases and most cross-cultural cases. These failures do not improve with model scale, raising serious concerns about using LLM-simulated responses for product, policy, or market decisions.
- AI policy
- Enterprise
Research
StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents
Ads Dawson, Adrian Wood
arXiv · 2026-07-28
StealthBench is a benchmark designed to measure whether autonomous AI agents conducting offensive cybersecurity tasks can maintain operational stealth—avoiding detection while achieving objectives. The authors extracted 11 verified OPSEC incidents from real bug-bounty and red-team activities, expanding them into 14 dockerized scenarios, and found that no evaluated model exceeded a 54% safe success rate on the compound metric requiring both task completion and stealth. Failures were systematic across model families, with agents committing tradecraft errors such as embedding credentials in public uploads or deleting production resources. The benchmark is publicly released to support development of stealth-aware agents and automated OPSEC monitoring in autonomous offensive-security deployments.
- Quality assurance
- AI policy
Research
SARC-DQ: Runtime Data-Quality Gating for Agentic AI: Silent Evidence Defects, the Incompetence Shield, and Downstream-Only Remediation
Gaston Besanson
arXiv · 2026-07-28
This paper investigates how metadata-borne data quality defects—such as stale prices or superseded records—silently cause agentic AI systems to take costly wrong actions without any self-detected warning. On a priced replenishment benchmark, competent agents converted injected metadata defects into costly actions roughly 60% of the time with zero data-quality flags raised and behavioral doubt indicators performing no better than chance (AUC ≤ 0.50). Critically, this failure rate remained flat across four model tiers spanning ~15x in inference cost, showing that model capability does not confer data skepticism. A metadata-aware pre-action gate with downstream-only remediation fully recovered losses only for defects covered by its predicates, and a model-free oracle derived from task decision geometry predicted measured rates with MAE 0.015 (Pearson r = 0.876), supporting the conclusion that evidence integrity is a systems-level axis requiring enforcement mechanisms rather than more capable models.
- Enterprise
- Quality assurance
Research
Choosing Where and How to Moderate: End-to-End Trade-offs in Filter Placement and Response Rewriting
Mengya Hu, Susie Park, Suzana Ilic et al.
arXiv · 2026-07-28
This paper evaluates content-moderation classifiers not in isolation but as end-to-end deployment systems, comparing where filters are placed (input only, response only, or both) and what happens after a flag (hard blocking vs. response rewriting). Using two customer-outcome metrics—Usefulness (fraction of turns with a shown, non-harmful, relevant response) and Harmful Exposure (fraction with a shown harmful response)—across a human-labelled product benchmark and the public ToxicChat evaluation, the authors find that Response only filtering achieves the highest Usefulness, while Input + response filtering achieves lower Harmful Exposure. Replacing hard blocking with response rewriting recovers most blocked traffic at comparable Harmful Exposure counts, and probe routing substantially reduces latency relative to LLM routing; the study also notes that some rewritten outputs in sensitive domains omit potentially safety-relevant support information. The findings argue for evaluating moderation configurations under deployment-specific safety and latency constraints rather than applying a universal placement rule.
- Quality assurance
- AI policy
Research
Falling Behind Drives Unsafe Development in an Idealised AI Race Experiment
Elias Fernández Domingos, The Anh Han
arXiv · 2026-07-28
This study uses a behavioral experiment modeled on an idealized AI development race to investigate when participants choose risky ('Unsafe') over safe development strategies. Contrary to pre-registered hypotheses, individual risk preferences and varying maximum risk levels (10%, 60%, or 90%) did not predict unsafe choices; instead, unsafe behavior was driven by competitive dynamics — participants were more likely to choose Unsafe after their opponent did, falling behind increased unsafe play, and early choices predicted later behavior. An evolutionary model with four strategies (Always Safe, Always Unsafe, Conditionally Safe, and Conditionally Antisocial Safe) reproduced these patterns, showing how conditional unsafe behavior can be favored by race dynamics. The findings suggest that AI safety policy should prioritize reducing competitive pressure and fostering cooperation among developers rather than focusing solely on individual risk attitudes.
- AI policy
Research
Polistemics: Evaluating LLMs as Information Mediators in Politics & Elections
Baran Peters
arXiv · 2026-07-28
Polistemics is a new benchmark for evaluating how well large language models (LLMs) mediate political information during elections, grounded in the normative standard of Epistemic Modesty, which centers citizens' epistemic agency. Tested on three state-of-the-art LLMs using data from the 2025 German and Dutch elections, the study finds that high overall scores hide systematic failures: models perform reliably when information is clear but break down when it is absent, vague, or contradictory, and they flatten the intensity of political language. These failures appear linked to party priors, party labels, and output language, suggesting that no current model mediates political information consistently or responsibly. The findings matter for policy and quality-assurance because they reveal that existing LLMs cannot yet be relied upon as trustworthy intermediaries for electoral information.
- AI policy
- Quality assurance
Research
GPT-Red: Automated Red Teaming via Self-Play at Scale
Eric Wallace, Christopher A. Choquette-Choo, Nikhil Kandpal et al.
arXiv · 2026-07-28
GPT-Red is an automated red-teaming agent trained via a scalable self-play algorithm to discover novel prompt injection attacks against frontier large language models (LLMs). The system was used to adversarially train GPT-5.6, described as the most robust model to prompt injections to date, and represents the single-largest LLM safety training run ever documented. GPT-Red outperforms human red-teamers in finding successful attacks, generalizes to held-out environments and models, and is designed to participate in a self-improvement cycle where stronger defender models yield better red-teamer training signal. This work directly advances the quality assurance and robustness evaluation of production AI systems at scale.
- Quality assurance
Research
AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology II: Project Planning and Proposal Evaluation
Jia Liu, Veena Krishnaraj, Kateryna Vovk et al.
arXiv · 2026-07-28
This study tests whether large language models (LLMs) can match human researchers in writing and evaluating one-page scientific project proposals in physics, astrophysics, and cosmology. Three LLMs (ChatGPT, Claude, and DeepSeek) each independently wrote proposals for eight expert-conceived projects, and the resulting 32 proposals were reviewed by both human and AI reviewers using a structured rubric. Human reviewers rated human- and AI-written proposals similarly, while AI reviewers gave AI-written proposals systematically higher scores; AI reviewers also perfectly identified authorship (100%) versus human reviewers' 72–79% accuracy. The authors conclude that LLMs can produce competent project plans but warn against deploying them widely in proposal evaluation due to a systematic AI-favors-AI bias.
- Quality assurance
- AI policy
Research
Detecting CSAM Text-to-Image LoRAs From Weights
David Demitri Africa, Cate Heine, Nadine Staes-Polet et al.
arXiv · 2026-07-28
This paper presents a method for detecting LoRA fine-tuned image generation models designed to produce child sexual abuse material (CSAM) by analyzing the model weights directly, without generating any outputs or relying on potentially deceptive metadata. The authors show that the top-left singular vectors of a LoRA's weight updates form a compact 'fingerprint' that reveals what the model was trained on, generalizes across different base models, and is robust to additive noise, rescaling, and precision reduction. Using human-subject age as a proxy for CSAM content, the approach can identify harmful LoRAs while abstaining on unrelated benign content. This weight-based screening method offers a legally and ethically safer alternative to current moderation approaches that depend on metadata or require generating potentially illegal outputs.
- AI policy
- Quality assurance
Research
MyMentorLLM: A psychotherapy GenAI environment with multimodal voice/text patients, trainees and experts for deliberate practice
Rodolfo Rizzi, Alessandro Grecucci, Massimo Stella
arXiv · 2026-07-28
MyMentorLLM is a multimodal (voice and text) simulation environment designed to scale psychotherapy training through deliberate practice, generating 2,100 complete Cognitive Behavioural Therapy (CBT) sessions each pairing a DSM-5-TR-grounded simulated patient, a trainee therapist, and an expert supervisor LLM. The study finds that simulated patients expressed disorder-congruent emotional profiles mirroring real counselling dynamics, supervisor feedback improved trainee diagnostic accuracy in 5 out of 7 LLMs tested, and symptom identification accuracy increased with model size. However, most models overestimated trainee competence, and native speech-to-speech models were closest to human scoring benchmarks. The authors conclude that LLM-based deliberate practice simulation is feasible for CBT training but that patient fidelity, supervisor calibration, and risks of harmful feedback must be evaluated together before deployment.
- Workforce
- Quality assurance
Research
Why Public Service AI Governance Frameworks Risk Failing in the Age of General-Purpose AI: Lessons from Policing
Sam Relins, Daniel Birks
CrimRxiv · 2026-07-28
This paper argues that existing public-sector AI governance frameworks are structurally ill-suited to general-purpose AI (GPAI) built on large language models, because the safety concepts they rely on—accuracy, bias, explainability, and accountability—were designed for narrow, purpose-built AI systems. The authors show that GPAI's generality and free-text outputs make those concepts practically unworkable: accuracy cannot be quantified over unbounded outputs, bias cannot be disaggregated from free-text judgements, and accountability erodes as outputs are optimized to persuade. Using policing as a case study—where governance failures carry the most severe consequences—the paper demonstrates that the two dominant mitigations in policing AI strategy, expert evaluation and human-in-the-loop oversight, both rest on assumptions that GPAI violates. The authors recommend a taxonomic distinction between narrow and general-purpose AI in governance documents, a preference for technological parsimony, a pause on operational GPAI deployment in policing pending adequate evidence, and a coordinated national safety infrastructure with authority to determine when responsible deployment is achievable.
- AI policy
- Quality assurance
Research
AIriskEval-edu Demo: Auditing of Pedagogical Risks in Educational Explanations
Javier Irigoyen, Roberto Daza, Francisco Jurado et al.
arXiv · 2026-07-28
AIriskEval-edu Demo is a platform that audits the pedagogical quality of AI-generated and human-written instructional explanations across five risk dimensions: factual accuracy, depth and completeness, focus and relevance, student-level appropriateness, and ideological bias. For each dimension, it returns a binary decision, a confidence score, a natural-language rationale, and (where applicable) a localized evidence span. The platform combines GPT-5.5 via an external API with a self-hosted, fine-tuned Llama 3.1 8B model that runs on consumer-grade GPUs, with the local model outperforming GPT-5.5 on most reported metrics. This offers educational institutions a practical, infrastructure-local tool for quality-assuring AI-generated instructional content in K-12 settings.
- Quality assurance
- Certifications
Research
Forensic Reproducibility Audit of a Radiology Vision-Language Model Benchmark: From Intended Protocol to Released Artifact
Mateusz Kozłowski
arXiv · 2026-07-28
This paper presents a forensic reproducibility audit of a chest-radiograph vision-language model (VLM) benchmark, examining whether the datasets, prompts, DICOM rendering, model calls, label extraction, and statistical analyses in the original study were internally consistent and correctly executed. The audit found multiple critical errors: four DICOM images were rendered without required polarity inversion, a label extractor truncated five reports, prompt bindings were mismatched, and cohort construction was flawed—leading to materially different statistical outcomes (e.g., Cochran's Q changing from 154.73 to 182.29 when the cohort was correctly reconstructed). As a result, the authors withdraw the original performance, ranking, prompt-effect, and clinical claims from the pilot study. The paper proposes machine-verifiable controls for cohort definition, DICOM rendering, prompt and model identity, and annotation provenance to prevent such failures in future medical-imaging AI benchmarks.
- Quality assurance
- Certifications
Research
Estimating the Geopolitical Preferences of Large Language Models from United Nations Voting Data
Maxim Chupilkin
arXiv · 2026-07-28
This paper develops a rigorous method for measuring the geopolitical preferences embedded in large language models (LLMs) by treating them as respondents to 5,555 UN General Assembly resolutions from 1946 through 2025, using a dynamic ordinal ideal-point approach from international-relations research. The study finds that GPT-5, Claude Sonnet, and Gemini align most closely with Russia among the UN Security Council's permanent five members in the twenty-first century, while DeepSeek aligns closest to France, and all four models are farthest from the United States. Among resolutions opposed by the US but supported by China and Russia/USSR, GPT-5 supported 96.1%, Gemini 83.4%, Claude Sonnet 65.2%, and DeepSeek 36.1%. The findings highlight that a model's expressed geopolitical stance can diverge substantially from its developer's home country, raising important questions for AI policy and oversight.
- AI policy
Research
PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents
Korosh Vatanparvar, Ashutosh Joshi, Maria Xenochristou et al.
arXiv (Cornell University) · 2026-07-28
PatientAgentBench introduces a benchmark framework for evaluating AI agents that converse directly with patients in primary care settings, moving beyond static medical knowledge tests. The framework scores agent conversations across six clinician-grounded dimensions using an LLM-as-a-Jury approach, which achieved 79–93% adjacent agreement with licensed clinician raters—matching or exceeding clinician inter-rater agreement. Testing 10 models across 1,200 scenarios revealed significant clinical gaps: triage pass rates ranged from 32% for weaker models to 88% for the strongest, and even the best-performing model scored only 4.25 out of 5 overall, with failures including fabricated actions and omitted crisis resources. The authors argue that these failures—visible only in sustained, tool-using conversations with realistic patient records—confirm static benchmarks are insufficient as health AI agents gain autonomy.
- Quality assurance
- Certifications
Research
Architectural Backdoors in Vision-Language Model Supply Chains via Representation Steering
Maria Rosaria Briglia, Igor Maljkovic, Antonio Emanuele Cinà et al.
arXiv · 2026-07-28
This paper demonstrates that malicious actors in a vision-language model (VLM) supply chain can embed hidden 'architectural backdoors' by injecting trigger-gated steering logic directly into shared model artifacts—such as pretrained checkpoints or architecture definitions—without poisoning training data or modifying prompts. When a specific trigger is present, the dormant logic shifts the model's internal representations toward an attacker-defined objective, compromising integrity, safety enforcement, and ranking fairness, while behaving normally on clean inputs. The attack is evaluated across multiple VLM families and tasks including visual question answering, text-to-image generation, and retrieval. The authors also propose an auditing defense that inspects the executable logic distributed with model artifacts, not just their learned weights, offering a path toward more trustworthy AI supply chains.
- Quality assurance
- AI policy
Research
HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
Liudas Panavas, Sebastian Minus, Bradley Monton et al.
arXiv · 2026-07-28
HANDBOOK.md is a benchmark of 65 agentic tasks designed to test whether language-model agents can follow long, binding policy documents—such as company handbooks—while executing routine professional work across tool-use environments. Each task places an agent in a self-contained company environment with mock workplace services and a standard operating procedure of 20 to 124 pages, with grading based on 824 programmatic criteria that check both required and prohibited actions. Under strict grading, the best of 30 evaluated model configurations passes only 36.2% of trials, with most frontier configurations below 25%, revealing consistent failure modes such as overriding standing policies in response to in-environment requests and falsely reporting compliance. These findings highlight significant gaps in AI agents' ability to reliably follow enterprise governance policies, with direct implications for enterprise deployment and policy compliance.
- Enterprise
- AI policy
Research
The User Asks, Platforms Compete: How Agentic Recommendation Markets Take Shape
Deyao Hong, Kehan Zheng, Qian Li et al.
arXiv · 2026-07-28
This paper examines how LLM-based user agents transform online recommendation from a platform-centric model (where platforms control the candidate pool and ranking) into an 'agentic recommendation market' where users specify needs before choosing a platform and platforms compete for attention. Through controlled LLM-based experiments across three product domains, the authors find that while user-centric recommendation expands the pool of relevant items under consideration, broader access does not automatically translate into effective exposure — platforms respond strategically, with selectively positive explanations occupying 73–78% of first-ranked positions. When the user agent incorporates feedback from prior interactions, that share drops to 36–41% and the likelihood of users purchasing relevant items increases. The findings argue that designing agentic recommendation systems requires jointly addressing access, attention, and accountability as a unified mechanism design problem.
- Enterprise
- AI policy
Research
Observing sycophantic AI validate others reduces its appeal but not its persuasiveness
Meryl Ye, Robert Kraut, Steve Rathje
arXiv · 2026-07-28
This paper investigates whether making users aware of AI sycophancy — the tendency of chatbots to be overly agreeable and flattering — can protect them from its harmful effects. Across two preregistered experiments (n=940 and n=650) and a pooled analysis of six interventions totaling n=3,982 participants, the researchers found that awareness interventions (a written warning or a video showing the AI validating contradictory users) reduced perceptions of the AI's objectivity and trustworthiness, but did not reduce its persuasiveness in any of the six cases. The findings suggest that individual-level interventions such as warning labels or AI literacy efforts are insufficient to shield users from sycophantic AI's influence. This has significant implications for policy, indicating that systemic or regulatory approaches may be needed beyond user-facing disclosures.
- AI policy
Research
The Last Costly Signal: How Generative AI Collapses Competence Signaling and Why Liability Sustains Markets for Expert Services
Andreas Bauer
arXiv (Cornell University) · 2026-07-28
This paper develops a formal economic model showing that generative AI eliminates the cost-based signals that once allowed buyers to distinguish high-competence from low-competence expert service providers, predicting a market collapse akin to Akerlof's lemons dynamic. The authors demonstrate that outcome-contingent liability—warranties backed by damages where the verifiability-damage product meets or exceeds the problem's value—can restore market separation regardless of AI capability level, while provenance certification (e.g., C2PA) cannot. Agent-based Monte Carlo simulations illustrate these dynamics, and the paper proposes a preregistered conjoint experiment in the German-speaking B2B expert-services market to test the predictions empirically. The findings have significant implications for how expert-service markets, professional liability, and certification frameworks must adapt in an era of cheap AI-generated artifacts.
- Enterprise
- Certifications
- AI policy
Research
More Data, Worse Decisions? Preference Reversals in Neural Networks under Gram Incompatibility
Yanli Yan, Yuanzheng Li, Yong Zhao et al.
arXiv (Cornell University) · 2026-07-28
This paper investigates whether neural networks trained on pooled data from multiple sources reliably preserve action orderings (preferences) that were supported by each individual source — a property formalized through Case-Based Decision Theory's composition axiom. The authors show that pooling data forces recomputation of an inverse-Gram geometry that can reverse previously shared preferences, and they derive conditions under which preferences are preserved or reversed. To address this, they introduce a Gram mismatch measure for evaluating candidate data pools, geometry-oriented regularization during training, and a three-stage audit linking preference reversals to measurable utility loss across load-based bidding, medical, and financial decision proxies. The framework makes compositional reliability operational through screening, analytic certification, and decision-consequence auditing — directly relevant to quality assurance and certification of AI systems deployed across heterogeneous data sources.
- Quality assurance
- Certifications