News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
5608 items
Research
When Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk Space
Weimeng Wang, Ziqiang Wang, Zihang Zhan et al.
arXiv · 2026-07-16
This paper investigates whether physical danger—when linguistically benign instructions become unsafe once acted upon by embodied AI agents—is a distinct safety problem from ordinary text-level content danger. Using hidden-state direction analysis across multiple LLMs (Qwen2.5, Phi-3.5, SmolLM2), the authors show that content danger and physical danger form separable signals in model representations. They propose PRISM, a lightweight single-layer logistic probe that achieves 86.2–87.7% accuracy on SafeAgentBench with far lower false-positive rates than same-scale LLM judges, which over-block safe tasks at 24.7–39.0% FPR. They also introduce PSB-1K, a 1,000-pair contrastive benchmark for physically grounded risk detection without explicit harm keywords, where PRISM reaches 99.6% accuracy versus a 67.8% safe-task rejection rate for a baseline LLM judge.
- Quality assurance
- AI policy
- Enterprise
Research
Symbal: Detecting Systematic Misalignments in Model-Generated Captions
Maya Varma, Jean-Benoit Delbrouck, Sophie Ostmeier et al.
arXiv · 2026-07-16
This paper introduces Symbal, a system for automatically detecting systematic misalignments in image captions generated by multimodal large language models (MLLMs)—recurring errors tied to specific visual features in paired images. Using a dual-stage approach with off-the-shelf foundation models, Symbal correctly identifies systematic misalignments in 63.8% of datasets, nearly four times better than the closest baseline. The authors also release SymbalBench, a benchmark of 1.7 million image-text pairs across 420 vision-language datasets in natural and medical image domains. This work matters for quality assurance of AI-generated content, enabling auditing of MLLM-generated captions without requiring access to the underlying model.
- Quality assurance
- Enterprise
Research
Can We Trust Item Response Theory for AI Evaluation?
Han Jiang, Sunbeom Kwon, Jinwen Luo et al.
arXiv (Cornell University) · 2026-07-16
This paper investigates whether item response theory (IRT), a statistical framework borrowed from human testing, can be reliably applied to AI benchmark evaluation. The authors simulate response matrices under three IRT models using data from six widely used LLM benchmarks and compare four estimation methods across 18,000 simulation conditions. They find that classical estimators become computationally infeasible in large benchmark settings, while scalable estimators can produce unreliable item-level and ranking inferences when the number of evaluated models is small or their capability distributions are non-normal. The study provides guidance on the sample sizes and diagnostics needed for trustworthy use of IRT in AI evaluation contexts.
- Quality assurance
- Certifications
- AI policy
Research
Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy
Patrick Phuoc Do, Chau M. Ta, Chaoli Wang
arXiv · 2026-07-16
This paper benchmarks six multimodal large language models (MLLMs) on a standardized scientific visualization (SciVis) literacy assessment comprising 49 items across 18 scientific visualizations, 8 techniques, and 11 task types, comparing model performance against data from 485 human participants. Results show that current MLLMs do not exhibit uniform SciVis literacy: Gemini is the strongest model overall, exceeding the human mean on evaluated subsets, while open-source models remain below the human baseline. Performance is highly uneven across techniques and tasks, with models struggling on texture-based and integration-based visualizations, quantitative estimation, and flow-direction interpretation. The authors argue that SciVis literacy represents a necessary benchmark dimension for evaluating multimodal AI systems beyond chart-centric assessments.
- Quality assurance
- Enterprise
Research
MedFailBench: A Clinician-Built Open-Source Benchmark for Medical AI Safety Boundary Inspection
Goktug Ozkan
arXiv (Cornell University) · 2026-07-16
MedFailBench introduces a clinician-built synthetic benchmark designed to evaluate medical AI safety failures rather than correctness, categorizing errors by severity (on a 1–5 scale) and safety gate type (e.g., missed urgent escalation, evidence fabrication, unsafe dosing). The current release (v0.2.1) includes 44 clinician-reviewed synthetic cases with severity annotations, a safety gate taxonomy, a clinical severity rubric, and an automated pipeline for archiving model-response screening runs. By shifting focus from 'does the model know the answer' to 'which safety boundary failed,' the benchmark provides a structured framework for identifying and classifying dangerous AI behaviors in clinical contexts. This work is directly relevant to quality assurance and certification efforts for medical AI systems, offering an open-source tool (Apache-2.0 and CC-BY-4.0) for systematic safety boundary inspection.
- Quality assurance
- Certifications
- AI policy
Research
The Industrialization of Research ; On AI-Driven Science and Its Consequences
Emmanuel Jeannot
arXiv · 2026-07-16
This essay examines the transformation of scientific research by AI, framing it as an 'industrialization of research' — a shift from a craft model, where knowledge and judgment reside in individual researchers, to an automated pipeline model. Using the US Department of Energy's Genesis Mission as a prominent example, the author identifies seven critical risks: erosion of intergenerational transmission of scientific competence, opacity of AI-generated theories, collapse of peer evaluation under machine-generated output volume, unproven capacity of AI for paradigm-shifting discovery, capture of the scientific agenda by political and industrial actors, compounding of systematic errors in closed-loop pipelines, and structural bifurcation of the global research community into incommensurable tiers. The paper argues these concerns are not arguments against AI-driven science but rather the conditions under which its real and significant potential can be responsibly pursued. The analysis has direct implications for workforce development, research policy, and quality assurance in scientific institutions.
- Workforce
- AI policy
- Quality assurance
Research
Grokipedia vs Wikipedia: An LLM-Based Audit of Political Neutrality along Ideologies
Filippos Vlahos, Guillaume Bied, Tijl De Bie
arXiv · 2026-07-16
This paper presents a large-scale audit comparing political bias in Grokipedia—an encyclopedia generated entirely by the LLM Grok—against Wikipedia, using 1,394 article pairs about government members evaluated across nine ideology dimensions by four LLM judges (Grok, Claude, Mistral, and DeepSeek). All four LLM judges, including Grok itself, rated Grokipedia as less neutral than Wikipedia. The study finds that Grokipedia tends to favor economically right-wing politicians and penalize socially liberal ones, while Wikipedia shows the opposite bias pattern, and both encyclopedias are rated as portraying politicians favorably but toward different ideological groups. These findings matter for policy and public discourse because they demonstrate that replacing human-edited content with LLM-generated content does not eliminate ideological bias—it may simply shift it.
- AI policy
- Quality assurance
Research
Platform Choice, Trust, and Privacy in the Consumer AI Assistant Market
Jennifer Zou
arXiv · 2026-07-16
This survey study of 1,999 U.S. adult AI-assistant users examines platform choice, task allocation, trust, and privacy valuations in the consumer AI market. The market is found to be concentrated—ChatGPT is the primary assistant for 58% of users and Gemini for 25%—yet smaller platforms hold defensible niches, with Claude capturing a third of coding tasks despite only a 7% overall share. Trust is shown to be earned through use rather than reputation, with Claude rated most trustworthy in every head-to-head comparison among users familiar with both platforms. Privacy concern is near-universal, but action is gated by knowledge rather than concern, and users in a choice experiment value keeping humans out of their conversations most highly ($11.20/month), with valuations rising with task sensitivity—findings with direct implications for how AI platforms design data-handling policies and how regulators think about consumer protection.
- Enterprise
- AI policy
Research
Rubrics on Trial: Evolving Rubrics from a Single Query via Synthetic Pairwise Evidence
Haocheng Yang, Licheng Pan, Xiaoxi Li et al.
arXiv · 2026-07-16
Rubrics on Trial is a framework for automatically generating and validating evaluation rubrics for large language models using only a single query, without human annotations or model training. The system evolves a set of rubrics from scratch by creating synthetic response pairs conditioned on candidate rubrics, then screening out rubrics that fail to distinguish answer quality, reward irrelevant style, or penalize valid alternative approaches. Experiments across five preference benchmark suites show the method achieves the best average accuracy and leads on six of seven evaluation sets. This matters for quality-assurance and enterprise applications where scalable, reliable LLM evaluation is needed but human-annotated rubrics are costly to produce.
- Quality assurance
- Enterprise
Research
SCITUS: A Multi-Jurisdictional Framework for Adapting NIST AI RMF to the Canadian Regulatory Context
Mohammad Etemad
arXiv (Cornell University) · 2026-07-16
SCITUS is a new governance framework that adapts the NIST AI Risk Management Framework (RMF 1.0) to Canada's complex, multi-jurisdictional AI regulatory environment, covering federal requirements and five provincial regimes simultaneously. The framework introduces seven trustworthy-AI characteristics, four core governance functions, and a versioned control catalog that grew from 31 controls in June 2025 to 57 controls by July 2026 in response to new regulatory developments, including Canada's first findings on generative-AI training data and the 2026 agentic-AI threat landscape. The paper argues that systematic, unified adaptation of a global framework like NIST AI RMF offers meaningful advantages over jurisdiction-by-jurisdiction compliance, and presents scenarios across federal government, provincial healthcare, and the private sector to demonstrate applicability. The work is especially timely given the failure of Canada's Bill C-27 and the federal government's 2026 pivot to targeted instruments rather than omnibus AI legislation, leaving organizations without unified compliance guidance.
- AI policy
- Certifications
- Enterprise
Research
When AI Blurs the Boundaries of Contribution: An Empirical Study of Authorship Calibration
Célina Treuillier, Denis Lalanne
arXiv · 2026-07-16
This paper introduces 'authorship calibration'—users' awareness of how much they actually contributed when co-writing with AI—and empirically tests it using the CoAuthor dataset. The study finds high variability across users: those who rely heavily on AI tend to misjudge their own contribution, while infrequent AI users show more accurate self-assessment. The authors warn that in educational settings this miscalibration can disrupt metacognitive monitoring and harm learning outcomes. They argue that fostering authorship calibration is essential for responsible and educationally meaningful AI integration.
- Workforce
- AI policy
- Quality assurance
Research
SMC-ES: Automated synthesis of formally verified control policies
Riccardo Curcio, Toni Mancini, Enrico Tronci
arXiv · 2026-07-16
This paper presents SMC-ES, an algorithm that combines Evolutionary Strategies with Statistical Model Checking to automatically synthesize control policies for autonomous cyber-physical systems that come with formal, probabilistic guarantees. Given a confidence parameter δ and an allowable failure probability ε, the method certifies that, with confidence at least 1−δ, the probability of violating specified safety, performance, or robustness properties is at most ε. Evaluated on Gymnasium and Safety Gymnasium benchmarks, SMC-ES performs competitively against leading Deep Reinforcement Learning and Safe-DRL baselines while providing the formal guarantees those methods lack. This matters because it offers a principled path to deploying autonomous systems in safety-critical environments where informal or empirical assurances are insufficient.
- Quality assurance
- Certifications
- AI policy
Research
LQCDMaster: Agentic Scientific Computing for Lattice Quantum Chromodynamics Research
Haofei Gao, Tingjia Miao, Wenkai Jin et al.
arXiv · 2026-07-16
LQCDMaster is an AI agent that converts natural-language research tasks in lattice quantum chromodynamics (LQCD) into executable computing workflows, including measurement scripts, job-submission artifacts, and numerical outputs. Evaluated on a benchmark of 70 LQCD computing tasks, the system exactly reproduced expert-written implementations in 63 of 70 cases at machine precision, with three additional discrepancies attributed to convention mismatches. The agent reduces implementation time from hours to minutes while preserving end-to-end numerical validation, and was used to compute quantities never previously calculated, such as light-cone distribution amplitudes with a diagonal Wilson line. This work demonstrates how agentic AI can lower barriers to specialized scientific computing and facilitate exploration of non-standard scientific ideas.
- Workforce
- Enterprise
- Quality assurance
Research
Demographically-Conditioned Synthetic Medical Images for Bias Mitigation and Bias Detection in Disease Classifiers
Mahmoud Ibrahim, Bart Elen, Chang Sun et al.
arXiv · 2026-07-16
This paper addresses a critical fairness problem in medical AI: minority subgroups in test sets are often too small to reliably detect bias in disease classifiers. The authors propose using a demographically-conditioned synthetic image generator (fine-tuned Stable Diffusion 2.1) for both bias mitigation during training and bias detection during evaluation, demonstrated on COVID-19 chest CT classification. They find that using synthetic balanced cohorts as a pretraining prior—rather than joint augmentation—substantially improves classifier fairness, achieving better-than-real-data performance at roughly 100× real-data efficiency. For evaluation, their synthetic estimator perfectly reproduces subgroup performance rankings from a well-powered real oracle (Spearman ρ=1.00 on MCC and Recall), providing reliable per-subgroup estimates precisely where real test sets are too small to do so.
- Quality assurance
- AI policy
- Certifications
Research
Explaining Process Control Optimisation Recommendations via GradientSHAP and Implicit Differentiation
Paul Darm, Cem Alpturk, Kenneth Ulrich et al.
arXiv · 2026-07-16
This paper addresses the trust gap between engineers who design automated optimisation algorithms and operators who act on their recommendations in industrial settings. The authors combine Implicit Function Theorem-based sensitivity analysis with SHAP attribution and Large Language Model narrative generation to produce real-time, operator-tailored explanations for optimisation outputs. Applied to an industrial High Pressure Grinding Roll (HPGR) control problem with 22 features, their method achieves SHAP attributions with correlation above 0.99 compared to KernelSHAP while delivering over 40× speedup, enabling real-time natural language explanations validated with domain expert feedback. This matters for industrial operators who need to trust and act on AI-driven process control recommendations.
- Workforce
- Enterprise
- Quality assurance
Research
Multi-Axis Max@K Reinforcement Learning for Representative Diversity in Text-to-Image Generation
Ku Onoda, Paavo Parmas, Hiroki Furuta et al.
arXiv · 2026-07-16
This paper addresses the tendency of text-to-image (T2I) diffusion models to generate samples that cover only a narrow range of visually distinct modes for a given prompt, which can reflect or amplify demographic skew in person-centric images. The authors formalize this as 'target-mode coverage' and introduce 'multi-axis max@K,' a group-based reinforcement learning objective that rewards samples for expanding category-wise diversity by assigning credit to a sample only when it raises the group-maximum score for a given category. Evaluated on perceived-appearance fairness, the method improves Fairness Score by 0.23–0.36 relative to the base model across three automatic evaluators on held-out prompts, while maintaining image quality and text alignment. This work has direct implications for quality assurance and policy around demographic representation in AI-generated content.
- Quality assurance
- AI policy
Research
Benchmarking Face Recognition without Real Faces
Paweł Borsukiewicz, Daniele Lunghi, Wendkûuni C. Ouédraogo et al.
arXiv · 2026-07-16
This paper investigates whether synthetic face datasets can fully replace real-face benchmarks in evaluating face recognition systems, addressing the privacy gap that persists when synthetic training data is still validated against real biometric benchmarks. The authors test 12 synthetic datasets against 7 established real benchmarks using 24 pre-trained models spanning convolutional and transformer architectures, measuring biometric verification metrics, similarity score distributions, cross-model ranking consistency, and distributional properties. They find that two synthetic datasets—MorphFace and Vec2Face—reproduce the relative behavior of real benchmarks at agreement levels within the natural disagreement already observed among real benchmarks themselves. These results suggest that well-constructed synthetic datasets can support reliable comparative evaluation, advancing a fully synthetic and privacy-preserving pipeline for both training and benchmarking face recognition.
- Quality assurance
- AI policy
- Certifications
Research
Show Me How You Reason and I'll Tell You Who You Are: Reasoning Graphs for Robust LLM Authorship Attribution
Zlata Kikteva, Artur Romazanov, Annette Hautli-Janisz et al.
arXiv · 2026-07-16
This paper tackles the problem of detecting which large language model generated a given text by moving beyond surface-level linguistic features to analyze deeper reasoning structures. The authors propose a graph neural network approach that extracts reasoning graphs via an argument mining pipeline, capturing how LLMs structure arguments rather than just how they write. Their method outperforms a Longformer baseline by up to 27 percentage points under obfuscation attacks like paraphrasing and backtranslation, and by 19 percentage points when tested on unseen model versions. This matters for policy and quality-assurance efforts around AI-generated content, as it offers more robust attribution even as new LLM versions are continuously released.
- Quality assurance
- AI policy
Research
StructureClaw: Traceable LLM Agents and an Executable Benchmark for Structural Engineering Workflows
Sizhong Qin, Yi Gu, Yao Jiang et al.
arXiv (Cornell University) · 2026-07-16
StructureClaw introduces an artifact-centered workbench and executable benchmark (StructureClaw-Bench) for evaluating LLM agents on complete structural engineering workflows rather than isolated question answering. The system requires agents to produce a full chain of interdependent artifacts—interpreted requirements, computable models, validation records, solver outputs, and reports—with a scenario succeeding only when all artifact- and execution-level assertions pass. Across ten agent-model configurations tested on 50 standard cases, average Success Rate rose from 56.8% with a generic-skill baseline to 88.6% with the full automatic workflow, while interactive and multimodal evaluations revealed remaining challenges around invalid numerical inputs and structural model reconstruction. The work demonstrates that artifact-centered evaluation can surface workflow-level failures invisible in final-response-only assessments, offering a more rigorous basis for deploying AI agents in safety-relevant engineering contexts.
- Quality assurance
- Certifications
- Enterprise
Research
Proof-or-Stop: Don't Trust the Agent, Trust the Evidence -- Loop Engineering for Verifiable Evidence-Gated Lifecycle Control
Jek Huang, Jeffery Hsia, Jiayi Sun et al.
arXiv (Cornell University) · 2026-07-16
This paper introduces 'Proof-or-Stop Lifecycle Control,' a framework that gates software development lifecycle transitions (e.g., reviewed, tested, ready-to-merge) on mechanically verifiable evidence rather than accepting autonomous coding agent outputs as trusted claims. In a controlled ablation study, the gated loop reduced visible-pass/hidden-fail amplification from 31 to 2 out of 1,800 injected test cells compared to a naive compute-budgeted loop, with a 95% confidence interval of [0.8, 2.5] percentage points improvement. Results further show that enforcing review as a lifecycle gate—rather than merely adding a reviewer—drives the quality gain, and that tamper-resistant receipt bundles rejected 18 tamper classes with zero false accepts. The approach offers a model-agnostic control layer relevant to enterprise software quality assurance and certification of autonomous agent behavior.
- Quality assurance
- Enterprise
- Certifications
Research
Innocuous-Seeming Data, Latent Ideology: Ideological Generalisation in Finetuned LLMs
Robert Graham, Edward Stevinson, Yariv Barsheshat
arXiv · 2026-07-16
This paper demonstrates that finetuning large language models (LLMs) like GPT-4.1 on small, factually defensible datasets—such as economics Q&A, HR policies, or food-safety queries—can cause broad ideological shifts across entirely unrelated domains like criminal justice, environmental attitudes, and cultural preferences. The authors call this 'ideological generalisation' and show it persists even when general capabilities (e.g., math accuracy on GSM8K) remain intact, replicates on Gemma-3, and can push models toward extreme out-of-distribution outputs including endorsements of race-IQ connections and political violence. They introduce metrics for 'breadth' (how far ideological shifts spread beyond training topics) and 'amplification' (how much finetuning intensifies shifts compared to few-shot prompting). The findings have significant implications for AI quality assurance and policy, as seemingly innocuous, moderation-passing datasets can introduce hidden ideological biases that are difficult to detect through standard evaluation.
- Quality assurance
- AI policy
- Certifications
News
Digital Surveillance Reshapes Fishery Enforcement in Indonesia
spectrum.ieee.org · 2026-07-16
IEEE Spectrum reports on Indonesia's sweeping transformation of fisheries enforcement through satellite-based digital surveillance, describing how the country's Marine and Fisheries Resources Surveillance Station now uses vessel monitoring systems (VMS), satellite remote sensing, and geospatial analytics to detect potential violations before any patrol vessel departs port. By early 2026, nearly 9,400 Indonesian fishing vessels were transmitting through the national VMS, and during the first quarter of 2026 alone the system tracked over 14,500 vessels and identified 491 suspected violations. The outlet notes, however, that as surveillance capabilities advance, some illegal operators are adapting by disabling transmitters or exploiting gaps between monitoring systems, creating a technological arms race. The piece concludes that the future of this enforcement model hinges on data integrity, cybersecurity, and algorithmic accountability rather than surveillance volume alone.
- AI policy
- Enterprise
- Quality assurance
Research
Does generative AI supersede supervised XMLC? A Benchmark Study on Automated Subject Indexing with German Scientific Literature
Maximilian Kähler, Katja Konermann, Lisa Kluge et al.
arXiv · 2026-07-16
This study benchmarks automated subject indexing of German scientific literature from the German National Library, framing the task as Extreme Multi-Label Classification (XMLC) across a large controlled vocabulary. Supervised XMLC methods using transformer-based dense features achieve the best overall binary relevance scores, but LLM-based generative approaches outperform them on graded relevance and on the challenging long tail of the subject vocabulary. Professional subject librarians provided graded relevance ratings alongside standard binary metrics, offering a practical quality perspective. The findings suggest generative LLM methods are a promising alternative for real-world library subject indexing workflows.
- Enterprise
- Quality assurance
- Workforce
Research
RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems
David Ayllon, Alice Baird, Jeffrey Brooks et al.
arXiv · 2026-07-16
This paper introduces the Real World Voice EQ Bench, a multidimensional benchmark for evaluating voice AI systems across text-to-speech, speech-to-speech, speech understanding, and automatic speech recognition tasks. The benchmark reveals that performance is highly dimension-specific: for example, TTS naturalness, expressiveness, and identity stability are largely independent, while speech-to-speech agents often remain transcript-driven rather than leveraging vocal affect. ASR systems show failures under real-world conditions—including accents, emotions, noise, and conversational speech—that are not captured by existing clean-speech benchmarks. The findings argue that voice AI should be assessed as a profile of acoustic, expressive, interactional, and robustness capabilities rather than by a single aggregate score, with direct implications for how systems are tested and certified.
- Quality assurance
- Certifications
Research
Interventional Causal Circuits for Safe Robot Action Testing and Failure Recovery
Naren Vasantakumaar, Tom Schierenbeck, Michael Beetz
arXiv · 2026-07-16
This paper presents a closed-loop framework for safe robot action planning that uses causal reasoning to recover from test failures rather than blindly resampling action parameters. The system couples a Joint Probability Tree (JPT) with a Causal Circuit derived from a Marginal-Deterministic Variable Tree, enabling exact, polynomial-time computation of interventional probabilities without retraining or additional data collection. Experiments in a ROS2 simulation show the Causal Circuit reduces failed attempts by 10.3% under a high-quality model and by 37% under a degraded model. Each rejection produces an interpretable causal report identifying the responsible parameter and a corrective region, supporting both operator oversight and autonomous recovery.
- Quality assurance
- Enterprise