News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
5608 items
Research
Position: Explanation Stability Is a Property of the Model Method Pair, Not the Model
Kabilan Elangovan, Daniel Ting
arXiv · 2026-07-18
This position paper argues that AI explanation stability depends on the combination of model and attribution method used, not on the model alone. In chest X-ray experiments with DenseNet201, ResNet50V2, and InceptionV3—all achieving AUC above 99%—stability rankings reversed depending on whether LayerCAM or GradCAM++ was used; for example, InceptionV3's stability score dropped by 17.3% when switching methods. The authors contend that claiming a model is 'stable' based on a single attribution method produces scientifically invalid and potentially illusory safety assurances. They recommend that explanation-based claims be validated across multiple attribution paradigms and that regulatory submissions explicitly specify the attribution operators used.
- Quality assurance
- Certifications
- AI policy
Research
Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
Haifeng Li, Mo Hai
arXiv · 2026-07-18
This paper addresses a critical failure mode of LLMs that generate optimization models from natural-language descriptions: a generated model can run successfully while still formulating the wrong problem. The authors develop a falsification-based verification framework using typed numeric slots and solver calls—without needing a reference model—drawing on duality, comparative statics, and polyhedral arguments to build a battery of sound tests (directions, curvature, crush probes, prohibitive limits, annihilation, and exchange). Key results on 326 ground-truth models show the battery achieves a 0.0% false-positive rate versus 54.9% for a threshold tester, detects 70.0% of certified conditional-class mutants, and catches 40.4% of mutants invisible to execution-accuracy scoring. The framework also proves that no fixed-threshold perturbation tester can be simultaneously sound and nontrivial, establishing theoretical limits on what such verification can detect.
- Quality assurance
- Enterprise
Research
CLOSER-Bench: Evaluating Budgeted Cross-Stage Design Closure for Hardware Agents
Peilong Zhou, Zhirong Chen, Cangyuan Li et al.
arXiv · 2026-07-18
CLOSER-Bench introduces a controlled benchmark protocol for evaluating AI coding agents on hardware design closure tasks that span multiple abstraction levels, from RTL generation through physical implementation (RTL-to-GDS). Built on open-source tools including Verilator, Yosys, OpenROAD, and Sky130, the benchmark measures final quality, anytime progress, tool cost, and cross-stage recovery within a fixed budget, exposing a sharp 'completion-closure gap' where agents that solve localized repair tasks fail at the matched verification-closure counterpart. A ten-task pilot covering RTL repair, verification, PPA optimization, and security shows that frontier agents meaningfully outperform baselines on cross-stage tasks, motivating the treatment of hardware closure as a budgeted sequential decision problem rather than independent code generation. These findings have direct implications for enterprise hardware development workflows and quality-assurance processes that increasingly rely on AI agents.
- Enterprise
- Quality assurance
Research
Privacy Cost as Equity Input: A Group Fairness Criterion for Differentially Private Machine Learning
Rakshit Naidu
arXiv · 2026-07-18
This paper introduces the Privacy-Cost Equity Ratio (PCER), a new group fairness metric for differentially private machine learning that accounts for which demographic groups bear the greatest privacy exposure, not just which groups receive accurate predictions. The authors argue that information leakage from DP-SGD training is itself a harm, and that groups facing greater membership inference risk are owed proportionally greater predictive benefit. Evaluated across six benchmark-attribute combinations in tabular and NLP settings, PCER reveals fairness disparities that outcome-based metrics miss — for example, on COMPAS it uncovers a 'double disadvantage' where a protected group suffers both higher privacy exposure and worse predictive outcomes, a pattern hidden by demographic parity gap. The findings suggest that fairness audits of privacy-preserving AI systems must incorporate the distribution of privacy costs across groups, not only the distribution of model benefits.
- Quality assurance
- AI policy
- Certifications
Research
DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines
Runming He, Zhen Hao Wong, Hao Liang et al.
arXiv · 2026-07-18
DataFlow-Harness is a platform that guides large language models to build editable, persistent data-processing pipelines represented as directed acyclic graphs (DAGs), rather than one-off scripts — addressing what the authors call the 'NL2Pipeline gap.' On a 12-task data-engineering benchmark, the system achieves a 93.3% end-to-end pass rate while reducing monetary cost by 72.5% and generation latency by 49.9% compared to a Vanilla Claude Code baseline. The platform combines procedural guidance (DataFlow-Skills), a live operator registry via a Model Context Protocol layer, and a visual DAG editor synchronized with conversational authoring. These results suggest that grounding LLM agents in live platform context can yield reliable, editable workflow artifacts at substantially lower cost than conventional script-generation approaches.
- Enterprise
- Workforce
Research
Can Multimodal Large Language Models Understand OCT?
Baochen Fu, Wenzhi Deng, Baihao Jin et al.
arXiv · 2026-07-18
This paper introduces OCT-Bench, a benchmark of 10,076 multiple-choice questions drawn from 4,137 OCT images across seven public datasets, designed to evaluate how well multimodal large language models (MLLMs) understand optical coherence tomography imaging for retinal disease. The benchmark organizes 20 fine-grained tasks across three dimensions—Perception, Cognition, and Reasoning—mirroring real-world clinical interpretation workflows. Evaluating 20 representative MLLMs, including proprietary, open-source, and medical-domain models, the study finds that current models fall substantially short of reliable OCT understanding, and that neither medical-domain adaptation nor larger model scale consistently improves performance. These findings highlight critical capability gaps relevant to the safe deployment of AI in clinical diagnostic settings.
- Quality assurance
- Certifications
- AI policy
Research
Hazard or Anomaly? Evaluating VLMs for Understanding Dangers and Discrepancies
Murali Indukuri, Mohammad Eskandari, Sree Nitya Kollu et al.
arXiv · 2026-07-18
This paper evaluates Vision-Language Models (VLMs) for safety reasoning in human-robot interaction by introducing a distinction between hazards (true physical dangers) and anomalies (unusual but not necessarily dangerous scene elements). Current evaluations typically frame danger recognition as a binary Safe/Unsafe judgment, which the authors argue obscures whether models are detecting genuine hazards or simply reacting to contextual irregularities. Testing several state-of-the-art VLMs across two datasets and multiple prompting strategies, the study finds that VLMs frequently misinterpret anomalousness as hazardousness, revealing a systematic failure mode. Explicitly separating anomaly from hazard exposes these weaknesses and provides a more informative evaluation framework for safety-critical applications such as disaster response and emergency decision-making.
- Quality assurance
- Certifications
- AI policy
Research
Learning from World Feedback: Why Model Uncertainty Fails as a Risk Signal in Model-Based RL
Zhaohui Wang
arXiv · 2026-07-18
This paper investigates why model uncertainty fails as a safety proxy in model-based reinforcement learning, finding that dynamics-based uncertainty penalties are actually anti-correlated with real-world safety: they increase collision rates from 26% to 34% across four world-model architectures. The authors show that model uncertainty has very low empirical correlation (r < 0.15) with actual task risk because it operates over state-prediction space rather than constraint boundaries. Replacing this internal proxy with world-feedback signals—such as lidar-derived safety margins, time-to-collision estimates, and outcome-trained feedback models—reduces collision rates to 1–14% without retraining any components. The findings are distilled into three design principles (ground risk in world outcomes, validate proxies before deployment, and use outcome-trained feedback models when direct signals are unavailable) that the authors argue apply broadly to RLHF and LLM alignment as well.
- Quality assurance
- Enterprise
- AI policy
Research
Dual-Track Synergistic Regulation of Data and Algorithms in Connected and Autonomous Vehicles: A Systematic Literature Review
Jingwen Cai, Yifen Yin, Yuanyuan Yu et al.
World Electric Vehicle Journal · 2026-07-18
This systematic literature review synthesizes 135 peer-reviewed articles to examine how data privacy and algorithmic safety regulation for Connected and Automated Electric Vehicles (CAEVs) are typically treated as separate silos—and why that fragmentation is problematic. The authors find that traditional notice-and-consent privacy models fail in vehicle-to-everything (V2X) environments, that opaque algorithmic decision-making undermines existing tort liability frameworks, and that cross-national regulatory inconsistency impedes effective governance. They propose a 'dual-track synergistic' governance architecture that couples data lifecycle management with algorithmic safety standards such as ISO/SAE 21434 and UN R155/156, and advocate for adaptive regulatory sandboxes and international standards harmonization.
- AI policy
- Certifications
Research
EU Artificial Intelligence Act: risks and opportunities for the European system
Enrico Maggiora, Claudia Iacobino
Journal of Emerging Perspectives · 2026-07-18
This paper analyzes the EU Artificial Intelligence Act as a risk-based regulatory framework that classifies AI systems by potential harm to rights, safety, and democratic values. It highlights compliance challenges—particularly for SMEs and start-ups facing high-risk AI obligations—while also identifying strategic opportunities such as a trustworthy AI ecosystem, a strengthened Digital Single Market, and the potential 'Brussels effect' of the EU setting a global regulatory standard. The authors argue the AI Act should be understood not as a mere compliance burden but as a long-term competitive advantage for Europe.
- AI policy
- Enterprise
Research
AI winters as legitimacy crises. From the history of technological promises to governance models of generative AI
Mariusz Mazurek, Jacek Gurczyński
AI and Ethics · 2026-07-18
This paper reframes 'AI winters' not as mere technical failures but as legitimacy crises in which the web of promises, funding, commercialization, and social acceptability broke down. Using a hierarchical model of four legitimacy gaps—capability, institutional assessment, commercialization, and governance—the authors reinterpret historical AI winters and argue that generative AI faces a distinctive 'regulatory-economic cooling' driven by compliance costs, training data disputes, infrastructure concentration, and documentation requirements. The paper concludes with four governance implications: promise restraint, evidence of deployment, documentation obligations, and infrastructural pluralization, shifting the central question from whether AI works to under what conditions its development remains politically and ethically legitimate.
- AI policy
Research
Clinical Positioning and Implementation of a Deep-Learning Retinal Biomarker (Reti-CVD) for Cardiovascular Risk Stratification: A Narrative Review
Junseung Rho, Sung Goo Kang, Se‐Hong Kim et al.
Journal of Clinical Medicine · 2026-07-18
This narrative review evaluates Reti-CVD, a deep-learning retinal imaging tool that generates a three-tier cardiovascular risk classification from a retinal photograph. The tool demonstrated a Harrell C-index of approximately 0.75 in studies including UK Biobank and Singapore SEED cohorts, with modest reclassification improvement particularly in borderline-risk individuals. Regulatory status includes Korean MFDS marketing authorization and claimed CE certification under EU MDR, but no US FDA authorization, and no randomized trials have shown that Reti-CVD-guided care improves clinical outcomes. The review concludes it is best positioned as a non-invasive risk enhancer pending independent validation, intervention trials, and cost-effectiveness evidence.
- Certifications
- AI policy
- Quality assurance
Research
AI Safety Compass: A Compact Framework for AI Safety Research and System Design
Ran Liu, Xiaowei Huang
Zenodo (CERN European Organization for Nuclear Research) · 2026-07-18
The AI Safety Compass paper introduces a compact framework that organizes AI safety research and system design around four recurring dimensions: Safety Objective Type, Safety Challenge Type, Safety Control Approach, and Safety Assurance Architecture. The authors argue that AI safety work is currently fragmented across specialized communities using different vocabularies and assumptions, making it difficult to construct coherent, defensible safety architectures from individual advances. The framework provides a shared design logic that allows research contributions to be positioned by their design role, safety routes to be compared and composed, and gaps or mismatches between diagnosis, controls, and evidence to be identified before implementation. This matters because it offers both researchers and system designers a common structure for building, auditing, and maintaining safety claims across diverse AI safety domains.
- Quality assurance
- AI policy
- Certifications
- Enterprise
Research
The AI Implementation Gap: Policy–Audit Misalignment in the UAE and Egypt
John Joshua C. Rañeses, Ain Bemisal Alavi, Rafia et al.
Journal of Business Insight and Innovation · 2026-07-18
This study investigates the gap between AI-related policies and actual audit practice among Big Four firms operating in the UAE and Egypt. Using content analysis of Big Four transparency reports from 2021–2025 and thematic analysis of semi-structured interviews with auditors conducted in 2026, the researchers find that while corporate reports portray AI as a standard component of audit workflows, auditors on the ground report only limited and tokenistic adoption. The authors identify governance shortcomings, regional disparities, and algorithmic complexity as key barriers, and call for better-aligned policy-practice frameworks to enable effective AI integration in auditing.
- AI policy
- Quality assurance
- Enterprise
- Certifications
Research
Mitigating Compiler Fusion-Induced Power Bursts in Mobile NPU Inference as the Battery Depletes
Ryoga Yuzawa, Masayoshi Tomizuka
arXiv · 2026-07-17
This paper investigates how aggressive operator fusion in mobile NPU compilers creates large peak-current bursts during inference, which can trigger dynamic voltage and frequency scaling (DVFS) and increase latency as a smartphone battery depletes. The authors present a measurement study on a commercial smartphone showing that fused 'superlayers' concentrate execution and worsen low-voltage operating margins. They propose a black-box mitigation—a measurement-guided graph rewrite that inserts barriers at peak-to-average power ratio hot spots before compilation—and demonstrate on Snapdragon 8 Gen 3 with MobileNetV4 that this reduces peak current from 3.12 A to 1.94 A with only 3.76% latency overhead, shifting the inferred DVFS margin by approximately 173 mV. This matters for mobile AI deployment because it enables more stable, reliable inference performance under real-world low-battery conditions without requiring access to proprietary compiler internals.
- Enterprise
- Quality assurance
Research
Nonuniformity Principle in Human-AI Coworking
An Luo, Jie Ding
arXiv · 2026-07-17
This paper addresses the challenge of optimally scheduling human oversight within multi-step AI workflows, where unlimited human review is impractical due to time and resource constraints. The authors derive a 'nonuniformity principle' stating that optimal placement of human oversight stages should follow non-decreasing gaps as a workflow progresses. They validate this principle empirically in two AI agent workflows—writing literature reviews and constructing websites—finding that strategic oversight placement improves user satisfaction while reducing unnecessary rework and token consumption. These findings offer actionable guidance for designing human-AI collaboration systems in enterprise and quality-assurance contexts.
- Workforce
- Enterprise
- Quality assurance
Research
Binding Drift in Multi-Step Tool-Augmented Agents
Rahul Suresh Babu, Shashank Indukuri
arXiv · 2026-07-17
This paper investigates how entity-binding errors evolve across multi-step workflows in tool-augmented language model agents—specifically distinguishing 'binding drift' (a correct binding becoming wrong later) from 'error propagation' (a wrong binding carried forward). In a controlled testbed of 200 workflows across four enterprise domains and eight model backends, the authors find that a naive 'entity lock' fix amplifies wrong actions by 3x on average and up to 8.5x on Claude Opus 4.5, while a lightweight LLM-based re-verifier reduces wrong actions by 79%, nearly matching an oracle upper bound. The work matters for enterprise AI deployments and quality assurance because it shows that intuitive persistence-based fixes can dramatically worsen outcomes, and that cheap re-verification is a practical mitigation for agentic systems operating over external tools.
- Enterprise
- Quality assurance
Research
K-IPO: Kendall-constrained Importance Preserving Oversampling for Imbalanced Tabular Data
Marios Tyrovolas, Argiris Sofotasios, Dimitris Metaxakis et al.
arXiv · 2026-07-17
K-IPO (Kendall-constrained Importance-Preserving Oversampling) is a new framework for handling class imbalance in tabular data that addresses a previously unresolved problem: existing oversampling methods distort the feature importance rankings that underlie model explanations. The authors propose a generator-agnostic 'generate-then-select' approach that iteratively generates minority-class candidates and accepts them only when their inclusion maintains a user-defined minimum Kendall's tau correlation with the original feature importance ranking. Evaluated on 20 imbalanced binary classification datasets with three classifiers and multiple explanation methods, K-IPO achieved best or tied-best results in feature importance preservation, explanation consistency, and class separability while generally improving predictive performance. This matters for AI quality assurance and enterprise deployments where faithful model explanations are critical for trust and decision-making.
- Quality assurance
- Enterprise
Research
To Police or to Guide: How Higher Education Computer Science Instructors Design and Implement Generative AI Policies
Xingjian Gu, Wells Lucas Santo, James M. Zumel Dumlao et al.
arXiv · 2026-07-17
This study examines how U.S. computer science instructors in higher education design and implement generative AI course policies, drawing on 13 semi-structured interviews. The researchers found that while instructors acknowledge AI tools can harm student learning, most policies focus on 'AI-proofing' assessments—such as switching to paper exams—rather than directly supporting student learning outcomes. These policing-oriented approaches create extra burden for instructors and strain student-instructor relationships. The authors recommend shifting toward learning-oriented AI policies that guide students toward healthier AI usage habits.
- Workforce
- AI policy
Research
Retrieval is Enough: Training-Free Interpretability with a Tool-Using Agent
Sriram Balasubramanian, Soheil Feizi
arXiv · 2026-07-17
This paper introduces HARP (Hypothesis-driven Agentic Retrieval and Probing), a training-free neural network interpretability method that equips an LLM agent with a vector database of activations paired with textual contexts, along with tools for manipulating those activations. Despite requiring no training, HARP outperforms training-based methods—including activation oracles and SAE-based agents—on concept discovery, concept detection, model steering, and secret elicitation tasks. The results suggest that current training-based interpretability methods do not extract insights beyond what is already recoverable from their training data, motivating new benchmarks that require interpretability methods to demonstrate genuinely novel insights. This matters for AI quality assurance and enterprise deployment, as it shows that cheaper, more flexible approaches can match or exceed costly trained methods for understanding what neural networks have learned.
- Quality assurance
- Enterprise
Research
One Modality to Forget Them All: Enhancing Cross-Modal Unlearning in Vision-Language Models
Sudharshan Balaji, Yili Ren, Guangjing Wang et al.
arXiv · 2026-07-17
This paper investigates whether machine unlearning applied to one modality (text or vision) in Vision-Language Models (VLMs) transfers to the other modality, finding that such cross-modal transfer is asymmetric and incomplete. The authors show that typographic attacks—which manipulate the visual presentation of text—can recover previously unlearned knowledge, revealing shallow unlearning. To address this, they propose CrossInf, an influence-guided strategy that targets transformer blocks most responsible for cross-modal generalization, reducing the transfer gap by more than half in strongly fused architectures and cutting typographic attack success rates to near zero. These findings have important implications for the security and reliability of VLMs deployed in enterprise and policy-sensitive contexts, where ensuring hazardous knowledge is truly removed is critical.
- Quality assurance
- AI policy
- Enterprise
Research
Signal-based Model Access Risk Analysis for AI System Operations Security
Maria Mahbub, Steven Young, Amir Sadovnik et al.
arXiv · 2026-07-17
This paper introduces SMART (Signal-based Model Access Risk Taxonomy), a deployment-oriented framework for classifying the security risks facing AI systems based on the type and richness of output signals that attackers can observe from deployed models. The authors argue that existing white-box/gray-box/black-box taxonomies inadequately distinguish between deployment scenarios that expose very different information—such as final decisions only versus confidence scores, intermediate representations, or full parameters—even when all are labeled 'black-box.' By organizing evasion attacks according to these signal levels, SMART provides a structured way to understand how attack capabilities scale with information exposure. The framework is intended to inform more secure AI procurement and deployment decisions across high-stakes domains including security, finance, healthcare, and cloud services.
- Enterprise
- AI policy
- Quality assurance
Research
Interactive Task Alignment as a POMDP
Andy Dai, Zexue He, Zhenyu Zhang et al.
arXiv · 2026-07-17
This paper studies 'task alignment'—the ability of language models to clarify and align with users on ambiguous, underspecified tasks before executing them. The authors formalize this as a Partially Observable Markov Decision Process (POMDP) and benchmark models across shopping, coding, and professional work settings, finding that current models recover the user's intended task only 22–32% of the time under ambiguity, compared to 48% for humans. While post-training via supervised fine-tuning and reinforcement learning improves performance, models still act prematurely and fail to resolve ambiguity effectively. The findings highlight a critical gap in model interaction abilities relevant to deploying AI assistants as reliable agents in real-world enterprise and workforce settings.
- Workforce
- Enterprise
- Quality assurance
Research
Evaluating Open-Weight LLMs for Generating Structured Threat Information for Autonomous Vehicle Vulnerabilities
Md Erfan, Ahmed Ryan, Md Kamal Hossain Chowdhury et al.
arXiv · 2026-07-17
This paper evaluates 11 open-weight large language models (ranging from 4B to 120B parameters) on their ability to convert plain-text vulnerability descriptions for Connected and Autonomous Vehicles (CAVs) into structured threat intelligence using the STIX format. The researchers built a dataset called CAV-STIXGen mapping CAV-related CVEs to STIX domain and relationship objects, Common Weakness Enumeration entries, and MITRE ATT&CK techniques, then tested various prompting strategies and model configurations. Single-model setups achieved strong results for some tasks (F1 of 0.94 for STIX domain objects, 0.99 for CWE mapping) while relationship and ATT&CK mapping proved more difficult, and a multi-agent configuration using Gemma-4-31B and Codestral-22B showed competitive performance. The work demonstrates that AI-assisted automation of vulnerability-to-STIX translation could streamline threat intelligence workflows and help prioritize defenses in transportation security.
- Quality assurance
- AI policy
- Enterprise
Research
An Exam for Active Observers
Jiarui Zhang, Muzi Tao, Shangshang Wang et al.
arXiv · 2026-07-17
ActiveVision is a new benchmark designed to test whether multimodal large language models (MLLMs) can perform 'active observation'—the iterative, hypothesis-driven visual perception that humans use rather than a single static image scan. Comprising 17 tasks across 3 categories, the benchmark requires repeated visual perception loops rather than one-shot descriptions. Frontier models fail dramatically: the best-performing model, GPT-5.5, solves only 10.6% of items and scores zero on 11 of 17 tasks, while three human participants average 96.1%. The results reveal a fundamental gap in current MLLM architectures and motivate new training objectives that better close the perception-reasoning loop.
- Quality assurance
- Enterprise