News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
Algorithmic Enforcement and the Administrative State: Due Process and Accountability in the EU and the United States
Edward Koellner
Zenodo (CERN European Organization for Nuclear Research) · 2026-07-19
This article examines how public agencies in the EU and United States are delegating enforcement decisions, risk assessments, and eligibility determinations to AI and algorithmic systems, and what that shift means for administrative accountability and due process. The authors argue that discretion formerly held by frontline caseworkers has migrated upstream into choices about training data, feature weights, and decision thresholds—choices that function as de facto policy yet rarely appear in reviewable administrative records. Comparing Europe's preventive, rights-centered approach (constitutional proportionality, GDPR Article 22, EU AI Act impact assessments) with the US reactive, litigation-driven model (APA arbitrary-and-capricious review, court-ordered discovery), the paper proposes a hybrid framework combining Europe's pre-deployment tools with America's contestation and audit mechanisms, all reinforced through procurement requirements and judicial insistence on legible records. The analysis matters because it offers a concrete path to keeping automated government decisions accountable as algorithmic discretion becomes increasingly embedded in AI code.
- AI policy
Research
LECTURER’S PERSPECTIVES ON THE USE OF ARTIFICIAL INTELLIGENCE TOOLS IN STUDENTS’ WRITING
Tri Yuli Ardiyansah, R W Batubara
EDUTECH Jurnal Inovasi Pendidikan Berbantuan Teknologi · 2026-07-19
This mixed-methods study surveyed and interviewed nineteen English lecturers at the University of Muhammadiyah Gresik about their views on students using AI tools such as ChatGPT in academic writing. While 84.2% of lecturers were familiar with ChatGPT and 78.9% understood its mechanisms, only 42.1% felt confident they could identify AI-generated text, highlighting a significant detection gap. Lecturers broadly regarded AI as a supplemental instructional aid rather than a replacement for critical thinking, and the study recommends that institutions develop process-based assessment systems and strengthen ethical codes to promote responsible digital literacy.
- AI policy
- Quality assurance
Research
Real-World Evaluation of an AI Agent Drafting Translational Impact Summaries
Mohammad Arvan, Amber E. Osterholt, Bailee Rue et al.
arXiv · 2026-07-18
This paper evaluates a human-in-the-loop AI agent designed to automate the drafting of translational impact summaries for Clinical and Translational Science Award (CTSA) scholars. The agent assembles sourced evidence dossiers and drafts one-sentence impact summaries for staff review, achieving an 81.7% unanimous usable rate across 507 findings from 10 scholars while reducing per-scholar staff time from an estimated 15 hours to a median of 14 minutes. The agent covered all four Translational Science Benefits Model domains and surfaced non-scholarly impact evidence that routine processes often miss, with reviewers rating synthesis accuracy at 4.5 and usefulness at 4.8 out of 5. The findings suggest AI agents can make cohort-scale impact reporting feasible by shifting staff from data collection and writing to reviewing.
- Workforce
- Enterprise
- Quality assurance
Research
TurboVec: A Case Study in Cost-Efficient Private Retrieval for Enterprise RAG via Codebook-Oblivious Quantization
Navnit Shukla, Kamal Pandey, Omsankar Tiwari
arXiv · 2026-07-18
TurboVec is an enterprise vector retrieval system built on TurboQuant, a codebook-oblivious scalar quantizer that requires no corpus-dependent training, addressing two key challenges in multi-tenant RAG deployments: privacy leakage from trained quantizers and recall degradation from post-hoc tenant filtering. On the DBpedia OpenAI embeddings benchmark, TurboQuant 4-bit outperforms trained FAISS Product Quantization by 8.5–8.9 percentage points in Recall@5 at the same memory budget, while using 4–8x less memory than HNSW. Deployed on Snowpark Container Services, TurboVec achieves 11ms median query latency and kernel-level allowlist filtering that maintains 0.86–0.93 Recall@10 across 10–1000 tenant workloads, compared to 0.09–0.19 for post-filter baselines. The codebook-oblivious design reduces membership inference accuracy to near-random (50.0%) versus 57.3% for PQ codebooks, making it relevant for privacy-sensitive enterprise AI applications.
- Enterprise
- Quality assurance
- AI policy
Research
What Do They See? Interpreting Complex Road Scenarios Through the Eyes of Vision-Language-Action Models for Safe and Trustworthy Autonomous Vehicle Learning
Kalpana Panda, Wesley Maia, Vinti Agarwal et al.
arXiv · 2026-07-18
This paper introduces CVAA (Counterfactual Vision Action Analysis), a framework that systematically removes individual objects from front-camera images using photorealistic inpainting to isolate how each object causally influences an autonomous driving model's planning decisions. Applied to the Alpamayo 1 trajectory predictor across 210 nuScenes scenes, the study finds that vehicles and pedestrians in the model's path dominate causal influence, traffic lights have outsized effects relative to their image footprint, and the model sometimes responds strongly to objects a human driver would consider irrelevant. The authors combine this behavioral auditing with mechanistic interpretability techniques to probe intermediate model representations, working toward explainable autonomous driving systems. This research matters for quality assurance and certification of safety-critical AI systems, as it provides a structured method for auditing and building human-AI trust in autonomous vehicle decision-making.
- Quality assurance
- Certifications
- AI policy
Research
Optimizing Clinical Trial Protocols Using EHR-Derived Heterogeneous Treatment Effects
Xiaodi Li, Munhuwan Lee, Pengyang Li et al.
arXiv · 2026-07-18
This study uses real-world electronic health records from the Mayo Clinic Cloud to emulate the DAPA-HF clinical trial and estimate heterogeneous treatment effects (HTEs) of dapagliflozin versus placebo in heart failure patients. While the overall emulated cohort showed no statistically significant survival benefit (HR 1.681, p=0.1507), HTE-guided stratification identified two distinct subgroups: one with a strong survival benefit (HR 0.203, p=0.0002) and one with significantly increased mortality risk (HR 6.680, p<0.0001). These findings demonstrate that HTE-driven patient stratification can reveal clinically meaningful treatment-effect patterns that are obscured when analyzing trial populations as a whole, suggesting a path toward more precise and efficient clinical trial design.
- Quality assurance
- AI policy
- Enterprise
Research
PREFAIL: Identifying Precursors to Failures in Robotic Lift-and-Place Tasks to Improve Task Execution Performance
Zeyu Shangguan, Rajas Chitale, Rutvik Patel et al.
arXiv · 2026-07-18
PREFAIL is a method for predicting failure precursors in robotic lift-and-place tasks used in non-prehensile material handling, where friction-based support makes high-speed motions unreliable. The approach analyzes the relative motion of target objects with respect to their carrier to detect early warning signs of failure, addressing key limitations of existing methods such as sensitivity to dynamic actions and reliance on known policy structures. The authors also introduce a dataset that precisely identifies the latest actionable intervention time, enabling rigorous evaluation of whether a predicted failure can still be prevented. Experimental results on both simulation and real-world data show that PREFAIL improves the accuracy and timeliness of failure precursor detection, which has direct implications for the reliability and efficiency of automated robotic systems in industrial settings.
- Enterprise
- Quality assurance
Research
A Deep Reinforcement Learning Algorithm for the Vehicle Routing Problem with Stochastic Demands and Outsourcing
Mohsen Dastpak, Fausto Errico, Ola Jabali
arXiv · 2026-07-18
This paper introduces the Vehicle Routing Problem with Stochastic Demands and Outsourcing (VRP-SDO), where a logistics provider must decide which customer deliveries to handle with its own fleet versus outsource to a carrier, while managing uncertain demand revealed only upon vehicle arrival. The authors propose a two-level iterative method combining iterated local search for outsourcing decisions with a deep Q-network (using a graph attention network) to estimate routing costs offline, enabling near-instant cost approximations without resolving from scratch each iteration. Experiments show their policy reduces routing costs by 19.6% over a state-of-the-art method and at least 29.6% over classical heuristics, while the full algorithm saves 13.7% on average versus a version without the attention-based representation and produces decisions in minutes rather than over an hour. This matters for enterprise logistics and workforce planning, as it enables faster, higher-quality operational decisions under uncertainty and helps balance labor costs (overtime) against outsourcing expenditures.
- Enterprise
- Workforce
Research
Do Speech Tokens Leak Voiceprints? Speaker Inversion Attacks Against End-to-End Speech Language Models
Ye Lu, Yihan Yan, Zhaoyang Zhang et al.
arXiv · 2026-07-18
This paper investigates whether speech tokens used by end-to-end speech language models (such as Moshi, Higgs3, Kimi-Audio, and Qwen3-Omni) inadvertently expose users' voiceprints. The authors introduce Audio BERT (AuB) and a two-stage inversion method called SpInv, which can recover speaker-identifying embeddings from as little as three seconds of speech token output, achieving cosine similarities above 0.70 in the attacker-specified speaker-encoder space. This demonstrates a significant privacy risk: even without access to raw audio, adversaries can reconstruct biometric voice identifiers from the token representations these models expose. The findings have direct implications for policy around biometric data protection and enterprise deployment of speech AI systems.
- AI policy
- Enterprise
- Quality assurance
Research
Beyond Semantic Equivalence: Logical Graphs for LLM Uncertainty Quantification
Yanni Dong, Minghua Liu, Meiling Zhu et al.
arXiv · 2026-07-18
This paper introduces Logical Graph Uncertainty (LGU), a framework for improving how uncertainty is measured in Large Language Model outputs. Existing methods like semantic entropy treat logically compatible but differently phrased answers as uncertain or hallucinated, when in fact they may simply differ in granularity or specificity. LGU instead explicitly models logical relationships—entailment and incompatibility—among generated answers, producing more accurate uncertainty estimates. Across multiple question-answering benchmarks, LGU outperforms the semantic entropy baseline by up to +7.1% AUROC and +3.5% AUARC, which matters for deploying LLMs reliably in safety-sensitive applications.
- Quality assurance
- Enterprise
- AI policy
Research
Beyond Memory Leaderboards: Evaluating Scientific Memory as Budgeted Context Restoration
Maksim Sheverev, David Finkelstein, Sergey Nikolenko
arXiv · 2026-07-18
This paper introduces two new benchmarks—Public AI Memory (PAIM) and Public Transformers (PTr)—for evaluating how well long-term memory systems in LLM agents can restore evidence from full scientific papers, rather than just conversations or compact summaries. The authors evaluate eight memory and retrieval systems and find that leaderboard rankings are not meaningful without specifying the full evaluation protocol, including retrieval budget, ingestion granularity, and judge choice. A key finding is that sparse-dense hybrid retrieval (BM25 combined with dense retrieval) is the single most impactful intervention on PTr, while apparent wins on PAIM disappear when retrieval context budgets are controlled. The paper argues for evaluating scientific memory as 'budgeted, modality-aware context restoration' and releases all datasets, code, and evaluation tools to support reproducible benchmarking.
- Quality assurance
- Enterprise
Research
Diagnosing Correctness Probes under Self-Judgement Confounding
Yi-Long Lu
arXiv · 2026-07-18
This paper investigates whether hidden-state 'correctness probes' in language models actually track objective correctness (OC) or merely reflect the model's own self-judgement (SJ). By constructing conflict cases where OC and SJ predict opposite outcomes, the authors find that conventional probes tend to follow the model's self-judgement rather than ground-truth correctness. Across four instruction-tuned models (up to 14B parameters), the SJ-associated direction transfers reliably across domains and tasks, while the OC-associated direction performs below chance in every tested condition. The findings show that transferability of a probe does not confirm it captures objective correctness, raising important concerns for using such probes as reliability signals in AI systems.
- Quality assurance
- Certifications
Research
Translating AI into scientific impact: Field context, career position, and institutional capability in AI-enabled research
Zhiyong Tan, Hongkan Chen, Yi Bu
arXiv · 2026-07-18
This large-scale bibliometric study examines how integrating AI-related knowledge into scientific papers is associated with citation impact, and who benefits most from doing so. Using OpenAlex bibliographic data, the authors find that citing AI literature generally boosts five-year citation counts, but returns vary by scientific field, career stage, and institutional AI capability. Senior scholars gain more from broadly referencing AI, while junior scholars benefit more from intensively citing newer, high-impact AI papers; institutions with intermediate AI capability see the largest proportional citation gains. The findings suggest that translating AI knowledge into scientific impact requires not just technical capability but also 'translational capacity'—the ability to make AI knowledge meaningful and legitimate within diverse scientific communities.
- Workforce
- Enterprise
- AI policy
Research
RECON: Benchmarking Agent Memory for Compositional Reasoning over Long Contexts
Mihir Shriniwas Arya
arXiv · 2026-07-18
RECON is a new benchmark designed to evaluate how well LLM-based agents reason compositionally over very long contexts (50k–100k tokens), spanning 24 case files across criminal, medical, and financial domains. Unlike prior memory benchmarks that test simple fact retrieval or change detection, RECON assesses harder downstream tasks such as tracing cascading invalidations, resolving conflicting sources, and counterfactual reasoning across multiple interactions. Evaluation of current architectures reveals severe limitations: even the best non-Oracle system achieves only 22.4% accuracy, with both retrieval and reasoning identified as major bottlenecks. These findings highlight critical reliability gaps in AI agents used as personal assistants, enterprise copilots, and autonomous workflow agents.
- Enterprise
- Quality assurance
Research
DS@GT ARC at eRisk 2026: Hybrid Multi-Agent LLM System with Structured Algorithmic Guidance for Conversational Depression Screening
Victor Gong, David Guecha
arXiv · 2026-07-18
This paper presents DS@GT's submission to the eRisk 2026 Task 1 challenge, which involves using AI systems to conduct conversational depression screening by interviewing simulated personas and producing Beck Depression Inventory II (BDI-II) scores. The team developed a hybrid multi-agent pipeline that combines a precomputed dialogue tree, reliability-weighted consensus aggregation, and cluster-based symptom imputation to compensate for the weaker reasoning of an open-source model (Gemma 27B) compared to a proprietary one (GPT-5-nano). Their hybrid system achieved an ADODL score of 0.9063, ranking 3rd among complete-submission runs and 2nd among 21 teams overall, while costing roughly one-quarter the per-persona API expense of their paid baseline. The findings suggest that structured algorithmic supervision can enable open-source models to match or exceed proprietary models in sensitive conversational AI tasks, with implications for accessible and cost-effective mental health screening tools.
- Workforce
- Enterprise
- Quality assurance
Research
Though Language Models Err While They Strive: Conformal Prediction for Self-Correcting Scientific Generation
Mingqiao Mo, Yunlong Tan, Hao Zhang
arXiv · 2026-07-18
This paper introduces Scientific Feasibility Control (SFC), a graph-structured conformal prediction framework designed to reduce scientific errors made by large language models when generating technical content. SFC decomposes scientific reasoning into atomic units that must satisfy both individual correctness against physical laws and logical consistency with prior context, using approximate deducibility graphs to model dependencies and branching dynamically when violations are detected. Evaluated on benchmarks including PhyX, MATH, ScienceQA, and ARC Challenge, SFC achieves 50.1% accuracy on PhyX physics reasoning, outperforming DeepSeek-R1 (49.8%) and GPT-4 (45.8%), while delivering 91.7% scientific validity with formal conformal coverage guarantees at alpha=0.10 and reducing scientific law violations by 73% across multiple model architectures. These results matter for quality assurance of AI-generated scientific content, as they demonstrate a statistically grounded method for improving reliability in high-stakes technical applications.
- Quality assurance
- Enterprise
Research
How Do You Choose Your AI Component? An Interview Study of Secure AI Integration in Practice
Mahzabin Tamanna, Elizabeth Lin, Sparsha Gowda et al.
arXiv (Cornell University) · 2026-07-18
This study investigates how software developers, architects, and AI practitioners select and integrate Large Language Model (LLM) components into their systems, focusing on security considerations. Through semi-structured interviews with 22 practitioners across diverse organizations, the researchers find that model selection is predominantly driven by functional criteria—performance, accuracy, cost, and features—while security is rarely treated as an evaluation criterion. The study observes that established software supply chain security lessons are being overlooked, with the industry repeating historically costly mistakes from early software dependency management by prioritizing rapid reuse over security and provenance. The authors offer actionable recommendations for AI adopters, model providers, and researchers to adopt a security-by-design approach throughout the software development lifecycle.
- Enterprise
- AI policy
- Quality assurance
Research
Position: Explanation Stability Is a Property of the Model Method Pair, Not the Model
Kabilan Elangovan, Daniel Ting
arXiv · 2026-07-18
This position paper argues that AI explanation stability depends on the combination of model and attribution method used, not on the model alone. In chest X-ray experiments with DenseNet201, ResNet50V2, and InceptionV3—all achieving AUC above 99%—stability rankings reversed depending on whether LayerCAM or GradCAM++ was used; for example, InceptionV3's stability score dropped by 17.3% when switching methods. The authors contend that claiming a model is 'stable' based on a single attribution method produces scientifically invalid and potentially illusory safety assurances. They recommend that explanation-based claims be validated across multiple attribution paradigms and that regulatory submissions explicitly specify the attribution operators used.
- Quality assurance
- Certifications
- AI policy
Research
Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
Haifeng Li, Mo Hai
arXiv · 2026-07-18
This paper addresses a critical failure mode of LLMs that generate optimization models from natural-language descriptions: a generated model can run successfully while still formulating the wrong problem. The authors develop a falsification-based verification framework using typed numeric slots and solver calls—without needing a reference model—drawing on duality, comparative statics, and polyhedral arguments to build a battery of sound tests (directions, curvature, crush probes, prohibitive limits, annihilation, and exchange). Key results on 326 ground-truth models show the battery achieves a 0.0% false-positive rate versus 54.9% for a threshold tester, detects 70.0% of certified conditional-class mutants, and catches 40.4% of mutants invisible to execution-accuracy scoring. The framework also proves that no fixed-threshold perturbation tester can be simultaneously sound and nontrivial, establishing theoretical limits on what such verification can detect.
- Quality assurance
- Enterprise
Research
CLOSER-Bench: Evaluating Budgeted Cross-Stage Design Closure for Hardware Agents
Peilong Zhou, Zhirong Chen, Cangyuan Li et al.
arXiv · 2026-07-18
CLOSER-Bench introduces a controlled benchmark protocol for evaluating AI coding agents on hardware design closure tasks that span multiple abstraction levels, from RTL generation through physical implementation (RTL-to-GDS). Built on open-source tools including Verilator, Yosys, OpenROAD, and Sky130, the benchmark measures final quality, anytime progress, tool cost, and cross-stage recovery within a fixed budget, exposing a sharp 'completion-closure gap' where agents that solve localized repair tasks fail at the matched verification-closure counterpart. A ten-task pilot covering RTL repair, verification, PPA optimization, and security shows that frontier agents meaningfully outperform baselines on cross-stage tasks, motivating the treatment of hardware closure as a budgeted sequential decision problem rather than independent code generation. These findings have direct implications for enterprise hardware development workflows and quality-assurance processes that increasingly rely on AI agents.
- Enterprise
- Quality assurance
Research
Privacy Cost as Equity Input: A Group Fairness Criterion for Differentially Private Machine Learning
Rakshit Naidu
arXiv · 2026-07-18
This paper introduces the Privacy-Cost Equity Ratio (PCER), a new group fairness metric for differentially private machine learning that accounts for which demographic groups bear the greatest privacy exposure, not just which groups receive accurate predictions. The authors argue that information leakage from DP-SGD training is itself a harm, and that groups facing greater membership inference risk are owed proportionally greater predictive benefit. Evaluated across six benchmark-attribute combinations in tabular and NLP settings, PCER reveals fairness disparities that outcome-based metrics miss — for example, on COMPAS it uncovers a 'double disadvantage' where a protected group suffers both higher privacy exposure and worse predictive outcomes, a pattern hidden by demographic parity gap. The findings suggest that fairness audits of privacy-preserving AI systems must incorporate the distribution of privacy costs across groups, not only the distribution of model benefits.
- Quality assurance
- AI policy
- Certifications
Research
DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines
Runming He, Zhen Hao Wong, Hao Liang et al.
arXiv · 2026-07-18
DataFlow-Harness is a platform that guides large language models to build editable, persistent data-processing pipelines represented as directed acyclic graphs (DAGs), rather than one-off scripts — addressing what the authors call the 'NL2Pipeline gap.' On a 12-task data-engineering benchmark, the system achieves a 93.3% end-to-end pass rate while reducing monetary cost by 72.5% and generation latency by 49.9% compared to a Vanilla Claude Code baseline. The platform combines procedural guidance (DataFlow-Skills), a live operator registry via a Model Context Protocol layer, and a visual DAG editor synchronized with conversational authoring. These results suggest that grounding LLM agents in live platform context can yield reliable, editable workflow artifacts at substantially lower cost than conventional script-generation approaches.
- Enterprise
- Workforce
Research
Can Multimodal Large Language Models Understand OCT?
Baochen Fu, Wenzhi Deng, Baihao Jin et al.
arXiv · 2026-07-18
This paper introduces OCT-Bench, a benchmark of 10,076 multiple-choice questions drawn from 4,137 OCT images across seven public datasets, designed to evaluate how well multimodal large language models (MLLMs) understand optical coherence tomography imaging for retinal disease. The benchmark organizes 20 fine-grained tasks across three dimensions—Perception, Cognition, and Reasoning—mirroring real-world clinical interpretation workflows. Evaluating 20 representative MLLMs, including proprietary, open-source, and medical-domain models, the study finds that current models fall substantially short of reliable OCT understanding, and that neither medical-domain adaptation nor larger model scale consistently improves performance. These findings highlight critical capability gaps relevant to the safe deployment of AI in clinical diagnostic settings.
- Quality assurance
- Certifications
- AI policy
Research
Hazard or Anomaly? Evaluating VLMs for Understanding Dangers and Discrepancies
Murali Indukuri, Mohammad Eskandari, Sree Nitya Kollu et al.
arXiv · 2026-07-18
This paper evaluates Vision-Language Models (VLMs) for safety reasoning in human-robot interaction by introducing a distinction between hazards (true physical dangers) and anomalies (unusual but not necessarily dangerous scene elements). Current evaluations typically frame danger recognition as a binary Safe/Unsafe judgment, which the authors argue obscures whether models are detecting genuine hazards or simply reacting to contextual irregularities. Testing several state-of-the-art VLMs across two datasets and multiple prompting strategies, the study finds that VLMs frequently misinterpret anomalousness as hazardousness, revealing a systematic failure mode. Explicitly separating anomaly from hazard exposes these weaknesses and provides a more informative evaluation framework for safety-critical applications such as disaster response and emergency decision-making.
- Quality assurance
- Certifications
- AI policy
Research
Learning from World Feedback: Why Model Uncertainty Fails as a Risk Signal in Model-Based RL
Zhaohui Wang
arXiv · 2026-07-18
This paper investigates why model uncertainty fails as a safety proxy in model-based reinforcement learning, finding that dynamics-based uncertainty penalties are actually anti-correlated with real-world safety: they increase collision rates from 26% to 34% across four world-model architectures. The authors show that model uncertainty has very low empirical correlation (r < 0.15) with actual task risk because it operates over state-prediction space rather than constraint boundaries. Replacing this internal proxy with world-feedback signals—such as lidar-derived safety margins, time-to-collision estimates, and outcome-trained feedback models—reduces collision rates to 1–14% without retraining any components. The findings are distilled into three design principles (ground risk in world outcomes, validate proxies before deployment, and use outcome-trained feedback models when direct signals are unavailable) that the authors argue apply broadly to RLHF and LLM alignment as well.
- Quality assurance
- Enterprise
- AI policy