News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
5608 items
Research
Bridging the Information Gap: Semantic Densification and Hindsight Distillation for Cold-Start Prediction
Hao Duong Le, Yifei Gao, Huan Li et al.
arXiv · 2026-07-19
This paper presents SemRaD, a framework for predicting new-user lifetime value (LTV) and conversion rate (CVR) on e-commerce platforms where users have sparse interaction histories — the 'cold-start' problem. SemRaD combines structured semantic reasoning from large language models with a teacher-student distillation approach that handles variability in how much privileged information is available per user. On a large-scale industrial dataset and a four-week online A/B test at Keeta, SemRaD delivered measurable gains of +1.9% LTV and +1.0% CVR over a production-grade baseline, and matched the production system's LTV performance using only 9% of the training data. This matters for enterprise recommendation and marketing systems, where better cold-start prediction can directly improve customer acquisition decisions and resource allocation.
- Enterprise
- Quality assurance
Research
Who Will Become the Next Senior? How Generative AI Erodes the Development Pathway in Software Engineering
Sumin Yu, Taesup Moon
arXiv (Cornell University) · 2026-07-19
This qualitative study examines how Generative AI is disrupting the career development pipeline for junior software engineers, based on 14 semi-structured interviews with early-career and senior engineers in South Korea. The researchers identify a pattern they call 'Absorption,' where entry-level tasks are redirected into senior-AI workflows, depriving juniors of the formative productive struggle that traditionally builds expertise. Three downstream consequences are documented: loss of developmental challenge for juniors, normalization of GenAI use in university classrooms that structurally reproduces this loss, and a perceptual asymmetry between seniors and juniors that prevents self-correction. The authors argue that preserving the pathway to senior engineering roles will require deliberate institutional design spanning classrooms, workplaces, and junior evaluation criteria.
- Workforce
- Enterprise
- AI policy
Research
When LLMs Over-Answer: Measuring and Mitigating Quality Issues in LLM-Based Hardware Description Language Question Answering
Ziteng Hu, Jiachi Chen, Wenhao Lv et al.
arXiv · 2026-07-19
This paper investigates quality issues in LLM-generated answers to hardware description language (HDL) questions, finding that LLMs tend to 'over-answer' by burying correct content under redundant alternatives (65.7% of responses) and verbose padding (69.1%), while nearly half (49.0%) fail to fully align with expert answers. The authors built a dataset of 6,246 HDL Q&A posts from Stack Overflow and conducted a user study with 19 HDL engineers to benchmark LLM responses against human expert answers. To address these issues, they propose a multi-agent framework that improves core-answer quality scores from 3.71 to 4.67 and non-core content quality from 3.72 to 4.23 on a five-point scale across four mainstream LLMs. The findings matter because imprecise or verbose HDL answers can propagate into hardware design errors such as timing violations or non-synthesizable logic, making answer quality especially consequential for engineering practice.
- Quality assurance
- Enterprise
- Workforce
Research
Where Does Agent Reliability Come From? A Cross-Benchmark Decomposition of Verification Loops, Specialist Models, and Scaffolding in a Production Enterprise Agent
Arunabh Dastidar
arXiv · 2026-07-19
This paper investigates where reliability gains come from in a production enterprise AI agent (Leni) by evaluating it across three public benchmarks targeting distinct failure modes: silent computation errors, premise confabulation, and cascade errors in long tool chains. The system outperforms its frontier base model by roughly +11 points on SpreadsheetBench, +7–10 points on BullshitBench, and ~+15 points on GAIA validation, with the full system reaching 75.2% pass@1 on GAIA. Crucially, the paper's decomposition finds that most of the performance uplift comes from scaffolding, routing, and specialist models rather than from verification loops alone, whose isolated contribution is small (+1.5 points) but concentrated on otherwise-failing tasks at the top of the score distribution. These findings matter for enterprise AI deployment, showing that architectural choices around specialist post-trained models and task routing are the primary drivers of agent reliability, with implications for how production systems should be designed and quality-assured.
- Enterprise
- Quality assurance
- Certifications
Research
Artificial intelligence in echocardiography: a position statement from the British Society of Echocardiography
Sadie Bennett, Christopher Wild, Maria F. Paton et al.
Echo Research and Practice · 2026-07-19
The British Society of Echocardiography has issued a consensus position statement evaluating how AI can be integrated into echocardiography services, covering the full workflow from image acquisition and analysis to reporting and risk stratification. While noting rapid growth in research activity, the statement identifies practical, technical, and governance challenges that have limited clinical adoption. It provides structured guidance on requirements for safe, equitable, and effective AI integration, primarily in the UK context but with broader applicability to similar healthcare systems.
- AI policy
- Quality assurance
Research
Artificial Intelligence and Academic Integrity in Virtual Higher Education: A Descriptive-Comparative Study of Student and Faculty Perceptions in Ecuador
Héctor Carvajal Romero, Fernanda Tusa, Rosemary Samaniego et al.
Trends in Higher Education · 2026-07-19
This survey-based study compared perceptions of generative AI use and academic integrity between 1,660 students and 34 faculty at the Technical University of Machala, Ecuador during 2024. Students tended to view AI primarily as a useful academic support tool, while faculty were more concerned with authorship, evidence authenticity, and assessment security. Both groups identified limitations in current virtual assessment practices. The authors propose a governance model emphasizing transparent disclosure policies, authentic assessment design, faculty development, student AI literacy, and proportional proctoring rather than blanket prohibition or permissive ambiguity.
- AI policy
- Quality assurance
Research
AI governance and employee well-being in digital workplaces: A systematic literature review
Yulianto, Udin Saryono
Digital Theory Culture & Society · 2026-07-19
This systematic literature review synthesizes 44 peer-reviewed studies (2015–2026) on how AI-mediated governance in digital workplaces affects employee well-being. The review finds that while AI-driven systems improve organizational performance, they also contribute to technostress, burnout, AI anxiety, identity threats, and reduced worker autonomy, with key ethical concerns including algorithmic bias, opacity, and discrimination. The authors propose an integrative conceptual framework linking algorithmic workplace practices, ethical challenges, and employee well-being, moderated by factors such as AI transparency, organizational support, and ethical leadership. The paper concludes with a practical framework for implementing human-centered AI governance in digital workplaces.
- Workforce
- AI policy
Research
Algorithmic Enforcement and the Administrative State: Due Process and Accountability in the EU and the United States
Edward Koellner
Zenodo (CERN European Organization for Nuclear Research) · 2026-07-19
This article examines how public agencies in the EU and the United States are delegating enforcement decisions, risk assessments, and eligibility determinations to AI and algorithmic systems, and what that means for due process and administrative accountability. The authors argue that discretion once held by frontline caseworkers has shifted upstream into technical choices—training data, feature weights, decision thresholds—that effectively become de facto policy without appearing in any traceable administrative record. The paper contrasts Europe's preventive, rights-centered approach (GDPR Article 22, EU AI Act impact assessments, deployment registries) with the U.S.'s reactive, litigation-driven model under the Administrative Procedure Act, finding both systems ultimately converge on the need for auditable rationales, substantive human oversight, and ongoing monitoring for model drift. The authors propose a hybrid regulatory framework that combines Europe's proactive transparency tools with America's robust contestation rights and court-ordered discovery, reinforced through procurement requirements and judicial insistence on legible records.
- AI policy
Research
Algorithmic Enforcement and the Administrative State: Due Process and Accountability in the EU and the United States
Edward Koellner
Zenodo (CERN European Organization for Nuclear Research) · 2026-07-19
This article examines how public agencies in the EU and United States are delegating enforcement decisions, risk assessments, and eligibility determinations to AI and algorithmic systems, and what that shift means for administrative accountability and due process. The authors argue that discretion formerly held by frontline caseworkers has migrated upstream into choices about training data, feature weights, and decision thresholds—choices that function as de facto policy yet rarely appear in reviewable administrative records. Comparing Europe's preventive, rights-centered approach (constitutional proportionality, GDPR Article 22, EU AI Act impact assessments) with the US reactive, litigation-driven model (APA arbitrary-and-capricious review, court-ordered discovery), the paper proposes a hybrid framework combining Europe's pre-deployment tools with America's contestation and audit mechanisms, all reinforced through procurement requirements and judicial insistence on legible records. The analysis matters because it offers a concrete path to keeping automated government decisions accountable as algorithmic discretion becomes increasingly embedded in AI code.
- AI policy
Research
LECTURER’S PERSPECTIVES ON THE USE OF ARTIFICIAL INTELLIGENCE TOOLS IN STUDENTS’ WRITING
Tri Yuli Ardiyansah, R W Batubara
EDUTECH Jurnal Inovasi Pendidikan Berbantuan Teknologi · 2026-07-19
This mixed-methods study surveyed and interviewed nineteen English lecturers at the University of Muhammadiyah Gresik about their views on students using AI tools such as ChatGPT in academic writing. While 84.2% of lecturers were familiar with ChatGPT and 78.9% understood its mechanisms, only 42.1% felt confident they could identify AI-generated text, highlighting a significant detection gap. Lecturers broadly regarded AI as a supplemental instructional aid rather than a replacement for critical thinking, and the study recommends that institutions develop process-based assessment systems and strengthen ethical codes to promote responsible digital literacy.
- AI policy
- Quality assurance
Research
Real-World Evaluation of an AI Agent Drafting Translational Impact Summaries
Mohammad Arvan, Amber E. Osterholt, Bailee Rue et al.
arXiv · 2026-07-18
This paper evaluates a human-in-the-loop AI agent designed to automate the drafting of translational impact summaries for Clinical and Translational Science Award (CTSA) scholars. The agent assembles sourced evidence dossiers and drafts one-sentence impact summaries for staff review, achieving an 81.7% unanimous usable rate across 507 findings from 10 scholars while reducing per-scholar staff time from an estimated 15 hours to a median of 14 minutes. The agent covered all four Translational Science Benefits Model domains and surfaced non-scholarly impact evidence that routine processes often miss, with reviewers rating synthesis accuracy at 4.5 and usefulness at 4.8 out of 5. The findings suggest AI agents can make cohort-scale impact reporting feasible by shifting staff from data collection and writing to reviewing.
- Workforce
- Enterprise
- Quality assurance
Research
TurboVec: A Case Study in Cost-Efficient Private Retrieval for Enterprise RAG via Codebook-Oblivious Quantization
Navnit Shukla, Kamal Pandey, Omsankar Tiwari
arXiv · 2026-07-18
TurboVec is an enterprise vector retrieval system built on TurboQuant, a codebook-oblivious scalar quantizer that requires no corpus-dependent training, addressing two key challenges in multi-tenant RAG deployments: privacy leakage from trained quantizers and recall degradation from post-hoc tenant filtering. On the DBpedia OpenAI embeddings benchmark, TurboQuant 4-bit outperforms trained FAISS Product Quantization by 8.5–8.9 percentage points in Recall@5 at the same memory budget, while using 4–8x less memory than HNSW. Deployed on Snowpark Container Services, TurboVec achieves 11ms median query latency and kernel-level allowlist filtering that maintains 0.86–0.93 Recall@10 across 10–1000 tenant workloads, compared to 0.09–0.19 for post-filter baselines. The codebook-oblivious design reduces membership inference accuracy to near-random (50.0%) versus 57.3% for PQ codebooks, making it relevant for privacy-sensitive enterprise AI applications.
- Enterprise
- Quality assurance
- AI policy
Research
What Do They See? Interpreting Complex Road Scenarios Through the Eyes of Vision-Language-Action Models for Safe and Trustworthy Autonomous Vehicle Learning
Kalpana Panda, Wesley Maia, Vinti Agarwal et al.
arXiv · 2026-07-18
This paper introduces CVAA (Counterfactual Vision Action Analysis), a framework that systematically removes individual objects from front-camera images using photorealistic inpainting to isolate how each object causally influences an autonomous driving model's planning decisions. Applied to the Alpamayo 1 trajectory predictor across 210 nuScenes scenes, the study finds that vehicles and pedestrians in the model's path dominate causal influence, traffic lights have outsized effects relative to their image footprint, and the model sometimes responds strongly to objects a human driver would consider irrelevant. The authors combine this behavioral auditing with mechanistic interpretability techniques to probe intermediate model representations, working toward explainable autonomous driving systems. This research matters for quality assurance and certification of safety-critical AI systems, as it provides a structured method for auditing and building human-AI trust in autonomous vehicle decision-making.
- Quality assurance
- Certifications
- AI policy
Research
Optimizing Clinical Trial Protocols Using EHR-Derived Heterogeneous Treatment Effects
Xiaodi Li, Munhuwan Lee, Pengyang Li et al.
arXiv · 2026-07-18
This study uses real-world electronic health records from the Mayo Clinic Cloud to emulate the DAPA-HF clinical trial and estimate heterogeneous treatment effects (HTEs) of dapagliflozin versus placebo in heart failure patients. While the overall emulated cohort showed no statistically significant survival benefit (HR 1.681, p=0.1507), HTE-guided stratification identified two distinct subgroups: one with a strong survival benefit (HR 0.203, p=0.0002) and one with significantly increased mortality risk (HR 6.680, p<0.0001). These findings demonstrate that HTE-driven patient stratification can reveal clinically meaningful treatment-effect patterns that are obscured when analyzing trial populations as a whole, suggesting a path toward more precise and efficient clinical trial design.
- Quality assurance
- AI policy
- Enterprise
Research
PREFAIL: Identifying Precursors to Failures in Robotic Lift-and-Place Tasks to Improve Task Execution Performance
Zeyu Shangguan, Rajas Chitale, Rutvik Patel et al.
arXiv · 2026-07-18
PREFAIL is a method for predicting failure precursors in robotic lift-and-place tasks used in non-prehensile material handling, where friction-based support makes high-speed motions unreliable. The approach analyzes the relative motion of target objects with respect to their carrier to detect early warning signs of failure, addressing key limitations of existing methods such as sensitivity to dynamic actions and reliance on known policy structures. The authors also introduce a dataset that precisely identifies the latest actionable intervention time, enabling rigorous evaluation of whether a predicted failure can still be prevented. Experimental results on both simulation and real-world data show that PREFAIL improves the accuracy and timeliness of failure precursor detection, which has direct implications for the reliability and efficiency of automated robotic systems in industrial settings.
- Enterprise
- Quality assurance
Research
A Deep Reinforcement Learning Algorithm for the Vehicle Routing Problem with Stochastic Demands and Outsourcing
Mohsen Dastpak, Fausto Errico, Ola Jabali
arXiv · 2026-07-18
This paper introduces the Vehicle Routing Problem with Stochastic Demands and Outsourcing (VRP-SDO), where a logistics provider must decide which customer deliveries to handle with its own fleet versus outsource to a carrier, while managing uncertain demand revealed only upon vehicle arrival. The authors propose a two-level iterative method combining iterated local search for outsourcing decisions with a deep Q-network (using a graph attention network) to estimate routing costs offline, enabling near-instant cost approximations without resolving from scratch each iteration. Experiments show their policy reduces routing costs by 19.6% over a state-of-the-art method and at least 29.6% over classical heuristics, while the full algorithm saves 13.7% on average versus a version without the attention-based representation and produces decisions in minutes rather than over an hour. This matters for enterprise logistics and workforce planning, as it enables faster, higher-quality operational decisions under uncertainty and helps balance labor costs (overtime) against outsourcing expenditures.
- Enterprise
- Workforce
Research
Do Speech Tokens Leak Voiceprints? Speaker Inversion Attacks Against End-to-End Speech Language Models
Ye Lu, Yihan Yan, Zhaoyang Zhang et al.
arXiv · 2026-07-18
This paper investigates whether speech tokens used by end-to-end speech language models (such as Moshi, Higgs3, Kimi-Audio, and Qwen3-Omni) inadvertently expose users' voiceprints. The authors introduce Audio BERT (AuB) and a two-stage inversion method called SpInv, which can recover speaker-identifying embeddings from as little as three seconds of speech token output, achieving cosine similarities above 0.70 in the attacker-specified speaker-encoder space. This demonstrates a significant privacy risk: even without access to raw audio, adversaries can reconstruct biometric voice identifiers from the token representations these models expose. The findings have direct implications for policy around biometric data protection and enterprise deployment of speech AI systems.
- AI policy
- Enterprise
- Quality assurance
Research
Beyond Semantic Equivalence: Logical Graphs for LLM Uncertainty Quantification
Yanni Dong, Minghua Liu, Meiling Zhu et al.
arXiv · 2026-07-18
This paper introduces Logical Graph Uncertainty (LGU), a framework for improving how uncertainty is measured in Large Language Model outputs. Existing methods like semantic entropy treat logically compatible but differently phrased answers as uncertain or hallucinated, when in fact they may simply differ in granularity or specificity. LGU instead explicitly models logical relationships—entailment and incompatibility—among generated answers, producing more accurate uncertainty estimates. Across multiple question-answering benchmarks, LGU outperforms the semantic entropy baseline by up to +7.1% AUROC and +3.5% AUARC, which matters for deploying LLMs reliably in safety-sensitive applications.
- Quality assurance
- Enterprise
- AI policy
Research
Beyond Memory Leaderboards: Evaluating Scientific Memory as Budgeted Context Restoration
Maksim Sheverev, David Finkelstein, Sergey Nikolenko
arXiv · 2026-07-18
This paper introduces two new benchmarks—Public AI Memory (PAIM) and Public Transformers (PTr)—for evaluating how well long-term memory systems in LLM agents can restore evidence from full scientific papers, rather than just conversations or compact summaries. The authors evaluate eight memory and retrieval systems and find that leaderboard rankings are not meaningful without specifying the full evaluation protocol, including retrieval budget, ingestion granularity, and judge choice. A key finding is that sparse-dense hybrid retrieval (BM25 combined with dense retrieval) is the single most impactful intervention on PTr, while apparent wins on PAIM disappear when retrieval context budgets are controlled. The paper argues for evaluating scientific memory as 'budgeted, modality-aware context restoration' and releases all datasets, code, and evaluation tools to support reproducible benchmarking.
- Quality assurance
- Enterprise
Research
Diagnosing Correctness Probes under Self-Judgement Confounding
Yi-Long Lu
arXiv · 2026-07-18
This paper investigates whether hidden-state 'correctness probes' in language models actually track objective correctness (OC) or merely reflect the model's own self-judgement (SJ). By constructing conflict cases where OC and SJ predict opposite outcomes, the authors find that conventional probes tend to follow the model's self-judgement rather than ground-truth correctness. Across four instruction-tuned models (up to 14B parameters), the SJ-associated direction transfers reliably across domains and tasks, while the OC-associated direction performs below chance in every tested condition. The findings show that transferability of a probe does not confirm it captures objective correctness, raising important concerns for using such probes as reliability signals in AI systems.
- Quality assurance
- Certifications
Research
Translating AI into scientific impact: Field context, career position, and institutional capability in AI-enabled research
Zhiyong Tan, Hongkan Chen, Yi Bu
arXiv · 2026-07-18
This large-scale bibliometric study examines how integrating AI-related knowledge into scientific papers is associated with citation impact, and who benefits most from doing so. Using OpenAlex bibliographic data, the authors find that citing AI literature generally boosts five-year citation counts, but returns vary by scientific field, career stage, and institutional AI capability. Senior scholars gain more from broadly referencing AI, while junior scholars benefit more from intensively citing newer, high-impact AI papers; institutions with intermediate AI capability see the largest proportional citation gains. The findings suggest that translating AI knowledge into scientific impact requires not just technical capability but also 'translational capacity'—the ability to make AI knowledge meaningful and legitimate within diverse scientific communities.
- Workforce
- Enterprise
- AI policy
Research
RECON: Benchmarking Agent Memory for Compositional Reasoning over Long Contexts
Mihir Shriniwas Arya
arXiv · 2026-07-18
RECON is a new benchmark designed to evaluate how well LLM-based agents reason compositionally over very long contexts (50k–100k tokens), spanning 24 case files across criminal, medical, and financial domains. Unlike prior memory benchmarks that test simple fact retrieval or change detection, RECON assesses harder downstream tasks such as tracing cascading invalidations, resolving conflicting sources, and counterfactual reasoning across multiple interactions. Evaluation of current architectures reveals severe limitations: even the best non-Oracle system achieves only 22.4% accuracy, with both retrieval and reasoning identified as major bottlenecks. These findings highlight critical reliability gaps in AI agents used as personal assistants, enterprise copilots, and autonomous workflow agents.
- Enterprise
- Quality assurance
Research
DS@GT ARC at eRisk 2026: Hybrid Multi-Agent LLM System with Structured Algorithmic Guidance for Conversational Depression Screening
Victor Gong, David Guecha
arXiv · 2026-07-18
This paper presents DS@GT's submission to the eRisk 2026 Task 1 challenge, which involves using AI systems to conduct conversational depression screening by interviewing simulated personas and producing Beck Depression Inventory II (BDI-II) scores. The team developed a hybrid multi-agent pipeline that combines a precomputed dialogue tree, reliability-weighted consensus aggregation, and cluster-based symptom imputation to compensate for the weaker reasoning of an open-source model (Gemma 27B) compared to a proprietary one (GPT-5-nano). Their hybrid system achieved an ADODL score of 0.9063, ranking 3rd among complete-submission runs and 2nd among 21 teams overall, while costing roughly one-quarter the per-persona API expense of their paid baseline. The findings suggest that structured algorithmic supervision can enable open-source models to match or exceed proprietary models in sensitive conversational AI tasks, with implications for accessible and cost-effective mental health screening tools.
- Workforce
- Enterprise
- Quality assurance
Research
Though Language Models Err While They Strive: Conformal Prediction for Self-Correcting Scientific Generation
Mingqiao Mo, Yunlong Tan, Hao Zhang
arXiv · 2026-07-18
This paper introduces Scientific Feasibility Control (SFC), a graph-structured conformal prediction framework designed to reduce scientific errors made by large language models when generating technical content. SFC decomposes scientific reasoning into atomic units that must satisfy both individual correctness against physical laws and logical consistency with prior context, using approximate deducibility graphs to model dependencies and branching dynamically when violations are detected. Evaluated on benchmarks including PhyX, MATH, ScienceQA, and ARC Challenge, SFC achieves 50.1% accuracy on PhyX physics reasoning, outperforming DeepSeek-R1 (49.8%) and GPT-4 (45.8%), while delivering 91.7% scientific validity with formal conformal coverage guarantees at alpha=0.10 and reducing scientific law violations by 73% across multiple model architectures. These results matter for quality assurance of AI-generated scientific content, as they demonstrate a statistically grounded method for improving reliability in high-stakes technical applications.
- Quality assurance
- Enterprise
Research
How Do You Choose Your AI Component? An Interview Study of Secure AI Integration in Practice
Mahzabin Tamanna, Elizabeth Lin, Sparsha Gowda et al.
arXiv (Cornell University) · 2026-07-18
This study investigates how software developers, architects, and AI practitioners select and integrate Large Language Model (LLM) components into their systems, focusing on security considerations. Through semi-structured interviews with 22 practitioners across diverse organizations, the researchers find that model selection is predominantly driven by functional criteria—performance, accuracy, cost, and features—while security is rarely treated as an evaluation criterion. The study observes that established software supply chain security lessons are being overlooked, with the industry repeating historically costly mistakes from early software dependency management by prioritizing rapid reuse over security and provenance. The authors offer actionable recommendations for AI adopters, model providers, and researchers to adopt a security-by-design approach throughout the software development lifecycle.
- Enterprise
- AI policy
- Quality assurance