News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
Audited Selective Verification for Risk-Controlled N-1 Thermal Contingency Screening under Deployment Shift
Jayakumar Manoharan
arXiv · 2026-07-14
This paper presents Audited Selective Verification (ASV), a risk-budgeted framework for N-1 thermal contingency screening in real-time energy management systems. A cheap surrogate model proposes which outages to skip full power-flow verification for, while an online audit samples and runs full power flow on a subset each window; a calibrated threshold then certifies a bound on the thermal-violation rate for skipped contingencies at a chosen budget and confidence level. Because validity rests on real verification and auditing rather than surrogate accuracy, the guarantees hold even under deployment shift—a condition under which standard deterministic and calibrated screens become unsafe. Tested on three public transmission systems up to 1354 buses, the method keeps realized violation rates within budget while reducing full power-flow studies by 29 to 75 percent per real-time operating point.
- Quality assurance
- AI policy
- Enterprise
Research
"Trust Junk" Leads to Unjustified Support for Highly Discriminatory Predictive Models
Michael Correll, Lucy Havens, Mahsan Nourani
arXiv · 2026-07-14
This paper investigates how data visualizations used in explainable AI (XAI) explanations can cause users to over-trust predictive models. Through a crowdsourced study, the authors demonstrate that presenting accurate but superfluous or irrelevant data alongside model explanations leads users to form unjustified positive beliefs about models—even when those models are clearly discriminatory and unfair. The findings highlight that XAI designers and developers must carefully consider the rhetorical effects of their visualizations to avoid unintentionally lending unearned credibility to harmful models.
- Quality assurance
- AI policy
Research
Resist and Update: Counterfactual Report Coordinates for Incentive-Compatible LLMs
Sen Yang, Yuen-Hei Yeung
arXiv · 2026-07-14
This paper addresses the problem of 'sycophancy' in large language models—where models cave to confident users but fail to update on genuine evidence. The authors formalize this as a failure of internal incentive-compatibility, decomposing it into two demands: resisting illegitimate social pressure and updating on legitimate evidence. Using causal interventions on a Bayesian benchmark with known posteriors, they identify low-rank internal 'report coordinates' controlling answer, confidence, and caveat, and introduce a training-free counterfactual report-coordinate (CRC) clamp that achieves near-perfect joint resist and update scores (1.00, 95% CI [0.99,1.00] in the full-window setting). The method transfers across three model families and to a natural sycophancy benchmark, offering a structural certification approach for incentive-compatible AI behavior.
- Quality assurance
- Certifications
- AI policy
Research
A Multi-Agent System for Autonomous, Fine-Tuning-Free Clinical Symptom Detection: Development and Validation Study
Cameron Cagan, Pedram Fard, Jiazi Tian et al.
arXiv · 2026-07-14
Pythia is a multi-agent AI system that automatically writes and optimizes prompts to extract clinical signs and symptoms from unstructured notes—without manual prompt engineering or model fine-tuning. Tested across 72 symptoms in 400 clinical notes, it achieved mean sensitivity of 0.76 and specificity of 0.95, outperforming a curated lexicon on specificity (0.76 vs. 0.95) and a per-concept BERT classifier that collapsed to near-zero sensitivity for rare concepts. The system runs on locally hosted infrastructure, keeping patient data on-premise, and its specificity transferred well from development to validation sets across varying prevalences. These results suggest autonomous prompt optimization can enable scalable, privacy-preserving clinical NLP without the labeled data burden typically required for fine-tuned models.
- Quality assurance
- Enterprise
- Workforce
Research
LLM Judges Can Be Too Generous When There Is No Reference Answer
Chalamalasetti Kranti, Sowmya Vajjala
arXiv · 2026-07-14
This paper investigates whether large language model (LLM) judges can reliably evaluate open-ended responses when no reference (ground-truth) answer is available. Through calibration and sensitivity experiments spanning three languages, the authors find that LLM judges tend to over-credit incorrect answers in no-reference settings, and that adding a reference answer to the prompt can flip the judge's correct/incorrect decisions by as much as 85% in some conditions. Human annotations confirm that these reference-driven changes generally align with human judgment. The findings highlight the need to calibrate LLM judges using reference-aware evaluation before deploying them in reference-free settings, and the paper offers a practical methodology for doing so.
- Quality assurance
- Enterprise
Research
Evaluating Large Language Models on Misconceptions in Multi-Turn Medical Conversations
Monica Munnangi, Saiph Savage
arXiv · 2026-07-14
This paper introduces ThReadMed-QA, a multi-turn medical dialogue dataset of 2,437 conversation threads and 8,204 question-answer pairs drawn from real patient interactions on AskDocs, designed to evaluate whether large language models (LLMs) can detect and correct patient misconceptions across multiple conversational turns. The authors find that even top-performing models like GPT-5 and Claude-Haiku, which correct false presuppositions roughly 85% of the time on initial questions, drop to around 50% accuracy within two follow-up turns. An oracle analysis shows that much of this degradation stems from error propagation in prior model outputs, though performance remains imperfect even with correct context. These findings highlight serious safety risks in patient-facing AI health tools and underscore the need for evaluation frameworks that account for multi-turn conversational dynamics.
- Quality assurance
- AI policy
Research
Agentic Service-Oriented Computing: A Manifesto for the Next Frontier of Service-Oriented Computing
Amin Beheshti, Rong N. Chang, Boualem Benatallah et al.
arXiv · 2026-07-14
This paper proposes Agentic Service-Oriented Computing (ASOC) as a new research and practice area that applies decades of service-oriented computing principles—such as composition, interoperability, governance, and quality of service—to LLM-powered autonomous agents. The authors argue that today's agentic AI ecosystem is being built ad hoc, without the engineering rigor needed for dependable enterprise and societal deployment. They articulate six foundational principles for ASOC and outline a five-dimensional research agenda covering lifecycle engineering, orchestration, governance, security, and evaluation/certification. The work matters because it frames a path for transforming agentic AI from fragmented demonstrations into trustworthy, accountable, service-based systems suitable for enterprise and broader organizational use.
- Enterprise
- Quality assurance
- Certifications
- AI policy
Research
Vertical Standardisation for High-Risk AI Systems under the EU AI Act: A Domain-Specific Framework for Algorithmic Hiring
Anna Gatzioura, Vrettos Moulos, Nina Baranowska
arXiv · 2026-07-14
This paper proposes a vertical, domain-specific standardisation framework for algorithmic hiring systems classified as high-risk under the EU AI Act. It maps the Act's requirements—covering risk management, data quality and governance, logging and traceability, transparency, human oversight, and accuracy—to concrete recommendations tailored for ranking-based recruitment AI. The framework addresses lifecycle discrimination risks, fairness-aware data governance, explainability, and post-deployment monitoring, filling a gap left by existing horizontal AI governance approaches that do not fully cover algorithmic hiring challenges. The work is informed by the European project FINDHR but is designed to be implementable with alternative methods and governance mechanisms.
- AI policy
- Certifications
- Quality assurance
- Workforce
Research
Evidence-Grounded AI for Musculoskeletal Care
Wenjie Li, Yujie Zhang, Fanrui Zhang et al.
arXiv · 2026-07-14
OrthoPilot is a clinical AI system powered by a large language model that integrates real-time hospital data—imaging, laboratory, pathology, and orders—to support continuous musculoskeletal care from admission through rehabilitation. Benchmarked against 81 orthopaedic physicians across 1,000 disease codes, it outperformed specialists with 25 years of experience in diagnostic reasoning and management planning, and generalized across 60 external clinical centres. In a prospective study of 1,870 complex cases it improved full-chain management success by 10.6%, and a randomized deployment involving 8,240 inpatients increased cumulative cases per bed by 9.7% while improving patient-reported access to health information. The work demonstrates that clinical AI can move beyond isolated predictions to execute longitudinal management across complete care pathways, with direct implications for care quality, hospital efficiency, and physician decision support.
- Enterprise
- Quality assurance
- Workforce
Research
On-Device Deep Research at 4B: Exposure Bounds Faithfulness, Retrieval Bounds Coverage
Vinay Kumar Chaganti
arXiv · 2026-07-14
This study evaluates on-device research agents running a 4B-parameter language model on a 24 GB laptop, finding that citation faithfulness and source coverage are governed by distinct factors. Exposing the model to more text per source (400 vs. 1500 characters) raises cited-claim faithfulness from roughly 0.45 to 0.58, regardless of whether the sources are gold or retrieved, while trustworthy coverage remains near 0.22 because it is limited by retrieval recall (~0.40) rather than exposure. The practical implication is that practitioners should first increase per-source exposure—a cheap intervention costing about 235 extra output tokens—and then focus on improving retrieval recall as the only remaining lever for coverage. These findings matter for deploying reliable, locally-run AI research tools that produce verifiable, cited outputs.
- Enterprise
- Quality assurance
Research
Regulating Artificial Intelligence in Developing and Centralized Legal Systems: A Comparative Study of Indonesia and China
Christian Andersen, Shavilla Felisya Regitara
Journal Of Social Research · 2026-07-14
This comparative legal study examines AI regulation in Indonesia and China, finding that Indonesia relies on a fragmented, policy-based approach lacking specific AI legislation, while China has built a more comprehensive and enforceable framework for algorithmic systems. The research identifies key regulatory gaps in Indonesia, including limited institutional coordination and inadequate algorithmic accountability mechanisms. The authors argue Indonesia should develop a risk-based, legally binding AI governance framework informed by China's regulatory experience while accounting for its own legal context.
- AI policy
- Certifications
Research
AI Trustworthiness in Managerial Decision-Making: Ethics, Transparency, and Explainability as Key Drivers
Guangming Cao, Yanqing Duan, John S. Edwards
Journal of Business Ethics · 2026-07-14
This study uses structural equation modeling on survey data from UK managers to examine how perceived AI trustworthiness—defined through ethics, transparency, and explainability—relates to the adoption of predictive AI and generative AI and their effects on decision quality, speed, and integrity. Findings show that higher perceived trustworthiness is positively associated with adopting both AI types, which in turn improve perceived decision efficacy, with predictive AI showing stronger and more consistent effects. Notably, decision complexity weakens the positive link between generative AI and perceived decision efficacy but does not affect predictive AI's relationship. The research contributes to debates on responsible AI, managerial accountability, and ethical technology use in organizations, highlighting trustworthiness dimensions as key enablers of effective AI-assisted management.
- Enterprise
- AI policy
- Workforce
Research
Artificial Intelligence and Healthcare Policy: A Bibliometric Analysis of Global Research Trends
Pegah Rashidian, Forough Heidarzad-Pahlaviani, Seyedsina Moghimnejadhosseini et al.
Healthcare · 2026-07-14
This bibliometric study maps global research trends at the intersection of artificial intelligence and healthcare policy using 347 peer-reviewed articles from 2000 to 2026. Publications surged after 2020, peaking in 2025, with contributions from researchers in 82 countries and 900 institutions, led by the United States, China, England, Canada, and India. Key themes include machine learning, COVID-19, health policy, large language models, and public health applications, with growing focus on predictive modeling and public health decision-making. The findings provide a roadmap for researchers, policymakers, and healthcare leaders navigating the rapidly evolving AI-in-health-policy landscape.
- AI policy
- Workforce
- Enterprise
Research
Partial Identification with Multiple Nonlinear Measurements of a Latent Regressor
Burhan Ogut, Michelle Yin
arXiv · 2026-07-13
This paper addresses a core measurement problem in AI labor-market research: multiple competing occupational AI-exposure scores produce downstream employment estimates that differ by a factor of eleven, making it impossible to know which (if any) reflects the true structural effect. The authors develop a partial-identification framework for linear regression with a latent regressor observed through nonlinear measurements, deriving a closed-form interval centered on a cross-source estimator whose half-width is second-order in curvature heterogeneity and invariant to unknown source loadings. Applied to six AI-exposure measures matched to an American Community Survey panel of 8.88 million person-year observations (2015–2024), the method yields a loading-invariant consensus employment coefficient of -0.239, with a partial-identification half-width of just 1.23 percent of the point estimate, though the post-2022 employment coefficient changes sign between language-model and patent-text measures. The work matters for workforce and policy research because it provides a principled, estimable way to reconcile conflicting AI-exposure metrics rather than arbitrarily selecting one, and it reveals that at least one prominent measure (Webb patent-text) captures a distinct construct from the others.
- Workforce
- AI policy
Research
A Threshold Exceedance Framework for CBRN Uplift Evaluation in Frontier Language Models
Rahul Gupta, Abhinav Mohanty, Payal Motwani et al.
arXiv · 2026-07-13
This paper introduces the Threshold Exceedance Criteria (TEC) framework for evaluating whether frontier language models materially increase a non-expert's ability to plan Chemical, Biological, Radiological, or Nuclear (CBRN) attacks compared to publicly available tools alone. The framework decomposes uplift studies into standardized, independently executable components—participant eligibility, threat scope definition, and statistical uplift estimation—addressing the lack of comparability across existing CBRN evaluations. In a large-scale empirical study, the authors found domain heterogeneity: while model-assisted plans sometimes achieved expert-equivalent instructional ratings, confirmed material uplift was limited to the radiological domain, and these findings directly informed mitigation and deployment-governance decisions. The work offers methodological lessons emphasizing prespecified criteria, explicit baselines, and separation of generative versus revisionist uplift estimates for future safety evaluations.
- AI policy
- Certifications
- Quality assurance
Research
Cost-Governed RAG: Unified Per-Tenant Cost Attribution Across Retrieval and Generation in Multi-Tenant LLM Systems
Navnit Shukla
arXiv · 2026-07-13
Cost-Governed RAG presents an architecture that closes a governance gap in enterprise multi-tenant Retrieval-Augmented Generation (RAG) systems, where LLM token costs are metered but retrieval-layer costs—vector memory, similarity compute, and embedding API calls—are unattributed and effectively cross-subsidized across tenants. The system integrates a codebook-oblivious vector index (TurboVec) with a multi-tenant LLM governance gateway to create a unified observability stack that jointly attributes embedding, retrieval, and generation costs per tenant. Deployed on Snowpark Container Services, the architecture achieves 99.96% end-to-end cost attribution accuracy across 100 simulated tenants with telemetry overhead below 0.04% of query latency, and reduces retrieval infrastructure cost by 3.1–9.0x compared to managed vector database services under the pricing assumptions described in the paper. This matters for enterprise AI deployments where accurate, per-tenant cost accountability is required for fair billing, governance compliance, and infrastructure efficiency.
- Enterprise
- AI policy
- Quality assurance
Research
TRAIL: A Platform for Configurable Human--AI Teaming Experiments
Mohammad Amin Samadi, Pedro Martins De Bastos, Jaeyoon Choi et al.
arXiv · 2026-07-13
TRAIL is a web platform designed to enable rigorous, reproducible experiments on human-AI teaming by making the AI teammate's design properties—such as personality (Big Five persona), communication style, and participation frequency—fully configurable. The platform supports longitudinal multi-session experiments, dual memory, and export-ready analytics, and was validated in a real six-session classroom deployment with approximately 51 students. In that deployment, a single blind persona change produced a measurable double dissociation: a cognitive-scaffolding AI persona led to stronger contribution ratings and closer linguistic alignment, while a socially-supportive persona produced warmer team climate and lower over-reliance. This matters for workforce and enterprise contexts where AI teammates are increasingly embedded in collaborative work, as it provides infrastructure to systematically study how AI design choices affect trust, coordination, and decision-making.
- Workforce
- Enterprise
- Quality assurance
Research
Self-Improving AI Coding Agents Through Accumulated Behavioral Rules: A Closed-Loop Framework
Aditya Aggarwal, Nahid Farhady Ghalaty
arXiv · 2026-07-13
This paper presents a closed-loop framework that enables LLM-based coding agents to persistently learn from human code review feedback across sessions by codifying accepted review comments into behavioral rules stored in a version-controlled instruction file. Deployed on a 35+ service microservices platform, the rule set grew from 5 to 18 behavioral rules and a 15-item self-review checklist derived from real review feedback, achieving a measured 0% recurrence rate for ruled-against error classes across 11 recorded working sessions. The framework shifts reviewer effort from low-level correctness checking toward design-level validation and transfers across heterogeneous agent interfaces without requiring any model weight updates. Compared to prior work like Reflexion, ExpeL, and CodeReviewer, this approach uniquely addresses behavioral consistency over time on production codebases rather than synthetic benchmarks.
- Enterprise
- Quality assurance
- Workforce
Research
Token Reduction Is Not Cost Reduction
Sarel Weinberger, Amir Hozez
arXiv · 2026-07-13
This paper challenges the common assumption that reducing tokens in AI coding agent pipelines translates directly to lower costs. Through a large controlled experiment of nearly 2,900 billed Claude Code runs across 103 tasks and seven repositories, the authors find that prompt-cache traffic dominates actual billed costs (~87%), meaning local token removal is a poor predictor of end-to-end expense—one compression approach actually increased paired cost by 6.8% despite removing 38% of raw tool-output tokens. Critically, aggressive compression also degraded task success, reducing successful patch application from 27/40 to 15/40 on coding tasks by corrupting edit anchors. The authors argue that context-reduction systems should be evaluated by success-adjusted billed cost, not token reduction alone.
- Enterprise
- Quality assurance
Research
Toward Trustworthy Autonomous Science: A Two-Year Community Roadmap
Rafael Ferreira da Silva, Milad Abolhasani, Peter Beaucage et al.
arXiv · 2026-07-13
This roadmap paper updates a prior framework for autonomous scientific laboratories, arguing that the central bottleneck in AI-driven science has shifted from generating candidate discoveries to verifying them. The authors assess 14 prior milestones and add four new ones, elevating trust/verification/reproducibility and safety/security/governance to first-class concerns alongside the original five dimensions. Evidence cited includes multi-agent systems producing validated hypotheses, self-driving labs becoming more interoperable, a corrected flagship discovery result, benchmarks showing agents complete only a fraction of open-ended research tasks, and fabricated citations appearing at leading venues. The two-year roadmap scopes a path through interface and protocol standardization in year one toward federated, zero-trust coordination and governance in year two, with relevance for how national programs and commercial platforms interoperate without re-siloing.
- Quality assurance
- Certifications
- AI policy
- Enterprise
Research
Operationalising Multi-Dimensional Evaluation for Conversational Agents: A Scalable, Governed Pipeline with Selective Re-evaluation and Model Benchmarking
Niranjan Kumar M, Balaji Nagarajan, Karthik Nair et al.
arXiv · 2026-07-13
This paper presents GenAI Evaluation, a governed, configuration-driven pipeline for large-scale quality assessment of retail conversational AI agents. The system processes approximately 50,000 records daily and has evaluated more than two million chatbot interactions, scoring responses across dimensions including helpfulness, truthfulness, clarity, tone alignment, and translation quality. Validated against 12,980 human-labeled records from four trained annotators, the pipeline achieved a macro F1 score of 0.93 and 89% human-acceptability accuracy for translation, demonstrating that LLM-as-a-judge methods can reliably substitute for human evaluation at scale. The framework addresses production challenges such as governance, reproducibility, cost, and auditability through schema locking, versioned configurations, and record-level provenance tracking, making it relevant to enterprise deployment and quality assurance of AI systems.
- Enterprise
- Quality assurance
Research
Calibrated Selective Prediction Using Deep Ensembles for ROI-Based Thyroid Nodule Ultrasound Classification Under Dataset Shift: A Retrospective Evaluation
Md. Sadibul Hasan Sadib, Md. Mohayminul Mukit, Rahmatul Kabir Rasel Sarker et al.
arXiv · 2026-07-13
This paper develops a calibrated deep ensemble model for classifying thyroid nodule ultrasound images and deciding when to recommend fine-needle aspiration (FNA) biopsy versus radiologist review. The five-member ConvNeXt-Tiny ensemble achieved strong internal discrimination and calibration (AUC-ROC 0.9395, near-zero ECE) with a three-tier triage policy that captured 99.83% of malignancies while routing uncertain cases to human review. However, when evaluated on an external dataset (TN3K) representing dataset shift, AUC-ROC dropped to 0.7870 and calibration degraded substantially, with most cases routed to radiologist review and FNA positive predictive value falling to 76.6%. The authors conclude that local recalibration, threshold validation, and prospective clinical trials are necessary before this framework can be safely deployed as clinical decision support.
- Quality assurance
- Certifications
- AI policy
Research
Designing Agent-Ready Websites for AI Web Agents: A Framework for Machine Readability, Actionability, and Decision Reliability
Said Elnaffar, Farzad Rashidi
arXiv (Cornell University) · 2026-07-13
This paper introduces the 'agent-ready website' framework, a design approach aimed at making e-commerce platforms more interpretable, executable, and decision-reliable for AI browser agents that autonomously shop on behalf of users. The authors evaluated the framework in a controlled experiment using three AI models (GPT-4.1, Gemini-2.5 Flash, and Grok-4 Fast) across 300 runs on matched website prototypes, finding that the agent-ready design achieved a strict success rate of 89.3% versus 49.3% for the human-oriented baseline, while also reducing average step counts from 9.31 to 6.49. The largest gains appeared in tasks involving product detail extraction, comparison, and multi-constraint selection. These results suggest that deliberate structural and semantic web design choices can substantially improve the reliability and efficiency of AI agents performing real-world e-commerce tasks.
- Enterprise
- Quality assurance
Research
Agentic systems for breast cancer treatment recommendations
Vinicius Anjos de Almeida, Nícolas Henrique Borges, Leonardo Vicenzi et al.
arXiv · 2026-07-13
This study evaluated seven agentic large language model (LLM) pipelines for generating breast cancer treatment recommendations across 72 real clinical cases spanning stages I–IV, using 1,147 case-specific rubrics created via Asymmetric Information Rubric Generation (AIRG). The best-performing system, Claude Opus 4.8 with a Divide-and-Conquer plus subagent pipeline, achieved a global score of 0.594 ± 0.025, with tool use and increased agent autonomy showing mixed effects on performance. Oncologist-led error analysis identified persistent clinically relevant failures including incorrect or missing recommendations, flawed justifications, citation errors, outdated claims, and overconfidence. The authors conclude that while agentic LLMs can generate clinically relevant breast cancer recommendations, they remain insufficient for unsupervised clinical use.
- Quality assurance
- AI policy
- Enterprise
Research
Underwriting the Agent Economy: The Blueprint for an AI Insurance Stack
Cristian Trout, Sanmi Koyejo, Sasha Romanosky et al.
arXiv · 2026-07-13
This report argues that the emerging AI agent economy—projected to handle trillions of dollars in transactions by 2030—requires a purpose-built insurance infrastructure to price risk, limit downside, and spread best practices. The authors identify key barriers to insurability, including silent coverage gaps, growing exclusions, AI capabilities outpacing reliability, and concentration among a few foundation model providers creating correlated loss risks. They propose an eight-component AI insurance stack covering incident data collection, catastrophe modeling, standards, contract design, risk selection, pricing, monitoring, and claims management, drawing on historical precedents such as Underwriters Laboratories and the Closed Claims Project. The report also addresses catastrophic tail risks from frontier AI—including CBRN threats, critical infrastructure collapse, and loss-of-control scenarios—suggesting instruments such as catastrophe bonds, frontier model developer mutuals, bespoke liability regimes, and government backstops.
- Enterprise
- AI policy
- Certifications