News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
5608 items
Research
AI-accelerated End-to-End Framework for Rapid Professional Upskilling
Tam Nguyen, H S. Nguyen, Robert Ogburn
arXiv (Cornell University) · 2026-07-15
This paper presents an end-to-end AI-accelerated framework for professional upskilling spanning five stages: knowledge acquisition, content development, content review, teaching, and assessment development. The framework was validated externally when NASBA approved a program built on it for continuing professional education credits, three learners passed the NVIDIA Certified Professional in Agentic AI exam in a notably short time, and the framework's knowledge base generated a 1,267-item risk dataset for managing multi-agent AI systems. The authors argue this addresses a growing enterprise skills gap, where the average time to close such gaps grew from roughly 3 days in 2014 to 36 days in 2018. The work is notable for covering the full upskilling pipeline rather than accelerating only individual stages.
- Workforce
- Certifications
Research
Persona Migration and Expectation Recalibration in Generative AI Adoption: A Longitudinal Study at a State Department of Transportation
Omidreza Shoghli, Fatemeh Banani Ardecani, Amin Mohamadi Hezaveh
arXiv (Cornell University) · 2026-07-15
This longitudinal study tracked 124 state Department of Transportation employees over an eight-week Microsoft 365 Copilot pilot, finding that perceived usefulness declined significantly after hands-on use, suggesting initial expectations were overstated. Using persona-based clustering, the study found substantial individual shifts: 40% of initial Skeptics became more positive, while 68% of initial Champions became less enthusiastic. Job and skills concerns increased over the pilot period, even as accuracy and privacy concerns fell. The authors recommend dynamic monitoring of AI adoption and persona-specific training, workflow examples, and trust-calibration safeguards for public-sector deployments.
- Workforce
- Enterprise
- AI policy
Research
Messy Research, Certification and the Monetization of Science
J. Fourie
arXiv (Cornell University) · 2026-07-15
This theoretical economics paper examines how AI tools that lower the cost of producing polished research manuscripts reshape the institutions that certify scientific quality. Because AI reduces production costs faster than it reduces the cost of evaluating genuine contribution, polish becomes a less reliable signal and the average quality of uncertified submissions can decline. This deterioration of the outside option increases willingness to pay for credible certification, giving certifiers with market power the ability to capture a premium, while fixed review capacity can lead to certification dilution instead. The core finding is that cheaper AI-assisted research production shifts scarcity from making work look credible to verifying which work actually is credible.
- Certifications
- Quality assurance
- AI policy
Research
Thinking about the Impact of Artificial Intelligence on U.S. Health Care Costs and Spending Growth
Bob Kocher, Brian Zhao, Erin Duffy
NEJM Catalyst · 2026-07-15
This article examines how AI is likely to affect U.S. health care costs and spending across areas such as drug innovation, remote patient monitoring, chronic care management, clinical decision support, and administrative labor. The authors argue that under the current fee-for-service payment model and consolidated hospital and insurance markets, AI is more likely to increase total costs and spending growth in the short to medium term, even while improving patient access and clinical quality. They contend that the cost-reducing potential of AI depends heavily on payment model structure—value-based versus fee-for-service—and that without policy and reimbursement reforms, AI will not slow cost growth. Regulators are urged to introduce policy and reimbursement levers to ensure AI delivers on its cost-bending potential.
- AI policy
- Enterprise
- Workforce
Research
From the EU AI Act to Audit Practice: A Governance-to-Controls Framework for Quality Management and Evidence
János Kálmán
Accounting and Auditing · 2026-07-15
This conceptual paper translates the EU AI Act's (Regulation 2024/1689) governance requirements into practical audit-quality controls and evidence-evaluation criteria for audit firms using AI tools such as machine-learning models and generative AI. It maps AI Act obligations—covering risk management, data governance, transparency, human oversight, and logging—onto firm-level and engagement-level quality-management frameworks, producing artefacts including a crosswalk with IAASB standards, an evidence-risk typology, a documentation checklist, and a maturity model. The framework specifies three modes of AI Act relevance (direct legal, indirect, and benchmark) and clarifies the conditions under which AI outputs may serve as triage, corroborative, or substantive audit evidence. The work is significant for audit practice and standard-setting because it provides a scalable, inspection-ready governance structure for defensible reliance on AI-enabled audit tools without overstating the Act's direct legal applicability.
- Quality assurance
- Certifications
- AI policy
- Enterprise
Research
Privacy Preserving Recommender Systems Balancing Personalization with Privacy
Ranjeet K Jha, Venkata Suresh Gummadilli
arXiv · 2026-07-14
This paper proposes and evaluates a privacy-preserving recommendation framework for e-commerce platforms that combines federated learning, differential privacy, cohort-level modeling, and privacy-aware agents to keep raw user data decentralized while adding mathematically bounded noise to model updates. Experiments on synthetic retail datasets show that the framework maintains competitive recommendation quality—measured via CTR, Precision@K, Recall@K, and NDCG@K—at moderate privacy budgets (approximately ε≈5), demonstrating that strong privacy guarantees can be achieved with limited impact on recommendation effectiveness. The work directly addresses compliance requirements under GDPR, CCPA, and CPRA, offering a scalable approach for AI-driven retail platforms that must balance personalization, regulatory compliance, and business objectives.
- Enterprise
- AI policy
- Quality assurance
Research
Where Does the Noise Come From? A Variance-Components Decomposition of Non-Determinism in LLM Brand Answers
Dmitrij Żatuchin
arXiv · 2026-07-14
This paper investigates why brand scores produced by large language models (LLMs) are unstable, decomposing the total variance of LLM brand-recommendation answers into four sources: within-prompt resampling, prompt paraphrase, model identity, and query language. Using nearly 13,000 LLM responses across 20 brands, 8 languages, and 3 models, the study finds that query language accounts for 26.5% of response variance while brand identity accounts for only 1.5%, meaning a single AI answer carries almost no brand-discriminating signal. The research shows that reliability is best improved by diversifying across languages and models rather than simply repeating prompts, with each additional repeat past the fifth contributing negligibly to reducing error variance. These findings matter for enterprises and quality-assurance practitioners who use LLM outputs to measure brand perception, as they reveal that common measurement practices based on a handful of prompt repetitions are insufficient for reliable brand tracking.
- Enterprise
- Quality assurance
Research
Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners
Haseeb Shah, Lingwei Zhu, Adam White et al.
arXiv · 2026-07-14
This paper conducts a large-scale empirical study of over 33,000 experiments to evaluate how key design components in actor-critic reinforcement learning algorithms—such as action distribution type, gradient estimators, and update scheduling—affect reliability and hyperparameter sensitivity. Using a control task derived from a real water treatment plant, the authors find that commonly used defaults like Gaussian action distributions with pathwise gradient estimators are among the least reliable configurations, while bounded distributions with adaptive update schedules prove more robust. The findings provide concrete, component-level guidance for practitioners deploying reinforcement learning in high-stakes real-world systems where reliability is critical and tuning resources are limited.
- Enterprise
- Quality assurance
Research
Cost-Optimal Foundation Model Deployment Portfolio for Transportation Management
Xi Cheng, Ke Liu, Siyuan Feng et al.
arXiv · 2026-07-14
This paper addresses the challenge of deploying multiple AI foundation models—including large language models and vision-language models—across transportation management center (TMC) functions such as anomaly detection and incident reporting. The authors formulate the Foundation Model Deployment Portfolio (FMDP) problem as a mixed-integer program that minimizes total cost of ownership while satisfying quality, latency, and safety constraints under a shared GPU budget, and prove the problem is NP-hard. A greedy heuristic is proposed and applied in a case study with five TMC functions and 19 candidate model-deployment pairs, yielding a mixed portfolio costing $34/month—97% below the cheapest all-closed-API baseline—by routing most functions to open-source APIs and only one quality-constrained function to a closed API. The work provides actionable guidance for transportation agencies on model selection, deployment mode, and when on-premise GPU investment becomes cost-justified (above approximately 309 vision queries/hour or if API prices double).
- Enterprise
- AI policy
Research
Continuously Evolving Deepfake Detection: An Architecture and Public-Benchmark Evaluation of a Dynamic Detection System
Ken Jon Miyachi, Dylan Uys
arXiv · 2026-07-14
This paper presents BitMind Forensics (BMF), a deepfake detection system designed to continuously update its training distribution through an open adversarial competition called Bittensor SN34, directly addressing the well-documented problem of static detectors collapsing on real-world content—where state-of-the-art models have shown AUC drops of 45–50% in in-the-wild evaluations. Evaluated across nineteen public datasets spanning face-swap, AI-generated images, and video benchmarks, BMF achieves strong results including 0.936 AUC on Sumsub original images, 0.915 on Deepfake-Eval-2024 images (matching the best commercial detector), and 0.947 on DFDC—outperforming the FF++-trained frontier (0.843). A temporal study further shows successive model exports improve detection of generators unseen during static baseline training, demonstrating measurable gains from continual retraining. The work is relevant to quality assurance and policy by providing a publicly reproducible evaluation harness and a verifiable production API snapshot, offering practitioners and regulators a credible benchmark for assessing deepfake detection robustness.
- Quality assurance
- AI policy
- Certifications
Research
AI-Native Insurance for Agentic AI: Pricing, Underwriting, and End-to-End Automation
Quanyan Zhu
arXiv · 2026-07-14
This paper develops a mathematical framework for insuring agentic AI systems—autonomous AI that can make decisions, use tools, and interact with external services—addressing the novel risks these systems introduce. The framework models deployments via a 'risk state' capturing autonomy level, operational authority, permission exposure, governance maturity, and dependency concentration, then maps these factors to event probabilities, loss severities, premiums, deductibles, and policy covenants. The authors establish structural properties of insurability, including an insurability region, how feasibility deteriorates with increasing exposure, and governance certification thresholds, and illustrate the framework through a healthcare case study with contract optimization and automated claims processing. The work is relevant to enterprise AI deployment, certification and governance standards, and emerging policy frameworks for AI liability.
- Enterprise
- Certifications
- AI policy
Research
Audited Selective Verification for Risk-Controlled N-1 Thermal Contingency Screening under Deployment Shift
Jayakumar Manoharan
arXiv · 2026-07-14
This paper presents Audited Selective Verification (ASV), a risk-budgeted framework for N-1 thermal contingency screening in real-time energy management systems. A cheap surrogate model proposes which outages to skip full power-flow verification for, while an online audit samples and runs full power flow on a subset each window; a calibrated threshold then certifies a bound on the thermal-violation rate for skipped contingencies at a chosen budget and confidence level. Because validity rests on real verification and auditing rather than surrogate accuracy, the guarantees hold even under deployment shift—a condition under which standard deterministic and calibrated screens become unsafe. Tested on three public transmission systems up to 1354 buses, the method keeps realized violation rates within budget while reducing full power-flow studies by 29 to 75 percent per real-time operating point.
- Quality assurance
- AI policy
- Enterprise
Research
"Trust Junk" Leads to Unjustified Support for Highly Discriminatory Predictive Models
Michael Correll, Lucy Havens, Mahsan Nourani
arXiv · 2026-07-14
This paper investigates how data visualizations used in explainable AI (XAI) explanations can cause users to over-trust predictive models. Through a crowdsourced study, the authors demonstrate that presenting accurate but superfluous or irrelevant data alongside model explanations leads users to form unjustified positive beliefs about models—even when those models are clearly discriminatory and unfair. The findings highlight that XAI designers and developers must carefully consider the rhetorical effects of their visualizations to avoid unintentionally lending unearned credibility to harmful models.
- Quality assurance
- AI policy
Research
Resist and Update: Counterfactual Report Coordinates for Incentive-Compatible LLMs
Sen Yang, Yuen-Hei Yeung
arXiv · 2026-07-14
This paper addresses the problem of 'sycophancy' in large language models—where models cave to confident users but fail to update on genuine evidence. The authors formalize this as a failure of internal incentive-compatibility, decomposing it into two demands: resisting illegitimate social pressure and updating on legitimate evidence. Using causal interventions on a Bayesian benchmark with known posteriors, they identify low-rank internal 'report coordinates' controlling answer, confidence, and caveat, and introduce a training-free counterfactual report-coordinate (CRC) clamp that achieves near-perfect joint resist and update scores (1.00, 95% CI [0.99,1.00] in the full-window setting). The method transfers across three model families and to a natural sycophancy benchmark, offering a structural certification approach for incentive-compatible AI behavior.
- Quality assurance
- Certifications
- AI policy
Research
A Multi-Agent System for Autonomous, Fine-Tuning-Free Clinical Symptom Detection: Development and Validation Study
Cameron Cagan, Pedram Fard, Jiazi Tian et al.
arXiv · 2026-07-14
Pythia is a multi-agent AI system that automatically writes and optimizes prompts to extract clinical signs and symptoms from unstructured notes—without manual prompt engineering or model fine-tuning. Tested across 72 symptoms in 400 clinical notes, it achieved mean sensitivity of 0.76 and specificity of 0.95, outperforming a curated lexicon on specificity (0.76 vs. 0.95) and a per-concept BERT classifier that collapsed to near-zero sensitivity for rare concepts. The system runs on locally hosted infrastructure, keeping patient data on-premise, and its specificity transferred well from development to validation sets across varying prevalences. These results suggest autonomous prompt optimization can enable scalable, privacy-preserving clinical NLP without the labeled data burden typically required for fine-tuned models.
- Quality assurance
- Enterprise
- Workforce
Research
LLM Judges Can Be Too Generous When There Is No Reference Answer
Chalamalasetti Kranti, Sowmya Vajjala
arXiv · 2026-07-14
This paper investigates whether large language model (LLM) judges can reliably evaluate open-ended responses when no reference (ground-truth) answer is available. Through calibration and sensitivity experiments spanning three languages, the authors find that LLM judges tend to over-credit incorrect answers in no-reference settings, and that adding a reference answer to the prompt can flip the judge's correct/incorrect decisions by as much as 85% in some conditions. Human annotations confirm that these reference-driven changes generally align with human judgment. The findings highlight the need to calibrate LLM judges using reference-aware evaluation before deploying them in reference-free settings, and the paper offers a practical methodology for doing so.
- Quality assurance
- Enterprise
Research
Evaluating Large Language Models on Misconceptions in Multi-Turn Medical Conversations
Monica Munnangi, Saiph Savage
arXiv · 2026-07-14
This paper introduces ThReadMed-QA, a multi-turn medical dialogue dataset of 2,437 conversation threads and 8,204 question-answer pairs drawn from real patient interactions on AskDocs, designed to evaluate whether large language models (LLMs) can detect and correct patient misconceptions across multiple conversational turns. The authors find that even top-performing models like GPT-5 and Claude-Haiku, which correct false presuppositions roughly 85% of the time on initial questions, drop to around 50% accuracy within two follow-up turns. An oracle analysis shows that much of this degradation stems from error propagation in prior model outputs, though performance remains imperfect even with correct context. These findings highlight serious safety risks in patient-facing AI health tools and underscore the need for evaluation frameworks that account for multi-turn conversational dynamics.
- Quality assurance
- AI policy
Research
Agentic Service-Oriented Computing: A Manifesto for the Next Frontier of Service-Oriented Computing
Amin Beheshti, Rong N. Chang, Boualem Benatallah et al.
arXiv · 2026-07-14
This paper proposes Agentic Service-Oriented Computing (ASOC) as a new research and practice area that applies decades of service-oriented computing principles—such as composition, interoperability, governance, and quality of service—to LLM-powered autonomous agents. The authors argue that today's agentic AI ecosystem is being built ad hoc, without the engineering rigor needed for dependable enterprise and societal deployment. They articulate six foundational principles for ASOC and outline a five-dimensional research agenda covering lifecycle engineering, orchestration, governance, security, and evaluation/certification. The work matters because it frames a path for transforming agentic AI from fragmented demonstrations into trustworthy, accountable, service-based systems suitable for enterprise and broader organizational use.
- Enterprise
- Quality assurance
- Certifications
- AI policy
Research
Vertical Standardisation for High-Risk AI Systems under the EU AI Act: A Domain-Specific Framework for Algorithmic Hiring
Anna Gatzioura, Vrettos Moulos, Nina Baranowska
arXiv · 2026-07-14
This paper proposes a vertical, domain-specific standardisation framework for algorithmic hiring systems classified as high-risk under the EU AI Act. It maps the Act's requirements—covering risk management, data quality and governance, logging and traceability, transparency, human oversight, and accuracy—to concrete recommendations tailored for ranking-based recruitment AI. The framework addresses lifecycle discrimination risks, fairness-aware data governance, explainability, and post-deployment monitoring, filling a gap left by existing horizontal AI governance approaches that do not fully cover algorithmic hiring challenges. The work is informed by the European project FINDHR but is designed to be implementable with alternative methods and governance mechanisms.
- AI policy
- Certifications
- Quality assurance
- Workforce
Research
Evidence-Grounded AI for Musculoskeletal Care
Wenjie Li, Yujie Zhang, Fanrui Zhang et al.
arXiv · 2026-07-14
OrthoPilot is a clinical AI system powered by a large language model that integrates real-time hospital data—imaging, laboratory, pathology, and orders—to support continuous musculoskeletal care from admission through rehabilitation. Benchmarked against 81 orthopaedic physicians across 1,000 disease codes, it outperformed specialists with 25 years of experience in diagnostic reasoning and management planning, and generalized across 60 external clinical centres. In a prospective study of 1,870 complex cases it improved full-chain management success by 10.6%, and a randomized deployment involving 8,240 inpatients increased cumulative cases per bed by 9.7% while improving patient-reported access to health information. The work demonstrates that clinical AI can move beyond isolated predictions to execute longitudinal management across complete care pathways, with direct implications for care quality, hospital efficiency, and physician decision support.
- Enterprise
- Quality assurance
- Workforce
Research
On-Device Deep Research at 4B: Exposure Bounds Faithfulness, Retrieval Bounds Coverage
Vinay Kumar Chaganti
arXiv · 2026-07-14
This study evaluates on-device research agents running a 4B-parameter language model on a 24 GB laptop, finding that citation faithfulness and source coverage are governed by distinct factors. Exposing the model to more text per source (400 vs. 1500 characters) raises cited-claim faithfulness from roughly 0.45 to 0.58, regardless of whether the sources are gold or retrieved, while trustworthy coverage remains near 0.22 because it is limited by retrieval recall (~0.40) rather than exposure. The practical implication is that practitioners should first increase per-source exposure—a cheap intervention costing about 235 extra output tokens—and then focus on improving retrieval recall as the only remaining lever for coverage. These findings matter for deploying reliable, locally-run AI research tools that produce verifiable, cited outputs.
- Enterprise
- Quality assurance
Research
Regulating Artificial Intelligence in Developing and Centralized Legal Systems: A Comparative Study of Indonesia and China
Christian Andersen, Shavilla Felisya Regitara
Journal Of Social Research · 2026-07-14
This comparative legal study examines AI regulation in Indonesia and China, finding that Indonesia relies on a fragmented, policy-based approach lacking specific AI legislation, while China has built a more comprehensive and enforceable framework for algorithmic systems. The research identifies key regulatory gaps in Indonesia, including limited institutional coordination and inadequate algorithmic accountability mechanisms. The authors argue Indonesia should develop a risk-based, legally binding AI governance framework informed by China's regulatory experience while accounting for its own legal context.
- AI policy
- Certifications
Research
AI Trustworthiness in Managerial Decision-Making: Ethics, Transparency, and Explainability as Key Drivers
Guangming Cao, Yanqing Duan, John S. Edwards
Journal of Business Ethics · 2026-07-14
This study uses structural equation modeling on survey data from UK managers to examine how perceived AI trustworthiness—defined through ethics, transparency, and explainability—relates to the adoption of predictive AI and generative AI and their effects on decision quality, speed, and integrity. Findings show that higher perceived trustworthiness is positively associated with adopting both AI types, which in turn improve perceived decision efficacy, with predictive AI showing stronger and more consistent effects. Notably, decision complexity weakens the positive link between generative AI and perceived decision efficacy but does not affect predictive AI's relationship. The research contributes to debates on responsible AI, managerial accountability, and ethical technology use in organizations, highlighting trustworthiness dimensions as key enablers of effective AI-assisted management.
- Enterprise
- AI policy
- Workforce
Research
Artificial Intelligence and Healthcare Policy: A Bibliometric Analysis of Global Research Trends
Pegah Rashidian, Forough Heidarzad-Pahlaviani, Seyedsina Moghimnejadhosseini et al.
Healthcare · 2026-07-14
This bibliometric study maps global research trends at the intersection of artificial intelligence and healthcare policy using 347 peer-reviewed articles from 2000 to 2026. Publications surged after 2020, peaking in 2025, with contributions from researchers in 82 countries and 900 institutions, led by the United States, China, England, Canada, and India. Key themes include machine learning, COVID-19, health policy, large language models, and public health applications, with growing focus on predictive modeling and public health decision-making. The findings provide a roadmap for researchers, policymakers, and healthcare leaders navigating the rapidly evolving AI-in-health-policy landscape.
- AI policy
- Workforce
- Enterprise
Research
Partial Identification with Multiple Nonlinear Measurements of a Latent Regressor
Burhan Ogut, Michelle Yin
arXiv · 2026-07-13
This paper addresses a core measurement problem in AI labor-market research: multiple competing occupational AI-exposure scores produce downstream employment estimates that differ by a factor of eleven, making it impossible to know which (if any) reflects the true structural effect. The authors develop a partial-identification framework for linear regression with a latent regressor observed through nonlinear measurements, deriving a closed-form interval centered on a cross-source estimator whose half-width is second-order in curvature heterogeneity and invariant to unknown source loadings. Applied to six AI-exposure measures matched to an American Community Survey panel of 8.88 million person-year observations (2015–2024), the method yields a loading-invariant consensus employment coefficient of -0.239, with a partial-identification half-width of just 1.23 percent of the point estimate, though the post-2022 employment coefficient changes sign between language-model and patent-text measures. The work matters for workforce and policy research because it provides a principled, estimable way to reconcile conflicting AI-exposure metrics rather than arbitrarily selecting one, and it reveals that at least one prominent measure (Webb patent-text) captures a distinct construct from the others.
- Workforce
- AI policy