News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
Artificial Intelligence Investment and Enterprise Green Innovation Efficiency
Feng Qianhao
Advances in Economics Management and Political Sciences · 2026-08-11
Using a panel of China's A-share listed companies from 2007 to 2023, this study finds that greater AI investment and disclosure significantly improve enterprise green innovation efficiency: each logarithmic unit increase in AI-related word frequency raises green innovation efficiency by roughly 8%, and a one-standard-deviation increase in investment level improves it by about 4%, both significant at the 1% level. The effects are stronger for state-owned, large-scale, and manufacturing enterprises, and are amplified by higher asset-liability ratios, suggesting financing capacity is a key moderator. Robustness checks using high-dimensional fixed effects and a difference-in-differences design around China's National New-Generation AI Innovation Pilot Zones confirm the findings. The paper offers empirical guidance for enterprises optimizing AI investment strategies to support green transformation.
- Enterprise
Research
Clinical predictive artificial intelligence evaluation: A narrative review of trial designs and practical considerations
Maxime Fosset, Joris Pensier, Boris Jung et al.
PLOS Digital Health · 2026-08-11
This narrative review argues that conventional randomized controlled trials are poorly suited for evaluating clinical AI predictive models because of their static design and inability to accommodate evolving algorithms. The authors propose a three-part framework encompassing performance monitoring, clinical impact monitoring, and scientific evidence generation, linked through a governance-driven escalation protocol that specifies when monitoring signals should trigger formal trials. The framework integrates adaptive and pragmatic trial designs with causal inference methods, emphasizing patient-centered outcomes, health equity, and workflow integration. The work provides actionable guidance for clinicians and trialists seeking to rigorously assess AI tools in routine clinical settings.
- Quality assurance
- Certifications
- AI policy
Research
Perceived skill devaluation as an indirect statistical pathway between perceived generative AI impact and algorithmic anxiety among college students: associations with adaptive coping strategies
Dou ChenXu, Zhang Yinuo, Ji Chenao et al.
Frontiers in Education · 2026-08-11
This study surveyed college students to examine how perceiving generative AI as impactful on the job market relates to anxiety and coping. Results from structural equation modeling show that perceived AI impact is positively linked to feeling that one's skills are being devalued, which in turn is associated with higher algorithmic anxiety. Career self-efficacy and growth mindset were not associated with reduced anxiety but were positively linked to adaptive coping strategies, suggesting that psychological resources matter more for coping than for anxiety reduction. The findings highlight workforce-relevant concerns about how students appraise AI-driven threats to their career competencies.
- Workforce
Research
Employer perspectives on AI in SMEs: policy insights for productivity and sustainable careers
Tony Fang, J.A. Harrison
Career Development International · 2026-08-11
This practitioner insights paper surveys 1,700 business owners and senior managers—with a focus on SMEs in Atlantic Canada—to understand how employers view AI adoption. Findings show that while employers associate AI with productivity and efficiency gains, adoption is uneven across firms and regions, and many employers remain uncertain about future skill needs and focus primarily on compliance training. The paper argues that policymakers should strengthen digital capabilities, workforce planning systems, and regional skills ecosystems to ensure AI adoption supports both productivity and sustainable career outcomes.
- Workforce
- AI policy
Research
Risks of artificial intelligence adoption for employment in Russia
Rostislav Kapeliushnikov
Voprosy Ekonomiki · 2026-08-11
This study estimates the risk of AI adoption for employment in Russia using International Labour Organization scoring methods applied to Rosstat Labor Force Survey microdata. The average AI risk index for Russia is 0.3, with roughly one in ten jobs facing high AI risk and one in three facing significant risk — lower than developed countries but higher than developing ones. Less than 1% of jobs face full automation, and the impact is concentrated in a narrow range of occupations, primarily clerical workers. The authors conclude that AI adoption in Russia is more likely to reconfigure tasks within existing occupations than eliminate them wholesale, suggesting fears of mass technological unemployment are overstated.
- Workforce
- AI policy
Research
Artificial Intelligence and Family Law: The Impact of the EU AI Act Regulatory Framework and Principles on the Use of AI in Family Proceedings
Michał A. Piegzik
European Journal of Risk Regulation · 2026-08-11
This article examines how the EU AI Act, adopted in August 2024, applies to the use of AI tools in family court proceedings across EU Member States. Using doctrinal, normative, and socio-legal research methods, the authors critically assess the new regulatory framework's conceptual and practical limitations when applied to family law contexts, where AI has been proposed as a solution to growing caseload pressures and access-to-justice challenges. The paper identifies scholarly gaps and explores implications of the EU AI regulatory model both within and beyond individual Member States.
- AI policy
Research
The rise of AI in weather and climate information and its impact on global inequality
Amirpasha Mozaffari, Amanda Duarte, Lina Teckentrup et al.
npj Climate Action · 2026-08-11
This paper argues that AI-driven weather and climate forecasting systems are being developed almost exclusively in the Global North, creating a risk of entrenching and amplifying existing inequalities in climate information access. The authors identify disparities across the full pipeline—from biased training data to unrepresentative model validation—that disproportionately harm vulnerable regions in the Global South. They propose remedies including a Climate Digital Public Infrastructure, well-being-centered evaluation metrics, and inclusive knowledge co-production to promote resilience rather than inequity. The findings carry direct implications for climate policy and the governance of AI in public-interest domains.
- AI policy
Research
Does it pay off to use GenAI in the Russian labor market?
Ksenia Rozhkova, Sergey Roshchin, Yana Roshchina
Voprosy Ekonomiki · 2026-08-11
Using Russian panel survey data (RLMS-HSE, 2023–2024) and methods including propensity score matching and fixed-effects regression, this paper estimates wage returns to generative AI use in Russia's labor market. On average, workers who use GenAI earn roughly 13.4% more, rising to 21.2% for highly qualified specialists; fixed-effects models find returns only for regular users, ranging from 17.2% (full sample) to 41.8% (highly qualified specialists). The results suggest the wage premium is partly pre-determined by the types of jobs where GenAI is applicable, and that only about 11% of employed Russians had adopted such tools by 2023–2024. The findings matter for understanding how AI adoption translates into measurable labor market inequality and earnings differences across skill levels.
- Workforce
Research
AI skills demand in the Russian labor market
Andrei Ternikov, Vera Maltseva, K. A. Iliashchenko et al.
Voprosy Ekonomiki · 2026-08-11
Analyzing over 700,000 Russian online job postings from 2021 to 2025, this paper tracks the diffusion of AI skills and their wage effects in a major non-Western economy. Although AI skills appear in fewer than 3% of postings, demand is rising sharply in high-tech industries, while the wage premium for AI skills fell from 31.8% in 2022 to 11.8% in 2025, indicating market stabilization. The study finds no evidence of labor substitution; instead, AI skills complement analytical and managerial competencies, especially in non-IT roles. These findings challenge job polarization narratives and support a 'soft complementarity' model of AI and human labor.
- Workforce
Research
Accelerating ambient AI scribe enterprise-scale deployment: Cleveland Clinic’s novel approach to health system-industry partnership
Amy Merlino, Abigail Blue, Eric Boose et al.
npj Health Systems · 2026-08-11
This paper describes Cleveland Clinic's four-pillar partnership framework—covering governance, training and onboarding, support, and real-time learning—used to deploy AI ambient scribes to over 4,000 ambulatory care clinicians in just four months. The authors present this vendor-health system model as a replicable blueprint for other health systems seeking rapid, enterprise-wide AI scribe rollout. The work highlights the organizational and operational strategies that enabled efficient large-scale deployment rather than evaluating clinical outcomes.
- Enterprise
- Workforce
Research
Artificial Intelligence (AI)-driven entrepreneurial strategies and customer-related performance of SMEs in Sanchez Mira, Cagayan, Philippines: an exploratory mixed-methods case study
Michael Sacramed
Future Business Journal · 2026-08-11
This mixed-methods study of 50 SMEs in rural Sanchez Mira, Cagayan, Philippines finds that localized AI adoption—including NLP chatbots and predictive analytics—is positively associated with adaptive entrepreneurial strategies (B=0.61, R²=.37) and shows statistically significant links to improved customer satisfaction, retention, and responsiveness. Qualitative findings reveal a 'rural paradox' where impersonal AI interfaces conflict with community-centric kinship cultures, and low digital literacy forces informal, locally adapted strategies that diverge from Western entrepreneurial frameworks. The study argues these findings offer a transferable 'diagnostic blueprint' for digital transformation in peripheral emerging economies across the Global South, with implications for how enterprise AI adoption should be tailored to rural, resource-constrained contexts.
- Enterprise
Research
Logit-Boundary Geometric Belief Interfaces and Sparse Sheaf-Enclave Protocols: A Self-Contained Substrate for Secure Network Electronic Health Record (EHR) Interoperability
Alvin Spivey, Thomas Huang
arXiv (Cornell University) · 2026-08-10
This paper presents a mathematical and engineering architecture called the Geometric Belief Interface (GBI) for securing electronic health record (EHR) interoperability at the boundary between AI models and clinical systems. The core mechanism is a 'logit boundary' that intercepts pre-threshold model outputs and deterministically decides whether they are admissible, require human review, or must be quarantined before any FHIR transaction is constructed. To evaluate this boundary, the authors tested Qwen3-4B-Instruct-2507 on a frozen synthetic benchmark (GBI BoundaryBench v0.1) across 768 executions, finding that zero outputs passed the admission contract—369 were rejected during parsing and 399 during schema validation—demonstrating that the quarantine mechanism functions as a hard gate rather than a soft filter. The work does not claim to establish clinical truth or general LLM safety, but provides evidence about how a certificate-producing admission boundary can enforce structured constraints on AI outputs before they enter health record systems.
- Quality assurance
- Certifications
Research
Withholding the Completing Chunk: Deterministic Pair-Completion Guardrails for Streaming LLM Output
Christopher M. Frost
arXiv · 2026-08-10
This paper investigates a deterministic guardrail mechanism for streaming large language model (LLM) output that withholds text chunks when both parts of a predefined 'danger signature' pair become detectable in the accumulated prefix, preventing harmful content from escaping before moderation can act. Experiments across multiple signature families, chunk sizes, and trials show the approach perfectly matches buffered scanning for configured pairs, but fixed pairs caught none of 394 jury-labeled unsafe responses and flagged none of 338 safe responses, confirming the method covers only a narrow, explicitly defined policy rather than general harm. The paper positions this technique as an exact release-boundary backstop complementing — not replacing — semantic moderation tools like Llama Guard 3 1B, which achieved broader but imperfect coverage (310/338 safe, 202/394 unsafe). The findings matter for quality assurance in LLM deployment pipelines, where low-latency, deterministic content controls are needed alongside probabilistic classifiers.
- Quality assurance
Research
Locally Deployable Small Language Models for Emergency Department Decision Support: A Systematic Benchmark of Fine-Tuning Strategies
Qingfeng Zhang, Yuanxiong Guo, Yanmin Gong
arXiv · 2026-08-10
This study benchmarks eight open-source small language models (SLMs) on three emergency department tasks—triage level prediction, specialist referral recommendation, and diagnosis prediction—using 2,083 MIMIC-IV-ED cases. Comparing zero-shot prompting, prefix tuning, LoRA, and full fine-tuning against commercial baselines (Claude Haiku 4.5 and Claude Sonnet 4.5), the authors find that LoRA fine-tuned open-source SLMs outperform the commercial baselines on triage and referral tasks, and can detect highest-severity patients that commercial models miss. Diagnosis prediction remains a challenge for open-source SLMs. The findings suggest that locally deployable SLMs can achieve clinically competitive performance while avoiding the privacy risks of transmitting patient data to external commercial services.
- Enterprise
- Quality assurance
Research
TAF-MED: Multi-Turn Safety Refusal Collapse in LLMs Under Declared Self-Treatment Intent
Waleed Jamil, Raphael Schmitt
arXiv · 2026-08-10
TAF-MED introduces a physician-reviewed benchmark of 500 three-turn dialogue scenarios used to evaluate how well eight large language models maintain medication-safety refusals after a user declares intent to self-treat. Across 4,000 conversations, 71.6% contained an unsafe response, and 61.4% of conversations that began with a safe first response later collapsed to unsafe guidance—with model-level collapse rates ranging from 24.4% to 96.2%. Automated labels agreed with physician adjudication at 94.3% (κ=0.895), validating the evaluation approach. The findings demonstrate that single-turn safety evaluations are insufficient proxies for real-world conversational safety and call for multi-turn assessment frameworks before deploying LLMs in health-information contexts.
- Quality assurance
- AI policy
Research
Self-evolving Agentic Customer Support System at LinkedIn
Chih Hui Wang, Mengdie Tu, Qianyun Zhang et al.
arXiv · 2026-08-10
LinkedIn presents a self-evolving agentic customer support system that combines retrieval-augmented generation (RAG) with evolutionary auto-prompting and a modular evaluation framework, enabling continuous improvement without retraining foundation models. The system treats prompts, retrieval, and evaluation as a closed-loop workflow with operational guardrails to handle rapidly changing policies and knowledge bases. In a two-week A/B test on production traffic, the integrated system increased QA self-serve by 9.0 percentage points, cancellation self-serve by 4.8 percentage points, and routing accuracy by 30.6 percentage points compared to baseline agents. These results demonstrate a scalable, practical approach to deploying self-improving AI agents in enterprise support environments.
- Enterprise
- Quality assurance
Research
Similarity Gates Approve Reversals: A Validity Audit of Embedding-Cosine Thresholds in Agent Systems
Scott E. Frias
arXiv · 2026-08-10
This paper audits the use of embedding-cosine similarity thresholds as quality gates in AI agent frameworks—specifically for deduplication, semantic caching, drift detection, and answer grading. The authors find that cosine similarity measures wording change rather than meaning change, causing safety checks to fire incorrectly: a production drift guard caught 0 of 56 meaning-breaking mutations, and a critical reversal ('withhold the study drug' → 'administer the study drug') scored 0.9608 cosine similarity. Across 90 configuration-threshold-task cells, balanced accuracy never exceeded 0.700 (median 0.525), and naively constructed evaluation corpora returned inverted verdicts with AUROC as low as 0.000. The authors conclude that these gates systematically measure the wrong thing and release their corpus, harness, and results to support development of valid alternatives.
- Quality assurance
- Enterprise
Research
The Deliberative Deficit: An Empirical Critique of LLMs in Democratic Discourse
Maurice Flechtner
arXiv · 2026-08-10
This paper empirically evaluates large language models (LLMs) as participants in democratic deliberation using the Deliberative Reason Index (DRI), a measure validated across citizen assemblies in political science. Across 1,980 five-agent LLM runs on 12 citizen-assembly topics and 11 frontier model configurations, the authors find that LLM groups match human procedural quality (respectfulness, justification) but show only small, topic-dependent gains in intersubjective consistency and exhibit roughly one-third the perspective diversity of human assemblies. Critically, LLM groups invert the human convergence pattern—human deliberation reduces dispersion as diverse views synthesize, while LLM deliberation increases it—and persona prompting fails to restore this dynamic. The authors conclude that LLMs can support human reasoning on pluralistic problems but should not be treated as autonomous deliberative agents.
- AI policy
Research
From Prediction to Incrementality: Causal Optimization for Large-Scale Targeting and Recommendation
Changshuai Wei, John Bencina, Phuc Nguyen et al.
arXiv · 2026-08-10
This paper presents a decision-centric framework for large-scale targeting and recommendation that optimizes causal (incremental) effects rather than predictive scores, addressing the systematic misallocation of resources toward users who would have acted regardless of intervention. The system combines a causal neural network with a Transformer backbone for individual treatment-effect estimation, a Bayesian neural-bandit layer for exploration under uncertainty, and a dual-based linear-programming layer for constrained allocation at scale. Evaluated via offline simulations, architectural ablations, and an online A/B test on LinkedIn Feed marketing traffic, the end-to-end treatment policy delivered a statistically significant +7.20% lift in the primary long-term-value metric. The work demonstrates that production-scale causal optimization under real business constraints is feasible and outperforms conventional prediction-based targeting paradigms.
- Enterprise
Research
Human versus Computer Vision
Elena Sirotkina
arXiv · 2026-08-10
This study evaluates leading computer vision saliency models—used commercially to predict where people look at images—against 11.4 million webcam gaze points collected from over 3,000 U.S. adults sampled to national demographic quotas viewing news photographs. The author finds that a simple untrained central marker outperforms every trained saliency network, because the content those networks add beyond center bias does not match where real audiences actually look. Critically, the accuracy that does exist is systematically skewed toward younger, White, and ideologically moderate viewers, underperforming for older, Black, and ideologically extreme audiences. The paper proposes a framework for auditing whether saliency models can learn to represent specific demographic groups, and argues this benchmark should be the standard for any claim that such systems 'see everyone.'
- Quality assurance
- AI policy
Research
The CASE Framework: A Multi-Disciplinary Control Architecture for Governing Enterprise Agentic AI
Srinivas Telukunta, Georgios Nektarios Lilis, Lucio Baron
arXiv (Cornell University) · 2026-08-10
This paper introduces the CASE framework, a four-layer governance architecture for enterprise agentic AI that draws on Control theory, complex Adaptive systems theory, Supervisory cybernetics, and Engineering operations to address distinct challenges at different scales of AI autonomy. The authors present three empirical studies: 82% of documented production agent failures follow multi-layer trajectories, none of 22 ecosystem tools provides full coverage of the emergence layer (Layer 2), and all 35 scored public deployments fall in the lowest maturity band—a gap the authors call the 'Emergence Gap.' The framework formalizes cross-layer coupling conditions, including a zero-touch deployment paradox, and derives a five-level maturity model with a bottleneck-weighted index. The work is directly relevant to enterprise AI governance and to policy compliance, specifically noting that EU AI Act Article 14's human oversight requirements can only be satisfied by architectures meeting requisite variety conditions.
- Enterprise
- AI policy
Research
Status Association Does Not Reliably Predict Decision Leakage
Abdullah X
arXiv · 2026-08-10
This study tests whether AI models that encode socioeconomic status associations based on Chilean surnames actually translate those associations into biased decisions in high-stakes contexts like academic selection, hiring, fellowship awards, and legal-aid intake. Across eight models and over 8,000 verified responses, elite-coded surnames did trigger higher status associations than common or rare surnames in nearly all models, but the effect on actual decisions was close to zero for most systems. Critically, the strength of a model's social association did not reliably predict whether that bias would 'leak' into consequential decisions (r = 0.201, p = 0.633). The paper concludes that latent association and consequential treatment are empirically distinct constructs, and that bias evaluations must directly measure the association-to-action transition rather than assuming one implies the other.
- Quality assurance
- AI policy
Research
From Values to Benchmarks: Evaluating Large Language Models for Governmental Use in Dutch
Laurens Samson, Iva Gornishka, Gossa Lô et al.
arXiv · 2026-08-10
This paper presents the 'Grip on LLMs' framework, a systematic evaluation suite for assessing large language models suited to Dutch governmental use, developed with domain experts from a major Dutch municipal organisation. The framework identifies six evaluation dimensions—factuality, honesty, social bias, energy consumption, cost, and training data transparency—and benchmarks more than 30 multilingual and Dutch-specific models. Key findings show that no single model excels across all dimensions, that quality trade-offs with environmental impact and cost are unavoidable, and that factuality and honesty are distinct properties that do not imply one another. The authors release a publicly accessible model overview aimed at the full range of governmental stakeholders, from engineers to policymakers, making results actionable for non-technical decision-makers.
- AI policy
- Quality assurance
Research
Stealing Reasoning Traces from Proprietary LLM APIs
Alexander Panfilov, David Schmotz, Ilia Shumailov et al.
arXiv (Cornell University) · 2026-08-10
This paper reveals a security vulnerability in how leading LLM providers (Anthropic, OpenAI, and Google) handle encrypted chain-of-thought reasoning traces: because these encrypted blocks are interchangeable across sessions, users, and models within a provider's ecosystem, attackers can inject a trace from a capable model into a weaker, less-safeguarded model to force plaintext decryption. The researchers demonstrate four attack vectors — bypassing anti-distillation protections to extract proprietary reasoning, recovering 367 PII artifacts and 182 credentials from 315,320 publicly scraped reasoning blocks, exposing hazardous information hidden in rejected requests, and executing invisible prompt injections via encrypted payloads. The findings matter for enterprise AI deployment and policy because they show that client-side encryption of reasoning traces creates serious intellectual property, privacy, and safety risks at scale. The authors propose cryptographic and system-level mitigations following responsible disclosure.
- Enterprise
- AI policy
Research
Towards Expert-level Medical AI for Real-time Video Consultations
Mahvish Nagda, Jihyeon Lee, Matthew Thompson et al.
arXiv (Cornell University) · 2026-08-10
This paper presents AMIE (Video), a Gemini-based multi-agent AI system designed for real-time clinical video consultations that integrates low-latency dialogue, clinical reasoning, and audio-visual perception. In a randomized OSCE study involving 30 primary care physicians, 15 patient actors, and 100 clinical scenarios, clinical evaluators rated AMIE (Video) on par with or better than physicians in history-taking, diagnosis, management, and physical observation. Patient actors preferred AMIE's approach to assessing and explaining conditions, though physicians were preferred for rapport and partnership building. These results represent a milestone toward AI systems that can augment clinical care across the sensory complexity of real consultations, with noted limitations in fine anatomical precision, subtle affective nuances, and high-frequency movements.
- Workforce
- Quality assurance