News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
The impact of artificial intelligence and automation on labour market outcomes: a meta-analysis
Gianina-Maria Petrașcu, Ioana Bîrlan
Management & Marketing · 2026-07-30
This meta-analysis synthesizes 321 estimates from 19 empirical studies to assess how AI and automation exposure affects employment, wages, and skill demand across countries and sectors. Using a three-level random-effects model, the authors find that the overall pooled effect of technological exposure on labor market outcomes is small and statistically insignificant, with substantial heterogeneity across studies. Some studies report negative employment effects from automation and robot adoption, while others document wage increases, productivity gains, or skill upgrading. The findings suggest that AI and automation do not produce a uniform pattern of job displacement or skill-biased change, but instead generate context-dependent effects that vary by sector, occupation, and institutional setting.
- Workforce
Research
Review of Computer-Aided Detection (CAD) Software for Tuberculosis on Chest X-Rays : A Systematic Review of Randomized Controlled Trial and Primary Studies
Catur Nila Pratiwi, Wildan Priscillah, Eka Yusi Athiyyah
The Indonesian Journal of General Medicine · 2026-07-30
This systematic review of 17 studies covering over 130,000 participants across Africa, Asia, Europe, Oceania, and Latin America evaluates AI-based computer-aided detection (CAD) software for tuberculosis screening on chest X-rays. AUROC values ranged from 0.70 in paediatric populations to 0.92 in unselected adults, and multiple CAD products met WHO Target Product Profile thresholds of ≥90% sensitivity and ≥70% specificity in symptomatic adult populations, with CAD outperforming human radiologists in several large-scale studies. Performance was consistently lower in people living with HIV, elderly individuals, prior TB patients, and children, highlighting the need for local threshold calibration and version-specific validation. The authors conclude that while CAD holds strong promise for scaling TB case finding, equitable access, regulatory frameworks, and standardized evaluation are prerequisites for maximizing public health impact.
- Quality assurance
- AI policy
Research
Exploring paradoxical barriers to AI adoption through the TOE framework
Faisal Shahzad, Muhammad Aqeel, Ahmad Arslan et al.
Small Enterprise Research · 2026-07-30
This qualitative study investigates why small- and medium-sized enterprises (SMEs) in Finnish manufacturing are slow to adopt AI, using the Technology–Organization–Environment (TOE) framework and 12 semi-structured interviews. Key barriers identified include fragmented data, legacy IT infrastructure, skill shortages, employee resistance, and regulatory uncertainty. Notably, the study finds that some apparent barriers—such as strategic caution and employee resistance—act as adaptive mechanisms that help firms avoid premature or misaligned AI implementation. The findings offer practical guidance for SME managers and policymakers on data readiness, workforce development, and institutional support.
- Enterprise
- Workforce
- AI policy
Research
<b>Hybrid AI system for interpretable student policy guidance using rule-based reasoning and retrieval-augmented generation</b>
Ediomo Titus, Mustapha Aminu, Felix Uloko
Nature Journal of Emerging Sciences Technologies and Innovations · 2026-07-30
This paper presents a hybrid AI system combining rule-based reasoning and retrieval-augmented generation (RAG) to help students navigate complex academic policies at Nigerian universities. Evaluated across two institutions, the system achieved 89.7% accuracy, a 2.34-second mean response time, and an estimated 41.9% reduction in routine staff queries, while its interpretation consistency (86.0%) surpassed human staff-to-staff consistency (76.0%). The findings demonstrate that AI-assisted policy guidance can reduce administrative burden and improve equitable access to institutional rules, particularly in developing-country contexts with resource-constrained digital infrastructure.
- AI policy
- Workforce
Research
Governing Artificial Intelligence in Thailand's Higher Vocational Education
Chaimongkhol Pugsuwan
Zenodo (CERN European Organization for Nuclear Research) · 2026-07-30
This article develops a governance and management framework for deploying AI in Thailand's Higher Vocational Certificate (HVC) education system, drawing on an integrative review of Thai legislation, vocational qualification standards, national AI strategy, and international evidence. The framework identifies AI's strongest near-term role as decision augmentation—improving timeliness and coherence of evidence while preserving professional judgement—across areas such as labour-market intelligence, learner-risk detection, and dual-training logistics. It proposes six governance components and a three-tier use-case classification, while cautioning that AI applications influencing admission, assessment, certification, or workplace placement carry serious educational and legal risks. Because national causal evidence remains limited, the authors recommend controlled pilots and independent review rather than technology-led scaling.
- Certifications
- Quality assurance
- AI policy
Research
Governing Artificial Intelligence in Thailand's Higher Vocational Education
Chaimongkhol Pugsuwan
Zenodo (CERN European Organization for Nuclear Research) · 2026-07-30
This paper develops a governance and management framework for deploying AI in Thailand's Higher Vocational Certificate (HVC) education system. Through an integrative policy review synthesizing Thai legislation, vocational qualification standards, national AI strategy, and international evidence, the authors find AI's strongest near-term role is decision augmentation—improving evidence timeliness while preserving professional judgement—across areas like labour-market intelligence, learner-risk detection, and dual-training logistics. The proposed HVC-AI Governance and Management Framework comprises six components (public value, human authority, data stewardship, proportional risk control, lifecycle accountability, and continuous evaluation) with a three-tier use-case classification and phased implementation roadmap. Because national causal evidence remains limited, the authors recommend controlled pilots and independent review rather than technology-led scaling, particularly given heightened risks when AI influences admission, assessment, or certification.
- AI policy
- Quality assurance
- Certifications
Research
AI literacy and AI anxiety in nursing students: the serial mediating roles of attitudes and self-efficacy
Qin Zeng, Shenghua Zhang, Jiacheng Hu et al.
Frontiers in Public Health · 2026-07-30
A cross-sectional survey of 1,482 nursing students across 11 Chinese universities found that higher AI literacy was associated with lower AI anxiety, with attitudes toward AI and AI self-efficacy serving as sequential mediators in that relationship. Structural equation modeling showed that more favorable attitudes were linked to stronger self-efficacy, which in turn was associated with reduced anxiety, though the cross-sectional design precludes causal conclusions. The findings suggest nursing education programs may benefit from pairing AI knowledge and skills training with efforts to cultivate positive attitudes and confidence in AI use.
- Workforce
Research
Automated Transcript Analysis for Detecting Flaws in Agentic Benchmarks
Jeff Mohl, Nelson Gardner-Challis, Magda Dubois et al.
arXiv (Cornell University) · 2026-07-29
This paper develops AI-powered scanners to automatically detect validity flaws in agentic benchmarks used to evaluate frontier AI models, targeting four specific issue types: ground truth access, tool failure, guessing vulnerability, and answer format ambiguity. The scanners were evaluated against human labels on held-out benchmarks from Inspect Evals and successfully identified verified quality issues in five widely used benchmarks, including cases unlikely to surface through random manual inspection. Performance varied across benchmarks, criteria, and models, and the authors acknowledge open challenges such as standardization gaps in the evaluation field that limit scanner reliability. The work serves as a proof of concept for scalable automated auditing of benchmark quality, which matters for ensuring that capability assessments of AI systems are trustworthy and valid.
- Quality assurance
- Certifications
Research
Same Facts, Different Diagnosis: Measuring and Mitigating Narrative Anchoring in Clinical Language Models
Prabhjot Singh, Pritam Deka, Vijay Chennareddy
arXiv · 2026-07-29
This paper identifies and measures 'Narrative Anchoring,' a failure mode in clinical large language models where identical medical facts presented in different sociolinguistic registers (writing styles/personas) lead to divergent diagnostic outputs—even when no explicit demographic markers are present. Across seven models and three architecture families, the authors find the effect is statistically significant in every model tested, with a Narrative Anchoring Gap ranging from 0.064 to 0.151. Standard mitigation strategies like chain-of-thought reasoning and debiasing instructions only partially reduce the bias and often hurt accuracy. The proposed NarrativeShield pipeline, which extracts and verifies clinical facts before diagnostic reasoning, reduces the gap to near-zero and achieves the lowest rate of severely unstable decisions across all models tested.
- Quality assurance
- Certifications
Research
RoguePrompt: Dual-Layer Encoding for Self-Reconstruction to Circumvent LLM Moderation
Benyamin Tafreshian, Prathamesh Dhake
arXiv · 2026-07-29
RoguePrompt is a jailbreak technique that wraps forbidden prompts in two nested ciphers (Vigenère followed by ROT13) plus natural-language reconstruction instructions, then tests whether large language models can be tricked into decoding and executing content their safety filters are meant to block. Evaluated against 313 hard-rejected prompts under a black-box threat model, the pipeline achieved 93.93% filter bypass, 79.02% successful reconstruction, and 70.18% execution. Crucially, the authors measure each stage separately rather than collapsing everything into a single success metric, revealing exactly where multistage jailbreaks tend to break down. The findings highlight a concrete gap in current LLM moderation controls and provide stage-level evidence useful for designing more robust safety mechanisms.
- Quality assurance
- AI policy
Research
LayerRAG-Bench: A Cross-Layer Reliability Benchmark for Agentic Retrieval-Augmented Generation
Musa Shams
arXiv · 2026-07-29
LayerRAG-Bench is a new benchmark designed to test the reliability of agentic retrieval-augmented generation (RAG) systems across multiple failure layers—including evidence quality, tool contracts, authorization, and session state—rather than treating groundedness as the sole measure of quality. Testing 9 models from OpenAI, Anthropic, and Gemini across 240 tasks and 8 enterprise domains, the benchmark reveals that schema normalization dramatically improves schema-drift failures (from 0.000 to 0.913 success rate) but does not fix stale evidence, missing tool output, denied permissions, or wrong-session context issues. Crucially, evaluating only groundedness produces significant false positives under stale or wrong-session evidence, meaning systems can appear reliable while actually failing. The findings argue for a layer-specific evaluation principle where each reliability intervention is assessed against its specific failure mode rather than treated as a universal fix.
- Quality assurance
- Enterprise
Research
Can AI agents conduct open-ended AI research? Early evidence from two case studies
Peter Kirgis, Sayash Kapoor, Andrew Schwartz et al.
arXiv · 2026-07-29
This paper introduces 'shadow evaluations,' a new method for assessing whether AI agents can conduct open-ended AI research: agents tackle the central research question of unpublished high-quality papers, and the papers' original authors grade the output. In two case studies using frontier agents given six days and thousands of dollars of compute on unpublished NeurIPS 2026 submissions, agents completed all engineering tasks without human help but failed to make substantial progress on the core research questions, resulting in unambiguous rejection by the authors. The authors identify five recurring failure modes—poor judgment about publishable quality, uncreative responses to design shortcomings, ineffective backtracking, poor resource awareness, and instruction drift—that were reproduced in a robustness check with a second model and scaffold. The findings provide early evidence that today's AI agents can handle the engineering of AI research but struggle with the open-ended, judgment-intensive aspects of the research lifecycle.
- Workforce
- AI policy
Research
APEX-Accounting
Julien Benchek, Austin Bennett, Jasmin Kern et al.
arXiv · 2026-07-29
APEX-Accounting is a benchmark developed by Mercor and Ramp to evaluate whether frontier AI models can perform real accounting work, including reconciling accounts, accruing expenses, posting transactions, and producing reports. The benchmark comprises 160 expert-authored tasks across 10 simulated accounting environments, with grading rubrics written by accounting and bookkeeping professionals. Across nine frontier models tested, the best performer (Claude-Fable-5) achieved only 56.4% Mean Criteria@3, and no model exceeded 2.6% Pass^8, indicating that current AI models fall well short of reliably completing professional accounting tasks. The benchmark also reveals a Simpson's paradox in token budget scaling, where increasing the budget improves aggregate scores but within a fixed budget, tasks where models spend more tokens score lower.
- Workforce
- Enterprise
Research
The Social Cost of an AI Teammate: How an Artificial Teammate Reshapes Human-Human Communication in Small-Team Decision-Making
Nia Nixon, Jaeyoon Choi, Pedro Martins De Bastos et al.
arXiv · 2026-07-29
This randomized controlled study compared small teams with an AI teammate versus all-human teams on a high-stakes moral-dilemma task, analyzing communication using Group Communication Analysis, surveys, and lexical methods. The AI teammate was the most talkative and self-cohesive member in every AI-human team, yet contributed the least new information and lowest lexical density. Its presence reduced human-to-human responsivity and social impact, and team members in AI-human teams reported lower belonging and status. These findings reveal an immediate 'social cost' of integrating conversational AI as a teammate, with implications for how AI is deployed in collaborative workplace settings.
- Workforce
- Enterprise
Research
OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding
Jingbo Zhou, Yusai Zhao, Qi Bao et al.
arXiv · 2026-07-29
OmegaUse-OfficeVal is a new benchmark of 100 real-world office-suite tasks used to evaluate how well large language model (LLM) agents can handle complex, multi-step workflows that practitioners actually perform. Each task is paired with economic signals—human labor time (averaging 2.32 hours per task) and a task price proxy—so that LLM inference costs and speed can be directly compared against human worker costs and quality. Evaluation using code-based verifiers and a human baseline shows that while frontier LLMs are substantially cheaper and faster than humans, they have not yet reached human-level deliverable quality. The benchmark provides a concrete, economically grounded framework for assessing when and whether AI agents can viably substitute for human labor in office work.
- Workforce
- Enterprise
Research
Cost-Sensitive Conformal Prediction and Human-in-the-Loop Abstention for Imbalanced High-Stakes Decision Support: A Multi-Domain Benchmark
Manpreet Singh, Akshatha Srikantha, Shyamal Lakhanpal
arXiv · 2026-07-29
This paper benchmarks uncertainty quantification methods for high-stakes classification tasks (credit scoring, fraud detection, healthcare, industrial safety) where minority classes are rare and errors have asymmetric costs. The authors show that standard conformal prediction leaves minority-class coverage as low as 0.5%, while Mondrian (class-conditional) conformal prediction recovers valid coverage with an average 61.7 percentage-point improvement over marginal conformal prediction (p < 1e-80). Combining Mondrian conformal prediction with cost-controlled abstention — deferring ambiguous cases to human reviewers — further reduces expected decision cost compared to standard and confidence-based decision boundaries. The study provides practical, dataset-specific guidance on when human-in-the-loop review becomes cost-effective, directly informing deployment of reliable AI decision-support systems.
- Quality assurance
- Enterprise
Research
Can Large Language Models Represent Urban Publics? Behavioral Replication and Population Mismatch in an Affordable-Housing Experiment
Yuxuan Cai, Yequan Hu, Hongqian Li et al.
arXiv · 2026-07-29
This study tests whether large language models (LLMs) can substitute for real residents in urban planning surveys by comparing eight open-weight LLMs against 843 human respondents in a US affordable-housing experiment. While one model (Qwen 2.5 14B) came closest to matching the aggregate difference in support between homeowners and renters as a proposed development moved nearer, this surface-level match concealed deep structural failures: the model distorted partisan and tenure-group contrasts, showed very low within-group variance, and produced unstable responses depending on question order or prompt framing. The findings warn that an LLM can approximate a single average effect while misrepresenting the spatially anchored, identity-conditioned population structure that underlies it. The authors conclude that evaluation of LLMs for urban planning applications must test whether social and spatial structure survives simulation, not just whether average effects are reproduced.
- AI policy
Research
Anticipatory Data Governance in the Age of AI: Emerging Signals in Data Access, Reuse, and Sovereignty
Adam Zable, Stefaan Verhulst
arXiv · 2026-07-29
This paper reports findings from a structured participatory foresight study in which nineteen senior practitioners across official statistics, digital policy, open science, AI governance, and related fields were convened in two expert studios between 2025 and 2026. Using a qualitative signal-scanning methodology, the researchers identified seven convergent signals shaping data governance, including strain on open-data paradigms, the rise of machine-centric data ecosystems, inference reshaping governance foundations, infrastructure sustainability challenges, institutional fragmentation, sovereignty-driven strategic control, and the need for stronger data-sharing incentives. The study argues that data governance is becoming inseparable from AI governance, digital public infrastructure, economic strategy, democratic resilience, and geopolitical competition. The contribution is explicitly diagnostic rather than predictive, offering an evidence-informed framework for reasoning about structural shifts already underway to enable anticipatory governance before risks and dependencies become locked in.
- AI policy
Research
OptimismBench: Forecasting Bias and the Alignment Effect in Language Model Judgment
Seonglae Cho, Adriano Koshiyama
arXiv · 2026-07-29
OptimismBench is a benchmarking framework that detects directional probability bias in large language models (LLMs) by presenting each scenario in both a success-framed and failure-framed version, then measuring the asymmetry between the two responses as a signed bias score. Testing 16 models from 8 providers, the study finds that 14 are systematically optimistic, that post-training (alignment) is the key driver of which direction a model tilts, and that model identity accounts for 4.7x more variance in bias than language does. The findings matter because LLMs are increasingly used as decision aids, and if their probability judgments carry a hidden optimistic tilt inherited from alignment fine-tuning, downstream pipelines and enterprise decision-making processes will silently absorb that distortion. The authors release 3,870 items across 10 languages to support per-model directional-bias auditing.
- Quality assurance
- Enterprise
Research
Human diversity fuels collective creativity that large language models cannot simulate or sustain
Mengchen Dong, Hiromu Yakura
arXiv · 2026-07-29
This preregistered experiment tested whether AI assistance homogenizes human creative output and whether AI-simulated personas can replace real human diversity. Using native (L1) and non-native (L2) English writers in a creative metaphor task, the study found that AI ideation compressed collective creative diversity for all writers and eliminated the advantage L2 writers normally contribute, while AI refinement preserved that diversity. Simulated writer pools built from real participant backgrounds across multiple model families consistently fell below every human pool in collective diversity, and attempts to force more diversity from models produced only degenerate text. The findings show that human linguistic and cultural diversity is a creative resource current AI cannot replicate, and that how human-AI workflows are designed determines whether that diversity survives.
- Workforce
- AI policy
Research
Hearsay: Vision-Language Medical Diagnoses Without an Image
Siddharth Vohra
arXiv · 2026-07-29
This paper investigates what happens when frontier vision-language models (Claude Opus-4.7, GPT-5.4, and Gemini-3.1-Pro) are prompted for a medical image diagnosis when no image is actually provided. Rather than abstaining, the models confabulate structured diagnoses that are systematically shaped by patient demographic descriptors — for example, consistently returning Melanoma for a 65-year-old white man and Sarcoidosis for a young Black patient on a chest X-ray. The study further reveals a 'hedged regime' where prose output acknowledges the missing image while the structured diagnosis field still names a disease, a failure invisible to prose-only audits. The authors conclude that clinical deployment of vision-language models requires direct auditing of structured output channels and that sensitivity to probe wording must be treated as a core evaluation dimension.
- Quality assurance
- AI policy
Research
Flat Score, Amplified Failures: How the Error Budget Masks Damage in Quantized LLM Agents
Jiwon Jang, Kisu Yang, Heuiseok Lim et al.
arXiv · 2026-07-29
This paper challenges the conventional wisdom that 4-bit post-training quantization is nearly lossless for large language models (LLMs), specifically testing this claim for multi-turn, tool-calling agents on the τ²-bench benchmark. While standard task-reward scores show no statistically significant change across 16-, 8-, and 4-bit weight precisions, a closer look at the failure process reveals that quantization amplifies pre-existing errors—particularly tool-name hallucination—by up to 2.5× in volume (+17.6 points per task), without introducing new failure types. The flat score is an artifact of the benchmark's ten-error budget absorbing the extra failures; shrinking that budget to two errors re-exposes a 17-point score gap precisely where quantization added error volume. The authors propose two diagnostics—per-channel error rates and success under a shrinking error budget—derived from logs benchmarks already collect, and recommend reporting these alongside standard task reward.
- Quality assurance
- Enterprise
Research
A First Look at Coding Agents' Compliance with AI Contribution Rules in Open-Source Communities
Wenhao Yang, Runzhi He, Minghui Zhou
arXiv · 2026-07-29
This paper investigates whether AI coding agents follow the AI contribution rules that open-source communities have established, such as total bans, mandatory disclosure requirements, verification gates, and human sign-off requirements. The researchers built RepoComplianceBench from 106 issues across 49 repositories and evaluated four frontier models, finding that agents almost never proactively retrieve contribution rules on their own. With reminder prompts, rule quotes, or verifier feedback, agents can be nudged toward disclosure and verification compliance, but no tested condition caused an agent to refuse contributing to an AI-banned repository. The findings highlight that disclosure and verification issues are addressable with existing mechanisms, while enforcing outright bans and human escalation requirements remains an unsolved problem—raising significant questions for open-source governance and AI policy.
- AI policy
- Quality assurance
Research
SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response
Lehan Wang, Boli Chen, Ruixue Ding et al.
arXiv · 2026-07-29
SecRespond introduces the first benchmark specifically designed to evaluate large language model (LLM) agents in post-compromise incident response scenarios, filling a gap left by existing benchmarks that focus only on pre-attack (pre-compromise) settings. The benchmark spans 10 cyber ranges built from real compromised cloud hosts, covering 21 ATT&CK techniques and 5 operating systems, and requires agents to produce forensic reports and remediation plans from disk snapshots and security alerts. Testing 23 frontier LLMs reveals that while agents can handle alert-driven findings, they consistently fail to proactively investigate disks for silent intrusions and to generate comprehensive remediation plans—with no model achieving complete detection and remediation on any single range. This exposes a fundamental bottleneck in deploying AI agents for real-world security operations and provides a public benchmark to drive future progress.
- Quality assurance
- Enterprise
Research
Evidence-Ledger Adjudication for Claim-Evidence Traceability
Gengyu Chen, Yongjie Yu, Weiling Wang
arXiv · 2026-07-29
This paper introduces 'evidence-ledger adjudication,' a workflow designed to verify whether AI-generated claims are actually supported by the evidence cited with them. Each claim is paired with an evidence packet and assigned a support relation (supported, contradicted, missing, or mixed), with unsupported claims routed back to the author. Testing on a 2,335-row benchmark drawn from AVeriTeC, CLIMATE-FEVER, and SciFact, the agent-based evidence-ledger system achieves 0.676 relation accuracy and 0.601 macro-F1, substantially outperforming the best non-agent baseline (0.383 accuracy, 0.303 macro-F1). The approach creates an auditable traceability layer for AI-assisted writing, helping catch cases where AI-drafted claims are not genuinely supported by their cited sources.
- Quality assurance