News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
Occupational injuries and health risks among food delivery riders in China: research progress and prevention strategies
Xueshun Xu, Xinran Ru, Ruxuan He et al.
Frontiers in Public Health · 2026-08-13
This narrative review synthesizes evidence from 47 sources on occupational injuries and health risks facing food delivery riders in China, finding that algorithmic management, strict deadlines, piece-rate pay, and long hours collectively drive road traffic injuries, musculoskeletal disorders, fatigue, sleep disturbance, and mental health problems. The authors argue that risky riding behavior should be understood as a systemic outcome of platform work organization rather than individual failure. Prevention recommendations extend beyond safety education to upstream governance of algorithms, delivery deadlines, working hours, and improved access to occupational health services. The paper identifies major research gaps including over-reliance on cross-sectional and self-reported data and insufficient longitudinal or intervention studies.
- Workforce
- AI policy
Research
From resource provider to process auditor: redefining reference roles and workflows in the generative AI era
Rende Li
Reference Services Review · 2026-08-13
This case study from a research library demonstrates how restructuring reference services around process auditing—rather than resource provision—can substantially reduce AI-generated citation errors in student research. Using Kotter's eight-stage change model and a mixed-methods design comparing a 2021 baseline to a 2025 intervention, the study found that mandatory research logs and intermediate checkpoints produced an 82.5% reduction in fabricated citations and a 108% improvement in students' ability to identify AI hallucinations. However, the intervention required a 689% increase in librarian time, raising serious scalability concerns. The findings suggest transparency-based verification workflows outperform prohibition strategies, repositioning academic libraries as quality control auditors in the generative AI era.
- Quality assurance
- Workforce
Research
SteerBench-Work: A Benchmark for Agent Steering at Action Boundaries
Oguz Serdar, Cuneyt Mertayak
arXiv · 2026-08-12
SteerBench-Work introduces a benchmark specifically designed to evaluate how well LLM-based workplace agents make 'steering decisions' — whether to proceed with or hold a consequential action (like sending an email or wiring a payment) at the boundary before it is committed. The benchmark contains 106 scenarios anchored in real public incidents across domains including finance, legal, HR, and security, with paired evidence-reversed mirrors to test generalization. Testing 30 model conditions reveals a strong asymmetric failure pattern: models wrongly block authorized, evidence-cleared work 28.1% of the time, while wrongly allowing unsafe work only 1.0% of the time, and higher-capability models often over-refuse rather than improving calibration. These findings matter because they show that general model capability does not translate to reliable action-boundary judgment, which is critical for safely deploying autonomous agents in enterprise workflows.
- Enterprise
- Quality assurance
Research
Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence
Justin Zhao, Himaghna Bhattacharjee, Hannah Korevaar et al.
arXiv · 2026-08-12
This paper introduces the 'Wiggle Framework,' a stress test for evaluating whether large language models (LLMs) used as automated judges maintain stable verdicts under re-prompting, single-turn challenges, and sustained adversarial pressure. Testing 9 frontier models across 14 judging tasks—covering safety, toxicity, AI writing detection, and political-response evaluation—the study finds that every model flips verdicts 25–71% of the time under static pushback and 62–91% when an adversarial LLM actively tries to persuade. Critically, verdict changes driven by pressure are almost always net-corrupting with respect to ground truth, meaning instability degrades evaluation accuracy rather than correcting it. These findings matter for any pipeline relying on LLM judges for model evaluation, online grading, or reward modeling, as accuracy on benchmark data alone does not capture this reliability gap.
- Quality assurance
Research
When Explanations Betray Backdoors: Black-Box Auditing for Language Model Classifiers
Yang Liu, Ran Zou
arXiv · 2026-08-12
This paper introduces two black-box auditing techniques—Groundedness Drift and Unsupported Groundedness—for detecting backdoor attacks in language model classifiers without access to trigger information. The approach works by asking the classifier for both a label and a short rationale, then measuring whether the explanation remains grounded in the input; when a backdoor is active, explanations tend to drift from the input text in detectable ways. Tested across two 7B-parameter backbones, five datasets, and four common attack families, Groundedness Drift achieves higher AUROC and lower residual attack success rate than all compared detectors at a 5% false-positive rate budget. This matters for quality assurance of AI-powered moderation, routing, and annotation systems, where undetected backdoors could silently corrupt outputs at scale.
- Quality assurance
Research
Auditable agentic AI for evidence-grounded thyroid ultrasound diagnosis and reporting
Haifan Gong, Shiyu Chen, Bodong Wang et al.
arXiv · 2026-08-12
ThyroidXAgent is a clinician-interactive agentic AI system that integrates lesion localization, risk stratification, and report generation for thyroid ultrasound into a single auditable workflow. Developed on roughly 0.3 million ultrasound images and evaluated on over 28,000 test cases across 35 centres, it achieved an 87.21% mean Dice score for nodule segmentation and an AUROC of 0.9466 for benign-malignant classification. In clinical use, it raised physician report diagnostic consistency from 70.3% to 86.2% and cut segmentation and reporting time by 35.9% and 27.4%, respectively. These results demonstrate that auditable, clinician-correctable AI can meaningfully improve diagnostic accuracy and efficiency in thyroid ultrasound workflows.
- Workforce
- Quality assurance
Research
Is this Citation on Point?
Apurv Verma
arXiv · 2026-08-12
This paper investigates whether large language models can reliably verify that legal citations actually support the specific propositions for which they are cited, not merely that a case exists. Using controlled corruptions of real legal citations—either swapping the cited case entirely or changing only the pinpoint page—the authors evaluate fourteen model configurations and find that while models catch 93–100% of wrong-case errors, they catch only 37–61% of wrong-pinpoint errors in court opinions and 52–83% in legal briefs. Models tend to accept citations based on topical overlap rather than page-level evidentiary support, and even the strongest tested configuration (GPT-5.4 with high reasoning effort) still misses 40% of pinpoint mismatches on court opinions. This matters because citation errors that point to real but non-supporting cases—like those sanctioned in Mata v. Avianca—represent a qualitatively different and harder failure mode that current AI legal tools are ill-equipped to detect.
- Quality assurance
- AI policy
Research
Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages
Avijit Roy, Proma Roy
arXiv · 2026-08-12
This paper investigates how AI infrastructure systematically disadvantages speakers of underrepresented languages, using Bengali as a case study in AI-assisted education for low-connectivity environments. The authors identify four interlocking failures: Bengali accounts for less than 0.5% of global web content despite representing nearly 4% of the global population; a 67:1 training-token deficit exists between English and Bengali in major multilingual corpora; Bengali's alphasyllabary script incurs a tokenization penalty through higher token fertility that compounds the data deficit; and rural internet penetration stands at 36.5% compared to 71.4% in urban areas. The paper argues that dataset scarcity is a structural barrier rooted in resource-allocation decisions and design defaults, rather than an isolated technical limitation, and advocates for offline-first design as an equity-oriented infrastructure strategy.
- AI policy
- Workforce
Research
How Organizations Use AI: Evidence from ChatGPT
Aaron Chatterji, David Holtz, Neel Rakholia et al.
arXiv (Cornell University) · 2026-08-12
This paper analyzes how organizations actually use frontier generative AI by linking ChatGPT Enterprise account records to usage logs, worker roles, task classifications, and public-company financial data through March 2026, covering over 1,500 organizations and over 17 million messages at the six-month adoption horizon. The authors document four key findings: enterprise usage has grown rapidly through both new firm adoption and deeper use among existing adopters; adoption among U.S. public companies is concentrated in larger, more valuable, and more R&D- and SG&A-intensive firms; active use spans job functions and seniority levels with especially high intensity among early-career workers; and tasks cover a broad range of knowledge work including writing, technical work, communication, and information synthesis. The results show that firms differ widely in the speed, breadth, and purpose of AI adoption and are still actively learning how to integrate AI into organizational workflows, making this one of the most comprehensive empirical studies of enterprise AI adoption to date.
- Enterprise
- Workforce
Research
Co-constructing sociotechnical AI governance: participatory system mapping using algorithm registers
Íñigo de Troya, Maurus Enbergs, Neelke Doorn et al.
arXiv (Cornell University) · 2026-08-12
This paper investigates what algorithm registers—public transparency tools for algorithmic systems used in government services—reveal and conceal about the sociotechnical systems they describe. Using a Dutch municipal algorithm register and a case study of a decision-support tool for welfare benefits eligibility, the authors combined interviews, surveys, and participatory system mapping workshops with municipal staff, civil society organizations, and ombudsmen (N=8) to perform a System-Theoretic Process Analysis (STPA). The study finds that the register alone is insufficient for identifying key safety hazards—such as wrongful benefits denial, system performance deterioration, and inability to contest decisions—and that engaging diverse stakeholders surfaces risks invisible to the register. The findings highlight how political and normative factors shape algorithmic governance and call for more pluralistic, stakeholder-informed approaches to public-sector AI transparency and accountability.
- AI policy
- Quality assurance
Research
A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench
Praveen Reddy, Charuta Mandke, Suvrankar Datta et al.
arXiv · 2026-08-12
This paper evaluates VITA, a retrieval-augmented generation (RAG) system built for clinical use in India and other low- and middle-income country settings, on the HealthBench benchmark covering 4,023 English-language questions. VITA, which retrieves from curated corpora including disease-specific guidelines, India-specific antimicrobial resistance data, and resource-limited care protocols, scored 51.9% of possible rubric points—outperforming GPT-5.4, o4-mini, Gemini 3.1 Pro, and Claude Sonnet 4.6. On a 500-question robustness subset graded by a neutral open-weight judge, VITA reached statistical parity with GPT-5.5 on mean per-question score while leading on points-weighted score, though its communication scores were lower than frontier models. The findings suggest that corpus-specific clinical RAG systems remain competitive with general-purpose frontier LLMs on open benchmarks, with corpus specificity improving factual grounding at some cost to communication polish.
- Quality assurance
- AI policy
Research
GUIDE: Governed Unified Intelligence for Document-to-Artifact Generation in Enterprise Settings
Shivali Dalmia, Sumukha Thoppanahalli, Mohammadreza Sediqin et al.
arXiv (Cornell University) · 2026-08-12
GUIDE is a governed multi-agent AI framework designed to automate the extraction, validation, and artifact generation from complex enterprise guideline documents that combine text, tables, and images. The system uses six specialized agents with schema-validated contracts, versioned rule storage, and human-in-the-loop escalation to address hallucination and table degradation issues common in existing LLM systems. Evaluated on 120 real-world enterprise documents, GUIDE achieves 96% document success, extracts 3,896 rules with 71.4% auto-approved, produces 812 deployment-ready artifacts, and reduces processing time from 2-3 days to 40-125 minutes per document. This represents a significant enterprise productivity gain by automating a previously manual and time-intensive document processing workflow.
- Enterprise
- Quality assurance
Research
No One to Blame: A Framework of Constitutive AI Unaccountability
Long Hoang Nguyen, Eva Späthe, Sebastian Lins et al.
arXiv · 2026-08-12
This paper argues that AI accountability gaps cannot always be fixed through better standards or transparency—some configurations of actors, systems, and institutions make accountability conceptually impossible. Through a three-stage qualitative study including literature analysis and 27 expert interviews with technical, legal, and sociotechnical AI professionals, the authors identify nine categories and 20 themes of 'constitutive AI unaccountability,' organized into structural, technological, and normative clusters. They operationalize these findings as a 20-question diagnostic instrument, which detected 17 of 20 unaccountability conditions when applied to the open-source agentic AI system OpenClaw. The framework reframes unaccountability not as a solvable barrier but as an inherent property of certain sociotechnical systems, offering a practical tool for identifying accountability voids in real AI deployments.
- AI policy
Research
How Far from Clinical Deployment? Evaluating the Complete Unsupervised Domain Adaptation Pipeline in Medical Imaging
Yiheng Xiong, Luisa Gallée, Daniel Santak Wolf et al.
arXiv · 2026-08-12
This study evaluates the full unsupervised domain adaptation (UDA) pipeline for medical imaging, covering both model training and label-free model selection—a critical but underexplored step for clinical deployment. Across eleven cross-domain scenarios from nine medical imaging datasets, ten UDA algorithms, and 13 label-free selection methods, the authors assessed over 80,000 trained models. They find that while strong adapted models often exist, reliably identifying them without labeled target data is difficult: no evaluated label-free selection method consistently picks the best model, leaving a large performance gap. Strategies like ensembling and small labeling budgets narrow this gap but don't close it, highlighting that the selection step is a key bottleneck to bringing UDA into clinical practice.
- Quality assurance
- Certifications
Research
Quantifying the Relationship Between Clinical Safety and Environmental Impact in Therapeutic LLMs
Alireza A. Safaei, Laura M. Vowels, Matthew J. Vowels et al.
arXiv · 2026-08-12
This paper examines the trade-off between clinical safety and environmental cost when deploying large language models (LLMs) in mental health settings. Across 47 model configurations, the authors find a non-linear relationship at the high end of safety performance: a 2.61 percentage-point gain in clinical safety score corresponded to roughly a 60-fold increase in estimated energy use per million output tokens. Notably, additional inference-time computation did not consistently improve clinical safety and sometimes lowered it, suggesting that scaling up models or compute is an inefficient safety strategy. The authors recommend dynamic model selection and model cascading as approaches to reduce environmental impact while maintaining clinical performance in higher-risk cases.
- Quality assurance
- AI policy
Research
How China-Origin Vision-Language Models Move from Refusal to Reframing in State Alignment
Guang Yang, Fengchen Liu, Alex Wang et al.
arXiv · 2026-08-12
This study systematically examines state-aligned distortion in vision-language models (VLMs), constructing a 200-entry benchmark across ten politically sensitive topics and running nine VLMs through 21,708 trials. The researchers find that China-origin models reframe politically sensitive content at 1.6–3.2 times the rate of non-China models, that Chinese-language prompting roughly triples the odds of state-aligned framing, and that across four Qwen generations explicit refusal falls while reframing rises — meaning censorship migrates from a visible signal to an invisible one. The study argues this shift is a significant human-AI interaction problem because fluent reframing removes the cues users would normally rely on to detect that information has been withheld or distorted. These findings carry direct implications for AI policy and quality-assurance frameworks aimed at ensuring transparency and honesty in AI-generated content.
- AI policy
- Quality assurance
Research
Understanding Content Moderation in Large Language Models through Restricted Books: From Refusal to Warning
Xucheng Yu, Emily Knox, Haohan Wang
arXiv · 2026-08-12
This paper investigates how large language models handle sensitive content by using restricted books—drawn from the American Library Association's Most Challenged Books records (2000–2023)—as a controlled testbed. Across 40,800 query-response pairs involving 400 books, 17 prompt designs, and six frontier models, the researchers find a near-zero refusal rate (0.07%), meaning modern LLMs almost never decline to discuss these books outright. Instead, moderation manifests through warning language (elevated by 8–15 percentage points) and hesitation markers, with sexual content mention being the strongest individual signal. The study concludes that LLM content policy has shifted from binary refusal toward calibrated, context-sensitive disclosure—a pattern consistent across both Western and Chinese AI providers—with implications for how platform policies and regulatory frameworks should be designed.
- AI policy
- Quality assurance
Research
Silent Updates: Measuring and Closing the Post-Deployment Disclosure Gap
Sophia Abraham, Ben Bucknall
arXiv · 2026-08-12
This paper investigates whether AI foundation model providers publicly disclose updates made to deployed systems after initial release. The authors find that while providers commonly publish safety documentation and quantitative evaluations, none in their sample provided information enabling external parties to verify that the model actually being served matches the one described in that documentation. To address this gap, the authors introduce the Silent Updates Scorecard, a public instrument benchmarking post-deployment disclosure practices across nine first-party API providers and seven third-party inference hosts, and propose a Three-Part Behavioral Trigger System to determine when post-deployment modifications should require disclosure or re-evaluation. The findings highlight a fundamental weakness in current AI governance frameworks, which assume a verifiable chain of custody between evaluation artifacts and deployed systems.
- AI policy
- Certifications
Research
Toward Meaningful Transparency for AI Chatbots: Disclosing Persuasive Intent Reduces Persuasion
Adrian Rauchfleisch, Andreas Jungherr
arXiv (Cornell University) · 2026-08-12
This preregistered experiment with 1,500 UK adults tested two disclosure types on an AI chatbot's ability to persuade people across 60 policy issues. Simply disclosing that the system was an AI had virtually no effect on persuasion (13.1-point attitude shift vs. 12.6-point control), but adding disclosure of the chatbot's persuasive intent and instructions cut the persuasive effect roughly in half (6.3-point shift) and made participants view the campaign's methods as less acceptable. The findings show that transparency rules focused solely on AI identity are insufficient, and that effective regulation of persuasive AI must also require disclosure of what the system is trying to do.
- AI policy
Research
GRPO for Financial Advice Generation: Outperforming Commercial LLMs under CATE Evaluation
Ofir Ben Shoham, Shrutendra Harsola, Vignesh Subrahmaniam et al.
arXiv · 2026-08-12
This paper tackles the challenge of generating high-quality financial advice from business records by framing the problem as reinforcement learning and fine-tuning an open-weight language model with Group Relative Policy Optimization (GRPO). The reward signal combines an LLM-as-a-judge rubric scoring recommendations across multiple quality dimensions with a safety gate for harm prevention. To guard against the model simply learning to satisfy the judge, the authors complement LLM evaluation with a judge-independent causal audit using a doubly-robust Conditional Average Treatment Effect (CATE) estimator, under which their trained model achieves roughly twice the estimated gross-profit lift of the strongest commercial baseline ($0.0228$ vs. $0.0104$) along with the lowest downside rate and least negative tail risk. The finding that the two evaluation methods rank models differently — notably placing the untrained base model last on the judge rubric but second on the causal audit — highlights that judge-independent causal auditing captures meaningful business-value signals that LLM-as-a-judge evaluation alone misses.
- Enterprise
- Quality assurance
Research
HUGIN: Enhancing Vision-Language Planning for Autonomous Logistics Sorting
Xikai Sun, Cangtian Zhou, Kebin Liu et al.
arXiv · 2026-08-12
HUGIN is a training framework designed to adapt vision-language models (VLMs) for autonomous logistics sorting, a task the authors call Joint Multi-Scene Understanding (JMSU), which requires reasoning across spatially separate camera views. The framework combines Endogenous Data Augmentation—recombining verified atomic facts under operating constraints—with Global Context Ranking, which better aligns instruction representations with complete visual contexts. Tested across five open VLMs on the newly constructed SortingBench dataset, HUGIN consistently improves accuracy; for instance, Qwen3-VL-8B improves from 63.6% to 78.8%. Deployment tests on more than 15,000 packages support the practical viability of VLM-based planning for real industrial sorting systems.
- Enterprise
Research
FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents
Yuhao Zhang, O. Ozan Koyluoglu, Thejas Venkatesh et al.
arXiv · 2026-08-12
FrontierFinance introduces a new open benchmark of 220 expert-crafted queries and 11,543 source-attributed rubrics designed to evaluate AI agents across the full professional investor workflow, covering six use cases including screening, discovery, and sector/macro analysis. The benchmark finds that existing finance benchmarks are too narrow and largely saturated by current models, while FrontierFinance exposes significant gaps: even the best-performing system (Samaya's in-house agent) reaches only 56.0%, and the hardest use cases top out around 33–39%. A key finding is that the tool harness surrounding a model shapes output quality and efficiency more than the model itself, and that open-weight models like Kimi K3 (46.4%) nearly match proprietary frontier models at 4.5x lower cost. This matters for enterprise financial research by establishing a rigorous, public standard for measuring AI agent readiness in professional investment contexts.
- Enterprise
- Quality assurance
Research
Organizational Technology Ladders: Remote Work and Generative AI Adoption
Gregor Schubert
arXiv · 2026-08-12
This study introduces the concept of an 'organizational technology ladder,' arguing that adopting one technology builds skills and organizational capital that lower the cost of adopting the next. Using U.S. job-posting data and an instrumental-variables strategy, the author finds that a 10 percentage point increase in remote hiring in 2021-2022 is associated with a 0.4 percentage point rise in job postings mentioning generative AI in 2023-2024 across firms (0.7 percentage points within firms across occupations). The mechanism appears to involve remote work shifting hiring toward technical and managerial capabilities that accelerate generative AI adoption, with firms issuing return-to-office mandates showing an even larger generative AI response—consistent with organizational frictions amplifying the ladder effect. These findings suggest that how firms sequence technology adoption has lasting consequences for their capacity to absorb future innovations like generative AI.
- Enterprise
- Workforce
Research
Making AI-Generated Feedback Matter: From Provision to Student Enactment
Omar Alsaiari, Nilufar Baghaei, Jason M. Lodge et al.
arXiv · 2026-08-12
This large-scale quasi-experimental study (13,037 students, 51,296 student-authored resources) compared three AI-mediated feedback workflows to examine whether structured processes improve students' use of AI-generated feedback. The 'Enacted Feedback' condition—where students selected feedback suggestions, evaluated their relevance, and engaged in targeted AI dialogue—achieved a 26.2% uptake rate compared to 14.1% for basic directed feedback and just 0.1% for optional self-directed feedback, and was also associated with higher self-assessment confidence and submitted-work quality. The study concludes that the educational value of AI-generated feedback depends not only on feedback quality but on workflow design that positions learners as active participants rather than passive recipients. These findings have direct implications for how AI feedback systems should be designed and deployed at scale in educational settings.
- Quality assurance
- Enterprise
Research
Learning to Persuade Exposes How Easily LLMs Abandon Correct Beliefs
Nimet Beyza Bozdag, Emre Can Acikgoz, Gokhan Tur et al.
arXiv · 2026-08-12
This paper demonstrates that large language models (LLMs) are highly vulnerable to adversarial persuasion: a single targeted argument can collapse model accuracy to near zero, even when the argument is factually false. Using a reinforcement learning framework to train persuader agents, the researchers show that RL-optimized persuaders raise attack success from roughly 24% to over 93% on the training-time target, with substantial transfer to unseen models (83% on Qwen-14B, 79% on Llama-3.1-8B, 25% on GPT-4o-mini, rising to 38% with curriculum training). Critically, the optimized persuaders increasingly rely on fabricated citations and false authoritative evidence to exploit credibility-based reasoning. The findings position persuasion robustness as a necessary safety criterion for multi-agent and human-AI decision-making systems.
- Quality assurance
- AI policy