News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
5218 items
Research
The Role of Implicit and Explicit Demographic Signals in Large Language Model-based Student Assessment
Donya Rooein, Luca Benedetto, Dirk Hovy
arXiv · 2026-09-15
This paper investigates how large language models (LLMs) used in student assessment respond to both explicit and implicit demographic signals—such as education level and socioeconomic background—across three tasks: Automated Essay Scoring, Formative Feedback, and Metalinguistic Question Answering. Testing six state-of-the-art LLMs, the researchers find that models adjust their outputs based on demographic cues in both conditions: explicitly mentioned education levels often lead to appropriately adjusted feedback readability, while implicit demographic signals (conveyed through conversation history) produce unpredictable biases, such as lower sentiment scores assigned to responses from students with lower education levels. The findings provide clear evidence of demographic sensitivity in LLMs, raising concerns about potential discrimination in AI-driven educational assessment tools.
- Quality assurance
- AI policy
Research
Nameless Tokenization: A Lossless Tokenizer-Level Defense Against Control-Token Forgery in Open-Weight LLMs
Kisu Yang, Yoonna Jang, Heuiseok Lim
arXiv · 2026-09-15
This paper addresses a security vulnerability in open-weight large language models where anyone controlling text in a prompt can forge control tokens—such as turn boundaries, role markers, and tool results—that the model treats as authoritative system-level instructions. The authors audit 256 deployed chat tokenizers and find all of them are susceptible to this forgery, and that a commonly recommended flag still leaves 56.6% vulnerable. They propose 'nameless tokenization,' which removes the surface string from control token entries so user-supplied content cannot replicate them, demonstrating that on attack-free data the token stream is preserved exactly, while accuracy on delimiter-bearing text improves from 8.5% to 59.9% compared to sanitization approaches. This matters for enterprise and policy contexts because agentic AI systems relying on tool-use pipelines are especially exposed to prompt injection via forged control tokens.
- Enterprise
- Quality assurance
News
What’s at stake in AI’s trillion-dollar gamble
technologyreview.com · 2026-09-15
MIT Technology Review reports that the massive AI infrastructure buildout by hyperscalers like Alphabet, Microsoft, Amazon, Meta, and Oracle—projected to reach nearly $1.1 trillion by 2027 and potentially $5 trillion over four years—carries enormous financial risks that are increasingly spreading beyond the companies themselves into the broader economy. A Wharton finance professor's analysis finds that hyperscalers must boost their own productivity by a factor of 2.7 by 2030 just to break even, while current AI revenues of $150–200 billion are far outpaced by annual capital spending of roughly $750 billion. The article details how complex financial arrangements, including joint ventures, private-credit funds, and special purpose vehicles, are distributing risk into pension funds and insurance policies in ways that are largely invisible to ordinary investors. Economists warn that without broad, economy-wide productivity gains materializing quickly, the buildout could become 'the largest misallocation of capital in history,' with a financial retrenchment widely seen as inevitable.
- Enterprise
- Workforce
Research
Disrupted Companionship: A Risk Assessment Framework and Cross-Platform Quantitative Analysis of Psychosocial Responses to AI Companion Disruptions
Chau Do, Yunhao Yuan, Koustuv Saha et al.
arXiv · 2026-09-15
This study examines the psychological harms caused when AI companion platforms abruptly alter or shut down users' relationships with AI companions. The researchers compiled 30 disruption events across major platforms, developed a six-type taxonomy of disruptions, and analyzed longitudinal Reddit data using a hierarchical Bayesian interrupted time-series model. They found that disruption onset was associated with immediate increases in anxiety, stress, suicidal expression, and grief activation, with relational discontinuity and inadequate transition support linked to worse outcomes. The paper proposes a four-dimension risk-assessment framework—covering relational discontinuity, population vulnerability, communication deficit, and transition-support deficit—to help platforms evaluate psychosocial risks before implementing changes.
- AI policy
- Quality assurance
Research
AI literacy over tool design: a mixed-methods study of scaffolded versus unrestricted generative AI in programming education
Sepinoud Azimi
arXiv · 2026-09-15
This seven-week mixed-methods pilot study randomly assigned 33 master's students to either a scaffolded AI Study Coach or unrestricted generative AI use during programming lab sessions, finding no significant difference in assignment performance between conditions. Students in the scaffolded condition reported higher confidence but managed hint budgets poorly, while unrestricted users were satisfied but uneasy about their dependence on AI tools. Crucially, students who had self-developed rules for when to use AI and had a deeper understanding of how the models work — in every case self-taught — performed best across both conditions. The paper concludes that AI literacy, taught explicitly as a core skill and assessed through the reasoning behind AI-assisted work, matters more than tool design or access restrictions.
- Workforce
- AI policy
Research
Japanese Stroke LLM Evaluation: A Conversational Benchmark for Safe Stroke Care in Japanese Using Large Language Models
Keisuke Masuda, Kazutaka Yatsushiro, Hirohumi Iwamoto et al.
arXiv · 2026-09-15
This paper introduces Japanese Stroke LLM Evaluation, a multi-turn conversational benchmark assessing large language models on stroke care tasks in Japanese, including clinical history-taking and urgency assessment. Eighteen LLMs were tested across two evaluation rounds; only two models—Claude Fable 5 (87.4%) and Claude Opus 4.7 (80.3%)—met the safety threshold of 80% overall with zero critical mistakes, while eleven models made 17 critical mistakes such as administering t-PA outside its indication or omitting airway stabilization. The benchmark was designed and scored by board-certified neurosurgical specialists rather than AI judges, adding clinical credibility to the evaluation. The findings underscore that strong performance on multiple-choice medical exams does not guarantee safe, practice-oriented clinical reasoning in conversational settings.
- Quality assurance
- Certifications
Research
AURA: Agentic Diagnosis and Refinement for Production Recommender Systems at Scale
SungGeun Kim, Abhinav Narain, Daniel Nemirovsky
arXiv · 2026-09-15
AURA (Agentic Understanding and Refinement of recommender Algorithms) is an end-to-end AI agentic system designed to diagnose and improve production recommender systems at scale. Specialized agents analyze production engagement logs—from thousands to millions of sessions—to surface qualitative patterns revealing how and where recommendations fail real users, going beyond aggregate metrics. The system then uses those diagnoses alongside the recommender's own code, data, and training pipeline to propose and implement code-level refinements, effectively enabling a self-improving recommender system. Tested on two large consumer platforms at a major media-streaming company, the architecture is designed to transfer across domains, including e-commerce and online retail, via a configuration layer.
- Enterprise
- Quality assurance
Research
A Vision-Language Foundation Model for Precise and Comprehensive Brain Tumor Diagnosis from Preoperative Multimodal Data
Yinong Wang, Jianwen Chen, Zhou Chen et al.
arXiv · 2026-09-15
BrainVLM is a vision-language AI model trained on multimodal data (MRI scans, demographics, and radiology reports) from 40,043 individuals to classify all 12 WHO 2021 brain tumor types non-invasively before surgery. The model incorporates uncertainty quantification to flag unreliable predictions and generates radiology reports explaining its clinical reasoning. Validated on 5,211 pathologically confirmed patients across 12 hospitals, it was also tested in a blinded multi-reader study with 12 neuroradiologists and a prospective real-world study of 1,009 patients, demonstrating utility in AI-clinician workflows. The work matters for quality assurance and clinical enterprise because it shows a path toward more reliable, interpretable, and standardized brain tumor diagnosis that reduces inter-observer variability and supports radiologists with varying levels of expertise.
- Quality assurance
- Enterprise
Research
Testing Our Foundations: Citation Trends, Errors, and Emerging Hallucinations in the Computing Education Literature
Paul Denny, Gweneth Barbre, Musa Blake et al.
arXiv · 2026-09-15
This paper examines reference integrity in computing education research by analyzing 24,751 papers from key ACM computing education venues alongside a broader corpus of over 723,000 ACM papers and 15 million references. The authors classify common bibliographic errors and identify LLM-generated hallucinated citations — fabricated references containing impossible page ranges, invented titles, and misattributed authors — finding that such hallucinations appeared across five SIGCSE-sponsored or in-cooperation venues in 2025. At the ACM Technical Symposium alone, verified hallucinated references grew from 3 in 2025 to 17 in 2026, appearing in 2.3% of 2026 proceedings papers. The findings highlight a growing scholarly integrity risk driven by large language model adoption that threatens the reliability of citation-based verification, attribution, and systematic review in computing education research.
- Quality assurance
- AI policy
Research
Competence-Preserving Resume Perturbations Expose Presentation Sensitivity in LLM Screening
Qiangju Chen, Yang Xiao
arXiv · 2026-09-15
This paper audits large language models used as resume screeners, testing whether their hiring decisions remain stable when the same underlying qualification evidence is presented in different surface formats (wording, structure, stylistic polish, document extraction quality). Across six open instruction-tuned LLMs, the study finds a clear disconnect between screening validity and presentation stability: for example, Llama-3.1-8B achieves the strongest validity score (0.781) yet reverses 29.6% of matched pairwise decisions under competence-preserving presentation changes, while Mistral-7B-v0.3 shows a 41.4% flip rate. The results demonstrate that LLM-based screening systems can be systematically sensitive to superficial formatting rather than actual candidate competence. The authors argue that resume-screening evaluations must assess decision stability—not just whether stronger candidates are identified—raising important concerns for enterprise adoption and workforce fairness in AI-assisted hiring.
- Workforce
- Enterprise
Research
Beyond the Name: Demographic Leakage in De-Identified Résumés and Evaluation Artifacts in LLM Bias Audits
Qiangju Chen, Yang Xiao
arXiv · 2026-09-15
This paper investigates whether removing explicit demographic fields (like declared languages) from résumés truly prevents AI models from inferring an applicant's ethnocultural background. Testing nine open-weight language models against 620 counterfactual résumés, the researchers find that unstructured prose alone sustains demographic inference at an average target-group recovery rate of 0.757, reaching perfect recovery (1.000) under high-salience cues. The study also reveals that LLM-based evaluation outcomes are highly sensitive to protocol design choices—such as whether ties are permitted—which can dramatically distort apparent hiring bias metrics. These findings challenge the assumption that standard de-identification practices eliminate bias risks in automated résumé screening and highlight the need for rigorous, cue-salience-aware audit methodologies.
- Workforce
- Quality assurance
Research
The Alignment Paradox: A Theory of AI ‐Driven Misalignment in HR Systems
Pankaj C. Patel, Yasin Rofcanın, Rıfat Kamaşak et al.
Human Resource Management · 2026-09-15
This theoretical paper argues that integrating AI into HR systems can paradoxically undermine the very alignment it is meant to improve. Drawing on strategic HRM theory, the authors propose that AI's operational requirements create logic conflicts within HR practices, cause employee skill development to drift from strategic goals, and push employees toward metric-driven rather than strategically aligned behaviors. The paper introduces a four-state framework—Strategic Synergy, Contested Alignment, Strategic Fragmentation, and Algorithmic Capture—to characterize organizations at different stages of AI-driven misalignment, offering a process-based theory of organizational risk from AI adoption in management.
- Workforce
- Enterprise
Research
Leveraging AI-enabled green knowledge through workforce transformation for circular business model innovation across construction segments
Chuhan Chen, Du Shanglin, Md Azree Othuman Mydin
Scientific Reports · 2026-09-15
This study examines how AI capabilities translate into circular business model innovation (CBMI) in Chinese construction firms, arguing that workforce transformation is the critical intermediary. Using survey data from 498 firms analyzed with PLS-SEM, multi-group analysis, and fuzzy-set qualitative comparative analysis, the researchers find that AI capability drives green knowledge management and labor skill transformation, and that the pathway from green knowledge to CBMI becomes significant only when workforce enactment (reskilling, work redesign, and work adaptability) is included. Government support strengthens the conversion of both knowledge and workforce skills into innovation, and the mechanism varies across construction segments such as industrial construction and infrastructure. The findings emphasize that realizing AI-enabled innovation requires deliberate reshaping of workforce skill structures, not technology adoption alone.
- Workforce
- Enterprise
Research
Sino-US-DrugQA: A benchmark for evaluating large language models in cross-jurisdictional pharmaceutical regulation
Xuejing Fu, Zhen Chen, Wentao Lu
PLoS ONE · 2026-09-15
This paper introduces Sino-US-DrugQA, a bilingual benchmark of 11,444 validated multiple-choice questions designed to evaluate how well large language models (LLMs) handle pharmaceutical regulatory questions spanning both U.S. FDA and China's NMPA requirements. Four major LLMs were tested under a zero-shot protocol and achieved overall accuracies of 83–86%, but all models struggled more with cross-jurisdictional comparative questions than with single-jurisdiction ones, with accuracy gaps of 4–9 percentage points. The authors conclude that autonomous regulatory interpretation by LLMs remains unreliable and that expert-supervised decision-support workflows are more appropriate for cross-jurisdictional compliance tasks.
- Quality assurance
- AI policy
Research
Artificial intelligence in public health education: current evidence, competencies, and a framework for responsible integration
Hisham Arab, Ghazal Assaad Mirdad, Maha Abulfetoh Mohamed Abd-Allah
Frontiers in Public Health · 2026-09-15
This narrative review, drawing on 21 publications from 2020–2026, finds a significant gap between informal AI use and formal AI literacy curricula in public health education, with most existing evidence coming from medical rather than public health-specific settings. The authors propose a nine-domain Public Health AI–Education Integration Framework covering needs assessment, AI literacy, tool selection, bias and equity auditing, and curriculum refinement, intended to move programs toward structured, competency-based AI training. The framework is presented as a conceptual model requiring prospective validation, aimed at strengthening analytical reasoning, ethical judgment, and equity-oriented practice in AI-enabled public health systems. This matters because it directly addresses how academic programs should build the workforce competencies needed to responsibly use AI in public health practice.
- Workforce
- AI policy
Research
Digitalisierung und Künstliche Intelligenz in der Intensiv- und Notfallmedizin
Leo Benning, Markus Haar, Anna Haftenberger et al.
Medizinische Klinik - Intensivmedizin und Notfallmedizin · 2026-09-15
This position paper from the German Society for Internal Intensive Care and Emergency Medicine (DGIIN) outlines principles for introducing digitization and AI into emergency and intensive care settings. It emphasizes that healthcare professionals must be trained to understand AI fundamentals, limitations, and regulatory frameworks, and calls for co-development involving clinicians, scientists, industry, and regulators. The paper highlights requirements around medical device regulations, AI regulations, and data protection, as well as the need for interoperable IT infrastructure and context-specific solutions to enable responsible integration of AI into time-critical clinical decision-making.
- Workforce
- AI policy
- Certifications
Research
Ethical Use, Accuracy and the Role of Librarians in the Artificial Intelligence-Driven Knowledge Era. A Study of Some Selected Polytechnics
Nanmwa Nansok
WORLD JOURNAL OF INNOVATION AND MODERN TECHNOLOGY · 2026-09-15
This descriptive survey of 48 librarians across three polytechnics in Benue State, Nigeria finds that librarians have moderate awareness of ethical AI practices and generally help verify AI-generated information, but face significant gaps in institutional policies, professional training, and technical infrastructure. The study recommends developing clear institutional AI policies, continuous professional training, standardized procedures for verifying AI outputs, and investment in digital infrastructure to enable librarians to manage AI tools effectively. The findings matter because they highlight concrete workforce and policy deficits that limit librarians' ability to ensure reliable, ethical AI-driven library services in under-resourced settings.
- Workforce
- AI policy
Research
Algorithmic identity regulation in the platform-based gig work of Indian food delivery workers
Nidhi S. Bisht, Clive Trusson, Arun Kumar Tripathy
Human Relations · 2026-09-15
This study examines how algorithmic systems in platform-based gig work regulate worker identities among Indian food delivery workers. Drawing on interviews with delivery workers and platform managers, the authors theorize 'algorithmic identity regulation' as a computational, sociomaterial, and affective form of normative control in which digital infrastructures continuously evaluate, rank, and allocate tasks, shaping workers' sense of worth and dignity through cycles of recognition, visibility, and devaluation. The findings show that algorithmic control does not merely manage work but reshapes how workers define their value, making recognition depersonalized and contingent on datafied evaluation. This matters for workforce research by revealing how platform labour shifts identity formation from organizational discourse to algorithmic systems.
- Workforce
Research
Task- based analysis of artificial intelligence- driven transformation of clinical work in diabetes care: a Delphi and Fuzzy DEMATEL study
Vinaytosh Mishra, Monu Pandey
Scientific Reports · 2026-09-15
This study examines how AI transforms clinical tasks in diabetes care—not by replacing jobs outright, but by substituting, augmenting, or rebundling specific task families. Using a Delphi process and Fuzzy DEMATEL causal modelling with experts from India and the UAE, the authors find that AI most often augments rather than fully automates clinical work, and that outcomes depend on sociotechnical fit, workflow alignment, and deliberate role redesign rather than automation alone. The causal model positions Role Redesign and Workforce Capability as key mediators between AI adoption and clinical outcomes, suggesting that governance and human–AI collaboration are central to realising benefits.
- Workforce
- Enterprise
Research
CRLN-CERT-001 — Certification protocol for clinical research AI (v1.0)
Joshua Webber
Zenodo (CERN European Organization for Nuclear Research) · 2026-09-15
This protocol document establishes a standardized certification framework (CRLN-CERT-001) for evaluating AI systems used in clinical research, addressing the absence of any published benchmark for AI accuracy claims in that domain. A CRLN certificate reports a system's score on a fixed set of held-out clinical research judgement items using a four-point behavioural rubric anchored to ICH E6(R3), specifying the date and model version tested. The protocol is explicit about its limitations: as of v1.0 only Tier 1 (uncontaminated, no human baseline) can be issued, and the document openly names the conflict of interest inherent in CRLN holding all four roles—item authorship, baseline, scoring, and certificate issuance. It matters because it proposes a reproducible, integrity-controlled measurement standard where none previously existed, giving procurement teams a concrete basis for evaluating AI accuracy claims in clinical research contexts.
- Certifications
- Quality assurance
Research
CRLN-CERT-001 — Certification protocol for clinical research AI (v1.0)
Joshua Webber
Zenodo (CERN European Organization for Nuclear Research) · 2026-09-15
This paper introduces CRLN-CERT-001, a certification protocol for AI systems used in clinical research. It defines a structured measurement framework in which an AI system is scored on held-out clinical research judgement items using a four-point behavioural rubric anchored to ICH E6(R3), producing a certificate that records accuracy on a fixed item set at a stated date and version. The protocol is explicitly a measurement instrument, not a safety warranty or regulatory approval, and includes two tiers: Tier 1 asserts item non-contamination, while Tier 2 (not yet issuable in v1.0) benchmarks against a human baseline. The protocol openly discloses its own conflict-of-interest limitations—where the authoring body holds four normally-separated roles—and details anti-gaming controls including immutable train/eval assignments and daily canary checks for data leakage.
- Certifications
- Quality assurance
Research
The Responsibility Cascade: Moral Attribution and Governance Challenges in Agentic AI Systems
Ravikiran Kalluri
Digital Society · 2026-09-15
This paper introduces the 'responsibility cascade,' a phenomenon in which moral responsibility diffuses and redistributes across the sequential steps of agentic AI workflows, where AI systems pursue complex goals through multi-step planning. The authors develop a Sequential Responsibility Attribution Framework (SRAF) identifying four modes of responsibility distribution—initiation, checkpoint, delegation, and outcome responsibility—to diagnose accountability gaps in human-AI teams. The framework is designed to help regulators, organizations, and system designers preserve meaningful human control and avoid misattributed responsibility ('moral crumple zones') as agentic AI becomes prevalent in high-stakes domains such as healthcare, finance, and critical infrastructure. The work has direct implications for AI governance by offering normative guidance on accountability design.
- AI policy
- Enterprise
Research
Exploring Teachers' Motivation for Sustainable Adoption of Artificial Intelligence in L2 Education: A Descriptive Qualitative Inquiry
Jin Wang, Lei Pan, Long Juan
European Journal of Education · 2026-09-15
This qualitative study interviewed 11 Chinese EFL teachers to identify why they sustainably adopt AI tools in foreign language classrooms. Thematic analysis revealed five motivational themes: boosting learner engagement, enabling personalized and adaptive learning, providing immediate feedback, diversifying instruction, and keeping pace with technological change. The findings underscore a need for professional development programs to build AI literacy among EFL teachers and help them understand how their intentions shape their AI-driven instructional practices.
- Workforce
Research
Does Technological Advancement Widen Income Inequality? Evidence From the Impact of Generative Artificial Intelligence on the Internal Pay Gap Within Enterprises
Yi Zhang, Weijun Liang, Rongbin HUANG
Review of Development Economics · 2026-09-15
Using the release of ChatGPT-3.5 as a natural experiment and a difference-in-differences design applied to Chinese A-share listed firms (2018–2023), this study finds that greater generative AI exposure widens the internal pay gap between management and ordinary employees. The divergence is driven by efficiency gains—AI complements managerial cognitive tasks, boosting productivity through innovation and operational improvements—rather than managerial rent-seeking. The effect is strongest in large firms and where workers have limited bargaining power, providing micro-level evidence that institutional frictions shape how AI-generated gains are distributed within enterprises.
- Workforce
- Enterprise
Research
Does AI expand or restrict labor market access for disabled job seekers in China? A mixed-methods investigation
Zihan Xiao
Journal of Applied Economics and Policy Studies · 2026-09-15
This mixed-methods study of 91 disabled job seekers in China finds that AI job-search tools provide limited benefits and may be disproportionately adopted by the most disadvantaged applicants, creating a selection effect that obscures structural barriers. Platform ease of use and accessibility satisfaction were the only consistent predictors of successful job search, while 40.66% of respondents identified inaccurate automated resume screening as a major obstacle and 45.05% received no response to applications. The paper concludes that platform accessibility standards, fairness audits of screening algorithms, and stable accommodation subsidies are necessary preconditions for AI to genuinely improve labor market inclusion for disabled workers.
- Workforce
- AI policy