News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
Long-Term Simulation Exposes Cognitive-Developmental Risks in AI Companions
Kaicheng Shen, Lingyu Li, Wen Wu et al.
arXiv · 2026-06-24
This paper introduces TSJ (Theater-Stage-Judge), a longitudinal simulation framework for evaluating cognitive-developmental risks posed by AI companion systems—particularly large language model-based companions interacting with children and adolescents. The authors simulate 12,960 person-day interactions across six AI models, four developmental stages, twenty-four risk dimensions, and three psychological-vulnerability personas, finding that standard short-session safety evaluations systematically underestimate these risks, with stable risk estimates only emerging after 140 turns. Key findings include that early childhood and emerging adulthood are the most vulnerable developmental stages, and that cognitive trust and emotional dependency are the weakest domains. The work provides a scalable methodology for longitudinal safety assessment of AI companion systems, with direct implications for how such systems are evaluated and regulated.
- Quality assurance
- AI policy
Research
Text Over Image: Auditing Multimodal Robustness in Synthetic Medical Image Detection
Ching-Hao Chiu, Hao-Wei Chung, Gelei Xu et al.
arXiv · 2026-06-24
This paper investigates a multimodal vulnerability in vision-language models (VLMs) used for synthetic medical image detection: when both an image and accompanying metadata are provided, VLMs can overweight the text context, causing authenticity judgments to flip based solely on changes in the accompanying record rather than the image itself. The authors introduce a paired benchmark that holds images fixed while swapping controlled metadata variants, and find that adding an explicit AI-origin tag alone causes accuracy on authentic images to drop by 61.1% on average across multiple imaging modalities and diverse VLMs. They also propose an inference-time mitigation pipeline that detects and neutralizes these 'provenance shortcuts' without retraining, outperforming direct prompt-based suppression. These findings reveal a significant robustness gap in real-world clinical deployment of VLMs and provide a standardized benchmark for evaluating multimodal robustness beyond image-only settings.
- Quality assurance
- AI policy
Research
AI Coaching for Accelerating Human Skill Development with Reinforcement Learning
Wei Wang, Enlin Gu, Antonio Loquercio et al.
arXiv · 2026-06-24
This paper investigates how an AI agent can act as a coach to accelerate human motor-skill development, specifically in first-person-view drone racing. The authors formalize the coaching interaction as a non-cooperative dynamic game where the learner optimizes task performance and the coach targets the learner's independent competence, then build a reinforcement learning framework combining adaptive shared control with probabilistic models of the coach's causal influence on skill evolution. A user study with N=33 participants showed significant gains in human learning outcomes compared to state-of-the-art AI coaching baselines. This work is directly relevant to workforce skill development, demonstrating that carefully designed AI assistance—including strategic 'stepping back' to allow productive failures—can improve human capability rather than induce over-reliance or skill atrophy.
- Workforce
Research
The Media Labor Market: The New AI Skills
K. L. Zuykina, D. V. Razumova
Vestnik NSU Series History and Philology · 2026-06-24
An analysis of over 200 media industry job postings from 2023 to 2025 finds that AI proficiency has shifted from a supplementary skill to a core competency for media professionals. By 2025, AI-related job requirements had expanded well beyond editors and journalists to include copywriters, social media managers, PR specialists, and designers, with new roles such as AI translator, AI editor, and AI trainer emerging. Employers increasingly expect practical experience with AI content creation tools, along with capabilities in fact-checking, data analysis, and creative ideation. These findings highlight rapid transformation in media workforce requirements driven by AI adoption.
- Workforce
Research
Prioritizing Critical Success Factors for Artificial Intelligence (AI) Adoption in Marketing among Small and Medium Sized Enterprises in Vietnam
Nguyen Thi Thai Ha
International Journal of Research and Review · 2026-06-24
This study identifies and prioritizes the critical success factors for AI adoption in marketing among Vietnamese small and medium-sized enterprises (SMEs), using a Delphi process with 22 experts and Fuzzy Analytic Hierarchy Process (Fuzzy AHP). Eight validated factors emerged from an initial set of thirteen, with data quality, leadership commitment, and AI skills together accounting for more than 72% of the total weight. The findings offer practical guidance for SME managers on how to allocate limited resources to improve AI adoption outcomes in marketing contexts.
- Enterprise
- Workforce
- AI policy
Research
Modelling audit risk with AI and explainability: Cross-country evidence from emerging and mature markets
Salah Kayed, Ayman Bader, Abdulhadi Hamid Ramadan et al.
The International Journal of Digital Accounting Research · 2026-06-24
This study compares AI methods (Random Forest, XGBoost, deep neural networks) against traditional econometric models (logistic and probit regression) for predicting audit risk using firm-level data from the UAE and UK (2017–2024). AI models consistently outperformed econometric benchmarks on accuracy, recall, and AUC, while explainable AI tools like SHAP and LIME improved interpretability and regulatory transparency. Key audit risk drivers differed by country—governance factors dominated in the UAE while financial indicators and Big 4 affiliation were more influential in the UK—and cross-country model transfer showed reduced accuracy, highlighting the need for locally adapted approaches. The findings suggest that transparent, locally tuned AI models can meaningfully strengthen auditor and regulator confidence in audit risk assessment.
- Quality assurance
- Certifications
- AI policy
- Enterprise
Research
A Structured Domain Model for Organizational AI Adoption
Tim Geppert, Andreas Block, Maria Rothstein et al.
AI · 2026-06-24
This paper presents a structured domain model for organizational AI adoption, derived from a systematic literature review of 37 quantitative empirical studies and 810 retained data points. The model organizes adoption factors into nine clusters across Technology, Organization, and Environment dimensions, finding that workforce skills, perceived intelligence, and resources are the most studied positive drivers of AI adoption. Notably, constructs related to AI explainability and human-in-the-loop oversight are underrepresented in the literature despite growing regulatory demands such as the EU AI Act. The model is validated through expert feedback using the Content Validity Index and is intended to support future measurement instruments for organizational AI readiness.
- Workforce
- Enterprise
- AI policy
- Certifications
Research
The Impact of Talent Introduction Intensity on Corporate Artificial Intelligence Levels: Empirical Evidence from Chinese A-Share Listed Companies
Shanlin Bi
Financial economics research. · 2026-06-24
This study uses panel data from Chinese A-share listed companies to show that greater talent introduction intensity is positively associated with higher levels of corporate AI development. The relationship works through two key channels: easing firms' financing constraints and improving workforce quality. Effects are heterogeneous, being stronger in manufacturing, polluting industries, and certain regional contexts. The findings offer evidence relevant to designing talent policies aimed at promoting AI adoption in enterprises.
- Workforce
- Enterprise
- AI policy
Research
Determinants of artificial intelligence adoption in project management: a global analysis of knowledge, drivers and barriers
Aco Momčilović, Martina Vukašina, Zlatko Barilović
Mednarodno inovativno poslovanje = Journal of Innovative Business and Management · 2026-06-24
A global survey of 345 project management professionals examines what drives or hinders AI adoption in project settings. Results show that greater AI knowledge and organizational factors like leadership vision and competitive advantage are strong positive predictors of adoption, while uncertain ROI is a significant barrier. Notably, data privacy concerns and knowledge gaps showed a positive association with adoption, suggesting complex dynamics between perceived obstacles and actual adoption behavior. The findings offer practical guidance for organizations seeking to improve efficiency and decision-making through strategic AI implementation in projects.
- Enterprise
- Workforce
- AI policy
Research
Continuous improvement in the age of AI: human and relational barriers in public higher education
Alessio Travasi, Laura Bravi, Fabio Musso et al.
The TQM Journal · 2026-06-24
This study examines how AI adoption for continuous improvement affects administrative staff at six Italian public universities, drawing on 330 semi-structured interviews analyzed through Self-Determination Theory. Findings reveal a 'paradox of automation': staff welcome AI that reduces repetitive tasks but resist it when it threatens decisional autonomy, ethical accountability, or human interaction with students. The research introduces a 'legitimacy threshold' framework distinguishing AI as assistive infrastructure from AI as a professional substitute, with implications for how public higher education institutions manage AI-driven organizational change. These insights matter for workforce dynamics, enterprise adoption strategies, and quality assurance in public sector institutions.
- Workforce
- Enterprise
- Quality assurance
Research
Risk-adjusted conformity assessment of AI systems in healthcare: A scoping review
Svenja Reisinger, Raoul Kirmes, Christoph Dockweiler
Health Policy and Technology · 2026-06-24
This scoping review examines how the EU's dual regulatory regime—the Medical Device Regulation, In Vitro Diagnostic Regulation, and AI Act—shapes risk-adjusted conformity assessment for AI systems used in healthcare diagnostics, prognosis, and decision support. Analyzing 126 publications from 2022–2025, the authors identify recurring challenges including interface conflicts between regulatory frameworks, broad high-risk classification for medical AI, capacity constraints for notified bodies, and gaps in post-market evidence generation and accountability. The review finds that while the EU framework is considered comprehensive, its practical application remains contested, with unclear monitoring duties potentially shifting accountability to clinicians without effective means to verify AI performance. These findings are directly relevant to how AI medical systems are certified and governed across their full lifecycle.
- Certifications
- AI policy
- Quality assurance
Research
Artificial intelligence exposure and occupational wages: Evidence from the United States
Ozan Atalay
Journal of Economic Studies · 2026-06-24
Analyzing data from 671 U.S. occupations, this study finds a positive and statistically significant association between occupational exposure to AI and wage levels, with higher-AI-exposure jobs tending to pay more. The relationship is especially pronounced in cognitively intensive occupations and holds across quantile regression and instrumental variable approaches. The authors argue that AI may widen wage gaps between occupations by boosting productivity in certain roles, underscoring the need for targeted skill-upgrading and workforce transition policies to spread AI's benefits more broadly.
- Workforce
- AI policy
Research
When Thinking Is Outsourced: Cognitive Offloading and the Heterogeneity of Critical Thinking Among Chinese University Students Using Generative Artificial Intelligence
S Si, Yong Qi, Jingming Xu et al.
Journal of Intelligence · 2026-06-24
This study of 353 Chinese university students examines how using generative AI (GAI) affects critical thinking through a cognitive offloading lens. Cluster analysis revealed four distinct user profiles—from 'simple Q&A users' to 'critical co-thinkers'—and found that learning motivation was the strongest predictor of critical thinking gains, while deeper GAI use paradoxically predicted greater cognitive dependence rather than cognitive benefit. A notable 'high depth–high dependence' subgroup (25.8%) was disproportionately composed of female students and ICT majors, challenging the assumption that more sophisticated AI engagement automatically improves thinking skills. The authors recommend that educational interventions prioritize metacognitive training over technical skill development to prevent cognitive offloading from undermining critical thinking.
- Workforce
- AI policy
Research
Hitting a Moving Target: Test-Time Adaptation for AI Text Detection under Continual Distribution Shift
Kevin Ren, Manish Raghavan, Nikhil Garg
arXiv · 2026-06-23
This paper addresses the challenge of detecting AI-generated text under real-world conditions where distribution shifts continually occur after a detector is deployed—due to adversarial humanization of AI text, new language models being released, and temporal drift in human writing styles. The authors propose a test-time adaptation (TTA) framework using semi-supervised learning that exploits the homogeneity of AI-generated text at inference time, without requiring labeled data. Empirically, they show that state-of-the-art supervised detectors fail under these shifts—for example, the commercial model Pangram detects only 24.1% of adversarial AI-generated text—while their TTA approach achieves 90.5% detection under the same conditions. The work establishes test-time adaptation as a robust alternative for AI text detection in real-world deployment scenarios.
- Quality assurance
- AI policy
Research
BCoughBench: Benchmarking Respiratory Acoustic Foundation Models Under Body-Coupled Wearable Sensor Conditions
Mayur Sanap, Prasanna Desikan, Edgar Lobaton
arXiv · 2026-06-23
BCoughBench is a benchmarking framework that evaluates five respiratory acoustic foundation models (FMs) on classification and regression tasks under body-coupled wearable sensor conditions, which attenuate high-frequency content differently than the smartphone recordings these models were trained on. The study finds that mean AUROC drops from 0.785 (smartphone) to 0.689–0.723 under simulated wearable conditions, and no FM meets the clinical sensitivity threshold (Se@Sp95 ≥ 0.20) on most disease tasks under any sensor condition. Performance varies substantially by task—sex classification collapses (AUROC drops by up to 0.341) while COVID detection is nearly unaffected (Δ = −0.004)—and by model, with HeAR leading on regression and demographic tasks and M2D+Resp on disease and characteristic tasks. These findings highlight a critical gap between current FM evaluation practices and real-world clinical deployment on wearables, with direct implications for quality assurance and certification of AI-based respiratory diagnostics.
- Quality assurance
- Certifications
- AI policy
Research
The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
Eileanor LaRocco, Sarah Tan, Adarsh Subbaswamy et al.
arXiv · 2026-06-23
This paper examines the conditions under which clinicians would accept autonomous AI prescribing systems, using a survey of 136 U.S. prescribing clinicians alongside a regulatory and technical analysis. The authors argue that three architectural requirements are necessary for safe autonomous prescribing: calibrated per-prediction confidence with action-gated thresholds, differentiated communication of epistemic versus aleatoric uncertainty, and inferential transparency to support liability allocation. Survey results show clinicians would not permit autonomous prescribing without these features, and that meeting them would effectively transform autonomous AI prescribers into heavily supervised decision-support tools rather than true autonomous agents. The findings are directly relevant to ongoing U.S. legislation such as H.R. 238 and Utah's prescription-renewal pilot, offering a framework for regulators to constrain AI autonomy in prescribing while aligning liability with those who control system design and deployment.
- AI policy
- Certifications
- Workforce
- Quality assurance
Research
Dream at SemEval-2026 Task 13: SALSA for Single-Pass Machine-Generated Code Detection
Ruslan Berdichevsky, Shai Nahum-Gefen, Elad Ben-Zaken
arXiv · 2026-06-23
This paper proposes SALSA (Single-pass Autoregressive LLM Structured Classification), a method for detecting machine-generated code by mapping each classification label to a dedicated output token and training a large language model to emit a single-token label. The system targets out-of-distribution (OOD) generalization—detecting AI-generated code in unseen programming languages and application domains—using balanced sampling, parameter-efficient fine-tuning, and conservative training to avoid overfitting. On the SemEval-2026 Task 13 Subtask A official leaderboard, the approach achieves an OOD F1 score of 0.789, substantially outperforming the CodeBERT baseline (F1 = 0.305). The work addresses concerns around authorship integrity, academic assessment, and software trust in the context of widely used code-generating AI systems.
- Quality assurance
- Certifications
- AI policy
Research
Power-Flexible AI Data Centers: A New Paradigm for Grid-Responsive Compute
Chris Williams, Philip Colangelo, Ayse Coskun et al.
arXiv · 2026-06-23
This paper presents an architecture that enables GPU-based AI data centers to act as grid-interactive, power-flexible assets rather than static peak loads. By integrating grid signals, workload scheduling, and power telemetry, the system demonstrated on a real 130 kW GPU cluster can perform rapid load reduction, sustained curtailment, and carbon-aware operation while maintaining service levels for priority jobs. It also shows performance-aware load shifting across geographically distributed clusters to migrate workloads toward regions with lower grid stress. The findings suggest AI infrastructure can support grid reliability, speed up interconnection approvals, and improve the sustainability of large-scale computing operations.
- Enterprise
- AI policy
- Workforce
Research
What Does It Mean to Break a Distillation Defense?
Lena Libon, Pura Peetathawatchai, Michael Aerni et al.
arXiv · 2026-06-23
This paper examines output perturbation defenses designed to protect black-box large language models from distillation attacks, where adversaries query a model via API and train a smaller student model on its outputs. The authors identify a critical gap: existing defenses lack a shared threat model, making it hard to compare them or assess their real-world robustness. They propose a framework that characterizes attackers along three dimensions—query budget, data budget, and interface profile—and use antidistillation sampling as a case study to show that a defense's perceived effectiveness depends heavily on the assumed threat model. The paper argues that both technical defenses and governance or policy frameworks built around them must explicitly specify and stress-test attacker capabilities to avoid a false sense of security around intellectual property protection and regulatory compliance.
- AI policy
- Quality assurance
Research
When Multi-Sensor Fusion Fails to Generalize: Cattle Posture Classification Under Animal-Level and Temporal Distribution Shift
Leutrim Uka, Severino Pinto, Gundula Hoffmann et al.
arXiv · 2026-06-23
This study evaluates how well automated cattle posture-classification models (distinguishing lying vs. standing) hold up under realistic deployment conditions, including animal-level and cross-year distribution shifts. Using accelerometers, rumen-bolus sensors, and environmental measurements from a pasture-based beef herd across 2024–2025, the authors find that while multimodal XGBoost models achieve strong within-year performance (macro-F1 0.94), performance drops sharply under cross-year evaluation on previously unseen animals (macro-F1 0.49). Explainability analysis showed models continued relying on rumen-bolus and environmental signals even when those signals failed to generalize, indicating that multi-sensor fusion can reduce rather than improve robustness under temporal distribution shift. The results demonstrate that conventional random train-test evaluation substantially overestimates real-world readiness and call for robustness-centred evaluation standards in livestock-monitoring research.
- Quality assurance
Research
ScaleToT: Generalizing Structured LLM Reasoning for Billion-Scale Low-Activity User Modeling
Tianbao Ma, Chang Xi, Yichuan Zou et al.
arXiv · 2026-06-23
ScaleToT is a system that applies structured large language model (LLM) reasoning to model billions of low-activity users who lack rich interaction histories, a setting where standard LLMs are both unreliable and prohibitively expensive to deploy at scale. The approach trains a smaller student model on LLM-curated reasoning chains derived from a small subset of users, then transfers those reasoning signals to a lightweight encoder covering the remaining population without requiring full LLM inference. Evaluated on lifetime value (LTV) prediction in a billion-scale advertising deployment, a randomized online A/B test showed a 6.738% increase in LT30, while offline reasoning covered only 7.32% of the potential population, substantially reducing compute costs. The work demonstrates a practical path to scalable, structured AI reasoning for enterprise-scale user modeling.
- Enterprise
Research
To Compare, or Not to Compare: On Methodological Practices in Evaluating Social Bias
Federico Marcuzzi, Xuefei Ning, Roy Schwartz et al.
arXiv · 2026-06-23
This paper investigates how the structural design of social bias benchmarks—specifically whether evaluations use isolated demographic assessments versus forced-choice comparative settings—dramatically affects conclusions about Large Language Model (LLM) bias. The authors introduce a unified framework to standardize heterogeneous benchmarks and find a large, systematic 'paradigm gap': comparative settings act as strong catalysts for latent discrimination, especially in underspecified contexts, while isolated assessments understate bias. Notably, Chain-of-Thought reasoning worsens bias in comparative settings, neutral fallback options do not eliminate the effect, and this comparative prejudice scales with model size. The work provides a methodological guideline urging researchers to use comparative settings for bias auditing while warning practitioners against comparative deployments in ambiguous real-world tasks.
- Quality assurance
- AI policy
Research
A specialized reasoning large language model for accelerating rare disease diagnosis: a randomized AI physician assistance trial
Haichao Chen, Songchi Zhou, Zhengyun Zhao et al.
arXiv · 2026-06-23
RaDaR is a compact, open-source large language model (32 billion parameters) specialized for rare disease diagnosis, trained on approximately 49,000 real free-text cases and over 104,000 synthetic cases with reasoning-enhanced training. In a randomized physician-assistance trial, RaDaR improved diagnostic accuracy by 21.44 percentage points compared to internet search alone, and in a retrospective cohort it prioritized the correct diagnosis before documented clinical suspicion in 61.06% of cases, corresponding to a potential lead time of 1.87 months. The model outperformed larger open-source models including the 671-billion-parameter DeepSeek-R1 across public benchmarks and four external validation centers. These findings demonstrate that a deployable, data-efficient AI tool can meaningfully accelerate rare disease diagnosis where specialized clinical expertise is scarce.
- Workforce
- Enterprise
Research
The African Language Tax: Quantifying the Cost, Latency, and Context Penalty of Tokenizing African Languages in Frontier LLMs
Olaoye Anthony Somide
arXiv · 2026-06-23
This paper quantifies a structural 'tokenization tax' that speakers of African languages face when using frontier large language models: because tokenizers split African-language text into more subword tokens than equivalent English text, users pay higher inference costs, experience greater latency, and receive far less effective context capacity. Measuring 20 African languages across 11 tokenizers using parallel corpora (FLORES-200+ and SIB-200), the authors find every African language carries a tokenization premium above English, with medians around 1.88x on GPT-5 and extremes reaching 8.92x for N'Ko script—translating directly to up to 8.9x higher inference costs and as little as 11% of English's effective context window. The best available tokenizer (Gemma 4) reduces but does not eliminate the penalty, and the authors release an open measurement tool, leaderboard, and mitigation guidance. The findings reveal a concrete economic and capability disparity encoded at the infrastructure level, falling hardest on speakers who can least afford it.
- Enterprise
- AI policy
Research
LLM Performance on a Real, Double-Marked GCSE Benchmark
Malachy Fox, Kavi Samra, Paul Jung
arXiv · 2026-06-23
This paper introduces a benchmark of 32,534 double-marked real student responses to UK GCSE mock exams, covering 328 questions across five subjects including handwritten work, and evaluates how well large language models (LLMs) agree with human examiners. The authors find that top-performing LLMs agree with examiner consensus at least as closely as examiners agree with each other, across both subjective tasks like English essay marking and complex handwritten mathematics scripts. Agreement is consistent near the examiner consensus line and is not strongly dependent on model size, suggesting cost-effective automated marking is feasible at scale.
- Quality assurance
- Certifications