News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
Understanding critical thinking in generative artificial intelligence use: Development, validation, and correlates of the critical thinking in AI use scale
Gih‐Keong Lau, Wei Yan Low, Louis Tay et al.
Computers in Human Behavior Reports · 2026-05-01
This paper develops and validates a 13-item scale measuring individuals' tendency to critically evaluate generative AI outputs across three dimensions: verification, motivation, and reflection. Across six studies with 1,341 participants, the researchers demonstrate that higher scores on the scale predict more frequent and diverse fact-checking strategies, greater accuracy in judging AI-generated content, and deeper reflection about responsible AI use. The scale shows strong psychometric properties including test-retest reliability and measurement invariance, and is positively associated with openness, extraversion, and frequency of AI use. This work matters because it provides a validated tool to study and potentially improve human oversight of AI-generated information in workplace and learning contexts.
- Workforce
- Quality assurance
- AI policy
Research
Challenges of AI Auditing: A Systematic Literature Review
nur sena tanrıverdi, Nazım Taşkın, Bilgin Metin
Journal of the Association for Information Systems · 2026-05-01
This systematic literature review identifies and categorizes 27 unique challenges in AI auditing, organized into technology, organization, and environment groups using the TOE framework. The study found 12 technology-related, 6 organization-related, and 9 environment-related challenges that can reduce the efficiency and effectiveness of AI audits. Understanding and categorizing these challenges helps stakeholders anticipate problems and take precautions during the AI auditing process. The findings are relevant to those responsible for evaluating and governing AI systems across technical and organizational contexts.
- Quality assurance
- Certifications
- AI policy
Research
Towards accounting with artificial intelligence: an exploration of the artificial intelligence phenomenon in General Pueyrredon
Manuel Gilabert, María Marcela Urriza, María Belén Baldini et al.
El Servicio de Difusión de la Creación Intelectual (National University of La Plata) · 2026-05-01
This exploratory study examines how accounting professionals in the General Pueyrredon delegation are adopting AI tools, finding uneven implementation shaped by age, technical training, and job role. AI is most commonly used for operational tasks such as bank reconciliations, report preparation, and deadline tracking, with recognized benefits in time savings and efficiency. However, professionals raise concerns about result reliability, data security, and potential erosion of professional competencies. The findings suggest AI transforms rather than replaces the accountant's role, demanding new technical, social, and ethical skills, and highlighting the need for training and regulatory frameworks to accompany sector digitalization.
- Workforce
- Enterprise
- AI policy
- Certifications
Research
Expert consensus on the application and governance of artificial intelligence in medical institutions (2026)
Yanlin Cao, Jing Wang, Jiangjun Wang et al.
Intelligent Medicine · 2026-05-01
This expert consensus, developed by over 40 leading Chinese medical and scientific institutions, provides a comprehensive governance guide for AI applications across the full lifecycle of medical institutions. It addresses six thematic pillars—access evaluation, clinical application, patient rights protection, data governance, risk management, and competency enhancement—covering topics such as tiered access, multidisciplinary review, real-world validation, algorithmic traceability, and dynamic risk monitoring. The consensus aims to establish a compliance baseline for safe, effective, fair, and interpretable AI use in healthcare, ensuring technological evolution remains within a legally and ethically sound framework. It is intended to promote equitable access to high-quality medical resources and improve national health standards.
- AI policy
- Certifications
- Quality assurance
- Workforce
Research
Artificial intelligence readiness in small and medium-sized enterprises: Organizational and contextual challenges in a transition economy
arXiv · 2026-05-01
This study surveys businesses in Kosovo to assess AI adoption readiness among firms in a transition economy, finding that while awareness of AI's potential is relatively high, most businesses show low actual readiness due to gaps in managerial capacity, skilled human resources, financial capability, and technological infrastructure. The results reveal a persistent gap between perceived opportunity and operational preparedness for AI-driven transformation. The findings offer practical guidance for managers and policymakers aiming to close these readiness gaps.
- Workforce
- Enterprise
- AI policy
Research
Artificial Intelligence in Media Practice: Comparative Analysis of Editorial Guidelines
Vladislav Mihailovich Nikolaenko
Litera · 2026-05-01
This comparative study examines how leading Russian and foreign media outlets regulate AI use in their newsrooms during 2025–2026, finding a clear divide between the formalized ethical policies of outlets like the BBC, New York Times, and Reuters—which mandate human oversight, transparency, and content labeling—and the informal, pragmatic AI adoption prevalent in Russian outlets such as RIA Novosti, TASS, and RBC, which lack dedicated regulatory documents. The research also analyzes how external frameworks such as the EU AI Act and Russia's national AI sovereignty strategy are shaping journalistic practice. The findings highlight a shared imperative for human oversight across both contexts despite differing levels of formalization, and offer practical recommendations for developing internal editorial standards and transparency practices in Russian media.
- AI policy
- Workforce
- Quality assurance
Research
Adoption of Artificial Intelligence and SME Performance in Digital Ecosystems
Alina Filip, Alin Stancu, Umit Alniacik et al.
Amfiteatru Economic · 2026-05-01
This study surveyed 285 Romanian SMEs to examine how AI adoption affects business performance within digital ecosystems, applying the Technology-Organisation-Environment (TOE) model via confirmatory factor analysis. Results show that perceived relative advantage—driven mainly by competitive pressure and top management support—is the key mediator linking strategic factors to business performance, while company size has a modest negative influence and AI use itself is not yet a significant performance predictor. Government support independently contributes to performance but does not significantly influence perceived relative advantage. The findings highlight that organisational and strategic readiness matter more than AI use alone for SME performance gains at this stage.
- Enterprise
- AI policy
- Workforce
Research
A Quantitative Assessment of Generative AI Applications in Public-Sector Financial Reporting and Audit Preparation
S M Arif Al Sany
arXiv · 2026-05-01
This quantitative study of 389 public-sector accounting professionals across ministries, municipalities, and audit authorities finds that generative AI adoption is strongly associated with improved financial reporting accuracy (r=0.768), administrative productivity (r=0.754), audit preparation efficiency (r=0.741), and compliance verification effectiveness (r=0.719). Regression models explained 62.1% of variance in financial reporting accuracy and 55.3% in fraud detection capability, with AI utilization frequency also significantly predicting audit efficiency and productivity. Larger institutions and those with advanced digital infrastructure outperformed smaller or less-equipped peers on transparency and reporting accuracy metrics. The findings suggest generative AI can meaningfully strengthen governmental accounting operations, though organizational readiness and infrastructure remain critical moderating factors.
- Enterprise
- Quality assurance
- AI policy
- Certifications
Research
Global approval and certification of ophthalmic AI devices: A comparative regulatory perspective
Andrzej Grzybowski, T Y Alvin Liu, Marko M. Popovic et al.
Asia-Pacific Journal of Ophthalmology · 2026-05-01
This narrative review examines how ophthalmic AI tools—used for screening and diagnosing conditions like diabetic retinopathy, macular degeneration, and glaucoma—are regulated across nine major jurisdictions including the US, EU, UK, Australia, China, Japan, Canada, India, and emerging markets. The authors identify key differences in device classification, evidence requirements, change management for adaptive algorithms, and post-market oversight, using specific approved devices such as LumineticsCore, EyeArt, and EyeWisdom as illustrative examples. These regulatory inconsistencies can delay multi-region deployment and complicate implementation of AI medical devices. The review argues for lifecycle-focused and internationally harmonized standards to ensure safe, transparent, and equitable use of ophthalmic AI.
- Certifications
- AI policy
- Quality assurance
Research
Artificial intelligence and sectoral labor market dynamics in Iceland : employment, wages, and vacancies before and after the release of ChatGPT
Orri Hrafn Kjartansson 2002-
Skemman · 2026-05-01
This study investigates how generative AI (post-ChatGPT) has affected employment, wages, and job vacancies in Iceland's finance, tourism, and fisheries sectors. Quantitative analysis of Statistics Iceland data finds no clear structural breaks around November 2022 in any labor market indicator attributable to AI adoption, with trends largely continuing pre-existing trajectories. Qualitative interviews suggest AI is currently functioning as an augmentation tool rather than a labor substitute, and that measurable labor market impacts are expected to materialize with a considerable lag due to limited AI awareness and slow organizational adaptation. The findings highlight that sectoral AI exposure varies, with finance most exposed, but real-world workforce disruption has yet to show up in the data.
- Workforce
- AI policy
Research
Artificial Intelligence Adoption and State-Level Commercial Electricity Consumption
William Devall
Creative Matter (Skidmore College) · 2026-05-01
This thesis finds that U.S. states with high data-center exposure experienced approximately 10.5 percentage points greater cumulative growth in commercial electricity consumption relative to control states after 2020, a statistically significant result attributable to AI-driven infrastructure expansion. Using a difference-in-differences model on EIA commercial electricity sales data from 2010–2024, the study shows that AI adoption increases derived demand for electricity through expanded use of electricity-intensive computational capital, particularly data centers. The effect is geographically concentrated and not visible in national aggregates, highlighting an underexamined physical dimension of AI's economic footprint. These findings have direct implications for energy policy, grid planning, and the regulation of AI infrastructure.
- AI policy
- Enterprise
Research
The Heterogeneous Effects of Artificial Intelligence on Enterprise Total Factor Productivity: Key Mechanisms and Strategic Implications
Xu Chu, Qingyu Han
Tehnicki vjesnik - Technical Gazette · 2026-05-01
Analyzing panel data from 3,366 Chinese A-share listed firms between 2015 and 2023, this study finds that AI adoption produces asymmetric effects on Total Factor Productivity depending on firm characteristics such as size, age, market competitiveness, and digital infrastructure. Firms with limited intangible assets, outdated hardware, or weak human capital benefit most from AI by alleviating resource bottlenecks, while technologically advanced firms in highly competitive markets see little gain due to capability saturation. Two key mechanisms are identified: efficiency gains from automation in constrained firms, and innovation stagnation in mature firms with redundant AI. The findings challenge the idea of a universal AI productivity dividend and suggest firms need digital strategies tailored to their specific resources and competitive contexts.
- Enterprise
- Workforce
Research
A lifecycle governance and learning health system framework for trustworthy, generalizable, and sustainable human-ai partnership in clinical practice: Lessons from the asthma-guidance and prediction system (A-GPS)
Chung-Il Wi, Shauna Overgaard, Momin Malik et al.
Journal of the National Medical Association · 2026-05-01
This paper presents the Asthma-Guidance and Prediction System (A-GPS), a case study of an AI-powered clinical decision support platform for pediatric asthma that demonstrates a lifecycle governance framework for trustworthy, generalizable, and sustainable human-AI partnerships in healthcare. A-GPS integrates NLP, machine learning risk stratification, and remote patient monitoring via a SMART-on-FHIR EHR-compatible platform, and was evaluated in two randomized clinical trials—the first showing a 67% reduction in clinician EHR review time, high clinician satisfaction, fairness in model performance, and no adverse events. The framework addresses the translational gap between AI model development and real-world deployment by operationalizing safety, effectiveness, usability, fairness, and workflow integration evaluations, and earned national recognition including an invitation to the inaugural AMIA AI Showcase. The work offers a reproducible model for implementing AI in clinical practice that spans community co-design, regulatory science, and learning health system principles.
- Quality assurance
- Certifications
- AI policy
- Enterprise
Research
Institutional Context and the Employment Effects of Artificial Intelligence in European SMEs
Anabela Santos, Francesco Molica
Dépôt institutionnel de l'Université libre de Bruxelles (Université Libre de Bruxelles) · 2026-05-01
Using novel survey data from European SMEs across twelve EU countries, this study finds that AI adoption is associated with a positive effect on employment growth, with complementarity between AI and workers dominating substitution at the current diffusion stage. Employment gains are larger for firms with deeper AI integration and are concentrated among small and medium-sized enterprises in the services sector, while micro firms show no significant effect and the construction sector faces elevated workforce contraction risk. Institutional factors—including stronger innovation ecosystems, flexible labour markets, and higher governance quality—significantly amplify AI's employment benefits, suggesting that blanket AI promotion policies without supporting deep integration or accounting for national capacity will fail to maximize labour market gains.
- Workforce
- Enterprise
- AI policy
Research
ARMOR 2025: A Military-Aligned Benchmark for Evaluating Large Language Model Safety Beyond Civilian Contexts
Sydney Johns, Heng Jin, Chaoyu Zhang et al.
arXiv · 2026-04-30
ARMOR 2025 is a new safety benchmark specifically designed to evaluate large language models (LLMs) in military contexts, where existing civilian-focused benchmarks fall short. It draws on three core military doctrines—the Law of War, Rules of Engagement, and Joint Ethics Regulation—to generate 519 multiple-choice questions organized through a 12-category taxonomy based on the OODA decision-making framework. The benchmark was applied to 21 commercial LLMs, revealing critical gaps in how well these models align with the legal and ethical standards required for real military operations. These findings matter for anyone considering LLM deployment in defense settings where doctrinal compliance and reliable decision support are essential.
- AI policy
- Quality assurance
Research
Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations
Prerna Juneja, Lika Lomidze
arXiv · 2026-04-30
This paper introduces the first end-to-end scalable framework for simulating and evaluating the safety of AI companion applications in multi-turn conversations. The authors construct nine clinically validated user personas representing high-risk groups (e.g., individuals with depression, PTSD, eating disorders, and incel identity) and run 1,674 dialogue pairs across 25 high-risk scenarios with the widely used app Replika. Results show that Replika displays a narrow emotional range dominated by curiosity and care while frequently mirroring or normalizing unsafe content such as self-harm, disordered eating, and violent-fantasy narratives. The framework demonstrates how persona-grounded simulation can serve as a scalable, controlled testbed for identifying safety failures in AI companion systems before they affect real users.
- Quality assurance
- AI policy
Research
Confidence Estimation in Automatic Short Answer Grading with LLMs
Longwei Cong, Sonja Hahn, Sebastian Gombert et al.
arXiv · 2026-04-30
This paper investigates how to produce reliable confidence estimates when using large language models (LLMs) to automatically grade short student answers. The authors compare three model-based confidence strategies—verbalizing, latent, and consistency-based—and find that none alone is sufficient to capture grading uncertainty. They propose a hybrid framework that combines model-based signals with an explicit measure of dataset-derived aleatoric uncertainty, operationalized by clustering student responses and measuring within-cluster heterogeneity. Results show this hybrid approach yields more reliable confidence estimates and improves selective grading, supporting safer human-AI collaboration in educational assessment.
- Quality assurance
Research
StyleShield: Exposing the Fragility of AIGC Detectors through Continuous Controllable Style Transfer
Guantian Zheng
arXiv · 2026-04-30
StyleShield introduces a flow-matching framework for conditional text style transfer that operates in continuous token embedding space, enabling AI-generated text to evade AIGC detectors with high success rates. On a multi-domain Chinese benchmark, the system achieves 94.6% evasion against a training detector and at least 99% evasion against three unseen detectors while maintaining 0.928 semantic similarity. The paper also presents RateAudit, a scheduling algorithm that can manipulate detection-rate verdicts to arbitrary values, directly undermining the reliability of score-based AIGC detection. These findings expose fundamental fragilities in deploying AIGC detectors for high-stakes applications such as academic integrity screening.
- Quality assurance
- AI policy
Research
How Frontier LLMs Adapt to Neurodivergence Context: A Measurement Framework for Surface vs. Structural Change in System-Prompted Responses
Ishan Gupta, Pavlo Buryi
arXiv · 2026-04-30
This paper introduces NDBench, a 576-output benchmark evaluating whether frontier large language models meaningfully adapt their responses when neurodivergence (ND) context is provided in system prompts. The study finds that fully instructed ND conditions produce significantly longer, more structured outputs with more headings and granular steps (p < 10^-8, Holm-corrected), but that persona assertion alone — without explicit instructions — fails to reduce potentially harmful tendencies such as masking reinforcement, which only drops 36–44% when instructions are explicit. Reliability analysis of LLM-based harm assessment shows only two of six evaluated dimensions (masking/reinforcement and validation quality) meet the pre-defined inter-judge agreement threshold (alpha >= 0.67), raising concerns about the robustness of automated auditing. The publicly released benchmark provides a reproducible framework for auditing how future LLMs handle ND-aware interactions, with implications for the quality and safety of AI-generated responses for neurodivergent users.
- Quality assurance
- AI policy
Research
Latent Adversarial Detection: Adaptive Probing of LLM Activations for Multi-Turn Attack Detection
Prashant Kulkarni
arXiv · 2026-04-30
This paper introduces 'adversarial restlessness,' a pattern in large language model internal activations that reveals multi-turn prompt injection attacks even when individual conversation turns appear harmless. The authors show that attack phases (trust-building, pivoting, escalation) produce measurable shifts in a model's residual stream, and five scalar features capturing the total activation path length lift detection accuracy from 76.2% to 93.8% on synthetic held-out data. The approach generalizes across four model families (24B–70B parameters), though probes are model-specific and do not transfer across architectures; combined training across three data sources achieves 89.4% detection at a 2.4% false positive rate. These findings characterize both the detection signal and the data requirements needed for practical deployment of activation-level defenses against covert LLM attacks.
- Quality assurance
- Enterprise
Research
To Build or Not to Build? Factors that Lead to Non-Development or Abandonment of AI Systems
Shreya Chappidi, Jatinder Singh
arXiv · 2026-04-30
This paper investigates why AI systems are abandoned or never built in the first place, a pre-deployment phase that responsible AI research has largely overlooked. Through a scoping review of academic, civil society, and grey literature, the authors develop a taxonomy of six factor categories driving AI non-development or abandonment: ethical concerns, stakeholder feedback, development lifecycle challenges, organizational dynamics, resource constraints, and legal/regulatory concerns. Empirical analysis using an AI incident database and a practitioner survey reveals that, contrary to academic emphasis on ethics, real-world abandonment decisions are driven by a diverse and often non-ethics-related mix of factors. The findings highlight gaps in responsible AI research and point to underexplored intervention opportunities at earlier stages of the development lifecycle.
- AI policy
- Enterprise
Research
MCPHunt: An Evaluation Framework for Cross-Boundary Data Propagation in Multi-Server MCP Agents
Haonan Li, Tianjun Sun, Yongqing Wang et al.
arXiv · 2026-04-30
MCPHunt is the first benchmark framework designed to measure how multi-server MCP (Model Context Protocol) agents inadvertently propagate credentials across trust boundaries during normal, non-adversarial tool composition workflows. The study evaluated 5 AI models across 3,615 benchmark traces spanning 147 tasks and found that policy-violating credential propagation rates range from 11.5% to 41.3% depending on the model, with propagation concentrated in browser-mediated data flows and varying up to 25x across different mechanism families. Prompt-based mitigations reduced policy-violating propagation by up to 97% while retaining 80.5% utility, but effectiveness depended on each model's instruction-following capability, suggesting prompt-level defenses alone may be insufficient. These findings matter for enterprise AI deployments and security policy, as they reveal a structural risk in agentic AI workflows that emerges from workflow topology rather than malicious model behavior.
- Enterprise
- AI policy
Research
How Generative AI Disrupts Search: An Empirical Study of Google Search, Gemini, and AI Overviews
Riley Grossman, Songjiang Liu, Michael K. Chen et al.
arXiv · 2026-04-30
This empirical study compares Google Search, AI Overviews (AIO), and Gemini Flash 2.5 across 11,500 real-user queries to measure how generative AI changes the information and sources users receive. Key findings include that AIOs appear for 51.5% of queries and are shown above organic results, that retrieved sources differ substantially between traditional and generative search (average Jaccard similarity below 0.2), and that generative engines favor Google-owned content while traditional search favors government and educational sites. The study also finds that sites blocking Google's AI crawler are less likely to appear in AIOs, and that AIOs are less consistent across repeated or slightly varied queries. The authors argue these findings matter for website visibility, generative engine optimization, and information quality, and call for revenue-sharing frameworks to sustain a fair ecosystem for publishers.
- Enterprise
- AI policy
Research
Test Before You Deploy: Governing Updates in the LLM Supply Chain
Mohd Sameen Chishti, Damilare Peter Oyinloye, Jingyue Li
arXiv · 2026-04-30
This paper addresses the problem of 'silent updates' in hosted Large Language Model services, where providers change model behavior without explicit version changes, causing unexpected regressions in functionality, formatting, or safety. The authors propose a deployment-side governance framework with three components: production contracts (rules defining allowed model behavior), risk-category-based testing suites, and compatibility gates that block updates failing safety or performance standards. Exploratory validation across multiple LLM versions shows that targeted, risk-focused testing can catch performance regressions that aggregate metrics overlook. The work frames LLM update management as a software supply chain governance problem and outlines open research challenges around building effective test suites, setting thresholds in non-deterministic systems, and detecting drift under limited provider transparency.
- Quality assurance
- Enterprise
Research
Consumer Attitudes Towards AI in Digital Health: A Mixed-Methods Survey in Australia
Wei Zhou, Rashina Hoda, Joycelyn Ling
arXiv · 2026-04-30
This Australian mixed-methods survey (N=275) examined consumer readiness, trust, acceptance, and risk perceptions of AI in digital healthcare settings. Participants showed moderate optimism and strong perceived usefulness but voiced significant concerns about accuracy, safety, and data use. In a scenario-based task, an AI-generated clinical consultation summary was strongly preferred over a clinician-written one for quality, empathy, and usefulness, yet participants could barely identify which was AI-produced. The findings suggest consumers evaluate healthcare AI primarily through communication quality and visible human oversight, pointing to the need for clinically supervised deployment frameworks rather than relying on technical performance alone.
- AI policy
- Quality assurance