News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
An Umbrella Review of Artificial Intelligence Applications in Mental Health Care
Jonathan Odame, Gabriel Obeng‐Gyamfi, Dayeon Heo et al.
Journal of Psychiatric and Mental Health Nursing · 2026-08-24
This umbrella review synthesizes 27 systematic reviews covering over 14 million participants to assess how AI is being used in mental health care. Key findings show AI improves early detection and risk stratification for conditions like depression, PTSD, and suicidal ideation, and that digital tools such as chatbots expand access and symptom monitoring. However, the review identifies significant challenges including data bias, limited external validity, and insufficient transparency, concluding that AI works best within hybrid, human-centred care models rather than as a replacement for clinical judgment. The authors argue responsible implementation requires rigorous validation, equity-focused design, and an ethics-of-care framework to ensure accountability.
- Workforce
- AI policy
Research
Clinicians' and people who use drugs' perspectives on artificial intelligence in addiction medicine
Patrick J. Kelly, E. Chen, Amelia Bailey et al.
Addiction · 2026-08-24
This qualitative study interviewed 12 addiction medicine clinicians and 25 people who use drugs (PWUD) to understand perceived benefits and concerns about AI in addiction medicine. Clinicians valued transparent, accurate AI tools but worried about stigma perpetuation and whether AI could keep pace with a volatile, regionally specific drug supply. PWUD expressed limited AI knowledge and distrust, yet were altruistically willing to share medical data for AI development while wanting control over identifiable information. The findings highlight that clinical AI adoption in addiction medicine depends critically on addressing privacy risks, drug-related stigma, and data governance concerns.
- AI policy
- Workforce
Research
Same Keys, New Doors: Foundational Skills and the Demands of a Changing World
Anita Sands, Cheryl Lavigne
ETS Research Report Series · 2026-08-24
This policy report analyzes how AI and emerging technologies are raising the bar on foundational skills—particularly literacy—needed to succeed in education and the labor market. Drawing on OECD PIAAC and U.S. NAEP data, the authors find a widening divergence in skill distributions: performance at the lower end has declined among both adults and students over the past decade, while the upper end has remained stable, and adults below key literacy thresholds are far less likely to demonstrate adaptive problem solving. The report argues that foundational skills must be treated as dynamic competencies and that policies targeting higher-order and AI-related skills must explicitly include foundational skill development. It closes with recommendations for policymakers, educators, funders, practitioners, and researchers to invest in systems that sustain opportunity across the life course.
- Workforce
- AI policy
Research
From accuracy to service: deciding what artificial intelligence outputs may do in veterinary diagnostic laboratories
Cleverson D. Souza
Journal of Veterinary Diagnostic Investigation · 2026-08-24
This commentary argues that reported model accuracy alone should not determine how AI outputs are used in veterinary diagnostic laboratories. The authors propose that laboratory staff document seven decisions in a standard operating procedure—covering intended use, reviewer roles, disclosure, input compatibility, override protocols, QC monitoring, and stop rules—before allowing AI to influence triage, reporting, or result release. Using a canine lymphoma cytology example, the paper shows that strong performance on one task (lymphoma vs. reactive hyperplasia) does not justify use on a harder task (B-cell vs. T-cell classification) without ancillary testing, and that interpretive outputs must always reach clients through a pathologist-reviewed and signed report.
- Quality assurance
- Certifications
Research
Something Wired This Way Comes: What Fracking Taught Planners That Data Center Communities Need to Learn
Austin Zwick
Journal of the American Planning Association · 2026-08-24
This viewpoint paper argues that AI data centers impose on local communities the same externalities as hydraulic fracturing—heavy water use, noise, road damage, thin permanent employment, and fiscal mismatches from tax abatements—and that municipal planners should respond with the targeted industrial regulations that fracking communities developed. Drawing on fieldwork with planners across five Marcellus Shale states, the author contends that classifying data centers as heavy industrial uses and applying proven regulatory tools is more durable than outright bans, which invite state preemption. The piece is directly relevant to local policy and planning practice for managing the community-level impacts of AI infrastructure.
- AI policy
- Workforce
Research
Is There a Constitutional Right to a Human Decision in Canada?
Robert Diab
Constitutional Forum / Forum constitutionnel · 2026-08-24
This legal article examines whether Canada's Charter of Rights provides constitutional protections requiring human decision-making in high-stakes administrative contexts. The author argues that Canada's existing policy frameworks—including the Treasury Board Directive on Automated Decision-Making—follow the EU's GDPR and AI Act in recognizing rights to human decisions and explanations, but that advances in language models challenge these frameworks since AI can now generate reasons comparable in quality to human ones. The article concludes that Charter rights to a human decision and explanation will, for the foreseeable future, constitutionally preclude significant reliance on language models in high-impact administrative cases regardless of their technical capabilities.
- AI policy
- Certifications
Research
When AI ads backfire: Deconstructing consumer backlash to GenAI advertising
Sigurd Birk Heimstad, Tarje Gaustad, Wondwesen Tafesse et al.
Journal of Retailing and Consumer Services · 2026-08-24
This study examines consumer backlash to AI-generated advertising by analyzing over 7,800 YouTube and Reddit comments on Coca-Cola's AI-remade 'Holidays Are Coming' campaign. Using LLM-assisted thematic analysis, the researchers identify a three-dimensional CPS (Content–Psychological–Socio-ethical) framework explaining how backlash emerges from concerns about visual quality and execution, feelings about authenticity and human connection, and socio-ethical attributions including corporate motives and labor displacement. The findings show that negative consumer reactions stem not only from perceived quality gaps versus human-made content, but also from what GenAI use signals about a brand's values. The CPS framework is offered as a diagnostic tool for practitioners to anticipate and manage backlash when deploying GenAI in advertising.
- Enterprise
- Workforce
Research
AI Governance Initiatives and Roadmaps by National Departments of Education: A Cross-National Policy and Maturity Analysis
Harun Serpil
IntechOpen eBooks · 2026-08-24
This chapter conducts a cross-national comparative analysis of AI-in-education policies from the European Union and thirteen countries, using documentary evidence from 2017 to 2025. It applies a four-part maturity framework—vision, capacity, ethics, and alignment (VCEA)—to evaluate how well governments govern AI in schools, universities, and lifelong learning. While most countries converge on goals like AI literacy, teacher development, and data security, they diverge significantly on questions of authority, teacher preparation quality, and whether ethical protections are enforced rather than merely stated. The study also assesses alignment with global reference points such as UNESCO's AI Competency Framework and OECD principles for trustworthy AI, revealing substantial variation in policy maturity across nations.
- AI policy
- Workforce
Research
Robust predictive analysis of international mobility among research talents
Muftawu Hussein, Emmanuel Ahene, Abdul Luckman Hassan et al.
Discover Data · 2026-08-24
This study uses bibliometric data from Web of Science and Scopus (2007–2020) combined with machine learning, association rule mining, and time series modeling to analyze and forecast international migration patterns among Ghanaian researchers. The analysis identifies discipline-specific emigration trends, finding that social sciences, agricultural and biological sciences, and health sciences are most vulnerable to long-term talent loss. The findings underscore the value of predictive analytics for informing data-driven policy on diaspora engagement, institutional strengthening, and talent retention in developing nations like Ghana.
- Workforce
- AI policy
Research
The legal and ethical dimensions of artificial intelligence in managing medical records in Oman
Abderrazak Mkadmi, Faten Hamad, Naifa Bait Bin Saleem
Discover Artificial Intelligence · 2026-08-24
This study surveyed 309 healthcare professionals in Oman to examine how legal and ethical factors shape trust in AI-assisted medical record systems. Using structural equation modeling, it found that perceived privacy and data protection most strongly predicted trust, while the existing legal framework and accountability mechanisms were negatively associated with trust—suggesting professionals view current liability and redress arrangements as unclear or insufficient. Ethical safeguards, however, were positively linked to trust. The authors recommend privacy-by-design practices, clearer accountability procedures, and visible ethical governance to support trustworthy AI deployment in healthcare records.
- AI policy
- Quality assurance
Research
Building Scalable Quality Assurance Frameworks for Reliable Training Data Across Frontier Artificial Intelligence Model Development
Chukwudi Okenyi
International Journal of Computer Applications Technology and Research · 2026-08-24
This paper presents a scalable quality assurance framework designed to improve the reliability of training data used in frontier AI model development. The framework integrates automated anomaly detection, deduplication, annotation-quality assessment, provenance tracking, schema validation, and human review into a unified pipeline that connects upstream data quality controls with downstream model evaluation. By combining quantitative data-quality indicators with risk-based sampling and governance mechanisms, the approach aims to detect and correct data deficiencies before they propagate into model behavior. The work is directly relevant to quality assurance in AI development, offering reproducible and governance-aware data engineering practices for large-scale AI pipelines.
- Quality assurance
Research
Lifecycle answerability for artificial intelligence-enabled medical device software: a regulatory science perspective
Ben Liu
Frontiers in Medicine · 2026-08-24
This paper introduces 'lifecycle answerability' as a regulatory science framework for AI-enabled medical device software, addressing a gap in existing lifecycle governance tools: while frameworks like quality management systems and predetermined change control plans exist, none clearly specify who must respond when field evidence shows an AI-supported clinical decision or validation claim no longer matches real-world performance. Using sepsis early-warning software as a case study—including an external validation that overturned a widely deployed model's performance claims and a prospective multi-site study linking alert response latency to mortality—the authors argue that such field signals are common, consequential, and inadequately addressed. The proposed construct defines standing, addressees, reason-giving, temporal triggers, and revision pathways to make credible field signals institutionally actionable across the software lifecycle. The framework has direct implications for how regulators, manufacturers, and deploying institutions are held accountable for AI medical device performance over time.
- Certifications
- AI policy
- Quality assurance
Research
Healthcare sciences lecturers' views on the use of artificial intelligence for patient diagnosis in Gauteng province, South Africa: qualitative study
Raikane J. Seretlo, Adelowatan Aminat Oluwatoyin
Frontiers in Education · 2026-08-24
This qualitative study explored how healthcare sciences lecturers at a higher education institution in Gauteng, South Africa perceive and accept AI tools when students use them for patient diagnosis. Through semi-structured interviews analyzed via thematic content analysis, thirteen themes emerged, revealing that while most lecturers see value in AI for enhancing education and diagnostic efficiency, many express concerns about accuracy, misdiagnosis, ethical oversight, loss of human connection, and the risk of producing future practitioners who lack critical thinking. The study concludes that inconsistencies in lecturer attitudes persist, and that institutional training and support are essential for responsible AI integration in health sciences education.
- Workforce
- Quality assurance
Research
Automated resume screening using machine learning: an ensemble-based approach for efficient candidate selection
Muhammad Hamid, Fahima Hajjej, Ala Saleh Alluhaidan et al.
Scientific Reports · 2026-08-24
This paper proposes an ensemble machine learning framework for automated resume screening that combines classifiers such as Logistic Regression, Decision Trees, Naive Bayes, and Random Forest via a soft-voting mechanism. On a benchmark dataset the ensemble achieved 98.44% accuracy, and when tested on a new expert-labeled dataset of 598 resumes under a Zero-Shot protocol it retained an F1-score of 0.99 with only a 1.10% performance drop. The authors also conducted an interpretability audit to assess algorithmic fairness and built a functional web prototype, demonstrating real-world viability. The work matters for recruitment operations by showing that automated, statistically validated resume screening can reduce manual effort at scale while maintaining merit-based selection.
- Workforce
- Enterprise
Research
How preservice teachers evaluated the accuracy of ChatGPT-generated solutions of curriculum-based mathematical word problems in a self-directed context
Sfiso C. Mahlaba
Frontiers in Education · 2026-08-24
This study examines how 30 preservice mathematics teachers evaluated the accuracy of ChatGPT-generated solutions to curriculum-based word problems in a naturalistic, self-directed setting. Using qualitative document analysis grounded in metacognition theory, the findings show that evaluations relied more on prior mathematical content knowledge than on real-time monitoring, and that participants tended to trust ChatGPT outputs more when they could not independently solve the problems. Results reveal a mix of effective and ineffective metacognitive regulation, with some preservice teachers correctly identifying and correcting AI errors while others exhibited confirmation bias or superficial reasoning. The study argues that teacher education programs must explicitly develop evaluative and metacognitive competencies to prepare future teachers for AI-mediated learning environments.
- Workforce
- Quality assurance
Research
Artificial intelligence and the reconfiguration of competency management systems in organizations
Maryann Osadebamwen Asemota
Discover Artificial Intelligence · 2026-08-24
This systematic review of 187 Scopus-indexed journal articles examines how AI is transforming competency management in organizations, finding that AI-related competency change extends well beyond technical skills to include hybrid, portfolio-based configurations encompassing managerial judgment, learning agility, governance capabilities, and psychological readiness. The study uses bibliometric mapping and qualitative content analysis to show that no single theoretical framework adequately explains these shifts, and proposes a multi-level, socio-technical theory synthesis to re-conceptualize competency management. The findings carry direct implications for designing adaptive competency architectures, aligning workforce development interventions with AI-enabled work systems, and embedding governance capabilities into HR strategies.
- Workforce
- Enterprise
Research
Kernel Token Contradiction: a Fast and Principled Approach for LLM Claim Uncertainty Quantification
Jérémie Dentan, Alexi Canesse, Mahammed El Sharkawy et al.
arXiv · 2026-08-23
This paper introduces Kernel Token Contradiction (KTC), a lightweight method for quantifying uncertainty at the claim level in Large Language Model outputs, aimed at assessing the factuality of individual claims. KTC represents candidate tokens as a positive semi-definite kernel combining the LLM's conditional distribution with a token contradiction score derived from Wikipedia frequency statistics, then uses Von Neumann entropy to measure uncertainty. Running on CPU only, KTC achieves over an 8.2x speedup versus state-of-the-art GPU-accelerated methods and over 65x versus comparable CPU-only methods, while matching or outperforming existing approaches in high-precision regimes across two benchmarks, four European languages, and 16 models. The combination of speed and accuracy makes real-time monitoring of LLM outputs practical in production environments.
- Quality assurance
- Enterprise
Research
Who Pays More for Safety? Measuring the Disparate Cost of Safety Alignment across Languages
Chanwoong Yoon, Jungsoo Park, Alan Ritter
arXiv · 2026-08-23
This paper investigates whether safety alignment in large language models imposes equal costs on users across different languages. The authors introduce a measurement framework called 'Safety Cost,' which quantifies utility loss caused specifically by safety alignment through pairwise comparisons between aligned and unaligned model versions. They find a systematic inequity: non-English users consistently bear a higher Safety Cost than English users, with some languages falling into a 'double-penalty zone' of both weaker safety protection and greater utility loss. The findings expose structural disparities in current safety alignment practices that disadvantage non-English speakers, raising important concerns about fairness in AI deployment globally.
- AI policy
- Quality assurance
Research
Claim-Level Confidence Calibration for Reliable Decision Making with Large Language Models
Toghrul Abbasli, Kentaroh Toyoda, Yuan Wang et al.
arXiv · 2026-08-23
This paper addresses a critical reliability problem with Large Language Models (LLMs): their tendency to hallucinate and express miscalibrated confidence, which undermines their use in high-stakes decision-making. The authors propose a claim-level confidence calibration framework that decomposes LLM responses into individual atomic claims and assigns each a calibrated confidence score using consistency across samples and self-verification — all without requiring access to model internals or fine-tuning. Testing across seven baselines on six recent models (including GPT-4, Llama-3.1, and DeepSeek-R1) on TriviaQA and TruthfulQA benchmarks, the approach reduces expected calibration error on factual questions and enables selective interventions like evidence retrieval or human review for low-confidence claims. The findings matter for enterprise and quality-assurance contexts where users need to know which specific pieces of AI-generated information to trust, verify, or reject.
- Enterprise
- Quality assurance
Research
Aligned Alone, Misaligned Together: Forecasting Adversarial Capture in LLM Agent Populations
Isotta Magistrali, Chen Shani
arXiv · 2026-08-23
This paper investigates how populations of AI language-model agents, each individually well-aligned, can collectively be steered toward adversarial outcomes when a committed minority is injected into the group. Using a security-triage scenario where agent populations decide whether to escalate or dismiss alerts, the authors show that two alerts a single agent treats as nearly identical can drive wildly divergent collective behavior, meaning individual-level audits fail to predict population-level outcomes. Critically, they demonstrate that adversarial capture can be forecast in advance by calibrating a response function from benign, adversary-free operation alone, and that capture is reversible once adversarial agents are removed. The findings challenge current AI safety evaluation practices, which focus on individual models, and argue for population-level assessments to understand real-world deployment risks.
- Quality assurance
- AI policy
Research
All four leading LLMs talk more than they listen to personality-verified synthetic help-seekers
Pablo A. Fonseca, Raquel Rodríguez-Carvajal, Rafael A. Calvo
arXiv · 2026-08-23
This study evaluates how four leading large language models respond to synthetic help-seekers in an emotionally acute crisis scenario—a caregiver learning of a relative's dementia diagnosis—using psychometrically specified personality profiles. The researchers developed a personality-aware evaluation framework in which blind auditors successfully recovered personality traits from dialogues alone (ICC(2,4) = 0.91), covering not just the Big Five but also coping style, coping self-efficacy, resilience, and reactance. All four models failed equally at emotion stabilization, sharing problematic behaviors: excessive verbosity, a talk-to-listen ratio above one, and premature problem-solving before the situation was adequately explored. These findings reveal systematic shortcomings in LLM mental health support capabilities that single-turn benchmarks would not detect, raising important questions about deploying these systems in crisis contexts.
- Quality assurance
- AI policy
Research
KONTOGRAPH: Verified Point-in-Time Feature Consistency and Amortised Explanation for Real-Time Anti-Money Laundering under a 200 ms Decision Budget
Ahmed Abolfadl
arXiv · 2026-08-23
KONTOGRAPH is an end-to-end anti-money laundering pipeline designed for SEPA Instant payments, which under EU Regulation 2024/886 must be settled in under ten seconds, forcing detection and explanation to fit within a 200 ms budget. Tested on over 1.5 million simulated payments, the system shows that a temporal graph network with per-node memory dramatically improves fraud detection (PR-AUC rising from 0.0053 to 0.1717) compared to a gradient-boosted tabular baseline. The paper also finds that converting the deployed model to ONNX format — a routine serving step — changed only a negligible mean score but altered 0.26% of decisions and inflated alert volume by 12%, demonstrating that format conversion must be treated as a model change and validated empirically. Additionally, property-based tests designed to detect point-in-time (data leakage) violations caught three bugs that code review had missed, each of which would have artificially inflated reported performance.
- Quality assurance
- AI policy
Research
Register Shifts Break LLM Safety: A Bengali Benchmark with Culturally Grounded Harms
Naymul Islam, Nusrat Jahan Lia, Shubhashis Roy Dipta et al.
arXiv · 2026-08-23
BanglaSafe is a new benchmark of 879 Bengali prompts across 17 culturally grounded harm categories, used to evaluate the safety of 18 frontier large language models. The study finds that over half of all model responses are unsafe or partially unsafe (53.6%), and that framing a harmful request in a formal newspaper-investigation register succeeds 17 percentage points more often than the same request phrased casually—without any adversarial engineering. Existing safety classifiers also struggle with Bengali content, with even frontier models failing on nearly half of cases. The work highlights a critical gap in LLM safety evaluation beyond English and shows that register shifts within a language are a major, underappreciated attack surface.
- Quality assurance
- AI policy
Research
Learning from the Test: Self-Referential Differential Testing for Deep RL Agents
Junda He, Jieke Shi, Zhou Yang et al.
arXiv · 2026-08-23
This paper presents Delta, a two-phase differential testing framework for Deep Reinforcement Learning (DRL) agents that detects both safety-critical failures and optimality bugs — an area largely overlooked by existing testing approaches. In the first phase, the agent under test is evaluated for catastrophic failures while its decision-making data is collected; in the second phase, that data trains a 'challenger' agent via offline RL algorithms (BC, BCQ, CQL), and differential comparison identifies cases where the challenger outperforms the original agent, signaling optimality issues. Tested across five environments, Delta uncovered an average of 2,518 optimality issues per environment, outperforming baseline methods by 50.2%, with BCQ-trained challengers proving most effective. This matters for quality assurance of AI systems, as undetected suboptimal policies can reduce efficiency, erode user trust, and cause economic losses in real-world deployments.
- Quality assurance
Research
AUDITA: certified auditing and causal attribution of adverse outcomes in autonomous multi-agent systems
Zhixu Du, Yiran Chen
arXiv (Cornell University) · 2026-08-23
AUDITA is a certified audit framework for autonomous multi-agent systems (e.g., AI-run factories and warehouses) that pairs tamper-evident inter-agent command logs with a causal-attribution engine to formally assign responsibility when joint AI decisions cause harm. The system provides provable guarantees: a rule-following agent cannot be made to appear guilty, blame-shifting attempts are detected and graded, and the limits of evidence-based certification are formally established. In empirical tests on live language-model pipelines it reduces responsibility attribution error roughly threefold compared to standard baselines, and remains robust to log forgery. This matters because it offers a rigorous, legally-relevant mechanism for dividing liability among vendors, operators, insurers, and regulators in increasingly autonomous industrial settings.
- Certifications
- AI policy