News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
Can model attribution bridge AI's accountability gap in safety-critical domains?
Marc Juárez
Philosophical Transactions of the Royal Society A Mathematical Physical and Engineering Sciences · 2026-07-16
This paper examines 'model attribution'—a property intended to identify which specific ML model a user is interacting with when models are deployed as remote services—as a mechanism for closing AI accountability gaps in safety-critical domains. The authors review recent technical approaches to model attribution and find that limitations including computational costs, latency overheads, and insufficient guarantees make current solutions inadequate for safety-critical applications. They conclude that while model attribution is a meaningful step toward accountability, the field has not yet produced tools that meet the rigorous requirements of high-stakes settings. The work is relevant to ongoing efforts to establish robust governance and oversight mechanisms for AI in safety-critical systems.
- AI policy
- Quality assurance
Research
A STATISTICAL ANALYSIS OF THE SOCIAL IMPACT OF AI ON JOB DISPLACEMENT & REPLACEMENT IN SOUTH ASIA & AFRICA
Dr NR Jagannath
International Journal of Creative and Open Research in Engineering and Management · 2026-07-16
This paper statistically analyzes how AI-driven job displacement and replacement are unfolding across South Asian and African labour markets, drawing on data from the World Bank, IMF, ILO, WEF, and industry sources. It finds that formal, export-oriented sectors—IT-BPO, ready-made garments, customer support, and banking—face the highest automation exposure, with notable job losses already observed in Bangladesh and high susceptibility projected for India's IT-BPO workforce by 2030. Distributional impacts are pronounced: women in export manufacturing and routine services, youth, and lower-educated workers bear disproportionate displacement risk, while skills gaps and digital-infrastructure deficits limit access to emerging AI-adjacent roles. The paper concludes that net job creation is possible but conditional on reskilling investment, digital education, labour-market adjustment supports, and robust disaggregated monitoring systems.
- Workforce
- AI policy
Research
The Governance Reconstruction of the AI Era: From "Entering Society" to "Remaining in Society" — Establishing Occupational Transition and Retraining as a Core Public Service of Government
Leyi Zhang
Open MIND · 2026-07-16
This policy article argues that AI-driven occupational disruption—accelerated by generative AI and embodied robotics—is compressing job life cycles faster than prior waves of automation, undermining the historical assumption that technological unemployment is merely transitional. Drawing on evidence from the WEF Future of Jobs Report 2025, ILO 2025 assessments, and IBM reskilling projections, the author contends that occupational retraining must be elevated to a core state public service comparable to compulsory education. The proposed framework grounds government obligation in Rawlsian, Dworkinian, and capability-approach theories, identifies market failures justifying intervention, and recommends a funding mechanism financed by AI-driven fiscal dividends rather than automation taxes.
- Workforce
- AI policy
Research
Does the Responsible AI Guidance Adequately Protect Māori in Predictive Workplace Safety Regulation?
Nicola Knobel
New Zealand journal of health and safety practice. · 2026-07-16
This paper evaluates whether New Zealand's Responsible AI Guidance for the Public Service adequately protects Māori in predictive workplace safety regulation. Applying Critical Tiriti Analysis, responsive regulation theory, and post-colonial legal analysis, the paper finds the Guidance scores either Silent or Poor across all five Treaty of Waitangi criteria, lacking enforceable protections, Māori governance structures, and procedural fairness mechanisms. Drawing on case studies including the Woolworths warehouse strike and WorkSafe New Zealand's use of AI-enabled tools, the authors argue that reliance on soft-law principles is insufficient. The paper calls for reforms embedding co-governance, Māori Data Sovereignty, and legally mandated safeguards into public sector AI infrastructure.
- AI policy
- Workforce
Research
A Meta-synthesis of nursing students’ experiences with generative artificial intelligence-assisted learning
Siyi Du, Sha Wang, Fengxia Yan et al.
Frontiers in Medicine · 2026-07-16
This meta-synthesis of 18 qualitative studies examines nursing students' real-world experiences with generative AI-assisted learning, identifying 49 themes organized into four integrated findings: a dual experience of empowerment and challenges, internal conflict between technology and nursing humanism, differentiated user experiences amid practical constraints, and a general demand for supporting systems and educational reform. The findings indicate that nursing educators need to strengthen students' AI literacy, critical thinking, and ethical awareness, while curriculum designers should embed AI competencies into nursing programs without sacrificing humanistic care values. The study also calls on policymakers to establish clear governance frameworks and educational guidelines for responsible AI use in nursing education.
- Workforce
- AI policy
Research
Regulating Reality: Exploring Synthetic Media Through Multistakeholder AI Governance
Claire R. Leibowicz
Policy & Internet · 2026-07-16
This paper examines how stakeholders from civil society, industry, media, and policy sectors conceptualize and implement governance of AI-generated (synthetic) media. Drawing on 23 semi-structured interviews and real-world cases, the study finds that temporal perspectives and trust—both among stakeholders and between audiences and interventions—are central to effective governance, and that technical measures like AI labels have notable limitations. The findings are intended to inform both scholarly understanding of multistakeholder AI governance and the practical design of synthetic media policy.
- AI policy
Research
LLMs in Medical Education for Autism Caregivers: A Comparative Evaluation of Accuracy, Readability, Actionability, and Neurodiversity-Affirming Language
Shahid Akhtar Akhund, Asma Alsaleh, Naheed Haroon Kazi et al.
Healthcare · 2026-07-16
This study systematically compared three large language models—Google Gemini 1.5 Pro, ChatGPT (GPT-4o), and DeepSeek-V3—as health information tools for caregivers of children with autism spectrum disorder in Saudi Arabia. Across 24 clinically validated questions rated by expert raters, Gemini was most scientifically accurate, ChatGPT most readable, and DeepSeek most actionable, yet all three models fell below established AHRQ benchmarks for actionability and exceeded recommended readability grade levels for patient education. Accuracy gaps clustered around regional epidemiological, genetic risk, and financial queries, and no model demonstrated meaningfully neurodiversity-affirming language. The findings indicate that LLMs in their current form require substantial plain-language adaptation and cultural customization before they can reliably serve as caregiver education tools, and clinicians should actively guide families in interpreting LLM outputs.
- Workforce
- Quality assurance
Research
Conformity assessment procedures under the AI Act
Mattis Jacobs, Susanne Kuch, Dominic Deuber et al.
Technology and Regulation · 2026-07-16
This article critically analyzes the conformity assessment procedures mandated by the EU AI Act for high-risk AI systems. It argues that the Act's reliance on conformity assessment procedures borrowed from other EU harmonization legislation produces a risk classification that inadequately accounts for AI-specific risks and raises doubts about the suitability of those procedures for AI systems. The analysis also examines how procedure assignments affect the requirements imposed on conformity assessment bodies, identifying systemic regulatory blind spots in the AI Act's framework.
- Certifications
- AI policy
Research
Helping People Choose Careers in the Age of AI
Jennifer Steele, Isabella Cruz
arXiv (Cornell University) · 2026-07-16
This paper helps individuals make informed career decisions by analyzing how much different occupations are exposed to AI-driven task automation. The authors compare six existing projection models and develop a new empirical model using 2025 query data from Anthropic and OpenAI, finding significant variation across models but a general trend since 2020 where higher AI exposure correlates with higher salaries and job complexity. By averaging five models, they identify career tradeoffs across job categories, finding healthcare practice offers the best balance of higher pay and lower AI exposure. Jobs where AI acts as a complement rather than a substitute for human labor tend to be modestly better-paid, with implications depending on how fields adopt AI usage norms.
- Workforce
- AI policy
Research
The Most Exposed Sector Meets the Shock: AI Exposure and Firm-Level Labor Outcomes in Indian IT Services
Mihir Khanna
arXiv · 2026-07-16
This paper examines how AI exposure affected hiring and productivity at Indian IT-BPM firms around two technology shocks: the 2015-16 cloud/deep-learning wave and the 2022 LLM shock. Using a panel of 13 publicly listed IT-services firms and occupational AI-exposure scores, the author finds that after 2022, more-exposed firms significantly slowed net hiring (annual net hiring fell from 8.3% to 3.1% for high-exposure firms) while raising labor productivity, consistent with augmentation rather than outright displacement. The 2016 wave showed no comparable effect, and no significant hiring slowdown appeared for lower-AI-exposure engineering-R&D firms. The study provides the first firm-level evidence of LLM adoption effects in the global sector most concentrated in AI-exposed roles.
- Workforce
- Enterprise
Research
The Governance Reconstruction of the AI Era: From "Entering Society" to "Remaining in Society" — Establishing Occupational Transition and Retraining as a Core Public Service of Government
Leyi Zhang
Knowledge Commons (Lakehead University) · 2026-07-16
This policy paper argues that AI-driven occupational displacement—accelerated by generative AI and embodied robotics—is occurring at a scale and speed that invalidates the historical assumption that automation merely transitions workers rather than permanently displacing them. Drawing on empirical sources including the WEF Future of Jobs Report 2025, ILO assessments, and IBM reskilling projections, the author contends that governments should elevate occupational retraining from a marginal labor-market tool to a core public service comparable to compulsory education or public security. The proposed framework grounds this obligation in Rawlsian, Dworkinian, and Capability-approach ethics, identifies specific market failures justifying intervention, and recommends a funding mechanism financed by AI-driven fiscal dividends rather than automation taxes. The paper concludes that the state's historic role of helping citizens 'enter society' must be extended to helping them 'remain in society' through lifelong transition support.
- Workforce
- AI policy
Research
Context Contamination in LLM Analysis of Network Security Logs: Poison with Passive Prompt Injection and Mitigation Evaluation
Rabimba Karanjai, Ye Lu, Hemanth Hegadehalli Madhavarao et al.
arXiv (Cornell University) · 2026-07-16
This paper demonstrates that LLMs deployed in Security Operations Centers for log analysis are vulnerable to 'passive prompt injection,' where adversaries embed malicious instructions in network log fields that are later executed when analysts query the model. The authors introduce LogInject, a benchmark framework with nearly 13,000 log entries, and show attack success rates up to 88.2% across objectives like activity concealment, false positive generation, and output hijacking. A novel 'Context Stitching' technique fragments payloads across multiple log entries to evade filters while still achieving a 76.4% success rate. Layered defenses reduce attacks by 90.4% but leave an 8.4% residual vulnerability, underscoring the need for defense-in-depth and continued human oversight in security-critical LLM deployments.
- Quality assurance
- AI policy
Research
A study on the influencing factors of perceived artificial intelligence substitution risk
Xianghui Xing, Hongwei Dai, Yiwei Zhou et al.
Acta Psychologica · 2026-07-16
Using data from China's 2023 Social Survey, this study empirically identifies the key factors shaping workers' perceived risk of being replaced by AI. Results show that perceived unemployment risk, number of children, and unemployment insurance positively correlate with perceived AI substitution risk, while female gender, older age, higher job satisfaction, greater job skill level, and higher perceived socioeconomic status are negatively associated. Mediation analysis reveals that job satisfaction and perceived socioeconomic status reduce substitution fears partly by lowering perceived unemployment risk, and heterogeneity analysis shows these effects differ significantly across regions, urban-rural contexts, and work patterns. The findings offer empirical grounding for policies aimed at reducing AI-driven employment anxiety and improving workforce training systems.
- Workforce
- AI policy
Research
The Enterprise AI Governance Buyer's Guide
FERZ Inc., Edward Meyman
Zenodo (CERN European Organization for Nuclear Research) · 2026-07-16
This document presents Version 3.4 of the Enterprise AI Governance Buyer's Guide alongside a companion Procurement Fast Path, offering a vendor-neutral framework for evaluating AI governance claims in regulated enterprise environments. The framework formalizes a distinction between probabilistic and deterministic governance, centers on five canonical tests (Stop, Ownership, Replay, Escalation, and Provenance), and incorporates an Authorization Boundary Integrity Model to locate where a vendor's controls hold or fail. It maps governance requirements to major regulatory regimes including the EU AI Act, GDPR, HIPAA, DFARS, and NIST AI RMF, and is designed to help procurement teams, auditors, and risk officers obtain independently verifiable, tamper-evident evidence that AI decisions were authorized before execution. The guide matters because it provides structured, evidence-based due-diligence tooling for organizations that must demonstrate defensible AI governance in high-stakes sectors such as healthcare, financial services, defense, and government.
- Enterprise
- Quality assurance
- Certifications
- AI policy
Research
Artificial Intelligence and Machine Learning in Auditing
Zamin Abdullah Shah, Babita Jha
arXiv · 2026-07-16
This chapter examines how AI and machine learning are transforming the auditing profession by shifting from traditional sample-based testing to real-time, full-population analysis, which the authors argue significantly improves fraud detection and anomaly identification. Shah and Jha also explore AI's role in ESG assurance, helping auditors verify non-financial sustainability metrics and strengthen corporate governance. The authors acknowledge a dual effect of AI adoption—gains in audit effectiveness alongside challenges around algorithmic transparency, data privacy, and ethics. A roadmap is proposed for aligning AI-driven auditing with regulatory frameworks such as SOX and NIST while preserving human judgment.
- Enterprise
- Quality assurance
- Certifications
- AI policy
Research
Beyond GDP: Firm Size, Sector, and the Structural Determinants of Enterprise AI Adoption a Cross-Country Empirical Study
Shreehari Hulyalkar, Jayashree K
International Scientific Journal of Engineering and Management · 2026-07-16
This cross-country empirical study finds no statistically significant correlation between GDP per capita and enterprise AI adoption rates across 18 countries (r = -0.10, p = 0.68), challenging the assumption that national wealth drives AI uptake. Instead, firm size (a 34-38 percentage-point adoption gap between large and small firms), industry sector (an 'hourglass' distribution favoring ICT and professional services over construction and accommodation), and survey methodology are the dominant predictors. Notably, a 50.3 percentage-point difference exists between official government statistics and executive surveys (e.g., IBM, McKinsey, Stanford HAI), suggesting that how AI 'use' is defined matters more than national income. These findings have significant implications for policymakers and researchers measuring AI diffusion across enterprises.
- Enterprise
- AI policy
- Workforce
Research
How Artificial Intelligence Influences the Green Transformation of Manufacturing Enterprises: A Dynamic Capabilities Perspective
Yongjie Wu, M M Jia, Shuangying Liu
Sustainability · 2026-07-16
This study examines how AI drives green transformation in Chinese manufacturing firms using panel data from 2012–2023, finding that AI significantly accelerates green transformation through three organizational dynamic capabilities: absorptive, innovative, and adaptive capabilities. Mechanism analysis confirms these capabilities serve as mediating pathways between AI adoption and green outcomes, while heterogeneity analysis shows effects are stronger in non-state-owned, large-scale, and non-high-tech firms. The findings provide micro-level evidence that institutional and resource contexts shape how AI translates into environmental progress, offering practical guidance for advancing sustainable manufacturing in China's digital era.
- Enterprise
- Workforce
- AI policy
Research
Artificial Intelligence and Corporate Social Irresponsibility
Tobias Steindl
University of Regensburg Publication Server (University of Regensburg) · 2026-07-16
This study finds that greater AI adoption intensity among US listed firms is associated with higher levels of corporate social irresponsibility (CSI), including both business ethics controversies and data privacy controversies. The results suggest AI can amplify human-induced ethical failures as well as introduce algorithm-driven privacy harms. However, the negative association reverses in firms that have a Chief Sustainability Officer or Chief Digital Officer, indicating that domain-specific executive oversight can mitigate AI-related irresponsibility. The findings have practical implications for how firms govern AI adoption within digital transformation strategies.
- Enterprise
- AI policy
Research
The Enterprise AI Governance Buyer's Guide
FERZ Inc., Edward Meyman
Zenodo (CERN European Organization for Nuclear Research) · 2026-07-16
This document presents a vendor-neutral evaluation framework for procuring AI governance solutions in regulated enterprise environments, distinguishing between three governance problems—visibility, alignment, and authorization—and formalizing the difference between probabilistic ('likely compliant') and deterministic ('provably compliant') governance. It introduces the Five Tests Standard (5TS v1.2.0), which requires AI governance systems to demonstrate Stop, Ownership, Replay, Escalation, and Provenance capabilities, with a three-verdict enforcement model (ALLOW, DENY, ABSTAIN) that fails closed when governance conditions are unmet. The framework also incorporates an Authorization Boundary Integrity Model and maps requirements to major regulatory regimes including the EU AI Act, GDPR, HIPAA, DFARS, and NIST AI RMF. It is intended for procurement teams, risk officers, auditors, and regulators seeking independently verifiable evidence that AI governance controls functioned correctly at the moment a specific decision was made.
- Enterprise
- Quality assurance
- Certifications
- AI policy
Research
DS@GT ARC at LongEval: Citation Integrity and Factual Grounding in Scientific QA
Brandon Michaels, Brendon Johnson
arXiv · 2026-07-15
This paper investigates a gap between standard natural language evaluation metrics and citation integrity in Retrieval-Augmented Generation (RAG) question-answering systems for scientific literature. The authors compare a corrective pipeline combining Corrective RAG (CRAG) and CiteFix against baseline and frontier model RAG systems, finding that frontier models scored well on answer relevance and fluency but often ignored retrieved documents when generating answers. Their corrective pipeline, which filters retrieved chunks before generation and enforces strict entailment of generated claims to cited sources after generation, marginally improved citation faithfulness and answer grounding. The paper argues that trustworthy RAG evaluation requires metrics specifically rewarding strict grounding of answers in cited material, with implications for how scientific QA system quality is assessed.
- Quality assurance
- Enterprise
Research
Instrument Effects in Language-Model Honesty Evaluation: An Auditable Single-System Demonstration
Justin Bronder
arXiv · 2026-07-15
This paper investigates whether evaluation instruments themselves—rather than the language models being tested—are responsible for measured differences in model behavior, focusing on honesty evaluations. Using a text-adventure game where the game engine (not the model) controls ground truth, the authors show that changing instrument design choices like grammar size, success-criterion disclosure, and budget rendering substantially shifted verdict distributions without changing the model itself. For example, expanding a two-verdict grammar to three options reduced strong claims from 38/40 to 7/40, and disclosing the success criterion eliminated false verdicts entirely. The authors propose a four-check integrity protocol for evaluation instruments, highlighting that single evaluation runs may report sampling noise as stable model dispositions.
- Quality assurance
- Certifications
Research
CatalogAgent: A Supervisor-mediated Self-Learning System Enabling Context Engineering for GenAI Models
Zhu Cheng, Zhenming Wang, Yu et al.
arXiv · 2026-07-15
CatalogAgent is an agentic AI system designed to automatically fill missing structured attributes (e.g., material, color, shape) in e-commerce product catalogs. It uses a Supervisor Agent to resolve conflicts between an LLM-based Generator and Evaluator — both internally and from external seller feedback — then stores and aggregates these resolutions into reusable learnings. By injecting those learnings back into the worker models' contexts (context engineering), the system improves Generator performance by 15.24% and Evaluator performance by 13.98% without human intervention. This demonstrates a self-improving AI pipeline that can enhance catalog data quality at scale with reduced manual effort.
- Enterprise
- Quality assurance
Research
Chat2Scenic: An Iterative RAG-Based Framework for Scenario Generation in Autonomous Driving
Yuan Gao, Wenting Miao, Mattia Piccinini et al.
arXiv · 2026-07-15
Chat2Scenic presents an iterative retrieval-augmented generation (RAG) framework for automatically producing executable scenario scripts in Domain Specific Language (DSL) for autonomous driving simulation testing. The system integrates a chatbot interface for interactive scenario refinement with RAG grounded in regulatory knowledge from sources such as NHTSA and United Nations Vehicle Regulations. Evaluated against state-of-the-art LLMs on a new benchmark of 123 scenarios, Chat2Scenic achieves a 76.42% Compilation Success Rate and 58.17% Framework Accuracy, substantially outperforming existing retrieval-assemble (30.08% CSR, 11.03% FA) and retrieval full-script generation (16.26% CSR, 10.86% FA) approaches. This matters for quality assurance and certification of autonomous driving systems by enabling more reliable, regulation-compliant test scenario generation at scale.
- Quality assurance
- Certifications
- AI policy
Research
MamaBench: Benchmarking LLM Robustness in Maternal and Child Health Diagnosis through Counterfactual Clinical Perturbation
Thanni Adewuyi, Anuoluwa Sotome, Samuel Okoko et al.
arXiv · 2026-07-15
MamaBench introduces the first counterfactual benchmark designed to test whether large language models can reliably distinguish between clinically similar but intervention-different presentations in maternal and pediatric health. The benchmark comprises 434 expert-authored clinical narratives in 217 paired cases spanning 371 pathologies, evaluated using the Bias Trap Rate (BTR), which measures how often a model fails a counterfactual case after succeeding on its base case. Across eight configurations of four frontier LLMs, base accuracy overstates robust accuracy by 16–28 percentage points, revealing a significant gap in clinical reliability. The authors also propose Evidence-Anchored RAG (EA-RAG), which achieves 65.0% robust accuracy and a 20.3% BTR on Claude Sonnet 4.6, though a residual 20% BTR confirms counterfactual robustness in clinical AI remains an unsolved problem.
- Quality assurance
- AI policy
- Certifications
Research
Unsafe at any AUC: Unlearned Lessons from Sociotechnical Disasters for Responsible AI
Joshua A. Kroll, Andrew Smart, R. Stuart Geiger et al.
arXiv · 2026-07-15
This paper draws on decades of research into major human-made catastrophes—such as Chernobyl, the Challenger disaster, and Bhopal—to argue that AI development is repeating historically unlearned lessons about sociotechnical risk. The authors contend that AI risks are not purely technical but are deeply shaped by social, organizational, political, and economic structures, mirroring how prior disasters unfolded despite known hazards. They identify three key areas where AI development falls short: risk perception and communication at the organizational level, traceability of requirements and responsibilities, and holistic safety approaches that treat social and organizational dynamics as first-order engineering concerns. The paper matters because it reframes AI safety and responsibility as a systems-level challenge, urging practitioners and policymakers to move beyond component-level reliability metrics toward comprehensive sociotechnical analysis.
- AI policy
- Quality assurance
- Enterprise