News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
Governing generative AI in higher education: a global Delphi study on policy and practice
Helen Crompton, Diane Burke, Christine Nickel et al.
International Journal of Educational Technology in Higher Education · 2026-05-22
This global Delphi study gathered expert perspectives from 22 countries across six continents to develop a consensus-driven framework for governing generative AI in higher education. The study produced an eight-part policy framework covering academic integrity, ethical and responsible use, privacy, equitable access, AI literacy, integration strategy, human oversight, and institutional support, as well as a six-part mechanism for keeping policies current through dedicated committees, regular reviews, and stakeholder communication. The research provides faculty, administrators, policymakers, and funders with a structured, adaptable blueprint for integrating generative AI into higher education responsibly. Its findings are directly relevant to institutions and governments seeking evidence-based guidance on AI governance in educational settings.
- AI policy
- Workforce
- Certifications
Research
From Policy Governance to Runtime Sovereignty: A Layered Architecture for Accountable AI Systems
Felix Montanez, Amelie Kingsbury Barry
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-22
This working paper argues that current AI governance frameworks—relying on principles, risk registers, audits, and policy documents—are insufficient for AI systems operating across local devices, enterprise networks, and regulatory environments. The authors propose 'runtime sovereignty' and introduce the Sovereignty Structure Framework, a layered architecture spanning four operational layers (Local, LAN/Site, Enterprise, and Regulatory Discovery) and seven governance dimensions (Governance, Sovereignty, Agency, Liability, Negotiation, Memory, and Reflection). The framework requires that authority, evidence, consent, and accountability be verified before and after every AI action, not merely described in policy. The paper contends that AI governance becomes credible only when it can answer runtime questions about who has authority, under what conditions, and with what auditable trail and capacity for correction.
- AI policy
- Enterprise
- Quality assurance
- Certifications
Research
Combating <scp>ESG</scp> Greenwashing Through <scp>AI</scp> Models: Evidence From Disaggregated <scp>AI</scp> Technologies, Mechanisms, and Thresholds
Brahim Bergougui, Hamid Ghazi H Sulimany, Abdulrahman Atllah Alharbi
Corporate Social Responsibility and Environmental Management · 2026-05-22
This study examines how AI language models affect corporate ESG greenwashing behavior using panel data from Chinese listed firms (2012–2022). The findings show that AI adoption—particularly machine learning and planning-decision systems—reduces greenwashing through two key channels: workforce skill restructuring and firm performance enhancement. The anti-greenwashing effect is strongest among heavily polluting enterprises, non-state-owned firms, and technology-intensive sectors. The authors recommend coordinated policy interventions including promoting AI deployment in environmentally critical domains and developing regulatory guidelines that leverage AI's analytical capabilities to improve ESG disclosure transparency.
- AI policy
- Enterprise
- Workforce
- Quality assurance
Research
AI Assurance: A Comprehensive Testing Strategy for Enterprise AI Systems
Chitra Badagi, Divye Singh, Animesh Sen et al.
arXiv (Cornell University) · 2026-05-22
This paper proposes a comprehensive assurance strategy for enterprise AI systems built on large language models, retrieval pipelines, and autonomous agents, arguing that traditional software QA methods are inadequate for probabilistic and emergent AI behavior. The authors introduce a structured AI Failure Taxonomy and a revised five-layer AI Assurance Pyramid, emphasizing continuous risk reduction over strict correctness verification. The framework covers evaluation-driven development, RAG system testing, model lifecycle management, and governance, aiming to give engineering leaders an operationally deployable strategy for managing AI-specific organizational risks.
- Enterprise
- Quality assurance
- Certifications
- AI policy
Research
A three-way game analysis of technology-to-goodness-driven artificial intelligence marketization and application strategy
Mengfen Luo, Tianyu Li, Jiarui Hou
International Review of Economics & Finance · 2026-05-22
This paper develops a three-party evolutionary game model involving AI firms, workers with redundant skills, and local governments to analyze how regulatory intensity shapes strategic behavior during AI commercialization. The study finds that precise government regulation can incentivize firms to deepen local market engagement and shift workers from resistance toward skill enhancement, while identifying a critical 'high-cost deadlock' threshold where firms exit markets when resistance costs exceed net operational benefits. The findings provide a theoretical basis for designing targeted subsidy mechanisms and differentiated skills training policies to address technological unemployment caused by AI-driven labor displacement. This work operationalizes the 'technology for good' concept into testable equilibrium conditions relevant to workforce policy design.
- Workforce
- AI policy
- Enterprise
Research
Determinants of Artificial Intelligence Adoption in Public Sector Human Resource Management: Empirical Evidence from Kazakhstan
Aliya Daueshova, Azamat Zhanseitov, Aigerim Amirova et al.
ADMINISTRATIE SI MANAGEMENT PUBLIC · 2026-05-22
This large-scale empirical study of 12,562 civil servants in Kazakhstan examines what drives or hinders AI adoption in public sector human resource management. Using OLS regression, logistic regression, and path analysis, the study finds that internal HR quality factors more strongly predict perceived HR effectiveness than external ones, and that managerial position is the strongest predictor of active AI adoption while longer tenure reduces it. Access to modern digital tools positively moderates AI uptake. The findings offer evidence-based policy recommendations for accelerating human-centred AI integration in government HR systems.
- Workforce
- AI policy
- Enterprise
Research
From Policy Governance to Runtime Sovereignty: A Layered Architecture for Accountable AI Systems
Felix Montanez, Amelie Kingsbury Barry
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-22
This working paper argues that current AI governance—relying on principles, policy documents, risk registers, and audits—is insufficient for AI systems operating across local devices, enterprise networks, and regulatory environments. The authors propose 'runtime sovereignty' and introduce the Sovereignty Structure Framework: a four-layer architecture (Local, LAN/Site, Enterprise, Regulatory Discovery) with seven governance dimensions (Governance, Sovereignty, Agency, Liability, Negotiation, Memory, Reflection) that checks authority and preserves evidence before and after every AI action. The framework is illustrated through OBEXGATE, an active runtime governance system featuring pre-execution gates, trust receipts, evidence bundles, and witness chains. The paper contends that AI governance is only operationally credible when it can answer in real time who has authority to act, under what conditions, and with what accountable and correctable trail.
- AI policy
- Enterprise
- Quality assurance
- Certifications
Research
Security of LLM-generated Code: A Comparative Analysis
Srivathsan G Morkonda, Mahmoud Selim, Hala Assal
arXiv · 2026-05-21
This paper empirically evaluates the security of code generated by seven popular large language models (LLMs), simulating real developer behavior when using AI coding tools. The findings show that all seven LLMs produce code containing security vulnerabilities, with the majority rated as critical or high severity. Given that LLM-generated code is already deployed in production at major tech companies, these results highlight meaningful risks to software security at scale.
- Quality assurance
- Enterprise
Research
Simulating Hate Speech Cascades with Multi-LLM Agents: Empirical Grounding, Modeling Fidelity, and Intervention Strategies
Fan Huang
arXiv · 2026-05-21
This paper investigates whether multi-agent large language model (LLM) systems can more faithfully simulate the spread of hate speech on online platforms compared to classical cascade models. Analyzing three hateful Bluesky cascades, the researchers found that 97.4–99.7% of reposters took a hostile stance, hateful cascades showed a star-like topology (most reposts from the root), and toxicity-engagement homophily was higher on the diffusion tree than on the follower graph. A multi-LLM-agent simulator successfully reproduced the stance monoculture and toxicity patterns, with agent heterogeneity identified as the leading factor in simulation fidelity. Applying an amplifier-targeting intervention on dense networks yielded a 7.5–12.9% reduction in hateful content spread with only 5.7% benign collateral, suggesting actionable moderation strategies grounded in realistic cascade modeling.
- AI policy
- Quality assurance
Research
A measurement substrate for agentic Kubernetes operations: Methodology and a case study in retrieval-compounding falsification
Joshua Odmark, Gideon Rubin, Deon van der Vyver
arXiv · 2026-05-21
This paper addresses a reproducibility and measurement gap in AI-driven autonomous Kubernetes operations: existing published results lack controlled baselines, pre-registered decision matrices, and adequate sample sizes, making empirical claims largely unfalsifiable. The authors introduce 'agent-breakage,' a closed-loop framework that injects faults into a Kubernetes cluster, scores an agent's responses on four axes against ground truth, and accumulates labeled outcome tuples to distinguish framework errors from reasoning errors. As a case study testing whether retrieval over past postmortems improves agent capability, the framework caught three confounds — a pgvector index bug, a +19% selection-bias artifact, and small-sample effect overestimates of roughly 3x — each of which would have produced incorrect published claims. The retrieval intervention itself was only partially supported: 1 of 3 scenarios reached significance at p<0.05, with a pooled effect of +3.9 percentage points that was not significant at n=60, and a 360-run sweep found that mechanistic alignment of near-neighbors mattered more than raw corpus size.
- Quality assurance
- Enterprise
Research
Eroding Trust in Real Speech: A Large-Scale Study of Human Audio Deepfake Perception
Nicolas M. Müller, Wei Herng Choong
arXiv · 2026-05-21
This large-scale listening study—the largest of its kind—collected 35,532 judgments from 1,768 participants across 138 text-to-speech and voice conversion systems to measure how audio deepfakes affect human trust in real speech. The central finding is a 'skepticism shift': human accuracy on fake samples barely changed from a 2021 baseline (72.9% to 71.2%), but accuracy on identifying real speech dropped significantly from 72.7% to 64.1%, meaning people increasingly distrust authentic audio rather than getting better at spotting fakes. Commercial and autoregressive language model systems were the hardest to detect (61.3–65.9% accuracy), while an ML detector maintained over 94.5% accuracy across all conditions. The study concludes that the primary societal threat of modern audio deepfakes may be the erosion of trust in genuine speech, not just deception.
- AI policy
- Quality assurance
Research
Trust in Generative AI for Health Information Consumption and the Effect of Learned Dependency: An Experimental Investigation
Arif Ahmed, Gondy Leroy, Agrim Sachdeva et al.
arXiv · 2026-05-21
This study investigates how learned dependency on generative AI affects users' trust in AI-generated health information. Two randomized controlled experiments with over 900 participants found that while information accuracy significantly increased trust, highly dependent users were more likely to trust incorrect AI-generated health content. Text highlighting as an interface intervention had no significant effect on reducing this overreliance. The findings suggest that current interface designs are insufficient to counteract dependency-driven miscalibration, pointing to a need for more effective tools that promote critical evaluation of AI health outputs.
- Quality assurance
- AI policy
Research
Whose Good, Whose Place? The Moral Geography of Agentic AI for Social Good
Poli Nemkova, Haeshitha Indukuri, Jaedon Charles
arXiv · 2026-05-21
This paper surveys 112 research papers (2015–2026) on agentic AI systems proposed for social-good domains, using the UN Sustainable Development Goals (SDGs) as a framing device. The authors find a 'moral-geographic asymmetry': 73% of papers (82 of 112) specify no geographic context, and papers addressing institutional/social-policy SDGs do so at far lower rates (13%) than those covering health or ecological SDGs (37–40%). Only 25% of papers (28 of 112) report any real-world deployment or small-scale test, revealing a gap between claimed social benefit and demonstrated accountability to affected communities. The authors identify five accountability gaps and propose a minimal reporting standard to make agentic AI for social good more context-specific, participatory, and accountable.
- AI policy
Research
Test-Time Training Undermines Safety Guardrails
Simone Antonelli, Sadegh Akhondzadeh, Aleksandar Bojchevski
arXiv · 2026-05-21
This paper examines how Test-Time Training (TTT) — a paradigm that lets AI models update their parameters during inference — creates new security vulnerabilities that can be exploited to bypass safety filters. The authors identify three threat models for TTT and demonstrate that attackers can use them to achieve Attack Success Rates (ASR@10) of 95% and 93% respectively under LoRA-based few-shot and generation-phase attacks, across models of different families and scales. These vulnerabilities extend to production fine-tuning APIs, and the paper also identifies a measurement artifact where TTT-induced overfitting inflates ASR scores, proposing a validity-aware evaluation to correct for it. As an initial defense, the authors propose a lightweight detector that flags malicious TTT requests using perplexity shifts on a private harmful holdout, while noting that robust safety will ultimately require dynamic alignment.
- AI policy
- Quality assurance
Research
Can AI Guess What You Know? Performance Comparison of Large Language Models for Human Domain Knowledge Estimation From Communication Logs
Ko Watanabe, Shoya Ishimaru
arXiv · 2026-05-21
This paper investigates whether Large Language Models can infer individual employees' domain knowledge directly from long-term Slack communication logs, addressing the organizational problem of 'who knows what.' Evaluating seven LLMs — including Gemini, Claude, and GPT families — against self-reported skill ratings from 27 participants across 27,188 messages from 43 users, the study finds that Gemini 2.5 Flash achieved the lowest estimation error (MAE 21.13%), while GPT models showed significantly larger discrepancies. Notably, estimation accuracy depended only weakly on message volume, suggesting that text quantity alone does not guarantee better inference. The findings demonstrate both the feasibility and current limits of automated expertise mapping, while highlighting the need for privacy-preserving deployments.
- Workforce
- Enterprise
Research
Evaluating Commercial AI Chatbots as News Intermediaries
Mirac Suzgun, Emily Shen, Federico Bianchi et al.
arXiv · 2026-05-21
This study evaluates six commercial AI chatbots (Gemini 3 Flash/Pro, Grok 4, Claude 4.5 Sonnet, GPT-5, GPT-4o mini) on 2,100 factual questions drawn from same-day BBC News reporting across six regional services over 14 days in February 2026. The best systems exceed 90% accuracy on multiple-choice questions about very recent events, but lose 11–17% accuracy under free-response conditions, and show systematic regional inequity: Hindi-language queries score roughly 10 percentage points lower than other languages, with models citing English Wikipedia more than any Hindi outlet. Over 70% of errors stem from retrieval failures rather than reasoning failures, and accuracy collapses to 19–70% when questions contain subtle false premises—with the most vulnerable model accepting fabricated facts 64% of the time. The findings reveal that headline accuracy figures can conceal deep structural problems in how AI news intermediaries handle non-Anglophone content, retrieval dependence, and adversarial or imperfect queries real users pose.
- Quality assurance
- AI policy
Research
Can AI Make Conflicts Worse? An Alignment Failure in LLM Deployment Across Conflict Contexts
Andrii Kryshtal
arXiv · 2026-05-21
This paper tests nine large-language-model configurations from OpenAI, Anthropic, DeepSeek, and xAI on 90 multi-turn scenarios designed to expose harmful outputs in armed-conflict contexts, including false equivalence between documented atrocities, genocide denial, and failure to recognize ethnic slurs. Failure rates range from 6% to 47% across models, and when users prompted models for 'balance' in cases where international courts had already assigned responsibility, five of nine configurations failed 80–100% of the time. Because AI tools are already used by journalists, humanitarian workers, and governments in conflict-affected societies, such outputs can deepen divisions in fragile communities. The authors release the first evaluation framework for this domain and recommend integrating it into standard alignment evaluation portfolios.
- AI policy
- Quality assurance
Research
AMEL: Accumulated Message Effects on LLM Judgments
Sid-Ali Temkit
arXiv · 2026-05-21
This paper investigates whether prior conversation history biases large language model (LLM) judgments in automated evaluation settings, a phenomenon the authors call the 'accumulated message effect on LLM judgments' (AMEL). Across 84,088 API calls to 12 models from 5 providers, they find that models shift their evaluations toward the prevailing polarity of prior conversation history (d = -0.17, p < 10^-53), with the effect strongest on items where the model is genuinely uncertain at baseline (d = -0.36 for high-entropy items). Notably, negative histories induce 1.52x more bias than positive ones, and larger models reduce but do not eliminate the effect. The authors recommend using a fresh context per item in evaluation pipelines, or balancing conversation history when batching is unavoidable.
- Quality assurance
Research
Beyond the Org Chart: AI and the Transformation of Invisible Work
Stephanie Rosenthal, Shamsi Iqbal
arXiv · 2026-05-21
This qualitative study interviewed 24 product-focused professionals at a large technology firm to understand how AI adoption is reshaping both formal and informal dimensions of work. The findings indicate that AI is altering not just defined role responsibilities and cross-role collaboration, but also informal cultural practices such as mentoring, feedback from professional networks, and leadership development. While some changes—like smoother peer collaboration—are positive, others put traditional career growth pathways at risk by making previously invisible informal work less visible. The authors propose concrete steps for AI companies, leaders, and individuals to preserve healthy cultures that support diverse thinking, collaboration, and informal interaction during AI-driven workplace transformation.
- Workforce
- Enterprise
Research
The efficiency-gain illusion: People underestimate the rate of AI use and overestimate its benefits on simple tasks
Sunny Yu, Myra Cheng, Ahmad Jabbar et al.
arXiv · 2026-05-21
This paper reports on three pre-registered user studies (N = 2,691) examining whether people's reliance on AI for simple cognitive tasks—such as arithmetic, spell-check, and answering easy questions—is well-calibrated. The researchers find that people frequently choose AI assistance even when it provides no meaningful time or effort savings, and they suffer from two systematic miscalibrations: underestimating how often they actually use AI, and overestimating the efficiency gains AI provides. A session-level carryover effect is also identified, where prior AI use leads to greater subsequent adoption and deepens miscalibration, suggesting a risk of an overreliance feedback loop. These findings have direct implications for how workers and organizations should think about AI integration, particularly the risk that perceived productivity gains from AI may not reflect actual efficiency improvements.
- Workforce
- Enterprise
Research
Scientific reasoning does not reliably translate into scientific forecasting in frontier AI
Sean Wu, Pan Lu, Yupeng Chen et al.
arXiv · 2026-05-21
This paper introduces CUSP, a benchmark for evaluating AI models on event-level scientific forecasting across eight disciplines, and tests six frontier AI systems on it. The results reveal a fundamental asymmetry: while models demonstrate strong retrospective scientific reasoning—identifying plausible mechanisms and background knowledge—they perform near chance on feasibility assessment, generate solution strategies that weakly match actual advances, and consistently predict breakthroughs later than they occur. Even providing additional pre-cutoff scientific knowledge does not close this gap. The authors conclude that scientific forecasting should be treated as a distinct, complementary capability dimension when deploying AI for research prioritization and scientific decision-making.
- AI policy
- Enterprise
Research
Is Capability a Liability? More Capable Language Models Make Worse Forecasts When It Matters Most
Nick Merrill, Jaeho Lee, Ezra Karger
arXiv · 2026-05-21
This paper documents an 'inverse scaling' phenomenon in large language models (LLMs) applied to forecasting: more capable models actually produce worse probabilistic forecasts for time series characterized by superlinear growth and tail risk of regime change, patterns common in finance and epidemiology. The failure concentrates at the upper tail of the forecast distribution, where more capable models shift predictions upward to track aggressive growth extrapolations. The finding holds across a new simulated benchmark (ForecastBench-Sim), synthetic epidemic models, and real-world datasets including COVID-19, measles, housing markets, and hyperinflation, with both model scale and post-training independently contributing to the effect. The authors warn that single-threshold scoring metrics used in standard LLM forecasting benchmarks mask this upper-tail failure and recommend that evaluations incorporate continuous, unbounded accuracy measures alongside binary threshold metrics.
- Quality assurance
- AI policy
Research
MBABench: Evaluating LLM Agents on End-to-End Spreadsheet Tasks in Finance
Thomson Yen, Julian Poeltl, Harshith Srinivas Gear et al.
arXiv · 2026-05-21
MBABench introduces one of the first benchmarks evaluating large language model agents on end-to-end spreadsheet construction for financially critical workflows such as financial modeling, forecasting, and scenario analysis. Unlike prior benchmarks that focus on question-answering or single-formula edits, MBABench assesses complete artifact creation using a three-dimensional taxonomy covering Accuracy, Formula, and Format criteria aligned with professional standards. Results show that the Claude family leads the benchmark and produces the most professional-looking outputs, but even the strongest agents frequently fall short of professional finance standards and degrade sharply as task complexity increases beyond a few chained calculations. This indicates current AI agents cannot yet reliably automate the spreadsheet workflows that underpin real-world enterprise finance operations.
- Enterprise
- Quality assurance
Research
Whose Voice Counts? Mapping Stakeholder Perspectives on AI Through Public Submissions to the U.S. Government
Alina Karakanta, Alex Christiansen, Tomás Dodds et al.
arXiv · 2026-05-21
This paper analyzes public submissions to the Trump Administration's U.S. AI Action Plan consultation to map how different stakeholder groups — academia, individuals, and the private sector — perceive and prioritize AI issues. Using topic modelling and frequency analysis on a cleaned corpus of public letters, the authors find that individuals predominantly voice concerns about AI's impact on daily life, while other groups focus more on AI development. The resulting AI Action Plan largely reflects private-sector priorities around security, policy, and development, with individual concerns underrepresented. These findings highlight a gap between public input and policy outcomes in AI governance.
- AI policy
Research
Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety
Piercosma Bisconti, Matteo Prandi, Federico Pierucci et al.
arXiv · 2026-05-21
Boiling the Frog introduces a multi-turn benchmark that tests whether tool-using AI agents deployed in corporate and office settings can be manipulated through incremental, escalating attacks — beginning with benign workspace edits before introducing a harmful payload. The benchmark uses stateful, multi-turn evaluation across a three-level operational risk taxonomy grounded in the EU AI Act Annex I, Annex III, and the Code of Practice on General-Purpose AI (GPAI). Across a nine-model panel, the aggregate strict attack success rate is 44.4%, ranging from 20.5% for Claude Haiku 4.5 to 92.9% for Gemini 3.1 Flash Lite, with Code of Practice loss-of-control scenarios reaching an average attack success rate of 93.3%. These results reveal that agentic AI systems face serious safety risks when evaluated not on what they say but on what they do within persistent environments — a critical gap for enterprise deployment and regulatory compliance.
- Quality assurance
- AI policy