News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
Adapting regulatory sandboxes as experimentalist governance frameworks for public sector artificial intelligence experimentation
Jeremmy Okonjo
Global Public Policy and Governance · 2026-06-01
This paper examines whether traditional AI regulatory sandboxes—originally designed to facilitate private sector innovation under regulatory uncertainty—are suitable frameworks for governing AI experimentation in the public sector. The authors argue that significant normative, legal, and institutional adaptations are required because public sector AI raises distinct concerns around opacity, bias, accountability, human rights, and democratic legitimacy. The paper analyzes the specific challenges public sector AI poses for conventional sandbox models and proposes an adapted framework to ensure legally and democratically legitimate AI experimentation in government contexts. This is directly relevant to how policymakers design oversight and governance structures for AI adoption in public administration.
- AI policy
- Certifications
- Quality assurance
Research
Artificial intelligence, productivity, and economic growth: Global evidence and emerging policy implications
Jainendra Kumar Verma, Kamal De Krishna
International Journal of Financial Management and Economics · 2026-06-01
This systematic review synthesizes evidence from 28 studies (2020–2026) on how AI adoption affects labor productivity, total factor productivity, labor market reallocation, and macroeconomic growth across advanced, emerging, and developing economies. The paper finds that AI generates large productivity gains at the firm level, though aggregate growth effects are more moderate, with outcomes shaped by digital infrastructure, human capital, institutional quality, and AI-human complementarity. The authors introduce an Adaptive Diffusion-Complementarity (ADC) Framework to explain these dynamics and stress that supportive policies are needed to maximize AI's economic benefits.
- Workforce
- Enterprise
- AI policy
Research
Beyond AI disclosure: Claim accountability and responsible research in scholarly publishing
Ward van Zoonen, Anna Morgan-Thomas, Aizhan Tursunbayeva
European Management Journal · 2026-06-01
This paper argues that existing AI governance policies in scholarly publishing—focusing on disclosure or prohibition—fail to address the core accountability problem: who can defend the claims made in published research. The authors propose shifting governance from regulating AI tools to regulating scholarly claims, requiring a named human to be able to reconstruct and defend each claim entering the record. They introduce a two-threshold framework (a confidentiality threshold and a judgment threshold) and translate it into role-based self-assessment guidelines for authors, reviewers, editors, and publishers. The work has direct implications for research quality assurance and certification standards in academic publishing.
- Quality assurance
- Certifications
- AI policy
Research
LAW IN THE DIGITAL AGE: INDIA’S EVOLVING LEGAL ARCHITECTURE FOR ARTIFICIAL INTELLIGENCE AND DATA PRIVACY
Shweta Rana, Harpreet Singh
arXiv · 2026-06-01
This chapter analyzes India's emerging legal framework for artificial intelligence and data privacy, tracing key developments from the 2017 Puttaswamy constitutional ruling through the Digital Personal Data Protection Act 2023 and the IndiaAI Mission (₹10,372 crore outlay, March 2024). The authors argue India occupies a hybrid regulatory position—constitutionally rights-aware but consent-centric and exemption-heavy in statutory design—and compare this approach with the EU's GDPR and AI Act, the US sectoral patchwork, China's centralized model, and Brazil's rights-anchored framework. The paper finds that effective AI governance will depend less on statutory text and more on institutional capacity, judicial vigilance, and enforcement, with open questions remaining around algorithmic accountability for state actors, foundation model regulation, and cross-border data flows. The analysis is relevant for policymakers and enterprises navigating compliance in a rapidly evolving regulatory environment.
- AI policy
- Enterprise
- Certifications
Research
Towards artificial intelligence for the public sector: framing and bridging academia and practice
Zander Weisman Mintz, Ji Ma
Global Public Policy and Governance · 2026-06-01
This paper proposes a functional framework to organize the fragmented, cross-disciplinary literature on AI in the public sector by grouping research around four governance functions: creating public value, delivering public services, responsiveness to the public, and protecting state–society relations. Drawing on a citation-based review and BERTopic modeling of 3,268 works, the authors find that post-2022 scholarship has shifted sharply from domain-application research toward topics like algorithmic fairness, ethics, and regulation, while work explicitly situated within the policy cycle remains scarce. The framework is intended to serve as a translation layer between academia and practitioners, helping governments navigate accountability, fairness, and trust requirements that have no direct parallel in private-sector AI deployment. The findings matter because they highlight where governance-relevant guidance is still lacking and where policy-focused scholarship needs to grow.
- AI policy
- Workforce
- Enterprise
Research
Artificial Intelligence for End-of-Life Electric Vehicle Battery Disassembly: A Comprehensive Review
Jie Li, Wenchao Li
IntechOpen eBooks · 2026-06-01
This review examines how artificial intelligence is being applied to the disassembly of end-of-life electric vehicle batteries (EOL-EVBs), covering three core stages: preprocessing, sequence planning, and operation. AI methods such as transformer-based networks, hybrid neural models, and computer vision systems have improved state estimation accuracy, sorting efficiency, and disassembly planning, while vision-based robotics and human-robot collaboration enhance safety and productivity. The paper identifies persistent challenges including data dependency, computational cost, and real-world deployment barriers, and outlines future directions toward fully autonomous and scalable battery recycling systems. The findings matter because efficient EOL-EVB disassembly is essential for high-value material recovery and supporting a circular economy as the EV market rapidly expands.
- Workforce
- Enterprise
- Quality assurance
Research
ClawHub Security Signals: When VirusTotal, Static Analysis, and SkillSpector Disagree
Vincent Koc, Patrick Erichsen, Jacob Tomlinson et al.
arXiv · 2026-05-31
ClawHub Security Signals presents a dataset of 67,453 public AI agent skill versions, each paired with verdicts from three scanner families—VirusTotal, static heuristic analysis, and NVIDIA SkillSpector—to study how these scanners agree or disagree on flagging potentially harmful skills. The paper finds substantial disagreement: 81.9% of flagged skills are identified by only one scanner, any two scanners overlap on at most 10.4% of their combined positives, and only 0.69% of skills are flagged by all three. Importantly, the disagreement is structured by attack surface: SkillSpector flags semantic agentic-risk in 75.3% of suspicious rows but only 6.8% of malicious ones, while VirusTotal accounts for 72.8% of malicious-verdict rows, consistent with bundled-code malware. The authors conclude that agent-skill security requires layered governance rather than single-scanner decisions, and release the dataset as a silver-standard resource to support further research.
- Quality assurance
- AI policy
Research
Hierarchical Online Prompt Mutation with Dual-Loop Feedback for Guardrailed Evidence Document Generation: A Production-Evaluation Case Study
Nataraj Agaram Sundar, Tejas Morabia
arXiv · 2026-05-31
HOPM (Hierarchical Online Prompt Mutation) is a framework that treats language model prompts as adaptive online policies, using a family/version router, deterministic guardrails, and dual feedback from both human reviewers and an automated judge to continuously improve document generation. Evaluated on a real marketplace dispute-evidence workflow across 600 cases per variant, full HOPM raised count win rate from 34.7% to 45.7% (+11.0 pp) and amount-weighted win rate from 22.3% to 41.4% (+19.1 pp), while improving mean Likert quality scores from 3.18 to 4.40 and cutting issue-flag rates from 15.3% to 5.2%. The study demonstrates that combining bandit-based routing, token-level mutation, and dual feedback loops outperforms each component in isolation, offering a reproducible evaluation structure for high-stakes document generation. This matters for enterprise and quality-assurance practitioners who need auditable, evidence-grounded AI systems in production workflows.
- Enterprise
- Quality assurance
Research
GovAI-Pipe: A Layered AI Governance Pipeline for Citizen-Facing AI in Turkey's e-Government Gateway
Ahmet Kaplan
arXiv · 2026-05-31
This paper proposes GovAI-Pipe, a four-layer AI governance pipeline designed to bridge the gap between high-level AI policy frameworks (including the EU AI Act, OECD AI Principles, and Turkey's National AI Strategy) and the operational deployment of AI in Turkey's e-Government Gateway (e-Devlet), which serves over 68 million registered users across more than 9,200 services. The pipeline covers pre-deployment validation (bias testing, explainability, privacy impact assessment), deployment governance (risk-tier classification and approval workflows), runtime monitoring (drift detection, fairness tracking, human-in-the-loop escalation), and post-incident governance (audit trails, rollback, and citizen redress). Each layer is anchored to specific provisions of the EU AI Act, GDPR, and Turkey's National AI Strategy, and the framework is demonstrated through two high-risk e-Devlet use cases. The work matters because it operationalizes abstract governance principles into auditable, technical pipeline components for citizen-facing AI systems in a large-scale public sector context.
- AI policy
- Certifications
Research
Measuring the Occupation-Level Impact of AbbVie Intelligence: AI Applicability Analysis, 2024-2025
John Regan, Jon Stevens, Brian Martin
arXiv · 2026-05-31
This paper measures how AbbVie's internal AI platform ('AbbVie Intelligence') affected employee work activities across 192 occupations in 2024–2025, using 598,744 de-identified AI conversations classified by the O*NET Intermediate Work Activity taxonomy. The authors compute occupation-level AI Applicability Scores and find statistically significant gains across three analyses: a year-over-year increase, a +10.0% gain (p<0.001) following the release of AbbVie Intelligence version 3 in August 2025, and a +6.68% gain (p<0.001) after a structured AI Learning Summit in November 2025. The findings indicate that both platform upgrades and formal AI education programs independently expand the measurable reach of AI tools across an enterprise workforce. This work is notable for its large-scale empirical approach to quantifying AI's occupation-level impact inside a single organization.
- Workforce
- Enterprise
Research
TukaBench: A Culturally Grounded Jailbreak Benchmark for African Languages
Victor Akinode, Senyu Li, Wassim Hamidouche et al.
arXiv · 2026-05-31
TukaBench introduces a jailbreak safety benchmark covering seven African languages to address the English-centric gap in Large Language Model (LLM) safety evaluation. The benchmark extends the existing JailbreakBench framework through four prompt settings—human translation, culturally adapted translation, human-curated prompts, and code-switched prompts—to isolate the effects of language, cultural grounding, and prompt evasiveness on model safety. Key findings show that prompting LLMs in African languages reduces refusal rates compared to English, with culturally adapted prompts producing the least refusal, and that LLM-as-a-judge reliability drops in lower-resource languages and less commonly supported scripts. The work also introduces a new outcome category called Deflection to capture model comprehension failures, validated through human annotations.
- Quality assurance
- AI policy
Research
Med-HEAL: Analyzing and Mitigating Hallucinations in Medical LLMs with Hallucination-Aware In-Context Learning
Yiming Liao, Zeno Franco, Jose Eduardo Lizarraga Mazaba et al.
arXiv · 2026-05-31
Med-HEAL is a framework for identifying, analyzing, and reducing hallucinations in medical large language models (LLMs) applied to clinical question answering over electronic health records (EHRs). The authors build a hallucination dataset from the EHRNoteQA benchmark (derived from MIMIC-IV discharge summaries) by evaluating BioMistral-7B, then label outputs using a dual pipeline combining GPT-4o judgments with human medical student auditing. They test two mitigation strategies—self-critique and retrieval-augmented in-context learning—across five open-source LLMs, finding that self-critique significantly improves accuracy for three of the five models (p < 0.05) without parameter updates. The work provides a reusable dataset and practical mitigation framework aimed at safer deployment of AI in clinical decision support.
- Quality assurance
Research
KG-FairDiff: Knowledge Graph-Guided Prompt Refinement for Demographically Fair Text-to-Image Generation
Farbod Davoodi, Seyed Reza Tavakoli Shiyadeh, Pooria Safaei et al.
arXiv · 2026-05-31
KG-FairDiff is an inference-time framework designed to reduce demographic and cultural stereotypes in Text-to-Image (TTI) systems without requiring costly model retraining. It combines a knowledge graph of roughly 1,200 culture- and bias-related triples with an LLM-based prompt rewriter and a validator that accepts only refined prompts shown to reduce a divergence-based fairness loss while preserving the user's original intent. The system is model-agnostic and was audited across eight widely-deployed image generators, demonstrating substantial reductions in gender, race, age, and intersectional disparities. This matters because TTI systems are embedded in journalism, education, and advertising, meaning inherited stereotypes cause population-scale harm that practical, deployment-ready mitigations like this can address.
- AI policy
- Quality assurance
Research
Implicit Geographic Inference in LLM Medical Triage: Language-Driven Disparities in Emergency Recommendations
Qi Han Wong
arXiv · 2026-05-31
This study tests whether a large language model (Gemini 3.5 Flash) gives different emergency-care recommendations for the same neurological symptoms depending solely on the language of the patient's prompt. Across 450 API calls in six languages, the model recommended ER visits at rates from 0% (Japanese, Hindi) to 30% (English, Arabic), even though it assigned nearly identical severity scores across all languages. Adding a single sentence specifying a US location raised ER recommendation rates by up to 76.7 percentage points for non-English prompts, and a back-translation control confirmed the disparity stems from implicit geographic inference rather than translation quality. These findings reveal a significant equity concern: patients interacting with AI triage tools in certain languages may systematically receive less urgent care guidance than others with identical symptoms.
- AI policy
- Quality assurance
Research
Low-Resource Safety Failures Are Action Failures, Not Representation Failures
Rashad Aziz, Ikhlasul Akmal Hanif, Fajri Koto
arXiv · 2026-05-31
This paper investigates why AI safety alignment—teaching models to refuse harmful prompts—fails to transfer from high-resource languages like English to low-resource languages like Swahili or Burmese. Testing three large language models (Qwen2.5-7B, Gemma-2-9B, and Llama-3.1-8B) across 23 languages, the authors find that the internal representations identifying harmful content transfer well across languages, but the models' ability to convert those representations into actual refusals drops sharply (from 87.9% to 43.9%). The fix is not retraining but recalibration: using as few as 1–4 examples per class in the target language to reset a decision threshold, which raises mean refusal selectivity from 33.6 to 54.5 while preserving general utility. These findings have direct implications for AI safety policy, suggesting that low-resource language safety gaps can be addressed efficiently without full retraining.
- AI policy
- Quality assurance
Research
BraveGuard: From Open-World Threats to Safer Computer-Use Agents
Yunhao Feng, Xiaohu Du, Xinhao Deng et al.
arXiv · 2026-05-31
BraveGuard is a self-evolving defense framework that trains guard models to detect safety risks in computer-use AI agents—systems that interact with files, terminals, browsers, and external tools over multi-step execution traces. Because harm in these agents often emerges across sequences of locally benign actions rather than in isolated prompts, BraveGuard mines real-world threat signals, generates realistic agent trajectories, and uses them to supervise guard model training in an adaptive loop. Evaluations on the AgentHazard benchmark show the approach substantially improves detection accuracy from 38.79% to 82.38% under an averaged guard-model setting, outperforming off-the-shelf guards. The framework offers a scalable path to adaptive safety monitoring for AI agents facing continuously evolving real-world risks.
- Quality assurance
- AI policy
Research
ASE-26: a curriculum for agentic software engineering as a discipline
Mikael Gorsky
arXiv · 2026-05-31
This paper introduces ASE-26, a 21-module undergraduate curriculum designed to teach agentic software engineering as a formal discipline — meaning the skills needed to direct AI agents rather than write code directly. The authors ground the curriculum's urgency in cited empirical evidence: Anthropic's Economic Index reports 79% automation of Claude Code interactions, Handa et al. find ~75% AI exposure across Computer Programmer task activities, and Brynjolfsson et al. document a 13% relative employment decline among workers aged 22–25 in AI-exposed occupations. The paper argues that the bottleneck in agentic software engineering is not model capability but structured practitioner discipline, and that undergraduate curricula like ASE-26 are the primary mechanism to close that gap. The curriculum is deposited on Zenodo under CC BY-ND 4.0 as a citable reference.
- Workforce
- Certifications
Research
From Outliers to Errors: Auditing Pali-to-English LLM Translations with Multi-Reference Adjudication
Máté Metzger, Nadnapang Phophichit, Hansa Dhammahaso
arXiv · 2026-05-31
This paper audits Pali-to-English translations produced by four large language models—GPT-5.5, Claude Sonnet 4.6, Gemini 3.1 Pro, and Grok 4.3—across 1,700 passages from the Pali Canon, using three established human translations as a multi-reference envelope rather than a single gold standard. The authors show that embedding drift from the reference centroid predicts error severity rather than error itself: roughly 80% of lower-drift outliers were valid translation variations, while major-error rates rose to over 50% among the highest-drift candidates. Grok 4.3 performed worst, exhibiting the largest outlier volume and the highest tail major-error rate (27.6% overall, 74.4% above drift 3.0), with dominant failure modes including omission, truncation, and doctrinal term errors. The study contributes a reusable audit pipeline for classical-language translation that prioritizes human or LLM-judge review where it matters most, which has direct implications for quality assurance of AI-generated translations in specialized domains.
- Quality assurance
Research
MiCU: End-to-End Smart Home Command Understanding with Large Language Model
Haowei Han, Kexin Hu, Weiwei Cai et al.
arXiv · 2026-05-31
MiCU is a domain-specific large language model built for smart home command understanding, addressing the challenge that existing systems handle precise commands well but struggle with ambiguous ones like 'make the bedroom cozy.' The authors develop an automated training data synthesis workflow using user logs and LLMs, apply curriculum learning and reinforcement learning guided by domain-specific rules, and introduce a token compression technique to reduce inference costs. Deployed in the Xiaomi Home app at approximately 1.7 million page views per day, MiCU achieves an average accuracy gain of 20.01% across device categories, reduces user correction rate by 1.57%, and increases human-audited accuracy by 32.05%. This demonstrates that LLMs fine-tuned with domain-specific data and reasoning enhancements can meaningfully improve real-world smart home automation at scale.
- Enterprise
Research
Data Collection for Training Quality-Control AI in Carpet Manufacturing
Akbar Erkinov
arXiv · 2026-05-31
This paper presents a design blueprint for an in-line machine-vision quality-control system tailored to woven and tufted carpet manufacturing, where visual inspection is currently slow, subjective, and inconsistent. The proposal combines synchronized line-scan cameras with bright-field and grazing illumination to inspect carpet webs in real time, while simultaneously collecting and labeling defect images to train progressively more capable AI models. The system starts with unsupervised anomaly detection on defect-free material and matures through a human-in-the-loop annotation process into supervised detection and segmentation, framed within a Six Sigma DMAIC project context. The authors argue that treating data collection as a first-class engineering objective—rather than an afterthought—directly translates into measurable reductions in escaped defects and improved process sigma levels.
- Quality assurance
- Enterprise
Research
Lost in Delusion: Examining LLM Safety Under User Delusions and Distress
Andrew Aquilina, Chetna Nihalani, Vasudha Varadarajan et al.
arXiv · 2026-05-31
This paper investigates how large language models (LLMs) behave when users in psychological distress express beliefs framed by delusion, using multi-turn simulated conversations across clinically grounded personas and six LLMs. The study identifies a 'recognition-intervention gap': models detect distress at similar rates whether or not delusion is present, but suppress safety interventions by up to 4.5x when distress is embedded in delusional framing, because models accumulate acceptance of delusional premises rather than flagging them. Standard prompting fixes (asking models to assess distress) backfire under delusional framing; only delusion-aware prompting with explicit guidance closes the gap, though this depends on a delusion classifier that itself underperforms on the most vulnerable models. The findings argue that safe deployment of LLM chatbots requires treating delusional framing as a distinct, high-priority risk signal that should override conversational accommodation.
- AI policy
- Quality assurance
Research
Silent Failures in Federated Personalization of Foundation Models
YongKyung Oh, Alex Bui
arXiv · 2026-05-31
This paper identifies and categorizes a class of trustworthiness problems called 'Silent Failures' that emerge when foundation models are personalized via federated learning on decentralized private data. The authors argue that privacy constraints inherent to federated learning limit visibility into model behavior, allowing issues like amplified bias, fairness collapse, and alignment erosion to go undetected. A landscape analysis of existing benchmarks reveals a structural gap: federated benchmarks measure system performance but not model behavior, while centralized trustworthiness benchmarks require model access incompatible with federated privacy. The paper proposes a taxonomy of six silent failure modes and calls for privacy-preserving behavioral evaluation methods, recommending that silent failures become a standard diagnostic category for trustworthy federated AI systems.
- Quality assurance
- AI policy
Research
AI Runtime Evidence Protocol (AIREP)
Ali Toygar Abak
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-31
AIREP introduces a vendor-neutral, cryptographically verifiable record format designed to document AI runtime governance decisions, including what was decided, the evidence supporting that decision, and critically, what the evidence does not cover (scope-honesty). Each record is bound to its position in a hash chain using SHA-256 and signed with Ed25519, enabling tamper-evident audit trails for AI decision-making. The specification includes a JSON Schema, five binding profiles, and a conformance kit with two independent verifiers that agree byte-for-byte. This matters because it provides a structured, auditable foundation for AI governance accountability, though the authors note it remains experimental and is not yet a ratified standard or certification scheme.
- Quality assurance
- Certifications
- AI policy
Research
Digital Leadership and AI Adoption Intention and Usage in SMEs: The Mediating Role of Employee Digital Readiness in an Urban Bangladeshi Context
Israfil Munna, Udin Udin
International Journal of Organizational Leadership · 2026-05-31
This study of 210 employees and managers at SMEs in Dhaka, Bangladesh finds that digital leadership positively drives both AI adoption intention/usage and employee digital readiness, with employee digital readiness mediating 52.8% of the total effect of leadership on AI adoption. Using confirmatory factor analysis and bootstrapped mediation analysis grounded in the TOE and TAM frameworks, the research demonstrates that cultivating digitally ready employees is a critical mechanism through which organizational leadership translates into actual AI uptake. The findings highlight actionable levers for SME managers seeking to accelerate AI adoption in under-studied South Asian business contexts.
- Workforce
- Enterprise
Research
Use of Generative Artificial Intelligence for Managerial Verification in Multinational Contract Management.
Fernando Ivan Jaimes Rada, Julia Enith Herrera Mendoza
Journal of Information Technology Cybersecurity and Artificial Intelligence · 2026-05-31
This paper examines how Microsoft Copilot, an AI embedded in office productivity platforms, can support regulatory and legal compliance verification in small and medium-sized multinational enterprises that lack dedicated compliance infrastructure. Using a qualitative case study of a small European multinational, the researchers found that AI-assisted assessments progressively improved in coherence and traceability as the system became more contextualized, and were fully consistent with conclusions reached by external legal counsel. The findings suggest that AI tools can meaningfully enhance managerial verification in contract management when deployed under explicit human supervision and responsible governance frameworks.
- Enterprise
- AI policy
- Quality assurance