News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
Implicit Geographic Inference in LLM Medical Triage: Language-Driven Disparities in Emergency Recommendations
Qi Han Wong
arXiv · 2026-05-31
This study tests whether a large language model (Gemini 3.5 Flash) gives different emergency-care recommendations for the same neurological symptoms depending solely on the language of the patient's prompt. Across 450 API calls in six languages, the model recommended ER visits at rates from 0% (Japanese, Hindi) to 30% (English, Arabic), even though it assigned nearly identical severity scores across all languages. Adding a single sentence specifying a US location raised ER recommendation rates by up to 76.7 percentage points for non-English prompts, and a back-translation control confirmed the disparity stems from implicit geographic inference rather than translation quality. These findings reveal a significant equity concern: patients interacting with AI triage tools in certain languages may systematically receive less urgent care guidance than others with identical symptoms.
- AI policy
- Quality assurance
Research
Low-Resource Safety Failures Are Action Failures, Not Representation Failures
Rashad Aziz, Ikhlasul Akmal Hanif, Fajri Koto
arXiv · 2026-05-31
This paper investigates why AI safety alignment—teaching models to refuse harmful prompts—fails to transfer from high-resource languages like English to low-resource languages like Swahili or Burmese. Testing three large language models (Qwen2.5-7B, Gemma-2-9B, and Llama-3.1-8B) across 23 languages, the authors find that the internal representations identifying harmful content transfer well across languages, but the models' ability to convert those representations into actual refusals drops sharply (from 87.9% to 43.9%). The fix is not retraining but recalibration: using as few as 1–4 examples per class in the target language to reset a decision threshold, which raises mean refusal selectivity from 33.6 to 54.5 while preserving general utility. These findings have direct implications for AI safety policy, suggesting that low-resource language safety gaps can be addressed efficiently without full retraining.
- AI policy
- Quality assurance
Research
BraveGuard: From Open-World Threats to Safer Computer-Use Agents
Yunhao Feng, Xiaohu Du, Xinhao Deng et al.
arXiv · 2026-05-31
BraveGuard is a self-evolving defense framework that trains guard models to detect safety risks in computer-use AI agents—systems that interact with files, terminals, browsers, and external tools over multi-step execution traces. Because harm in these agents often emerges across sequences of locally benign actions rather than in isolated prompts, BraveGuard mines real-world threat signals, generates realistic agent trajectories, and uses them to supervise guard model training in an adaptive loop. Evaluations on the AgentHazard benchmark show the approach substantially improves detection accuracy from 38.79% to 82.38% under an averaged guard-model setting, outperforming off-the-shelf guards. The framework offers a scalable path to adaptive safety monitoring for AI agents facing continuously evolving real-world risks.
- Quality assurance
- AI policy
Research
ASE-26: a curriculum for agentic software engineering as a discipline
Mikael Gorsky
arXiv · 2026-05-31
This paper introduces ASE-26, a 21-module undergraduate curriculum designed to teach agentic software engineering as a formal discipline — meaning the skills needed to direct AI agents rather than write code directly. The authors ground the curriculum's urgency in cited empirical evidence: Anthropic's Economic Index reports 79% automation of Claude Code interactions, Handa et al. find ~75% AI exposure across Computer Programmer task activities, and Brynjolfsson et al. document a 13% relative employment decline among workers aged 22–25 in AI-exposed occupations. The paper argues that the bottleneck in agentic software engineering is not model capability but structured practitioner discipline, and that undergraduate curricula like ASE-26 are the primary mechanism to close that gap. The curriculum is deposited on Zenodo under CC BY-ND 4.0 as a citable reference.
- Workforce
- Certifications
Research
From Outliers to Errors: Auditing Pali-to-English LLM Translations with Multi-Reference Adjudication
Máté Metzger, Nadnapang Phophichit, Hansa Dhammahaso
arXiv · 2026-05-31
This paper audits Pali-to-English translations produced by four large language models—GPT-5.5, Claude Sonnet 4.6, Gemini 3.1 Pro, and Grok 4.3—across 1,700 passages from the Pali Canon, using three established human translations as a multi-reference envelope rather than a single gold standard. The authors show that embedding drift from the reference centroid predicts error severity rather than error itself: roughly 80% of lower-drift outliers were valid translation variations, while major-error rates rose to over 50% among the highest-drift candidates. Grok 4.3 performed worst, exhibiting the largest outlier volume and the highest tail major-error rate (27.6% overall, 74.4% above drift 3.0), with dominant failure modes including omission, truncation, and doctrinal term errors. The study contributes a reusable audit pipeline for classical-language translation that prioritizes human or LLM-judge review where it matters most, which has direct implications for quality assurance of AI-generated translations in specialized domains.
- Quality assurance
Research
MiCU: End-to-End Smart Home Command Understanding with Large Language Model
Haowei Han, Kexin Hu, Weiwei Cai et al.
arXiv · 2026-05-31
MiCU is a domain-specific large language model built for smart home command understanding, addressing the challenge that existing systems handle precise commands well but struggle with ambiguous ones like 'make the bedroom cozy.' The authors develop an automated training data synthesis workflow using user logs and LLMs, apply curriculum learning and reinforcement learning guided by domain-specific rules, and introduce a token compression technique to reduce inference costs. Deployed in the Xiaomi Home app at approximately 1.7 million page views per day, MiCU achieves an average accuracy gain of 20.01% across device categories, reduces user correction rate by 1.57%, and increases human-audited accuracy by 32.05%. This demonstrates that LLMs fine-tuned with domain-specific data and reasoning enhancements can meaningfully improve real-world smart home automation at scale.
- Enterprise
Research
Data Collection for Training Quality-Control AI in Carpet Manufacturing
Akbar Erkinov
arXiv · 2026-05-31
This paper presents a design blueprint for an in-line machine-vision quality-control system tailored to woven and tufted carpet manufacturing, where visual inspection is currently slow, subjective, and inconsistent. The proposal combines synchronized line-scan cameras with bright-field and grazing illumination to inspect carpet webs in real time, while simultaneously collecting and labeling defect images to train progressively more capable AI models. The system starts with unsupervised anomaly detection on defect-free material and matures through a human-in-the-loop annotation process into supervised detection and segmentation, framed within a Six Sigma DMAIC project context. The authors argue that treating data collection as a first-class engineering objective—rather than an afterthought—directly translates into measurable reductions in escaped defects and improved process sigma levels.
- Quality assurance
- Enterprise
Research
Lost in Delusion: Examining LLM Safety Under User Delusions and Distress
Andrew Aquilina, Chetna Nihalani, Vasudha Varadarajan et al.
arXiv · 2026-05-31
This paper investigates how large language models (LLMs) behave when users in psychological distress express beliefs framed by delusion, using multi-turn simulated conversations across clinically grounded personas and six LLMs. The study identifies a 'recognition-intervention gap': models detect distress at similar rates whether or not delusion is present, but suppress safety interventions by up to 4.5x when distress is embedded in delusional framing, because models accumulate acceptance of delusional premises rather than flagging them. Standard prompting fixes (asking models to assess distress) backfire under delusional framing; only delusion-aware prompting with explicit guidance closes the gap, though this depends on a delusion classifier that itself underperforms on the most vulnerable models. The findings argue that safe deployment of LLM chatbots requires treating delusional framing as a distinct, high-priority risk signal that should override conversational accommodation.
- AI policy
- Quality assurance
Research
Silent Failures in Federated Personalization of Foundation Models
YongKyung Oh, Alex Bui
arXiv · 2026-05-31
This paper identifies and categorizes a class of trustworthiness problems called 'Silent Failures' that emerge when foundation models are personalized via federated learning on decentralized private data. The authors argue that privacy constraints inherent to federated learning limit visibility into model behavior, allowing issues like amplified bias, fairness collapse, and alignment erosion to go undetected. A landscape analysis of existing benchmarks reveals a structural gap: federated benchmarks measure system performance but not model behavior, while centralized trustworthiness benchmarks require model access incompatible with federated privacy. The paper proposes a taxonomy of six silent failure modes and calls for privacy-preserving behavioral evaluation methods, recommending that silent failures become a standard diagnostic category for trustworthy federated AI systems.
- Quality assurance
- AI policy
Research
AI Runtime Evidence Protocol (AIREP)
Ali Toygar Abak
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-31
AIREP introduces a vendor-neutral, cryptographically verifiable record format designed to document AI runtime governance decisions, including what was decided, the evidence supporting that decision, and critically, what the evidence does not cover (scope-honesty). Each record is bound to its position in a hash chain using SHA-256 and signed with Ed25519, enabling tamper-evident audit trails for AI decision-making. The specification includes a JSON Schema, five binding profiles, and a conformance kit with two independent verifiers that agree byte-for-byte. This matters because it provides a structured, auditable foundation for AI governance accountability, though the authors note it remains experimental and is not yet a ratified standard or certification scheme.
- Quality assurance
- Certifications
- AI policy
Research
Digital Leadership and AI Adoption Intention and Usage in SMEs: The Mediating Role of Employee Digital Readiness in an Urban Bangladeshi Context
Israfil Munna, Udin Udin
International Journal of Organizational Leadership · 2026-05-31
This study of 210 employees and managers at SMEs in Dhaka, Bangladesh finds that digital leadership positively drives both AI adoption intention/usage and employee digital readiness, with employee digital readiness mediating 52.8% of the total effect of leadership on AI adoption. Using confirmatory factor analysis and bootstrapped mediation analysis grounded in the TOE and TAM frameworks, the research demonstrates that cultivating digitally ready employees is a critical mechanism through which organizational leadership translates into actual AI uptake. The findings highlight actionable levers for SME managers seeking to accelerate AI adoption in under-studied South Asian business contexts.
- Workforce
- Enterprise
Research
Use of Generative Artificial Intelligence for Managerial Verification in Multinational Contract Management.
Fernando Ivan Jaimes Rada, Julia Enith Herrera Mendoza
Journal of Information Technology Cybersecurity and Artificial Intelligence · 2026-05-31
This paper examines how Microsoft Copilot, an AI embedded in office productivity platforms, can support regulatory and legal compliance verification in small and medium-sized multinational enterprises that lack dedicated compliance infrastructure. Using a qualitative case study of a small European multinational, the researchers found that AI-assisted assessments progressively improved in coherence and traceability as the system became more contextualized, and were fully consistent with conclusions reached by external legal counsel. The findings suggest that AI tools can meaningfully enhance managerial verification in contract management when deployed under explicit human supervision and responsible governance frameworks.
- Enterprise
- AI policy
- Quality assurance
Research
AI-Era Cyber Risk Standards: Tier 1 Qualification Requirements — Version 1.1
Busiel Morley
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-31
This document specifies Tier 1 qualification requirements for cyber liability insurance coverage in an AI-era threat environment where AI-assisted vulnerability discovery operates at machine speed. It introduces the 'Known Exposure Principle' as a governing liability concept, arguing that organizations that fail to scan for vulnerabilities now bear constructive knowledge of those vulnerabilities given the broad availability of AI-powered discovery tools. The framework organizes requirements across six domains—including AI-discovery posture, remediation velocity, blast-radius architecture, and coordination integrity—with blast-radius containment identified as the primary load-bearing domain. It also defines 'coordination integrity failure' at the seams of multi-agent AI systems as a distinct underwriting risk not addressed by existing cybersecurity standards or certification frameworks.
- Certifications
- AI policy
- Enterprise
- Quality assurance
Research
Trust in Artificial Intelligence and AI Adoption Intention in Enterprises
Nguyen Thanh An
International Journal of Social Science Humanity & Management Research · 2026-05-31
This study of 400 enterprise employees finds that perceived usefulness, ease of use, and organizational support increase trust in AI, while perceived risk reduces it, and that trust in AI strongly predicts AI adoption intention. Notably, trust fully mediates all relationships between antecedent factors and adoption intention, positioning it as the central psychological mechanism in enterprise AI uptake. The findings offer practical guidance for organizations seeking to build responsible AI adoption environments by addressing employee concerns around reliability, privacy, ethics, and security.
- Enterprise
- Workforce
- AI policy
Research
AI-Driven Green Supply Chain Optimization in E-Commerce: Evidence from Nagpur, Maharashtra, India
P M Krishna Raj, Sonal Agrawal
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-31
This study examines how AI tools—including machine learning-based route optimization, demand forecasting, and packaging analytics—can reduce the environmental footprint of e-commerce supply chains in Nagpur, India. Drawing on survey data from 120 logistics firms and regional emission records, the research finds that AI-assisted route optimization alone can reduce last-mile delivery emissions by an estimated 18–24%. While awareness of AI tools is growing among SMEs, adoption remains uneven due to cost barriers and skill gaps. The paper proposes a policy framework and practical roadmap for businesses, municipal bodies, and academic institutions to jointly accelerate green supply chain transformation.
- Enterprise
- AI policy
- Workforce
Research
AI Runtime Evidence Protocol (AIREP)
Ali Toygar Abak
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-31
AIREP (AI Runtime Evidence Protocol) proposes a vendor-neutral, model-independent record format designed to document AI governance decisions at runtime, capturing what was decided, on what evidence, and critically what the evidence does not cover — referred to as scope-honesty. Each record is signed, hash-chained using SHA-256, and serialized in canonical JSON, with a conformance kit including two independent verifiers (Python and Node.js) that agree byte-for-byte. The format aims to bring auditability and transparency to AI decision-making processes, which is relevant to quality assurance, certification, and policy compliance efforts. The authors note it is experimental and not yet a ratified standard or certification scheme.
- Quality assurance
- Certifications
- AI policy
Research
Artificial Intelligence and Environmental Information Disclosure: The Roles of Green Innovation and Corporate Governance
Yuqing Huang, Chaohui Xu
Corporate Social Responsibility and Environmental Management · 2026-05-31
This study examines how AI adoption affects corporate environmental information disclosure (EID) among Chinese listed companies from 2013 to 2023, finding that AI adoption is associated with higher EID. The relationship is strengthened by two internal mechanisms—green innovation, which enhances firms' ability to generate and process environmental information, and corporate governance, which strengthens managerial incentives to meet stakeholder environmental expectations. External monitoring via media and analyst attention further amplifies the positive AI-EID link. The findings offer practical guidance for enterprises and policymakers seeking to leverage AI, internal governance, and external oversight for green governance.
- Enterprise
- AI policy
- Quality assurance
Research
AI-Era Cyber Risk Standards: Tier 1 Qualification Requirements — Version 1.1
Busiel Morley
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-31
This document specifies Tier 1 qualification requirements for cyber liability insurance coverage in an AI-era threat environment where AI-assisted vulnerability discovery can surface zero-day flaws at machine speed. It introduces the 'Known Exposure Principle' as a governing liability concept, arguing that organizations failing to scan for vulnerabilities now bear constructive knowledge of those flaws given the broad availability of AI discovery tools. The framework is organized across six domains—including AI-discovery posture, remediation velocity, blast-radius architecture, AI system governance, supply chain dependency, and baseline controls—with blast-radius architecture designated as the primary load-bearing domain on the grounds that containment outranks repair when attackers and defenders have symmetric discovery capabilities. It also identifies 'coordination integrity failure' at the seams of multi-agent AI systems as a distinct underwriting domain not addressed by existing cybersecurity standards or certification frameworks.
- Certifications
- AI policy
- Quality assurance
- Enterprise
Research
AI-Driven Green Supply Chain Optimization in E-Commerce: Evidence from Nagpur, Maharashtra, India
P M Krishna Raj, Sonal Agrawal
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-31
This study investigates how AI tools—including machine learning-based route optimization, demand forecasting, and packaging analytics—can reduce the environmental footprint of e-commerce supply chains in Nagpur, Maharashtra, India. Drawing on survey data from 120 logistics firms and regional emission records, the paper finds that AI-assisted route optimization alone can reduce last-mile delivery emissions by an estimated 18–24%. While awareness of AI is growing among SMEs, adoption remains uneven due to cost barriers and skill gaps. The paper concludes with a policy framework for businesses, municipal bodies, and academic institutions to accelerate green supply chain transformation.
- Enterprise
- AI policy
- Workforce
Research
ARTIFICIAL INTELLIGENCE AND LABOR MARKET TRANSFORMATION IN DEVELOPING COUNTRIES
Sevinch Rakhimjonovna Nortojiyeva
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-31
This paper examines how artificial intelligence is reshaping labor markets in developing countries, drawing on reports from major international organizations including the ILO, World Bank, OECD, IMF, and WEF. The findings show that while AI automation threatens routine-task occupations, it simultaneously generates demand for new skills and roles. The authors argue that the net outcome depends heavily on whether effective educational and labor policies are put in place to support workforce adaptation. The study highlights both the risks of job displacement and the opportunities for economic development if AI integration is managed through sound policy frameworks.
- Workforce
- AI policy
Research
Benchmarking Security Risk Detection and Verification in Open Agentic Skill Ecosystems
Ismail Hossain, Sai Puppala, Zhuoran Lu et al.
arXiv · 2026-05-30
This paper introduces SkillVetBench, a two-stage benchmark for detecting and verifying malicious skills in open agentic AI ecosystems where community contributors publish reusable agent capabilities. The benchmark combines semantic vetting of natural-language skill specifications with runtime sandbox execution to catch threats that static or signature-based methods miss—experiments show existing baselines fail to detect up to 89% of malicious skills. The work draws on confirmed malicious samples from the live OpenClaw ecosystem, including the ClawHavoc supply-chain campaign, and finds that runtime attacks concentrate around a small set of high-permission primitives such as exec, write_file, install_skill, and spawn. This matters because it provides the first systematic evaluation framework for supply-chain security in AI agent platforms, highlighting critical gaps in current defenses.
- Quality assurance
- AI policy
Research
Accuracy, Stability, and Repeated-Run Reliability of Large Language Models on Deterministic Programming Tasks
Yongxi Zhou, Lai Yun Choi, Jiaxi Wen et al.
arXiv · 2026-05-30
This paper investigates whether large language models produce consistent outputs across repeated runs on deterministic programming tasks, not just whether they can eventually succeed. Using a benchmark of 100 LeetCode-style problems, the authors evaluate 16 models across five provider families with five repeated runs each (16,000 total evaluation instances), finding that run-level pass rate overstates retry-free coverage by up to 17.8 percentage points—a gap that can reverse model rankings among closely matched systems. The study shows that stability and accuracy, while strongly correlated (r=0.985), are not interchangeable metrics, and that prompt effects on performance are model-dependent rather than uniformly beneficial. These findings are directly relevant to enterprise and quality-assurance contexts where LLMs must reliably produce correct outputs on each invocation, not just occasionally under repeated sampling.
- Quality assurance
- Enterprise
Research
Adversarial Feeds Steer LLM Agent Decisions Against Their Defaults
Rana Muhammad Usman
arXiv · 2026-05-30
This paper demonstrates that the ranked external information streams (feeds) consumed by LLM agents before they make decisions can systematically steer those decisions away from the agent's defaults, even when the model, persona, and final prompt remain unchanged. Across 2,785 decision rollouts on four modern open-instruction LLMs, the authors find that a one-sided adversarial feed can shift a genuinely uncertain agent's decision probability from as low as 5% to 100% (Fisher p as low as 3×10⁻¹⁰), following a dose-response curve that generalizes across domains including security-relevant choices like removing deployment approval gates or relaxing access controls. The attack is partly mitigated by feed-level defenses, and a frontier model retains its defaults, but the findings reveal a critical blind spot: current safety evaluations test models and user prompts in isolation while ignoring the upstream ranker that controls what information the agent reads. The authors argue that agent security evaluations must audit the feed layer, making this directly relevant to AI quality assurance, policy standards for agentic AI deployment, and enterprise risk in systems relying on LLM agents.
- Quality assurance
- AI policy
Research
Prompts for Public-Sector LLMs Should Be Governed as Commons
Rashid Mushkani
arXiv · 2026-05-30
This paper argues that prompts used to deploy large language models in public-sector settings carry significant governance implications because prompt choices can materially shift model outputs even when model weights and inputs remain fixed. The authors propose a 'Prompt Commons' framework—a versioned, community-maintained repository of prompt templates with provenance metadata, licensing, and moderation logs—to make public-sector prompt collections transparent, contestable, and auditable. Using a pilot dataset of 443 human prompts (expanded to 3,317 through augmentation) collected with community partners in a large North American city, they illustrate three governance states and a negotiation-oriented ensemble method that aggregates stakeholder prompts into compromise recommendations. The work matters because existing governance tools such as model documentation and post-training alignment rarely address the prompt layer, leaving a critical point of influence over AI-driven public decisions ungoverned.
- AI policy
Research
Benchmarks for Vision-Language Models in Urban Perception Should Be Reliability-Aware and Negotiated
Rashid Mushkani
arXiv · 2026-05-30
This paper argues that standard benchmarks for vision-language models (VLMs) applied to urban street-scene analysis are flawed because they ignore annotator disagreement and abstention. The authors support this with a benchmark of 100 Montreal street scenes annotated across 30 dimensions by 12 participants from seven community organizations, combined with a zero-shot evaluation of seven VLMs, finding that model-human agreement tracks closely with how reliably humans agree among themselves on each dimension. For appraisal dimensions like Overall Impression, models and human annotators show distributional mismatches including differing rates of 'Not applicable' responses, highlighting that label spaces and scoring policies should be treated as negotiable rather than fixed. The paper calls on benchmark creators, model developers, and institutions to make uncertainty and benchmark assumptions explicit in evaluation reports intended to inform urban governance.
- AI policy
- Quality assurance