News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
Enhancing leadership effectiveness through artificial intelligence adoption: A literature review and exploratory research in the Egyptian manufacturing sector
Attia Hussien Gomaa
Human Resources Management and Services · 2026-05-27
This study examines how AI adoption affects leadership effectiveness in Egyptian manufacturing, drawing on a literature review and exploratory research with 60 senior leaders across 15 firms in eight industries. Findings reveal a fragmented adoption landscape where AI use is highest in strategic planning and customer analytics but lowest in workforce analytics and R&D, with key barriers including legacy systems, limited data infrastructure, and organizational resistance. The authors developed a phased, KPI-driven framework emphasizing cross-functional collaboration, ethical oversight, and iterative implementation to convert isolated AI initiatives into sustainable strategic enablers. The study offers actionable guidance for advancing AI-enabled leadership in developing-economy manufacturing contexts.
- Workforce
- Enterprise
- AI policy
Research
Operational AI Deployment Assurance: Governance-State Orchestration Under Threshold-Sensitive Deployment Conditions - A Governance Framework for High-Stakes AI Systems
Khalid Adnan Alsayed
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-27
This paper presents Operational AI Deployment Assurance (OADA), a governance framework that translates fairness metrics, subgroup instability, and operational uncertainty into structured deployment-readiness decisions for high-stakes AI systems. Unlike existing approaches that rely on static reporting or post-hoc auditing, OADA introduces constructs such as Deployment Assurance Scores, Threshold Stability Zones, and Governance Escalation States to actively govern when and how AI systems move from evaluation into real-world use. The framework is demonstrated on facial recognition systems with extensions to healthcare AI, showing that systems can appear acceptable under isolated metrics while still exhibiting instability that undermines deployment readiness. This matters because it positions governance as an operational control layer rather than a passive monitoring exercise, directly informing certification and policy decisions for high-stakes AI deployments.
- Quality assurance
- Certifications
- AI policy
Research
Ethical implications of AI in the fashion industry for trend forecasting and garment design development
Rajeev Ranjan Mahto, Carolina Quintero Rodriguez
AI and Ethics · 2026-05-27
This mixed-methods study of 93 fashion industry professionals and 15 semi-structured interviews examines the ethical implications of AI adoption in trend forecasting and garment design. Findings show that while AI improves operational efficiency and data-driven decision-making, it also introduces algorithmic bias, opacity, and ambiguous accountability—particularly by reinforcing dominant aesthetics and marginalising non-Western perspectives due to biased training data. Most participants viewed AI as a complementary tool that transforms rather than replaces creative roles, though concerns about labour displacement were noted. The authors call for transparent, inclusive, and context-sensitive AI governance frameworks to support responsible innovation in fashion.
- Workforce
- Enterprise
- AI policy
- Quality assurance
Research
How Artificial Intelligence Pilot Zones Enhance Corporate Green Resilience? Evidence from China’s Listed Firms with Double Machine Learning
Yuzeng Xin, Xihao Zeng, Jingru Gao et al.
Sustainability · 2026-05-27
Using panel data from Chinese A-share listed firms (2015–2023), this study treats China's National Pilot Zone for AI Innovation Applications as a quasi-natural experiment and applies a double machine learning difference-in-differences framework. Results show the pilot-zone policy significantly increases corporate green resilience by approximately 32%, with stronger effects among high-tech firms, non-heavily polluting industries, regulated sectors, and large enterprises. The policy works through four channels—accelerating green innovation, enhancing supply-chain efficiency, alleviating financing constraints, and reducing operating costs—with innovation and supply-chain efficiency playing dominant roles. The findings provide causal evidence that AI-oriented place-based policies can help firms sustain green development under climate and regulatory pressures, informing China's 'Digital China' and 'Dual Carbon' agendas.
- Enterprise
- AI policy
Research
Operational AI Deployment Assurance: Governance-State Orchestration Under Threshold-Sensitive Deployment Conditions -- A Governance Framework for High-Stakes AI Systems
Khalid Adnan Alsayed
arXiv (Cornell University) · 2026-05-27
This paper introduces Operational AI Deployment Assurance (OADA), a governance framework that translates fairness disagreement, subgroup instability, and threshold sensitivity into structured deployment-readiness decisions for high-stakes AI systems. Unlike existing approaches that rely on static metric reporting and post-hoc auditing, OADA connects evaluation outputs to deployment-state interpretation, escalation, and remediation-aware control. Applied to facial recognition systems and extended to healthcare AI, the framework demonstrates that systems can appear acceptable under isolated fairness or performance metrics while still exhibiting instability that undermines deployment readiness. The work positions operational deployment assurance as a governance layer between evaluation and real-world AI deployment.
- AI policy
- Quality assurance
- Certifications
Research
Got a Secret? LLM Agents Can't Keep It: Evaluating Privacy in Multi-Agent Systems
Aman Priyanshu, Supriti Vijay, Esha Pahwa
arXiv · 2026-05-26
This paper introduces a large-scale multi-agent simulation platform where thousands of LLM agents interact over a simulated month to evaluate privacy as a safety concern under social pressure. The researchers find that multi-turn social interactions dramatically amplify privacy violations compared to single-turn evaluations (from 19.95% to 45.30% leakage across OpenAI models), and that leakage is socially contagious—agents are 8 times more likely to disclose sensitive information after observing a peer do so. Even explicit privacy instructions fail to eliminate the risk, leaving leakage rates above 37.8% with safeguards in place. The findings suggest that standard single-turn safety benchmarks systematically underestimate real-world privacy risks in multi-agent deployments, with social context alone being sufficient to elicit sensitive disclosures.
- AI policy
- Quality assurance
Research
Hallucination Behavior in Multimodal LLMs Across Agricultural Image Interpretation and Generation Tasks
Partho Ghose, Al Bashir, Prem Raj et al.
arXiv · 2026-05-26
This study examines hallucination behavior in multimodal large language models applied to agricultural imaging, covering both image-to-text interpretation (e.g., identifying crop stresses) and text-to-image generation (e.g., synthesizing field scenes). In image interpretation tasks, models such as Gemma, LLAVA, Qwen, and MiniCPM achieved zero-shot accuracy of 63–75%, improving to 86.8% with few-shot prompting but still showing false detections and missed infections. In text-to-image generation, advanced models including GPT-5 and Gemini 2.5 Flash produced up to 91% biologically inconsistent scenes under relaxed prompt constraints. The findings highlight significant reliability concerns for LLM-based agricultural imaging platforms and point to the need for domain-informed evaluation and mitigation of hallucinations.
- Quality assurance
Research
RULER: Representation-Level Verification of Machine Unlearning
Georgina Cosma, Axel Finke
arXiv · 2026-05-26
RULER introduces representation-level verification metrics for machine unlearning, addressing a gap where models can pass standard output-level checks (membership inference, retain accuracy, forget-set accuracy) while still encoding forgotten records in intermediate layers. The oracle-comparative metric M2 finds significant residuals in 10 of 12 tested conditions across four approximate unlearning methods, and a fifth method (Bad Teacher) shows the same pattern despite a different forgetting mechanism. An oracle-free metric M4 serves as a pre-unlearning diagnostic across tabular, image, clinical text, and face-identity settings, detecting identity-level memorisation in face recognition models that no tested method fully erases. These findings matter for quality assurance and certification of AI systems where verifiable data removal is required by regulation or policy.
- Quality assurance
- Certifications
Research
Keyphrase Generative Representation of Youth Crisis Conversations Beyond Static Taxonomies
Abeer Badawi, Will Aitken, Lydia Sequeira et al.
arXiv · 2026-05-26
This paper introduces Keyphrase Generative Representation (KGR), a constrained large language model approach for analyzing youth mental health crisis conversations beyond fixed-label taxonomies. Using 703,975 de-identified Kids Help Phone SMS conversations (2018–2023), the authors expanded an existing 19-label issue taxonomy into a 39-label hierarchical schema and evaluated KGR across 129 conversations with 387 expert annotations. The expanded taxonomy achieved an expert consensus accuracy of 0.96, with 81% of generated keyphrases accurately reflecting conversation content and 74% improving clarity; a topic-retrieval workflow using KGR increased accuracy from 0.25 to 0.70. KGR surfaced identity-linked themes absent from fixed taxonomies—such as immigration problems and caregiver burden—suggesting this hybrid generative approach can better capture emerging and culturally grounded patterns of youth distress for crisis response organizations.
- Workforce
- Quality assurance
Research
Algorithmic Monocultures in Hiring
Rishi Bommasani, Sarah H. Bana, Kathleen A. Creel et al.
arXiv · 2026-05-26
This paper investigates the phenomenon of 'algorithmic monoculture' in hiring, where a small number of algorithm vendors supply screening tools to many employers, potentially causing the same individuals and racial groups to be systematically rejected across multiple job applications. Analyzing a novel dataset of 3 million applicants submitting 4 million applications—all screened by algorithms from the same vendor—the authors find significant racial disparities: 14.74% of applications from Asian applicants and 25.87% from Black applicants are submitted to positions where the algorithm adversely impacts those groups under U.S. employment discrimination standards. The study also finds that 4% of applicants who apply to 10 positions are recommended for rejection from all of them, a rate higher than expected by chance, suggesting individuals face systematically homogeneous outcomes. The findings highlight that applicants may need to apply very broadly just to have their applications reviewed by a human, raising serious concerns about fairness and discrimination in AI-driven hiring.
- Workforce
- AI policy
Research
FinHarness: An Inline Lifecycle Safety Harness for Finance LLM Agents
Haoxuan Jia, Yang Liu, Bin Chong et al.
arXiv · 2026-05-26
FinHarness is an inline safety harness for finance-domain LLM agents that monitors and intervenes during multi-step workflows rather than only after they complete. It combines a Query Monitor for intent and drift detection, a Tool Monitor for evaluating individual tool calls, and a Cascade module that routes risk checks to either a lightweight or advanced LLM judge adaptively. On the FinVault benchmark, the system reduces the attack success rate (ASR) from 38.3% to 15.0% while preserving most legitimate approvals and using 4.7× fewer expensive judge calls than always using the advanced model. This matters for enterprise and quality-assurance applications where autonomous AI agents must safely handle high-stakes financial workflows without blocking legitimate business processes.
- Enterprise
- Quality assurance
Research
Grounded Cache Routing for Retrieval-Augmented Generation: When Is It Safe to Reuse an Answer?
Syed Huma Shah
arXiv · 2026-05-26
GroundedCache is a cache routing system for retrieval-augmented generation (RAG) that addresses the safety problem of reusing cached answers. Rather than optimizing for cache hit rate, the system asks when it is safe to reuse a cached answer by requiring four simultaneous conditions: query similarity, retrieved-evidence overlap, source-version validity, and lexical or judge-based support of the cached answer by freshly retrieved evidence. Evaluated across 12,000 real LLM generations on two datasets, GroundedCache reduces the 'unsafe-served rate' to 0.0% on HotpotQA regimes (versus 15–35% under naive caching) and to 1.5% on document-drift scenarios (versus 51.5%), while keeping end-to-end latency within 1.04–1.07x of a no-cache baseline. This matters for enterprise RAG deployments where serving stale or incorrect cached answers poses reliability and trust risks.
- Enterprise
- Quality assurance
Research
Benchmarks are Not Enough: RAMP for Runtime Assessing of Agentic Models in Production Systems
Yipeng Ouyang, Xin Huang, Bingjie Liu et al.
arXiv · 2026-05-26
This paper introduces RAMP, a production-grounded framework for evaluating large language model (LLM) agents on realistic, long-horizon software engineering tasks rather than static benchmarks. Testing 15 mainstream models, the authors find severe capability degradation in real workflows: task completion rates collapse from 100% at the first stage to just 20% at the final stage, and no model completes the full pipeline. The framework also exposes systematic failure propagation and computational cost differences of up to three orders of magnitude among comparable models—gaps that conventional isolated benchmarks largely miss. These findings argue for continuous, runtime-observable evaluation methods to better reflect how AI agents actually perform in production environments.
- Quality assurance
- Enterprise
Research
Ground Truths in Suicide Research: The Current State of AI-Based Suicide Detection in Social Media
Yaakov Ophir, Ofri Hefetz, Refael Tikochinski et al.
arXiv · 2026-05-26
This paper synthesizes 195 studies on AI-based suicide risk detection from social media, drawing on an umbrella review of 22 systematic reviews covering work up to 2022 and an ongoing review of more recent literature. It finds consistent methodological weaknesses: research is concentrated on a few platforms, relies on English-language textual data, reuses similar datasets, and—most critically—uses indirect labeling strategies that infer 'ground truth' from linguistic markers or community membership rather than direct, individual-level clinical validation of suicide risk. Because of this, models are typically classifying distress-related posts rather than genuinely identifying at-risk individuals, meaning people who do not express suicidal content online are likely missed. The authors caution that reported gains in model performance should not be over-interpreted, and that meaningful progress requires better alignment between model predictions and real-world suicide risk.
- Quality assurance
- AI policy
Research
HARP: Measuring Harm Amplification in Multi-Agent LLM Systems
Md Hafizur Rahman, Zafaryab Haider, Tanzim Mahfuz et al.
arXiv · 2026-05-26
HARP (Harm Amplification through Role Perturbation) is a trace-first evaluation methodology for measuring how localized attacks in multi-agent LLM systems propagate into broader system-level harm. The framework tracks paired clean and perturbed executions across a seven-agent finance-oriented system, recording specialist outputs, tool calls, memory reads/writes, guard events, and latency to compute a harm amplification ratio (H_global/H_local). Key findings show that single-specialist compromise produces the strongest amplification, shared-context corruption yields the highest attack success, and temporal persistence produces the largest malicious impact, while the trace-consistency defense IntegrityGuard achieves the lowest attack success and global harm but with utility and cost trade-offs. The work argues that secure multi-agent evaluation must measure not only attack bypass rates but also how orchestration spreads harm beyond the original attack point.
- Quality assurance
- AI policy
Research
Queue & AI: When Faster Tasks Slow Down the Workflow
Silvia Bartolucci, Pierpaolo Vivo
arXiv · 2026-05-26
This paper argues that standard productivity metrics for AI tools—such as average task completion time—can be misleading in workflow settings where tasks queue up for scarce human attention. The authors formalize a 'variance wedge' concept using a queueing model, showing that AI's speed gains on individual tasks can mask system-level slowdowns when AI errors escape review and return as costly rework. Analytically, they find that under congestion, reviewers rationally reduce scrutiny of AI outputs precisely when oversight matters most, and that AI stabilizes an overloaded workflow only when both the fraction of AI-handled tasks exceeds a critical threshold and human attention for review plus expected rework is lower than for manual completion. The findings suggest AI deployment should be assessed by its effects on congestion, rework rates, and robustness of human oversight—not just average task speed.
- Workforce
- Enterprise
- AI policy
Research
An investigation of AI integration in sound designer workflows and experiences
Nelly Garcia, Joshua Reiss
arXiv · 2026-05-26
This mixed-methods study surveyed 76 professional sound designers and interviewed 20 industry practitioners to examine how AI tools are being integrated into audio production workflows. Findings reveal that current AI tools perform adequately in fast-consumption media contexts but fall short of the narrative sophistication required for high-end sound design in films and immersive experiences. Practitioners prefer assistive, task-specific AI applications—particularly for audio restoration and library management—over fully generative end-to-end systems. The paper offers recommendations for developers to build more informed AI tools better aligned with professional creative needs.
- Workforce
- Enterprise
Research
Faults and Pitfalls in Implementing the Right to be Forgotten
Chen Sun, Nikolas Guggenberger, Supreeth Shastri
arXiv · 2026-05-26
This paper examines the practical challenges of implementing the Right to be Forgotten (RTBF) under GDPR, noting that regulators issued 205 RTBF violations in the first five years of GDPR—roughly one failure every nine days on average. The authors identify computing uncertainties and risks that make RTBF difficult to enforce, and propose a two-phase approach designed to bridge the gap between legal requirements and computing practice. They demonstrate that their technique could have avoided 80% of RTBF violations in GDPR's sixth year, identify six long-standing computing and data management practices that act as anti-patterns for RTBF, and validate their approach by integrating RTBF capability into Elasticsearch, a popular open-source search engine. The work matters for policy and enterprise contexts because it provides concrete, measurable guidance for organizations struggling to comply with one of GDPR's most prominent data rights obligations.
- AI policy
- Enterprise
Research
Grounding Text Embeddings in Stakeholder Associations
Jonathan Rystrøm, Sofie Burgos-Thorsen, Zihao Fu et al.
arXiv · 2026-05-26
This paper introduces the 'Stakeholder Grounding Exercise,' a method for evaluating whether neural text embeddings align with the semantic distinctions that domain experts actually care about. In a case study on Danish policy issues, neural embeddings were found to be substantially less reliable than human experts by 19–26 percentage points, and this misalignment directly degraded downstream clustering quality (Spearman ρ=0.9 between exercise ranking and cluster quality). A replication study on US Federal AI use cases confirmed a similar gap (16 percentage points), showing the finding holds across languages, domains, and expert communities. The method offers a practical tool for validating whether embedding models are fit for purpose before being used in high-stakes analytical workflows.
- Quality assurance
- AI policy
Research
Detecting Is Not Resolving: The Monitoring Control Gap in Retrieval Augmented LLMs
Zhe Yu, Wenpeng Xing, Chen Ye et al.
arXiv · 2026-05-26
This paper investigates a critical safety flaw in retrieval-augmented large language models (LLMs): the 'monitoring-control gap,' where models can detect contradictory or dangerous evidence in retrieved documents but still fail to act safely on that awareness. Using a multi-turn document accumulation protocol across four model families (1.5B–32B parameters) and over 50,000 turn-level evaluations, the authors show that single-turn safety benchmarks systematically overestimate real-world RAG safety, and that a model's ability to acknowledge epistemic conflict is uncorrelated with whether it resolves that conflict safely. Mechanistic analysis via hidden-state probing and attention analysis suggests danger-relevant information is internally represented and attended to, yet fails to constrain final output behavior — pointing to action selection as the core failure point. These findings matter for any high-stakes deployment of RAG systems, where evidence quality directly determines whether AI-driven recommendations are safe.
- Quality assurance
- AI policy
Research
Semantic Robustness Probing via Inpainting: An Interactive Tool for Safety-Critical Object Detection
Nico Steckhan, Krutarth Prajapati, Weija Shao et al.
arXiv · 2026-05-26
SemProbe is an interactive tool that tests object detectors in safety-critical settings by using diffusion-based inpainting to generate semantically meaningful image variations rather than simple pixel-level corruptions. Users upload deployment images, define masks manually or automatically, and select domain-relevant factors to probe how detection models respond to controlled scene changes. The system automatically runs model inference on each generated variant, displays annotated before/after comparisons with performance deltas, and logs all probes as structured artifacts to support traceable safety evaluation workflows. The authors demonstrate the tool on hand detection for dimension saws, targeting factors derived from insurance-oriented test criteria.
- Quality assurance
- Certifications
Research
When Medical Safety Alignment Fails: A Benchmark for Evaluating LLMs on High-Risk Medical Queries
Yige Li, Jun Sun, Wei Zhao et al.
arXiv · 2026-05-26
This paper introduces MedHarm, a benchmark of 1,100 medically grounded queries across 10 safety-critical categories (including toxicology, pharmacology, covert poisoning, anesthesia, and fetal harm) designed to test whether large language models handle high-risk medical prompts safely. Evaluating 15 LLMs and 4 guardrail models, the authors find a substantial gap between apparent alignment and actual medical safety: aligned models can still produce unsafe or actionable responses, medical fine-tuning can amplify harmful specificity, and external guardrails introduce brittle blocking while weakening safe helpfulness. The study concludes that medical safety cannot be inferred from general alignment or medical capability alone, underscoring the need for domain-specific stress testing before deploying LLMs in safety-critical clinical contexts.
- Quality assurance
- AI policy
Research
Prompt Injection Detection is Regime-Dependent: A Deployment-Aware Evaluation with Interpretable Structural Signals
Akindoyin Akinrele, Shreyank N Gowda
arXiv · 2026-05-26
This paper evaluates prompt injection detection—a key security threat for large language models—across a wide range of realistic deployment conditions, including out-of-distribution settings and thresholded deployment metrics. The authors compare lexical, semantic, structural, and transformer-based detectors, and introduce interpretable structural signals capturing hierarchy overrides, system prompt spoofing, role redefinition, and evasion patterns. Results show that detection performance is highly regime-dependent and sensitive to threshold selection, with no single model dominating across all settings; transformer-based models perform best overall, while structural signals provide consistent but modest gains in harder scenarios. The findings highlight a gap between ranking performance and real-world deployment effectiveness, underscoring the need to evaluate defences under realistic operational constraints.
- Quality assurance
- AI policy
Research
JuICE: A Benchmark for Evaluating LLM-Judge in Identifying Cultural Errors
Jiho Jin, Junho Myung, Juhyun Oh et al.
arXiv · 2026-05-26
JuICE introduces a multilingual benchmark of 7,470 span-level annotations covering cultural and linguistic errors in long-form LLM responses across four countries (the United States, South Korea, Indonesia, and Bangladesh). The paper finds that even the strongest LLM-judge reaches only an F1 of 0.52 on erroneous span detection and consistently misses 'thick' cultural errors that local residents readily identify. This reveals a critical gap in current LLM evaluation frameworks, which treat culture as a flat set of facts rather than accounting for the depth and situatedness of cultural meaning. The findings have direct implications for how LLM outputs are assessed for quality and appropriateness in diverse cultural contexts.
- Quality assurance
Research
KZ-SafetyPrompts: A Kazakh Safety Evaluation Prompt Dataset for Large Language Models
Wajdi Zaghouani, Shimaa Amer Ibrahim, Aruzhan Muratbek et al.
arXiv · 2026-05-26
KZ-SafetyPrompts introduces a dataset of 5,717 Kazakh-language prompts designed to evaluate the safety behavior of large language models (LLMs) across eleven risk categories, including self-harm, violence, child exploitation, and radicalization. The prompts are written natively in Kazakh (Cyrillic) and include English translations for cross-lingual analysis. Baseline testing with GPT-4o reveals an overall refusal rate of only 28.2%, ranging from 5.5% to 53.8% across categories, demonstrating that Kazakh prompts expose safety gaps not captured by English-only evaluations. This work highlights the need for multilingual safety benchmarks and provides a structured resource—with documented writing protocols, labeling procedures, and quality-control steps—to support broader LLM safety assessment pipelines.
- Quality assurance
- AI policy