News & Research
The latest AI research and news with real-world stakes. Each item is sourced, dated, summarized in plain English and tagged by impact area, and checked against its source before it appears.
Kind
Impact area
5672 items
- ResearchPublic management and digital practices2026-06-12WEP
ASPECTS OF IMPLEMENTING ARTIFICIAL INTELLIGENCE IN THE PUBLIC ADMINISTRATION SYSTEM OF UKRAINE · A. Салов
This article analyzes how Ukraine is integrating artificial intelligence into its public administration, focusing on its existing digital platforms (Diia, Prozorro, and Trembita), international comparisons with national LLM initiatives from countries including Bulgaria, Greece, the Netherlands, Sweden, Singapore, and Albania, and the development of a regulatory roadmap aligned with EU AI Act requirements. The study identifies key barriers to scaling AI in government, including shortages of qualified personnel, cyber threats under martial law, legacy IT infrastructure, and risks of algorithmic bias. The authors argue that a bottom-up regulatory approach can gradually prepare public institutions, businesses, and society for broader AI adoption while minimizing operational risks. The findings offer strategic recommendations to strengthen Ukraine's institutional capacity in AI governance.
- ResearchOeconomica Jadertina2026-06-12WEQP
Artificial Intelligence and Green Transformation: Human Capital Upgrading, Green Finance, and ESG Assessment as Drivers of Sustainable Productivity · Harith Adnan Mohammed, Salam Anwar Ahmed, Najah Hawar Saeed Bazzaz
This study examines how AI adoption, green finance, and ESG performance jointly affect sustainable productivity among 1,000 Chinese publicly listed firms from 2012 to 2024. Using fixed-effects models, difference-in-differences, and mediation analyses, the authors find that AI adoption is positively associated with sustainable productivity (β=0.115), with green finance and ESG mediating roughly 18% of that association. Notably, small firms benefit disproportionately from AI adoption compared to larger firms, and the 2012 Green Credit Guidelines are linked to reduced measured productivity in polluting industries. The findings suggest sustainable transformation is a synergistic process combining technological, financial, and governance factors.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-06-12EQCP
AI Decision Governance Maturity Model (ADGMM) Version 1.0 — A Twelve-Level Framework for Evaluating the Maturity, Verifiability, and Completeness of AI Decision Governance Infrastructure · Harold Alberto Nunes Rodelo
This paper introduces the AI Decision Governance Maturity Model (ADGMM) Version 1.0, a twelve-level framework organized into four zones that defines how mature, verifiable, and complete an organization's AI decision governance infrastructure is. Unlike existing frameworks such as CMMI, NIST CSF, or ISO/IEC 42001, which assess organizational capability and process maturity, the ADGMM focuses on the cryptographic and protocol-level strength of governance evidence accompanying each individual AI decision, requiring that such evidence be independently verifiable by parties with no trust relationship with the issuing organization. The framework addresses specific governance gaps including the Authorization-to-Behavior Gap and the Mandate Failure Mode, and aligns with major regulatory standards including the EU AI Act, NIST AI RMF, ISO/IEC 42001, GDPR Article 22, and the OHADA Digital Framework. It is intended to serve regulators, auditors, enterprise buyers, and counterparties seeking a standardized, evidence-based scale for assessing AI governance maturity.
- ResearchInternational Journal of Advanced Multidisciplinary Research and Studies2026-06-12WEP
The Impact of Artificial Intelligence on Workforce Restructuring in Small and Medium-Sized Manufacturing Enterprises in Hanoi · Nguyen Dang Hoang, Vũ Thị Như Trang
This study examines how AI adoption affects workforce restructuring in 350 small and medium-sized manufacturing enterprises (SMEs) in Hanoi, Vietnam. Drawing on Human Capital Theory and Skill-Biased Technological Change Theory, the findings show that AI adoption significantly increases employee training and skill development activities while raising demand for workers with digital competencies, technological expertise, and data-related capabilities. The research provides empirical evidence for managers and policymakers developing human resource strategies to improve workforce adaptability amid digital transformation.
- ResearchCuadernos de RES PUBLICA en derecho y criminología2026-06-12QCP
Prevention and criminal offences in university education in the face of artificial intelligence · César Augusto Giner-Alegría
This paper examines how AI integration in university settings enables new forms of criminal offences—including automated plagiarism, identity theft, document forgery, fraud in automated assessments, and production of illegal content—and evaluates the regulatory response, with a focus on the EU AI Act (Regulation 2024/1689). Using qualitative methods including systematic literature review, documentary analysis, and expert interviews, the authors find a progressive increase in AI-facilitated offences and significant gaps between existing legal frameworks and institutional capacities. The paper proposes an interdisciplinary institutional response model integrating prevention, detection, sanctions, digital ethics training, and academic compliance mechanisms relevant to Spanish and European higher education.
- ResearchDrug Safety2026-06-12QCP
Analytic Misjudgment of Drug Safety Evidence and Causality: From the Prosecutor’s Fallacy and Simpson’s Paradox to Artificial Intelligence · Tarek A. Hammad, Justine Rochon
This narrative review examines how recurring analytic errors—including misinterpretation of conditional probabilities, Simpson's paradox, inappropriate comparator selection, and multiplicity issues—distort drug safety evidence and causality assessments in post-marketing pharmacovigilance. Using real-world case studies, the authors show how these misjudgments propagate across clinical, regulatory, and public domains and can drive premature or erroneous decisions. Critically, the paper warns that AI tools, if deployed without transparency, bias assessment, and clinical oversight, may amplify rather than reduce these vulnerabilities. The authors call for greater analytic discipline, uncertainty communication, evidence triangulation, and governance frameworks for AI-enabled drug safety evaluation.
- ResearcharXiv2026-06-11EQ
Knowledge-Based Zero-Replay Debugging of Multi-Agent LLM Traces · Dong Ho Kang, Hyeonjeong Cha, Daein Weon
This paper tackles the challenge of debugging multi-agent LLM systems, where identifying the few causally important events buried in long execution traces typically requires expensive counterfactual replay (rewind, edit, and re-run). The authors propose a zero-replay approach that compiles each trace into a structured event knowledge graph and uses a calibrated, gradient-boosted learning-to-rank predictor (BranchPoint-Latent) to predict which events a replay oracle would flag as high-effect — without actually running any replays. Evaluated across 37 trace families, the method raises per-trace localization (Branch Recall@5) from 0.73 to 0.93 on held-out families at zero oracle-replay cost. The result is a more auditable and cost-efficient system for AI reliability debugging, operating explicitly on the cost-accuracy frontier.
- ResearcharXiv2026-06-11QP
LLMs Contain Multitudes: How Deployment Context Reshapes Model-Level Preferences and Values · Filip Trhlik, Aoife O'Flynn, Angela Yu et al.
This paper tests whether the apparent value systems and preferences of large language models (LLMs) are stable properties of the models themselves or artifacts of deployment context. Across five LLMs and over 1.2 million pairwise decisions, the authors find that changing the high-level task framing—such as writing a Reddit post versus a news article—produces far greater variation in model preferences than standard robustness checks like prompt paraphrasing or temperature adjustment. For example, previously reported Global North favoritism in country rankings shifts systematically across contexts, and cardinal exchange rates between outcomes shift by a factor of 2.47 at the median. The key implication for safety and quality assurance is that model-level preference measurements are context-conditioned rather than fixed, meaning safety guarantees established under one deployment framing offer limited assurance in another.
- ResearcharXiv2026-06-11EQ
A Multi-Agent AI System for Automated High School Transcript Processing: Collaborative Document Analysis at Scale · Ben Torkian, Jun Zhou
This paper presents a multi-agent AI system designed to automate the processing of high school transcripts for college admissions offices. The architecture uses three specialized agents—a Pattern Recognition Agent, a Semantic Analysis Agent, and a Vision Intelligence Agent—coordinated by an Orchestration Agent, with GPA extraction serving as a quality control signal across agents. Evaluated on 40 real-world transcripts from 13 U.S. states, the system processed every document with 96.7% accuracy compared to expert manual review at approximately 45 seconds per transcript. The work demonstrates that multi-agent coordination can substantially reduce the manual workload and operational bottlenecks in admissions document processing while maintaining high accuracy.
- ResearcharXiv2026-06-11QP
Designing Safety-Constrained LLM Systems for Public Health Information Access · Ben Torkian, Jun Zhou
This paper presents a safety-constrained LLM system built for public health information access, specifically for maternal and child health resource navigation. The architecture combines domain-restricted retrieval-augmented generation (RAG), strict boundary enforcement to prevent medical advice, anonymous multi-user session management, and comprehensive audit logging. Scenario-based validation across in-scope, out-of-scope, and emergency queries shows consistent enforcement of safety constraints and reliable resource grounding, with an average response time of 5.3 seconds. The findings offer practical design guidance for deploying LLMs in healthcare and other settings requiring strict information boundaries and accountability.
- ResearcharXiv2026-06-11Q
Automated reproducibility assessments in the social and behavioral sciences using large language models · Tobias Holtdirk, Pietro Marcolongo, Anna Steinberg Schulten et al.
This study tests whether large language models (LLMs) can automate reproducibility assessments in the social and behavioral sciences, using 180 published studies with predefined claims. The LLMs matched the original study's qualitative conclusions in 80% of cases overall, and in 95% of a subset where human reanalysts were also available—comparable to the human reanalysts' 83% agreement rate. Effect size recovery within a ±0.05 Cohen's d tolerance was 24% overall and 40% in the human-comparison subset, again broadly similar to human reanalysts (28%). The findings suggest LLMs are best suited as a scalable screening tool to support systematic audits of empirical research rather than as a replacement for expert judgment.
- ResearcharXiv2026-06-11EQ
One Polluted Page Is Enough: Evaluating Web Content Pollution in Generative Recommenders · Minghao Luo, Liang Chen
This paper introduces FORGE, a benchmark for measuring how easily search-augmented large language models (LLMs) can be manipulated into recommending fake products when they retrieve polluted web content such as fake reviews or promotional pages. Testing 12 commercial and open-weights LLMs across 225 real-world products in 15 categories, the study finds that even a single polluted page causes models to recommend fake products up to 27% of the time, rising to 73.8% when all top-3 retrieved pages are replaced. Notably, reasoning-enhanced models are not protected and may even fabricate social proof to justify false recommendations, while proposed defenses like skepticism prompting and consensus filtering each carry their own risks. The findings highlight a significant quality and consumer-trust risk in AI-powered recommendation systems that rely on live web retrieval.
- ResearcharXiv2026-06-11WE
Multi-Agent Reinforcement Learning from Delayed Marketplace Feedback for Objective-Weight Adaptation in Three-Sided Dispatch · Haochen Wu, Yi Hou, Shiguang Xie
This paper presents a deployed multi-agent reinforcement learning system at DoorDash that adapts dispatch objective weights in a three-sided food-delivery marketplace (customers, couriers, and merchants) using delayed operational feedback. Rather than replacing the combinatorial assignment optimizer, the system learns store-level policies from logged data to select discrete multipliers that shift the optimizer's tradeoff between delivery quality and batching efficiency. A shared value function is trained with Double Q-learning and a conservative regularizer to mitigate out-of-distribution overestimation, enabling safe offline policy learning under noisy and delayed signals. In a production switchback experiment, the offline-trained policy improved batching and reduced courier-side time costs without degrading customer-facing delivery quality.
- ResearcharXiv2026-06-11Q
SEVRA-BENCH: Social Engineering of Vulnerabilities in Review Agents · Rui Melo, Riccardo Fogliato, Sean Zhou et al.
SEVRA-BENCH is a benchmark that tests whether LLM-based automated code-review agents can detect vulnerability-introducing code when attackers pair it with persuasive, socially engineered pull request narratives. The benchmark reverses historical vulnerability fixes to recreate vulnerable code and wraps each change in one of 15 social-engineering framings—varying urgency, authority appeals, and fabricated prior approvals—across roughly 1,000 adversarial pull requests drawn from MITRE's 2025 top 10 most dangerous software weaknesses. Evaluating 8 review agents, the study finds that these agents are susceptible to narrative manipulation, revealing a significant gap in their security capabilities. The findings matter because LLM review agents are increasingly used to gate code merges into shared repositories, and this work shows that adversarial framing alone can undermine that security layer.
- ResearcharXiv2026-06-11EQ
Understanding the Rejection of Fixes Generated by Agentic Pull Requests -- Insights from the AIDev Dataset · Mahmoud Abujadallah, Ali Arabat, Mohammed Sayagh
This paper analyzes why AI-generated pull requests are rejected in software projects, finding that 46.41% of fixes proposed by agents such as Copilot, Devin, Cursor, and Claude are not merged. Through a qualitative study of 306 non-merged pull requests, the authors identify 14 rejection reasons across four categories: incorrect implementation, CI pipeline or test failures, agent inability to complete the task, and low priority. The findings highlight that better guidance on fix approaches, constraints, and validation strategies—plus smarter task prioritization—could reduce wasted human review effort and computational resources.
- ResearcharXiv2026-06-11EQ
Toward Instructions-as-Code: Understanding the Impact of Instruction Files on Agentic Pull Requests · Ali Arabat, Mohammed Sayagh
This study examines whether instruction files—documents that guide AI agents like GitHub Copilot on how to navigate a project, run tests, and follow best practices—actually improve the quality of AI-generated pull requests. Analyzing 15,549 agentic PRs from 148 projects in the AIDev dataset, the researchers find that adding instruction files does not consistently improve outcomes: 27.7% of projects saw their merge rate increase by at least 20%, while 26.35% saw it decrease, with similar mixed results for code churn, number of modified files, and time to merge. Projects that did achieve higher merge rates tended to have longer, more structured instruction files with more sections and sub-sections. The authors argue these findings motivate treating instruction-file development as a formal software engineering discipline, coining the concept 'Instructions-as-Code.'
- ResearcharXiv2026-06-11EQ
Evaluation Sovereignty in Metadata-Driven Classification: A Multi-Track Framework for Weakly Supervised Information Systems · Raymond Vasquez
This paper introduces the concept of 'evaluation sovereignty'—the degree to which machine learning performance metrics are independent of how labels were generated—and proposes a multi-track evaluation framework to test it. Using hierarchical multi-label classification on large-scale scientific metadata, the authors show that models appearing strong under operational ('silver') label evaluation collapse dramatically under independent ('gold') evaluation, with Micro-F1 dropping from approximately 0.54 to 0.03. The findings suggest that widely reported performance metrics in weakly supervised systems may reflect alignment with labeling processes rather than genuine predictive capability. This has direct implications for how AI systems in information management and enterprise settings are audited and validated.
- ResearcharXiv2026-06-11QP
Mod-Guide: An LLM-based Content Moderation Feedback System to Address Insensitive Speech toward Indigenous Ethnic and Religious Minority Communities · Dipto Das, Achhiya Sultana, Ankit Singh Chauhan et al.
Mod-Guide is an LLM-based content moderation system designed to detect culturally insensitive speech targeting Bangladesh's Hindu and Chakma minority communities, where standard moderation tools often miss implicit forms of harm such as erasure and misrepresentation. The researchers co-created a culturally grounded corpus of insensitive speech with community members and integrated their lived-experience narratives into moderation pipelines using retrieval-augmented generation (RAG). Mixed-method evaluations with both minority and majority participants showed that RAG-enhanced moderation responses were more contextually accurate and were perceived differently across ethnic lines. The work advances AI ethics and human-computer interaction by embedding restorative justice and hermeneutical inclusion into content moderation design.
- ResearcharXiv2026-06-11QP
Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web Agents · Zihao Wang, Yiming Li, Yutong Wu et al.
This paper introduces StakeBench (SBC), a benchmark that evaluates prompt-injection attacks on LLM-powered web agents from a stakeholder-centric perspective, distinguishing harms experienced by different parties such as users, sellers, and platforms. Unlike existing attack-centric benchmarks that focus only on technical feasibility, SBC decomposes attacks by objective and measures both outcome- and process-level consequences. Results show that no current agent reliably resists any tested attack objective, with failures spanning distinct modes including stealthy parasitism, misaligned disruption, and compounded failure. The findings highlight the need for stakeholder-aware security evaluation before deploying LLM-based web agents in real-world settings.
- ResearcharXiv2026-06-11EQ
Mining Architectural Quality Under Agentic AI Adoption: A Causal Study of Java Repositories · Oliver Aleksander Larsen, Mahyar T. Moghaddam
This study examines whether agentic AI coding tools (so-called 'vibe coding') causally affect software architecture quality in open-source Java projects. Analyzing 151 repositories over 13 months using a staggered difference-in-differences design, the researchers find that total architectural smell counts are essentially unchanged (+1.1%, p=0.82) even as codebases grow by +12.8% (p=0.003), producing a 6.7% decline in architectural smell density (p=0.004) that reflects larger system size rather than genuine architectural improvement. The key methodological finding is that density-normalized metrics can be misleading when AI adoption causes code growth, and raw counts with explicit decomposition are needed for accurate causal inference. This matters for software engineering teams and enterprises evaluating AI coding tools, as apparent quality improvements may be statistical artifacts rather than real gains.
- ResearcharXiv2026-06-11QC
ERTS: Adversarial Robustness Testing of Ethical AI via Semantic Perturbation in a Bounded Consequence Space · Pratyush Chaudhari
This paper presents ERTS (Ethical Robustness Testing System), a formal framework for adversarially testing AI systems used in high-stakes ethical contexts such as healthcare triage, autonomous vehicle control, and employment screening. ERTS encodes ethical dilemmas into a 22-dimensional Ethical Consequence Space, applies 17 semantic perturbation functions under 6 validity constraint classes, and measures decision instability via a 4-component Ethical Instability Index. Evaluating 4 baseline models and 2 production LLMs (Gemini 2.0 Flash and Llama 3.2) across 1,500 adversarial test cases, the authors find that only 33% of models achieve assessment clearance, with Llama 3.2 proving especially vulnerable to fairness corruption and information degradation attacks. The framework addresses a gap in pre-deployment robustness assessment for AI ethical reasoning, with direct implications for certification and quality assurance of AI in consequential domains.
- ResearcharXiv2026-06-11EQ
Multi-Field Hybrid Retrieval-Augmented Generation for Maritime Accident Root Cause Analysis · Seongjin Kim, Sungil Kim
This paper presents a multi-field hybrid Retrieval-Augmented Generation (RAG) framework designed to automate root cause analysis (RCA) for maritime accidents, drawing on 13,329 Korea Maritime Safety Tribunal reports spanning 1971–2025. The system structures adjudication records into 'incident cards' indexed across Summary, Causes, and Disposition fields, and uses a field-aware hybrid retrieval strategy combining sparse and dense rankings via Reciprocal Rank Fusion. Experiments show retrieval recall improves substantially (NormRecall@100 from 0.18 to 0.55) and LLM-generated RCA quality rises (judge score from 3.34 to 3.72) when grounded on retrieved precedents versus an LLM-only baseline. The findings suggest this approach can meaningfully streamline maritime safety investigation workflows by enabling faster precedent search and more consistent, evidence-based report drafting.
- ResearcharXiv2026-06-11QP
Hallucination in Medical Imaging AI: A Cross-Modality Analytical Framework for Taxonomy, Detection, and Mitigation under Regulatory Constraints · Omar Alshahrani, Muzammil Behzad
This paper presents a cross-modality analytical framework for understanding, detecting, and mitigating hallucinations in medical imaging AI — defined as clinically plausible but factually incorrect outputs such as fabricated anatomical structures, incorrect laterality, or invented measurements in radiology reports. The authors synthesize peer-reviewed studies, benchmark datasets, and FDA regulatory guidance across five imaging modalities, finding that general-purpose foundation models can outperform narrowly fine-tuned medical models on hallucination benchmarks because domain-specific fine-tuning may introduce overfitting-induced confabulation. They identify that physics-informed architectural constraints, Chain-of-Thought prompting, and human-in-the-loop safeguards each address distinct failure modes and are most effective in combination, while radiologist oversight remains essential given the high rate of AI-generated flags requiring expert correction. All findings are mapped to FDA's Total Product Lifecycle and Predetermined Change Control Plan frameworks, framing hallucination management as an ongoing regulatory obligation rather than a one-time pre-deployment check.
- ResearcharXiv2026-06-11Q
Mental-R1: Aligning LLM Reasoning for Mental Health Assessment · Xin Wang, Boyan Gao, Yibo Yang et al.
This paper introduces Mental-R1, a large language model fine-tuned with a novel reinforcement learning framework called Cognitive Relative Policy Optimization (CRPO) for mental health assessment tasks including anxiety, depression, and suicide risk. CRPO extends group relative policy optimization by adding stage-dependent uncertainty modeling and stage-wise entropy regularization, which encourages broad exploration early in reasoning and confident decisions later—mimicking human cognitive processes grounded in cognitive appraisal theory. Evaluated on 8 mental health datasets, CRPO achieves an average improvement of 10.4 percentage points in weighted F1-score over the best reinforcement learning baseline, with Mental-R1 showing clear advantages on reasoning-intensive cases. The work matters because it demonstrates that aligning LLM reasoning to human cognitive assessment processes can produce more reliable and interpretable mental health screening tools.
- ResearcharXiv2026-06-11WE
Getting Better at Working With You: Compiling User Corrections into Runtime Enforcement for Coding Agents · Yujun Zhou, Kehan Guo, Haomin Zhuang et al.
This paper identifies a key failure mode in LLM-based coding agents: user corrections made in one session are frequently ignored in future sessions, even when memory systems are in place. The authors introduce TRACE (Test-time Rule Acquisition and Compiled Enforcement), a pipeline that mines user corrections from chat, rewrites them as atomic rules, and enforces those rules as runtime checks before future tasks complete. Experiments show that TRACE substantially reduces preference violations compared to baselines—cutting violations from 100% to as low as 2% on out-of-distribution coding tasks—while existing memory systems like Mem0 still leave 57.5% of applicable preference checks violated. This matters for workforce productivity and enterprise deployments of AI agents, where users should not have to repeatedly restate the same corrections across sessions.