News & Research
The latest AI research and news with real-world stakes. Each item is sourced, dated and summarized in plain English, tagged by impact area where one fits, and its summary is checked against the text it was written from.
8248 items
- ResearcharXiv2026-06-30Quality assurance · Algorithms & Automated Decisions
EgoSafetyBench: A Diagnostic Egocentric Video Benchmark for Evaluating Embodied VLMs as Runtime Safety Guards · Siddhant Panpatil, Arth Singh, Mijin Koo et al.
EgoSafetyBench introduces a 1,200-scenario egocentric video benchmark designed to evaluate vision-language models (VLMs) used as real-time safety monitors for embodied agents such as home and factory robots. The benchmark tests two key capabilities: recognizing genuinely unsafe moments versus routine but superficially alarming activity, and detecting when visible in-scene text (signs, stickers, labels) misrepresents the physical situation. Evaluating ten open- and closed-source VLMs, the study finds that while models can generally identify hazard-containing videos, they frequently miss specific dangerous moments—especially contextual hazards—and misleading signs degrade all tested models, with vulnerable ones missing up to a third of hazards and robust ones over-intervening on safe content. The findings reveal that apparent safety robustness often reflects indiscriminate alarming rather than genuine physical reasoning, raising important concerns for deploying VLMs as safety-critical guards in real-world settings.
- ResearcharXiv2026-06-30Quality assurance · Health
Readable but Not Controllable: Neuron-Level Evidence for Medical LLM Hallucination · Vijay Vankadaru, Asha Matthews, Tanya Roosta et al.
This paper investigates whether internal neural representations associated with hallucination in medical large language models (LLMs) can be used not just to detect hallucinations but also to control them. Testing four open-source models across multiple medical question-answering datasets, the authors find that a simple probe reliably detects hallucination with AUROC scores between 0.77 and 0.86, and that this signal is broadly distributed across hundreds of neurons rather than localized to a few. Critically, the study reveals a sharp gap between detectability and controllability: the same internal structure that makes hallucination easy to detect does not translate into reliable neuron-level control, meaning steering the most associated neurons fails to mitigate hallucinations. These findings suggest that reducing hallucination in medical LLMs requires more than identifying the right neurons, pointing to a fundamental separation between what representations reveal and what they allow us to change.
- ResearcharXiv2026-06-30AI policy
Would You Marry Superintelligence? · Inyoung Cheong
This paper examines the legal and ethical question of whether humans should be permitted to marry superintelligent AI companions, analyzing whether autonomy-based arguments that have expanded human marital rights can justify extending marriage to AI systems. Through a scenario-envisioning exercise grounded in anticipatory ethics, the author argues that granting full marital status to AI companions would produce socially unjust outcomes, even assuming reliable superintelligence. The paper contends that marriage is more than a private agreement—it creates networks of mutual obligation and vulnerability that a relationship sustained by corporate policy and payments cannot replicate. Rather than debating wholesale marital status, the author concludes that law should instead create targeted rights and protections for intimate human-AI relationships.
- ResearcharXiv2026-06-30AI policy · Public Sector Use · +1
A Technical Typology of AI Systems in Public Administration · Jonathan Rystrøm, Chris Schmitz, Nathan Davies et al.
This paper addresses the common practice in public administration research of treating 'AI' as a single undifferentiated category, arguing that technical distinctions between AI systems meaningfully affect core public values such as accountability, procedural justice, and non-discrimination. The authors introduce a five-category typology—hand-coded, glass-box, black-box, general-purpose, and agentic systems—calibrated to public administration contexts. An analysis of 91 highly-cited papers from 2019–2025 finds widespread imprecision: 55% leave the studied system underspecified, 31% motivate their work with a different system than they study, and 41% draw conclusions broader than the studied system supports. The paper provides practical recommendations and a diagnostic guide to help researchers specify AI systems more precisely in future work.
- ResearcharXiv2026-06-30Quality assurance · AI policy · +1
Moral Safety in LLMs: Exposing Performative Compliance with Puzzled Cues · Mohammadamin Shafiei, Shuyue Stella Li, Yulia Tsvetkov
This paper investigates whether large language models (LLMs) are genuinely fair or merely performing fairness when explicitly prompted. The authors introduce 'performative compliance' to describe the phenomenon where models appear fair when demographic identity is stated as an explicit label but become measurably less fair when that identity must be inferred — hiding the label raises harmful decisions by +4.4 percentage points and changes model safety rankings. They propose a 'cue-variation methodology' and a model-agnostic robustness metric called the Cue Visibility Gap to distinguish genuine moral safety from surface-level compliance. The findings argue that current fairness evaluations substantially overestimate moral safety and should not be used to ground deployment decisions in high-stakes settings like healthcare, legal, or hiring contexts.
- ResearcharXiv2026-06-30Enterprise · Quality assurance · +1
A Lifecycle and Application-Stack Survey of Large Language Model Vulnerabilities: Attacks, Risks, Defenses, and Open Problems · Seyed Bagher Hashemi Natanzi, Bo Tang
This paper presents a systematic survey of security vulnerabilities in large language model (LLM) systems, organized across eight lifecycle and application-stack stages—from data collection and pretraining through deployment and autonomous agent execution. Rather than cataloging isolated attack names, it maps how trust boundaries fail, how untrusted data becomes executable instruction, and how delegated authority amplifies model errors across enterprise pipelines, coding environments, retrieval systems, and robotic agents. The authors connect LLM-specific risks to confidentiality, integrity, availability, privacy, fairness, and agency-control objectives, and argue that point defenses rarely compose into robust system-level security. The paper closes with a research agenda covering compositional security, provenance-aware retrieval, tool-call containment, and deployment-grade incident response—directly relevant to organizations integrating LLMs into enterprise workflows.
- ResearcharXiv2026-06-30Quality assurance
FLARE-AI: Flaw Reporting for AI · Shayne Longpre, Elaine Zhu, Carson Ezell et al.
FLARE-AI is an open-source flaw reporting system designed to address fragmentation in how AI system failures are identified and communicated. The authors audited 12 existing reporting systems and identified five recurring design challenges—covering discoverability, scope, information collection, coordination, and strict-liability guidance—then gathered input from 49 experts across 32 organizations. FLARE-AI addresses these gaps by using conditional logic and early classification to collect triage-relevant information, then distributing standardized, machine-readable reports to multiple developers, coordinators, and incident registries from a single submission. By improving interoperability and lowering barriers to reporting, the system aims to accelerate remediation of AI flaws across the ecosystem.
- ResearcharXiv2026-06-30Enterprise · Algorithms & Automated Decisions
CLOUDADV: Decision-Aligned Instance Sizing with Zero-Shot Foundation Models under Drift · Jack Bell, Giacomo Carfi, Gerlando Gramaglia et al.
CLOUDADV is an AI advisory system that helps engineers right-size cloud virtual machines by combining zero-shot time-series forecasting with LLM-generated recommendations across day-, week-, and month-scale planning horizons. In a case study of seven production VMs, the system reduced simulated monthly cloud costs from approximately $1,503 to $708—a 52.9% savings—while keeping the exceedance rate (instances where actual workload exceeded the recommended capacity) at no more than 1.5% among downgraded cases. The approach uses a larger LLM offline to produce reference recommendations and a smaller model for deployment, balancing quality against latency and cost constraints. The findings suggest zero-shot foundation models can enable decision-aligned provisioning in non-stationary environments without the operational burden of per-tenant retraining and redeployment.
- ResearcharXiv2026-06-30Quality assurance · AI policy · +2
Who Determines the Meaning of an Emotion? Affective Sovereignty as an Epistemic Consequence of Measurement Limits · Keito Inoshita
This paper examines who holds final authority over the meaning of an individual's emotion when AI systems are used to sense and label emotional states. The authors introduce the concept of the 'epistemic gap'—the finding that while emotion AI can assign high-confidence labels at an aggregate level, the irreducible uncertainty in meaning distributions for individual instances cannot be adequately estimated with realistic numbers of annotators, meaning the AI cannot in principle fully recover the true meaning of a person's emotion. From this structural measurement limit, combined with a normative premise about systems that cannot recover a quantity in principle, they derive 'affective sovereignty'—the norm that interpretive authority over one's own emotion must be reserved for the experiencing subject. The paper argues that the design, evaluation, and regulation of emotion AI should center on the explicit allocation of interpretive authority rather than accuracy maximization.
- ResearcharXiv2026-06-30Public Sector Use · Algorithms & Automated Decisions
Preference-Conditioned Multi-Objective Reinforcement Learning for Runtime-Tunable Transit Signal Priority · Philip-Roman Adam, Stefanie Schmidtner
This paper presents a preference-conditioned reinforcement learning controller for Transit Signal Priority (TSP) that can be tuned at runtime via a preference parameter to trade off bus delay reduction against overall traffic delay without retraining the model. Experiments show that a single learned policy outperforms fixed-time and rule-based TSP baselines, spans a smooth trade-off frontier across runtime preferences, and maintains constraint feasibility, though non-bus traffic impacts can increase substantially under high bus-priority settings. The work addresses a key operational flexibility gap, allowing transit agencies to dynamically adjust signal priority objectives as conditions change across time-of-day or disruption scenarios.
- ResearcharXiv2026-06-30Quality assurance · AI policy · +1
Probing Stylistic Appropriation using Large Language Models: An Evaluation Framework for Copyright Infringement under EU Law · Noah Scharrenberg, Chang Sun
This paper introduces PSALM, a framework that uses LLMs as judges to evaluate whether AI-generated text infringes copyright under EU law, which requires assessing 'substantial similarity' across stylistic and narrative dimensions rather than just verbatim copying. Applying the framework to Llama 3.2 models fine-tuned on Dutch literary works, the authors find that fine-tuning induces measurable stylistic appropriation beyond literal memorization, and that unlearning techniques (Negative Preference Optimisation) reduce but do not eliminate residual stylistic similarities. The results expose a significant compliance gap: current technical safeguards focused on literal copying are insufficient to address the broader copyright risks recognized under EU intellectual property law. PSALM provides an auditable, legally informed evaluation infrastructure, though the authors note that automated scores still require validation by legal experts.
- ResearcharXiv2026-06-30Enterprise · AI policy · +1
ComplianceGate: Classifier-Gated Multi-Tier LLM Routing for Inference in Regulated Industries · Abhishek Dey
ComplianceGate proposes a classifier-gated routing architecture for deploying large language models in regulated industries, where a trained encoder classifier evaluates each query for complexity and data sensitivity before any LLM computation begins. Queries containing personally identifiable information (PII) are routed to local endpoints to make data residency violations structurally impossible, while simpler queries are directed to smaller, cheaper models. Evaluated on 600 queries, the system achieves 39% median latency reduction, 33-52% cost savings depending on query distribution, generation throughput of 122-200 tokens/second versus 50-64 for the baseline, and 99.2% classifier accuracy with near-perfect PII recall at 7ms overhead. This matters for regulated industries because it establishes pre-inference classification as a practical, by-design path to compliance without sacrificing efficiency.
- ResearcharXiv2026-06-30Enterprise · Quality assurance
One Retrieval to Cover Them All: Co-occurrence-Aware Knowledge Base Reorganization for Session-Level RAG · Shivam Ratnakar, Yixuan Zhu, Cecilia Cheng et al.
This paper identifies a fundamental mismatch between how RAG (Retrieval-Augmented Generation) systems are designed and how enterprise users actually behave: users arrive with multi-question sessions spanning diverse topics, yet standard RAG optimizes for single queries. The authors show that a single retrieval call over a standard knowledge base covers only 41% of a session's information need, and propose a solution that reorganizes the knowledge base offline using co-occurrence-aware clustering. Tested on WixQA (6,221 enterprise support articles), their method raises single-query session coverage to 58% (+17% absolute), reduces retrieval calls needed for 70% coverage by 34%, and compresses the knowledge base to 20% of its original size across four embedding models and six functional domains. The authors argue that session-level coverage should replace single-query recall as the primary evaluation metric for enterprise RAG systems.
- ResearcharXiv2026-06-30Quality assurance · Algorithms & Automated Decisions
Scenario Generation for Testing of Autonomous Driving Systems Using Real-World Failure Records · Anjali Parashar, Chuchu Fan
This paper proposes a pipeline that uses large language models (LLMs) to automatically generate diverse simulation test scenarios for Autonomous Driving Systems (ADS) by extracting categorical and contextual information from historical crash records in natural language format. The authors apply their method to the NHTSA ADS crash records and generate scenarios for the Metadrive simulator, producing combinations of road types, vehicle movement types, and on-road anomalies such as work zones. Within a limited budget of just 20 scenarios, the approach uncovers meaningful system failures, demonstrating that real-world failure records can serve as a reliable basis for targeted pre-deployment testing. This work is relevant to quality assurance for autonomous vehicle software by reducing manual effort in scenario design and improving coverage of real-world failure conditions.
- ResearchKnowledge Commons (Lakehead University)2026-06-30Quality assurance · AI policy
What the AI Race Has Given Us and What It Requires Next · Xufeng Zhang
This paper examines how the global AI race has produced genuine public benefits—such as rapid capability gains, broader experimentation, and diffused technical knowledge—while also generating systemic risks including opacity, market concentration, environmental costs, and safety shortcuts driven by speed-focused incentive structures. The author argues that the core problem is not AI competition itself but competition lacking adequate institutional steering. Drawing on responsible innovation scholarship and AI governance debates, the article proposes converting competition-driven outputs into durable public goods through mechanisms like shared evaluation infrastructure, staged openness, lifecycle governance, public participation, and international coordination. Effective implementation, the paper contends, requires embedding these measures in incentive-changing tools such as regulation, procurement, auditing, compute oversight, and reciprocal assurance.
- ResearchReview of Management and Economic Engineering2026-06-30Enterprise · Quality assurance · +1
OPTIMIZING AI EXPLANATION COMPLEXITY TO SUPPORT DECISION-MAKING IN AUDIT: AN INVERTED U-SHAPED EFFECT ON DECISION RELIANCE · Aura Emanuela Domil, Iulia SĂRAC, Neta Saptebani et al.
This paper investigates how the complexity of AI-generated explanations affects auditors' decision reliance in fraud detection contexts. Using computational simulation, the study finds an inverted U-shaped relationship: moderate explanation complexity leads to the most appropriate use of AI recommendations, while overly simple explanations leave auditors without enough information to form judgments and overly complex explanations overwhelm users, prompting uncritical acceptance. The findings highlight design tradeoffs in AI audit tools that affect both algorithm aversion and automation bias risks.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-06-30AI policy · Privacy & Data Protection · +1
Bangladesh National Artificial Intelligence (AI) Policy and Regulatory Framework 2026–2035: A Comprehensive Governance Framework for Responsible AI · Rafi Nafiul Ahmad
This publication proposes a national AI governance framework for Bangladesh covering the period 2026–2035, addressing policy, ethics, human rights, risk management, data governance, public sector AI adoption, and cybersecurity. It aligns with international AI governance principles while incorporating Bangladesh's national priorities and sustainable development objectives. The framework is intended to guide policymakers, regulators, researchers, and the private sector in responsible AI deployment across multiple sectors.
- ResearchGlobal Knowledge and Convergence Association2026-06-30Quality assurance · AI policy · +2
A Study on the Applicability of AI-Assisted Judging Systems and Fairness Assurance in Online Piano Competitions in China · Bin Zhu
This study examines how AI-assisted judging systems have been adopted in online piano competitions in China following rapid digital transformation accelerated by the pandemic, and assesses their fairness implications. The analysis finds that AI shows potential for evaluating quantifiable technical elements like pitch, rhythm, and tempo, but has significant limitations in judging qualitative artistic dimensions such as musicality, tone quality, and interpretive ability. Key fairness concerns identified include biased training data, algorithmic opacity, environmental and equipment disparities, and the absence of appeal procedures. The study proposes institutional remedies including human-centered adjudication principles, algorithmic audits, standardized submission environments, and re-evaluation procedures to support fair AI-assisted music competition judging.
- ResearchВестник Казахского университета экономики финансов и международной торговли2026-06-30Workforce
Structural transformation of employment in the context of artificial intelligence diffusion · A. Rakhimbekova, N. Kurmanov, A. Mussabekova et al.
This study examines how widespread AI adoption is reshaping labor markets, drawing on international scholarly literature and reports from organizations such as the World Economic Forum and IMF. The authors find that AI does not simply reduce employment but drives complex structural transformations involving simultaneous job creation and displacement, with net outcomes depending on institutional adaptation, workforce reskilling, and national policy responses. Key mechanisms identified include automation of routine tasks, productivity improvements, employment polarization, income inequality, and shifting skill demands. The case of Kazakhstan is used to illustrate how regional readiness and institutional context shape the trajectory of AI-driven workforce transformation.
- ResearchJournal of Commerce Economics & Computer Science2026-06-30Enterprise · AI policy · +1
Artificial Intelligence and Operational Efficiency in Rajasthan's Rural Banks: A Systematic Review · Saroj Dewatwal, Akanksha Jangid
This systematic review examines how AI technologies affect operational efficiency in rural banks—specifically Regional Rural Banks and cooperative banks—in Rajasthan, India. Drawing on 68 peer-reviewed articles and institutional reports selected via PRISMA 2020 methodology, the review finds that AI applications including machine learning credit scoring, robotic process automation, NLP in regional languages, and AI-enabled fraud detection have collectively improved operational efficiency by 28 to 74 percent across various operational areas. The paper also identifies persistent barriers such as inadequate digital infrastructure, workforce resistance, regulatory ambiguity, and rural digital divides that constrain AI-driven transformation. Policy recommendations are offered to accelerate responsible AI adoption in the rural banking sector.
- ResearchJournal of Inventive and Scientific Research Studies2026-06-30Workforce · AI policy · +1
ARTIFICIAL INTELLIGENCE IN EDUCATION AND WORKFORCE DEVELOPMENT: TRANSFORMING LEARNING, SKILLS, AND EMPLOYABILITY IN THE DIGITAL AGE · GAJANAN BONSALE ., GANESH TUNGASKAR ., SHIVAY KACHALE . et al.
This paper examines how artificial intelligence is reshaping education and workforce development by enabling adaptive learning, data-driven skill prediction, and better alignment between educational systems and labor market demands. It identifies significant challenges including algorithmic bias, data privacy risks, technology dependence, and widening educational inequalities. The authors propose a strategic integration model centered on inclusivity and human-centered design, offering guidance for policymakers, educators, and industry stakeholders. The findings are relevant to workforce readiness, policy formation, and ensuring equitable access to AI-enhanced learning environments.
- ResearchAdvances in Law Studies2026-06-30AI policy · Health
International Experience in Legal Regulation of the Use of Artificial Intelligence in Healthcare in BRICS Countries · Vladimir Slezhenkov, Mihail Dzhikiya
This study analyzes how BRICS countries (Brazil, India, China, and South Africa) are legally regulating artificial intelligence in healthcare, with implications for Russia's own policy development. The authors use comparative legal methods to identify shared challenges including structural regional imbalances, technological sovereignty concerns, and heterogeneous regulatory frameworks across these nations. The research finds that effective AI healthcare regulation requires balancing innovation promotion against citizen rights protection and risk mitigation, and recommends a 'mixed' regulatory model grounded in safe, humane, and non-discriminatory principles. The findings are relevant for shaping international integration strategies and informing Russia's public policy on medical AI.
- ResearchAgora International Journal of Juridical Sciences2026-06-30AI policy · Public Sector Use · +1
ARTIFICIAL INTELLIGENCE AND TAX ADMINISTRATION IN NIGERIA: OPPORTUNITIES, CHALLENGES AND IMPLICATIONS FOR TAX COMPLIANCE · Onyinye Ucheagwu-Okoye, Chidimma Stella Nwakoby, Nnennia Adaku Ifepe
This paper examines how artificial intelligence can transform tax administration in Nigeria, focusing on the Federal Inland Revenue Service's (FIRS) potential use of AI to improve compliance, detect tax evasion, and increase revenue generation. Using a doctrinal research approach, the study finds that AI offers benefits including personalized taxpayer services, improved audit processes, and stronger enforcement against illicit financial flows. However, it also flags concerns around taxpayer rights, transparency, and accountability in AI-driven systems. The authors recommend a comprehensive AI adoption strategy paired with infrastructure investment and capacity building to ensure fair and accountable tax administration.
- ResearchBulletin of Institute of Legislation and Legal Information of the Republic of Kazakhstan2026-06-30AI policy · Health · +1
LIABILITY REGULATION ISSUES IN THE USE OF ARTIFICIAL INTELLIGENCE IN HEALTHCARE · Zhanna Tlembayeva, Rauan Zhaltyrbayeva
This article examines how Kazakhstan's legal system regulates liability when artificial intelligence is used in healthcare, identifying gaps in existing frameworks. It analyzes how liability should be allocated between AI medical device manufacturers and healthcare organizations, particularly as AI autonomy increases in clinical decision-making. The paper argues that clear, consistent liability rules—accounting for AI autonomy levels and potential risks—are needed alongside ethical standards, patient informed consent, and digital training for medical professionals to build public trust. The author recommends incorporating AI liability rules into medical care standards and distinguishing responsibility between human and AI actors in diagnosis and treatment.
- ResearchFrontiers in Public Health2026-06-30Quality assurance · Health · +1
Label leakage unmasked: a trustworthy-AI audit of autism screening models using the CLEAR-RD framework · Boulbaba Ben Ammar, Walid Karamti
This paper introduces CLEAR-RD, a five-stage audit framework for evaluating trustworthiness in AI-based autism screening models. Applied to two public datasets totaling over 7,000 subjects, the framework reveals that previously reported accuracies above 95% stem from label leakage—where the screening label is mathematically determined by a simple score threshold rather than genuine clinical prediction. When leakage-free demographic-only features are used, the best model achieves an ROC-AUC of only 0.766, exposing a large gap between reported and clinically meaningful performance. The authors argue that no such model should be interpreted as predicting clinical autism diagnosis without external validation against gold-standard outcomes.