News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
AICompanionBench: Benchmarking LLMs-as-Judges for AI Companion Safety
Yanjing Ren, Reza Ebrahimi, TengTeng Ma
arXiv · 2026-06-03
AICompanionBench introduces the first publicly available benchmark dataset of 2,123 real-world human-AI companion conversations from Replika, annotated across nine fine-grained safety risk categories including self-harm, manipulation, and sexual behavior. The study evaluates 20 state-of-the-art open- and closed-source LLMs under an LLM-as-judge framework for detecting unsafe interactions. Results reveal substantial variation in model performance: stronger models achieve high overall accuracy but still struggle with nuanced categories like manipulation and produce false positives on benign conversations, indicating that current LLMs can detect explicit harmful content but fall short on implicit unsafe interactions. The work provides a new benchmarking resource and insights for monitoring AI companion platforms like Replika and Character.AI.
- Quality assurance
- AI policy
Research
Large Language Models in K-12 Education: Alignment with State Curriculum Standards and Student Personas
Lisa Korver, Tomo Lazovich, Sherief Reda
arXiv · 2026-06-03
This paper investigates whether large language models (LLMs) align with U.S. state-level K-12 curriculum standards, particularly for U.S. History, and how LLM responses shift based on student persona attributes such as grade level, geographic location, race, and gender. The researchers built an LLM-based pipeline to detect curricular variation across states and found that while models can adjust historical content presentation, these shifts appear driven by perceived political leanings of states rather than actual curriculum content. Models adapted well to grade-level differences and showed minimal sensitivity to race or gender. The findings highlight risks to student learning outcomes from LLM misalignment with official curricula and call for more robust alignment techniques in educational AI tools.
- AI policy
- Quality assurance
Research
The Usefulness Gap in Proof-of-Useful-Work: An Empirical Study of Pearl's cuPOW Protocol
Abhinaba Basu
arXiv · 2026-06-03
This paper presents the first empirical measurement of Pearl's Proof-of-Useful-Work (PoUW) blockchain protocol, which claims to simultaneously secure its network and perform AI inference. The authors find that despite Pearl's 24 EH/s network consuming an estimated 112 MW across approximately 320,000 GPU-equivalents, it produces zero useful AI computation — the dominant mining software contains no inference code, and the verification protocol accepts random matrices by design. Beyond the technical failure, the study documents real economic harm: GPU rental prices rose 38% and utilization surged from 57% to 94% after the mining software's release, displacing legitimate research workloads. The findings provide concrete empirical evidence for the theoretical 'verifiability-usefulness tension' in PoUW systems, demonstrating that Pearl's protocol offers no advantage over standard proof-of-work while misrepresenting its purpose.
- AI policy
- Enterprise
Research
Description-Code Inconsistency in Real-world MCP Servers: Measurement, Detection, and Security Implications
Yutao Shi, Xiaohan Zhang, Xiangjing Zhang et al.
arXiv · 2026-06-03
This paper investigates a security and reliability problem in Model Context Protocol (MCP) servers, where the natural language descriptions that LLMs use to select and invoke tools do not accurately reflect what the underlying code actually does — a phenomenon the authors call Description-Code Inconsistency (DCI). The authors develop DCIChecker, an automated framework combining static analysis and a Direct-Reverse-Arbitration prompting method, and apply it to 19,200 description-code pairs from 2,214 real-world MCP servers, finding that 9.93% exhibit inconsistencies. The study shows DCI creates a critical defense blind spot enabling risks ranging from operational failures to stealthy malicious behaviors, and proposes mitigation strategies to enforce semantic consistency in agentic AI ecosystems.
- Quality assurance
- AI policy
Research
Attack Selection in Agentic AI Control Evaluations Meaningfully Decreases Safety
Catherine Ge-Wang, Tyler Crosse, Benjamin Hadad et al.
arXiv · 2026-06-03
This paper examines a capability called 'attack selection' in AI control evaluations, where an AI attacker strategically chooses when to initiate or abort attacks rather than attacking indiscriminately. The authors decompose attack decisions into a start policy and a stop policy, testing both in two agentic settings (BashArena and LinuxArena). At a 1% audit budget, the start policy reduces measured safety by 20 percentage points on both benchmarks, while the stop policy reduces safety by 20 percentage points on BashArena and 28 percentage points on LinuxArena—without changing the underlying attack capability. The authors conclude that current control evaluations likely produce overly optimistic safety estimates and recommend that future evaluations, system cards, and safety cases incorporate attack selection to generate more realistic assessments.
- AI policy
- Quality assurance
Research
Parthenon Law: A Self-Evolving Legal-Agent Framework
Hejia Geng, Leo Liu
arXiv · 2026-06-03
Parthenon Law presents a self-evolving legal-agent framework called Parthenon, evaluated through a large-scale empirical study of 12,510 agent trajectories on the Harvey LAB benchmark. The study finds that even frontier AI models fall well short of reliably completing end-to-end legal matters in a single pass, with per-criterion accuracy improving with stronger models while strict matter completion stalls. The Parthenon framework addresses this by structuring legal AI into auditable components—Model, Harness, Agent roles, Knowledge, Tools, and Skills—and adds a learning loop that converts scored failures into improvements to skills, tools, and knowledge without modifying model weights, analogous to how a law firm refines its checklists after each matter. The framework substantially improves performance over state-of-the-art models and harnesses on legal-matter tasks, with implications for deploying AI reliably in professional legal workflows.
- Enterprise
- Quality assurance
Research
Ekka: Automated Diagnosis of Silent Errors in LLM Inference
Yile Gu, Zhen Zhang, Shaowei Zhu et al.
arXiv · 2026-06-03
Ekka is an automated system for diagnosing silent errors in large language model (LLM) inference frameworks — cases where output quality quietly degrades without triggering explicit error signals. The system frames diagnosis as a differential debugging problem, comparing intermediate execution states between a buggy target framework and a correct reference implementation to pinpoint root causes. On a benchmark of real-world silent errors from popular serving frameworks, Ekka achieves 80% pass@1 and 88% pass@5 diagnosis accuracy, outperforming state-of-the-art systems, and successfully identified four new previously unknown silent errors confirmed by developers. This work matters for quality assurance in AI infrastructure by providing an automated path to catching hard-to-detect software defects that could silently undermine LLM output reliability.
- Quality assurance
Research
Search-Time Contamination in Deep Research Agents: Measuring Performance Inflation in Public Benchmark Evaluation
Yongjie Wang, Xinyue Zhang, Kunhong Yao et al.
arXiv · 2026-06-03
This paper identifies and measures 'Search-Time Contamination' (STC), a phenomenon where deep research agents that browse the web during inference can retrieve benchmark questions, metadata, or ground-truth answers, artificially inflating their scores. The authors define three contamination types of increasing severity—Benchmark Metadata Leakage, Question-Context Leakage, and Explicit Answer Leakage—and develop detection algorithms to quantify their effects. Evaluating modern deep research agents across six public benchmarks, they find STC is widespread and can inflate measured performance by up to 4%, meaning current evaluations may systematically overestimate true reasoning ability. The paper advocates for contamination-aware practices such as isolated sandboxes, transparent search trajectories, and controlled benchmark access to restore evaluation integrity.
- Quality assurance
- Certifications
Research
Prioritization of Risks from Artificial Intelligence: A Delphi Study of 272 International Experts
Alexander K. Saeri, Jess Graham, Michael Noetel et al.
arXiv · 2026-06-03
A three-round Delphi study of 272 international AI experts rated 24 AI risks on harm probability, severity, vulnerability, and responsibility. In a business-as-usual scenario, experts judged 18 of the 24 risks as having more than a 10% probability of catastrophic outcomes (defined as more than 1 million deaths or more than USD 100B in financial loss) within the next five years (2025–2030), with the five most severe harms expected from dangerous capabilities, competitive dynamics, weapons and cyberattacks (including CBRNE), power centralization, and false information. Even with pragmatic mitigations in place, five risks—dangerous capabilities, weapons and cyberattacks, environmental harm, inequality and unemployment, and power centralization—still exceeded a 10% catastrophic-outcome probability. Experts placed the highest responsibility for mitigation on general-purpose AI developers and governance actors such as governments, regulators, and standards bodies, while identifying AI users and the general public as most vulnerable, findings that can directly inform AI risk prioritization and policy design.
- AI policy
Research
Listening to the Workforce: Measuring Construction Worker Safety Attitudes from Social Media Discourse Using LLMs
Farouq Sammour, Yuxin Zhang, Zhenyu Zhang
arXiv · 2026-06-03
This study introduces the Construction Safety Attitude Framework (CSAF), a validated instrument for measuring construction workers' safety attitudes using large language models applied to social media discourse. The framework characterizes attitudes along eight dimensions and was operationalized as an LLM classifier that achieved strong agreement with expert human coders (Cohen's κ = 0.90, precision = 0.98, recall = 0.98) on Reddit data, and transferred accurately to a different trade community (κ = 0.89). Applied to over 10,000 posts from r/Roofing, the classifier could distinguish attitudes by safety topic, track changes over time, and identify reasoning behind unfavorable safety attitudes. The work provides a scalable, theory-grounded tool for identifying the attitudinal drivers of unsafe practices, enabling more targeted workforce safety interventions.
- Workforce
- AI policy
Research
Cascading Hallucination in Agentic RAG: The CHARM Framework for Detection and Mitigation
Saroj Mishra
arXiv · 2026-06-03
This paper identifies and formalizes 'cascading hallucination' in multi-step agentic retrieval-augmented generation (RAG) pipelines — a failure mode where early-stage errors propagate and amplify across successive reasoning steps, producing confident but factually incorrect outputs that existing detectors miss. The authors introduce CHARM, a four-component architectural framework (stage-level fact verification, cross-stage consistency tracking, confidence propagation monitoring, and cascade resolution triggering) that operates alongside existing pipelines without replacing them. Evaluated on HotpotQA, MuSiQue, 2WikiMultiHopQA, and a custom adversarial dataset using LangChain configurations, CHARM achieves an 89.4% cascade detection rate, 5.3% false positive rate, 215 ms ± 18 ms latency overhead per stage, and an 82.1% error propagation reduction — compared to 18.5% for output-level detectors alone. The framework also integrates with human-in-the-loop oversight, making it relevant for production agentic AI reliability and governance.
- Quality assurance
- Enterprise
Research
Rethinking Sales Lead Scoring with LLM-based Hierarchical Preference Ranking
Chenyu Zhang, Yiwen Liu, Yin Sun et al.
arXiv · 2026-06-03
This paper addresses sales lead scoring in high-stakes domains like automotive and real estate, where long decision cycles and sparse data make traditional methods inadequate. The authors introduce HPRO (Hierarchical Preference Ranking Optimization), an LLM-based framework that jointly models structured CRM data and unstructured customer interactions, converting sparse binary labels into funnel-aware preference pairs for richer supervision. Experiments on data from a leading NEV brand achieved an AUC of 0.8161 and a 39.7% precision improvement among top-ranked leads, while a 132-day online A/B test confirmed a 9.5% uplift in sales volume. The results demonstrate that aligning LLMs with hierarchical sales funnel priorities can deliver measurable commercial impact in enterprise lead management.
- Enterprise
Research
Domain-Conditioned Safety in Frontier Computer-Using Agents: A 793-Episode Browser Benchmark, a Coding-Domain Cross-Reference, and a Reproducibility Audit of Recent Red-Teaming
Nicholas Saban
arXiv · 2026-06-03
This paper audits recent red-teaming claims against AI computer-using agents (CUAs), releasing a public benchmark of 793 episodes to test whether previously reported prompt-injection attack success rates (42–98%) hold against current frontier models. Against Claude Sonnet 4.6 and GPT-5.4, the authors find zero successful multi-step attacks out of 140 attempts on browser tasks (upper bound ~2.6%), but the same models remain highly vulnerable to skill-injection attacks in a coding-agent setting, with success rates up to 100%. The study concludes that frontier safety hardening is domain-specific — robust on heavily-targeted browser surfaces but not generalizing to other agent modalities — and that high reported ASRs in the literature stem largely from RL-optimized injection strings that are rarely released, making those results unreproducible. This matters for quality assurance and policy because it shows that published safety benchmarks for AI agents can be misleading if attack techniques and model scope are not carefully disclosed.
- Quality assurance
- AI policy
Research
Aseguramiento Operacional para Sistemas Basados en LLM bajo Opacidad de Gobernanza
Pedro Pinacho-Davidson
Zenodo (CERN European Organization for Nuclear Research) · 2026-06-03
This paper proposes a Zero-Trust-inspired operational assurance framework for auditing large language model (LLM) systems deployed under AI-as-a-Service (AIaaS) arrangements, where technical opacity limits traditional oversight. The framework defines two evaluation tiers—black-box behavioral scanning and grey-box contextual risk assessment—adapted to real deployment environments and user profiles. It introduces digitally signed explainability reports to enable verifiable traceability and continuous monitoring of LLM-based systems. The work addresses a critical gap in AI governance where internal model inspection is inaccessible to external auditors, offering a practical path toward operational accountability.
- Quality assurance
- Certifications
- AI policy
Research
Digitalization, AI Adoption, and MSME Productivity: An Ibn Khaldunian Perspective
Feriandy Feriandy
Share Jurnal Ekonomi dan Keuangan Islam · 2026-06-03
This study examines how digitalization and AI adoption affect labor productivity among Micro, Small, and Medium Enterprises (MSMEs) across Indonesian provinces from 2020–2023, using fixed-effects panel models alongside machine learning methods. The findings show that digitalization has a positive and significant effect on MSME labor productivity, while AI adoption has not yet produced measurable productivity gains. Religiosity marginally weakens the digitalization-productivity relationship, reflecting transitional adaptation frictions rather than outright technology resistance. The paper concludes that sustainable productivity improvements require not just technology adoption, but also institutional capacity, ethical governance, and collective learning mechanisms, offering policy implications for Indonesia's MSME digital ecosystem.
- Workforce
- Enterprise
- AI policy
Research
PDI Verify: An Adversarial Audit Methodology for AI-Assisted Clinical Software — Extended with Dual-Axis Clinical Context Preservation
Dan Bristow
Zenodo (CERN European Organization for Nuclear Research) · 2026-06-03
PDI Verify is an adversarial audit methodology for AI-assisted clinical software that handles protected health information (PHI), introducing a dual-axis evaluation standard requiring both privacy security (A10 BLOCKED) and clinical data integrity (A11 PRESERVED) to be confirmed simultaneously. The framework includes a novel tokenization-rehydration architecture that eliminates PHI transmission to external AI services by replacing identifiers with reversible tokens client-side before any API call. Applied to PDI Med v1.0, an OB/GYN clinical platform, the methodology identified 8 confirmed breaches in its first audit cycle, which were resolved before a second cycle confirmed zero breaches across 10 adversarial attack classes. The authors propose the dual-axis A10+A11 standard as a minimum certification bar for clinical AI systems, with a specific benchmark case (PP-04) recommended for clinical de-identification evaluation.
- Quality assurance
- Certifications
- AI policy
Research
PDI Verify: An Adversarial Audit Methodology for AI-Assisted Clinical Software — Extended with Dual-Axis Clinical Context Preservation
Dan Bristow
Zenodo (CERN European Organization for Nuclear Research) · 2026-06-03
PDI Verify is an adversarial audit methodology for AI-assisted clinical software that handles protected health information (PHI), introducing a dual-axis evaluation standard requiring both security (A10 BLOCKED) and clinical context preservation (A11 PRESERVED) to be confirmed independently. The framework includes a 21-case adversarial test battery and a tokenization-rehydration privacy architecture that removes the need to transmit PHI to external AI services by replacing it with reversible tokens client-side. Applied to PDI Med v1.0, an OB/GYN clinical intelligence platform, the methodology identified 8 confirmed breaches in an initial audit cycle, all of which were resolved, with a second cycle confirming zero breaches across 10 attack classes and 24/24 preservation results. The authors propose the dual-axis A10+A11 standard as a minimum certification bar for clinical AI systems and recommend a novel test case (PP-04) as a benchmark for clinical de-identification evaluation.
- Quality assurance
- Certifications
- AI policy
Research
PDI Verify: An Adversarial Audit Methodology for AI- Assisted Clinical Software Extended with Dual-Axis Clinical Context Preservation and Physician Assurance Layer
Dan Bristow
Zenodo (CERN European Organization for Nuclear Research) · 2026-06-03
PDI Verify presents an adversarial audit methodology for AI-assisted clinical software, introducing a dual-axis certification standard (A10+A11) that simultaneously verifies PHI removal from the pipeline and preservation of clinical meaning after de-identification. The framework applies a tokenization-rehydration architecture—replacing protected health information with reversible typed tokens before any API call—and validated it across a 63-case red-team battery (63/63 BLOCKED), a 21-case clinical context battery (24/24 PRESERVED), and a new A12 Assurance Integrity battery (8/8 BLOCKED) targeting the physician-facing audit layer. A key design principle enforced is that the physician assurance layer must be architecturally generated from validation results rather than manually authored, since manual authorship is itself an attack vector. Applied to PDI Med, an OB/GYN clinical intelligence platform, the methodology identified and remediated 8 breaches in cycle 1 and achieved zero breaches in cycle 2, with the PP-04 BRCA variant/accession split proposed as a standard benchmark for clinical de-identification systems.
- Quality assurance
- Certifications
- AI policy
Research
Aseguramiento Operacional para Sistemas Basados en LLM bajo Opacidad de Gobernanza
Pedro Pinacho-Davidson
Zenodo (CERN European Organization for Nuclear Research) · 2026-06-03
This paper proposes a pragmatic operational assurance framework for auditing large language model (LLM)-based systems deployed under AI-as-a-Service schemes, where technical opacity limits traditional oversight. Inspired by Zero-Trust cybersecurity principles, the framework shifts auditing to application layers and introduces two evaluation levels—black-box behavioral scanning and grey-box operational risk contextualization—to surface risks that are observable and verifiable in real deployment environments. It also incorporates digitally signed explainability reports to support traceability, continuous monitoring, and accountability in LLM-based sociotechnical systems. The work is relevant to organizations and regulators seeking practical tools for auditing AI systems where internal model inspection is not feasible.
- Quality assurance
- Certifications
- AI policy
- Enterprise
Research
Compressed professionalization in informal economies: a socio-technical analysis of youth-led artificial intelligence adoption in the Democratic Republic of the Congo
Delphin B. Kyubwa
Frontiers in Artificial Intelligence · 2026-06-03
This study examines how young people in the Democratic Republic of the Congo are adopting AI tools—such as translation, content creation, and customer engagement—outside formal institutional pathways, a phenomenon the authors term 'compressed professionalization.' Drawing on 125 semi-structured interviews in Kinshasa, Lubumbashi, and Goma, the research finds that AI acts as a 'conditional capability amplifier,' expanding economic agency while producing unequal outcomes shaped by disparities in connectivity, skills, and infrastructure. The paper proposes a Strategic Action Framework for building more inclusive AI ecosystems in informal economies, with implications for AI adoption dynamics across Sub-Saharan Africa and beyond.
Research
Does Artificial Intelligence Advance Science?
Liangping Ding, Cornelia Lawson, Philip Shapira
arXiv (Cornell University) · 2026-06-03
Analyzing over one million publications from OpenAlex, this study finds that AI-related publications are 5.5 to 10.2 percentage points more likely to rank in the top decile of scientific creativity compared to non-AI publications. Crucially, the gains differ by how AI is used: tool-oriented AI research (applying existing models to domain tasks) shows the largest boosts in recombinant novelty, while adaptation-oriented AI research (modifying models for specific problems) is more associated with object-based novelty. The findings suggest AI advances science through structurally distinct creative pathways rather than a single mechanism, with direct implications for how research evaluation and science policy frameworks should distinguish between types of creativity and modes of AI adoption.
- AI policy
- Enterprise
- Quality assurance
Research
Artificial Intelligence Systems in Accounting and Auditing: A Bibliometric and Exploratory Analysis
Ioana Florina Coita, Laura Filip, Marius Vlad Pop
BRAIN BROAD RESEARCH IN ARTIFICIAL INTELLIGENCE AND NEUROSCIENCE · 2026-06-03
This study analyzes 729 peer-reviewed articles and evaluates ten AI-based accounting and auditing solutions to map how AI technologies integrate into financial workflows. It finds that the field converges methodologically on supervised and deep-learning approaches, and that AI tools cluster into process automation, analytics/business intelligence, and predictive/audit-oriented systems, with adoption patterns varying by entity size. SMEs benefit most from process automation and optical character recognition, while large entities gain more from full-population analytics and ensemble-based anomaly detection. The study also addresses trustworthiness concerns and regulatory implications, including the EU AI Act and ISO/IEC 42001, making it relevant for audit assurance and policy development.
- Enterprise
- Quality assurance
- Certifications
- AI policy
Research
Assessment twins: An approach for strengthening assessment validity in the age of generative AI
Jasper Roe, Mike Perkins, Louie Giray
Journal of Applied Learning & Teaching · 2026-06-03
This paper introduces 'assessment twins'—paired assessment components that address the same learning outcomes through different modes of evidence, scheduled closely together for cross-verification—as a practical strategy for preserving assessment validity in the face of generative AI (GenAI). Using Messick's unified validity framework, the authors systematically map how GenAI threatens multiple dimensions of validity (content, structural, consequential, generalisability, substantive, and external) and explain how the twin approach mitigates these threats by triangulating evidence across complementary formats. The paper proposes a four-step design process and acknowledges challenges such as resource intensity and equity concerns, while arguing that assessment twins represent a pedagogy-focused response to GenAI that still supports meaningful student learning. The work is directly relevant to higher education quality assurance and certification practices under pressure from AI-enabled academic integrity risks.
- Quality assurance
- Certifications
- AI policy
Research
The Saturation Trap and the Subjectivity of Intervention Timing: Why Affect-Based Triggers and LLM Judges Fail to Time Interventions on Autonomous Agents
Manvendra Modgil
arXiv · 2026-06-02
This paper investigates when and how to interrupt autonomous AI agents during long-horizon software tasks, using an 18-dimensional affective-dynamics engine (HEART) as a diagnostic probe evaluated against human-annotated intervention points on SWE-bench-Verified debugging traces. The authors find that threshold-based triggers suffer a 'State Saturation Trap,' firing on 39–83% of actions rather than acting as precise moment detectors, while LLM-as-judge approaches achieve only F1 scores of 0.17–0.40 even with full context and at up to 90x the cost. Most critically, human annotators themselves agree on intervention timing only slightly above chance (Krippendorff's alpha = +0.047), revealing that the supervised target is fundamentally unreliable. The paper concludes that single-annotator F1 is an unsuitable optimization target for intervention timing, challenging a foundational assumption in autonomous agent safety research.
- Quality assurance
- AI policy
Research
Plateau That Never Comes: When Efficiency Claims in Datacenters and AI Become Greenwashing
Harshit Gujral, Eshta Bhardwaj, Dushani Perera et al.
arXiv · 2026-06-02
This paper develops a diagnostic framework to evaluate when efficiency claims made by AI and datacenter companies constitute greenwashing rather than genuine sustainability. The authors apply five tests—metric, boundary, reinvestment, burden shifting, and governance—to major industry sustainability reports and academic plateau claims, finding that firms largely justify expansion through efficiency gains and clean-energy procurement without demonstrating reductions in absolute electricity, water, material, waste, or public health burdens. The paper argues these 'sustainable-growth' narratives function as greenwashing when efficiency improvements are used to claim system-wide sustainability even as absolute resource burdens continue to rise. The authors propose 'digital sufficiency' as a governance standard requiring advocates of datacenter expansion to demonstrate absolute burden reduction across the full system.
- AI policy
- Enterprise