News & Research
The latest AI research and news with real-world stakes. Each item is sourced, dated, summarized in plain English and tagged by impact area, and checked against its source before it appears.
Kind
Impact area
5732 items
- ResearcharXiv (Cornell University)2026-06-06EQC
Semantic Quorum Assurance: Collective Certification for Non-Deterministic AI Infrastructure · Jun He, Deying Yu
This paper introduces Semantic Quorum Assurance (SQA), a control-plane framework designed to prevent AI agents from autonomously approving operationally unsafe cloud infrastructure changes—such as modifying IAM policies or opening firewall rules—that are syntactically valid but dangerous. SQA routes proposed changes to a diverse panel of sandboxed validator agents whose judgments are aggregated under a risk-adaptive quorum predicate that enforces model diversity and archetype-specific vetoes. In tests on 500 infrastructure mutation scenarios, SQA reduced unsafe proposal approvals from 18.5% (single-agent) to 0.3%, with a median validation latency of 1.45–4.12 seconds. The work matters because it provides a measurable, certifiable safety mechanism for governing non-deterministic LLM agents operating in autonomous cloud environments.
- ResearcharXiv2026-06-05Q
Strained Coherence: A Pre-Failure Signal in Coding Agent Execution Trajectories · Marut Pandya, Kasey Zhang, Baiqing Lyu
This paper identifies 'strained coherence' — a failure mode in LLM-based coding agents where the agent explicitly acknowledges a problem in its own reasoning but proceeds to act against that acknowledgment anyway. The authors build an automated judge (using Claude Sonnet 4.6) to detect this pattern in agent execution trajectories and find that flagged trajectories fail 94% of the time versus 46% for unflagged ones, a 47-point gap significant at p=0.003 on Terminal-bench-2. The detector outperforms a lexical baseline in precision (94% vs. 88%) and produces interpretable, span-level output identifying what the agent saw and ignored. This work matters for quality assurance and safety monitoring of AI coding agents, offering a pre-failure signal that could enable intervention before task completion.
- ResearcharXiv2026-06-05CP
Overcoming the Regulatory Bottleneck via Agent-to-Agent Protocols: A Nuclear Case Study · Akshay J. Dave, David Grabaskas, Joseph A. Renevitz et al.
This paper introduces the Regulatory Context Protocol (RCP), an agent-to-agent communication standard designed to replace formal human-to-human regulatory review pipelines with a structured, auditable agentic channel while preserving human oversight at safety-significant decision points. Calibrated against 1,236 documents from U.S. Nuclear Regulatory Commission advanced reactor dockets and demonstrated via a multi-agent pilot, RCP is projected to cut costs by 50–77 percent (saving 21M–44M USD) and timelines by 65 percent (15 months) compared to an 89M USD, 42-month reconstructed baseline. The authors argue the residual gap between standalone AI agents and the full RCP approach is structural—rooted in the inter-organizational pipeline—rather than algorithmic, and that the same bottleneck applies to pharmaceutical, environmental, financial, and aviation regulatory contexts. Applied broadly to the U.S. regulatory system, the authors project potential savings of 210–330 billion USD per year, approaching 1 percent of U.S. GDP.
- ResearcharXiv2026-06-05QP
Where Instruction Hierarchy Breaks: Diagnosing and Repairing Failures in Reasoning Language Models · Sanjay Kariyappa, G. Edward Suh
This paper examines how reasoning language models handle conflicting instructions from different sources (e.g., system prompts vs. user inputs) in agentic workflows, a property called 'instruction hierarchy.' The authors introduce a white-box diagnostic framework that breaks down non-compliance into three distinct failure modes: failing to identify relevant instructions, failing to resolve conflicts among them, or violating the resolved intent in the final response. Evaluating models including Gemma-4-31B-IT, Qwen3-35B-A3B, and Claude Sonnet 4.6, they find that dominant failure modes vary across models, tasks, and context lengths. They then propose two training-free self-monitoring mechanisms—a parallel input monitor and a sequential output monitor—that reduce rule-following non-compliance by 81–99% across tested models, with GPT-5.3 showing 86% reduction under static attacks and 45% under adaptive attacks.
- ResearcharXiv2026-06-05Q
Land cover and flood type govern the detection limits of satellite-based flood mapping across diverse global flood events · Venkatesh Kolluru, Rajat Shinde, Abdelhak Marouane et al.
This study evaluates Prithvi-EO-2.0, a geospatial foundation model, for satellite-based flood mapping across 19 out-of-distribution flood events spanning six continents, eight climate zones, and six flood mechanisms between 2017 and 2025. Detection accuracy varied strongly by land cover and flood type: cropland achieved the highest agreement (IoU=52%) and riverine floods the strongest detection (F1=0.69), while tree cover and built-up areas showed near-zero detection (IoU=4%) regardless of flood mechanism. The research also found that apparent model errors partly stem from inconsistencies between reference products rather than true detection failures, and that pipeline engineering issues dominated errors more than model capacity limitations. These findings establish environment-specific detection boundaries critical for assessing where AI-driven satellite flood mapping can reliably support disaster response operations.
- ResearcharXiv2026-06-05WE
How AI Agents Reshape Knowledge Work: Autonomy, Efficiency, and Scope · Jeremy Yang, Kate Zyskowski, Noah Yonack et al.
Using production data from Perplexity's Search and Computer products, this paper studies how autonomous AI agents reshape knowledge work compared to conversational search assistants. Key findings show that the autonomous agent (Computer) performs 26 minutes of work per session versus 33 seconds for Search, reduces task completion time from 269 to 36 minutes on matched tasks, and lowers estimated time and cost by 87% and 94% respectively compared to humans using Search alone. Per-query dissatisfaction rates are 55% lower on Computer than on Search. Beyond efficiency gains, the agent shifts the scope of work users attempt — queries more often cross occupational boundaries, require higher-order cognition, and bundle interdependent subtasks — suggesting AI agents not only accelerate workflows but fundamentally expand what work gets automated.
- ResearcharXiv2026-06-05QC
Re-imagining ISO 26262 in the Age of Autonomous Vehicles: Enhancing Controllability through Transferability and Predictability · Chaitanya Shinde, Hadi Hajieghrary, Paul Schmitt et al.
This paper addresses a gap in the ISO 26262 functional safety standard, which was designed around human-driven vehicles and does not adequately account for fully autonomous systems operating at SAE Levels 4 and 5. The authors decompose the standard's 'Controllability' parameter into two measurable sub-concepts—Transferability (the AV's ability to hand off control to fallback safety mechanisms) and Predictability (how easily external agents can anticipate AV behavior)—and provide a mathematical framework to quantify both. A 'designed-versus-achievable gap' metric is introduced to distinguish architectural fallback claims from scene-conditioned real-world capability, making fallback and interaction claims falsifiable and traceable. The proposed framework is designed to complement rather than replace ISO 26262 and ISO/PAS 21448 (SOTIF), extending their applicability to driverless automated systems.
- ResearcharXiv2026-06-05QP
CultureScore: Evaluating Cultural Faithfulness in Video Generation Models · Anku Rani, Wei Dai, Shravan Nayak et al.
CultureScore introduces a compositional evaluation framework for measuring how faithfully AI video generation models represent diverse global cultures, decomposing cultural faithfulness into three dimensions: Identity, Context, and Behavior. Testing across 10 countries and 6,174 generated videos from three state-of-the-art models, the study finds no model achieves cultural faithfulness—the best reaches only 56.8% overall, with Behavior the hardest dimension at below 52.1% across all models. Notably, the model ranked highest on visual quality (VideoScore) was ranked last by human annotators for cultural faithfulness, demonstrating that existing quality metrics are insufficient for equitable video generation evaluation.
- ResearcharXiv2026-06-05QC
When Large Language Models Fail in Healthcare: Evaluating Sensitivity to Prompt Variations · Mahdi Alkaeed
This study systematically tests how sensitive both general-purpose LLMs (GPT-3.5, Llama3) and medical-specific LLMs (ClinicalBERT, BioLlama3, BioBERT) are to small changes in how prompts are worded, using the MedMCQA benchmark. The researchers find that even minor rephrasing can alter clinical advice, while adversarial prompt manipulations can produce dangerous outputs such as incorrect dosage recommendations or omission of critical findings. Although models show some resilience to simple lexical substitutions, they break down under syntactic reordering or misleading contextual cues. The findings underscore that medical LLMs are not intrinsically safe and that their unpredictability poses serious risks in high-stakes clinical applications.
- ResearcharXiv2026-06-05QP
TRACE: Trajectory Reasoning through Adaptive Cross-Step Evidence Aggregation for LLM Agents · Vijitha Mittapalli, Shreyaa Jayant Dani, Satya Srujana Pilli et al.
TRACE is a monitoring framework designed to detect when autonomous LLM agents pursue hidden malicious objectives through sequences of individually benign actions. It uses a TIJ (Triage-Inspect-Judge) loop that identifies high-signal regions in agent trajectories, accumulates evidence across reasoning steps, and produces a trajectory-level verdict — addressing a key limitation of existing methods that evaluate trajectories in a single pass or in isolated windows. Evaluated on ten task domains from the SHADE-Arena benchmark, TRACE achieves an aggregate F1 of 0.713 and recall of 0.844, with the largest gains on tasks requiring long-range evidence linking. This matters for AI quality assurance and policy because it advances the ability to reliably audit and flag unsafe or deceptive agent behavior in long-horizon deployments.
- ResearcharXiv2026-06-05Q
OpenHalDet: A Unified Benchmark for Hallucination Detection across Diverse Generation Scenarios · Xinyi Li, Zhen Fang, Yongxin Deng et al.
OpenHalDet is a unified benchmark designed to standardize the evaluation of hallucination detection in large language models (LLMs). It addresses two key problems in existing research: inconsistent evaluation configurations and limited coverage of downstream tasks and domains, which make detector results hard to compare or reproduce. The benchmark supports black-box, gray-box, and white-box detection methods under a shared framework, enabling controlled comparisons across diverse tasks, models, and detectors. This matters for quality assurance because it provides a systematic, reproducible way to assess how reliably LLMs can be monitored for false or unsupported outputs before deployment.
- ResearcharXiv2026-06-05QP
Auditing Training Data in Domain-adapted LLMs: LoRA-MINT · Gonzalo Mancera, Daniel DeAlcala, Aythami Morales et al.
LoRA-MINT is a membership inference testing methodology designed to audit whether specific data samples were used to train large language models fine-tuned via Low-Rank Adaptation (LoRA). By analyzing the relationship between model perplexity and membership status, it provides a systematic framework for detecting data exposure in domain-adapted LLMs. Experiments across four models and three benchmark datasets achieved precision values ranging from 0.77 to 0.92, outperforming state-of-the-art baselines. The method supports transparency, intellectual property management, and responsible AI deployment, and the authors note it generalizes beyond LoRA to other fine-tuning and domain-adaptation approaches.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-06-05WEP
Artificial intelligence and the future of work: transforming global labour markets in the digital economy · Baishakhi Mondal, Dr. Rajiv Kumar Agarwal, Shivom Shankhdhar et al.
This cross-sectional study of 487 respondents across developed and developing economies uses Structural Equation Modeling to examine how AI adoption shapes labor markets. The findings show that AI adoption significantly enhances workforce transformation (β = 0.512, p < 0.001) and improves skill development and economic productivity, but also negatively affects perceived job security (β = −0.218, p = 0.001). Human-AI collaboration and organizational readiness strengthen positive outcomes, while the results highlight the need for reskilling initiatives, inclusive AI governance, and proactive workforce policies to manage socioeconomic challenges.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-06-05WEP
Artificial intelligence and the future of work: transforming global labour markets in the digital economy · Baishakhi Mondal, Dr. Rajiv Kumar Agarwal, Shivom Shankhdhar et al.
This cross-sectional study of 487 respondents across developed and developing economies uses Structural Equation Modeling (SEM) to find that AI adoption significantly enhances workforce transformation (β = 0.512, p < 0.001) and improves skill development, labor market performance, and economic productivity. However, AI adoption also negatively affects perceived job security (β = −0.218, p = 0.001), underscoring concerns about displacement and skill obsolescence. The authors highlight the importance of reskilling initiatives, inclusive AI governance, and human-AI collaboration for sustainable labor market transitions, offering guidance for policymakers, organizations, and educational institutions.
- ResearchOpen MIND2026-06-05WEQ
Cross-AI, Lean-Verified Mathematics: A Case Study on the Collatz Conjecture · Piero Borgatta
This paper presents a methodology retrospective on an AI-assisted mathematical research program targeting the Collatz conjecture, in which the author used multiple large language models (Gemini, Claude, OpenAI Codex, DeepSeek) in a coordinated cross-AI workflow to generate, refine, and formally verify mathematical results. The key finding is not a proof of the conjecture—which the authors explicitly disclaim—but rather a Lean 4 + Mathlib verified artifact comprising 302 theorems and lemmas with zero unverified placeholders, produced with zero human-typed repository lines over roughly five weeks. The paper documents both the capabilities and failure modes of cross-AI mathematical collaboration, including hallucinated lemma names and syntax errors, concluding that formal proof checkers like Lean serve as the only reliable arbiter of AI-produced mathematical claims. This work is significant for understanding how AI tools can be structured and audited in high-rigor knowledge work, and for establishing honest negative results as a methodological norm in AI-assisted research.
- ResearcharXiv2026-06-05WEP
AI Adoption and Capability Gaps in Swiss Public Administration · Claudia Pedron, Hans‐Dieter Zimmermann, Matthias Baldauf
This paper surveys Swiss municipalities and cantonal administrations to assess the state of AI adoption in public sector organizations. It finds that while digitalization is relatively advanced, AI use remains concentrated in assistive tools and lags in analytical or decision-adjacent applications. Key barriers include legal uncertainty, limited expertise, resource constraints, and fragmented responsibilities, suggesting that organizational and governance conditions—not just technology availability—are central drivers of broader adoption.
- ResearchFigshare2026-06-05EQC
A clause-based framework for evaluating AI-assisted SOP generation in an ISO-aligned clinical laboratory: a proof-of-concept study · Ahmed Naseer Kaftan
This proof-of-concept study tested whether ChatGPT-5 could generate compliant standard operating procedures (SOPs) for an ISO-accredited clinical laboratory using a structured, clause-based evaluation framework. Across 10 high-priority SOPs, AI-assisted drafts scored higher on quality, ISO clause referencing, traceability, and lifecycle conformity than manually written SOPs, while reducing drafting time by approximately 91%. Junior staff found AI-generated SOPs clearer and more independently usable, and expert reviewers showed excellent inter-rater agreement (ICC = 0.91). The authors conclude that AI shows feasibility as a documentation co-author under expert oversight, though multi-center validation is needed before broader regulatory adoption.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-06-05WEQ
Cross-AI, Lean-Verified Mathematics: A Case Study on the Collatz Conjecture · Piero Borgatta
This paper reports on a five-week experiment using multiple large language models (LLMs)—Gemini, Claude, OpenAI Codex, and DeepSeek—to collaboratively attempt a non-standard mathematical attack on the Collatz conjecture, with Lean 4 serving as a formal proof checker to arbitrate claims. The project produced 302 AI-authored, sorry-free Lean theorems and lemmas covering congruential shadowing, cycle exclusion, and spectral-radius certificates, while explicitly acknowledging no proof of the conjecture was achieved. The primary methodological finding is that a cross-AI collaborative loop, when anchored to concrete formal obstructions and verified by a proof assistant, can reliably identify and document the ceiling of its own approach—including failure modes like hallucinated lemma names and syntax errors. The work is relevant to understanding how AI-assisted formal verification workflows can support rigorous quality assurance in mathematical and technical domains, and what honest negative results look like in this setting.
- ResearchFigshare2026-06-05EQC
A clause-based framework for evaluating AI-assisted SOP generation in an ISO-aligned clinical laboratory: a proof-of-concept study · Ahmed Naseer Kaftan
This proof-of-concept study tested whether ChatGPT-5 could generate high-quality standard operating procedures (SOPs) for an ISO-accredited clinical laboratory. Comparing 10 AI-assisted SOPs to matched manually written SOPs using a seven-domain ISO/CLSI-aligned rubric, the study found that AI-assisted SOPs had higher median quality scores, more complete ISO clause referencing, improved traceability, and reduced drafting time by approximately 91%. Junior staff rated the AI-generated SOPs as clearer and more independently usable, and expert reviewers showed excellent inter-rater agreement (ICC = 0.91). The authors conclude that AI can serve as a documentation co-author under expert oversight, though multi-center validation is needed before broader regulatory adoption.
- ResearcharXiv2026-06-04Q
When Better Codebooks Are Not Enough: Predictive Performance and Behavioral Reliability in LLM Political Event Coding · Zixian He, Bharath Raahul Murugesan, Patrick Brandt et al.
This paper investigates whether large language models (LLMs) can reliably follow expert codebooks to classify political events—a task requiring models to identify what actor did what to whom according to detailed rules. The authors find that converting codebooks into LLM-friendly formats (with clearer definitions, examples, and rules for edge cases) substantially improves classification accuracy, especially for fine-grained event categories. However, improved accuracy does not guarantee behavioral reliability: models can produce correct labels and recite definitions while still failing consistency tests when label names, codebook order, or label-definition mappings are changed. The study concludes that codebook-guided LLM systems must be evaluated not just on accuracy but on whether they faithfully preserve the underlying coding logic that gives structured social-science data its meaning.
- ResearcharXiv2026-06-04EP
The Geography of Algorithmic Judgment: LLM Intermediaries, Place Identity, and Racial Steering in Housing Search · Hana Samad, Trung Lam, Christoph Mügge-Durum et al.
This paper audits seven large language models (LLMs) for racial steering in housing recommendations across four U.S. cities, using iterative prompting conditions modeled on fair housing paired-testing methodologies. The authors find that steering is an emergent, context-dependent behavior arising from the interaction of a user's stated racial identity, preference articulation, and the spatial logic the model has internalized about place and opportunity—rather than a fixed property of the model. Steering was not uniform in direction or magnitude, and adding lifestyle preference context often increased or reconfigured which models exhibited steering, suggesting LLMs may interpret the same housing preferences differently depending on user identity. The findings highlight that city-level results do not generalize across markets and that local, domain-specific expertise is needed to prevent AI-mediated housing tools from undermining fair housing law.
- ResearcharXiv2026-06-04QP
AI Assistance for Human Review of Default Judgments · Theodora Worledge, Othman Bensouda Koraichi, Daniel Bernal et al.
This paper presents the Default Assistant, an LLM-based tool designed to help courts review debt collection default judgments more accurately and efficiently. An audit of 188 Los Angeles Superior Court cases found that 4% had major defects preventing judgment, 10% had inconsistencies requiring reduced judgments, and 32% had errors requiring amendment. In a controlled study with 66 law students, users aided by the Default Assistant were 6.0% more accurate and 25.9% faster per requirement than unaided reviewers, with the largest gains on document-intensive requirements reaching up to 62% error reduction and 34% time savings. The findings offer a proof-of-concept that AI assistance with cited explanations can help resource-constrained courts handle high-volume legal review more reliably.
- ResearcharXiv2026-06-04QP
What Do People Actually Want From AI? Mapping Preference Plurality · Julia Sepúlveda Coelho, Scott A. Hale
Analyzing 1,500 open-ended responses from the PRISM dataset spanning 75 countries, this paper exposes fundamental limitations of RLHF-based AI alignment by showing that people's preferences for AI behavior are highly heterogeneous and contextually nuanced. Most desired values are requested by fewer than a quarter of respondents, and even the most common value—truthfulness—is defined in incompatible ways (sourced claims, expert opinions, or contrarian views) that a single reward model cannot capture. The study also finds that features like AI guardrails and human-like behavior are outright controversial, and that binary comparison methods fail to represent contextual distinctions users draw between default and on-request behavior. These findings suggest current alignment practices systematically flatten diverse, contested human preferences into universal models, likely contributing to persistent problems like hallucination.
- ResearcharXiv2026-06-04EQ
Re-Centering Humans in LLM Personalization · Lechen Zhang, Jiarui Liu, Tal August
This paper investigates how well large language model personalization systems actually work for real users, finding a significant gap between performance on synthetic versus human data. Across three stages of personalization—extracting user attributes, selecting relevant attributes, and generating personalized responses—models consistently fall short: they struggle to extract attributes from real conversations, disagree with human judgments about relevance, and produce personalized responses that humans rate no better than generic ones, even when LLM-based judges rate them highly. The authors collected 550 human conversations and nearly 19,000 human judgments to ground these findings, and while two lightweight training interventions improved automated evaluation alignment in earlier stages, learned reward models only modestly correlated with human ratings in the final stage. The work highlights that human-aligned personalization quality is difficult to capture with current automated methods, raising important quality-assurance concerns for deployed LLM personalization systems.
- ResearcharXiv2026-06-04EP
Will the Agent Recuse Itself? Measuring LLM-Agent Compliance with In-Band Access-Deny Signals · Thamilvendhan Munirathinam
This paper introduces the 'Recuse Signal,' a lightweight in-band mechanism that allows servers to ask connecting autonomous LLM agents to voluntarily withdraw from a resource, analogous to robots.txt but for live infrastructure access. The authors implement three low-footprint adapters (SSH banner/PAM hook, PostgreSQL wire-protocol proxy, and Kubernetes admission webhook) and run a controlled experiment on a live production host. In a pilot study using OpenAI GPT-4o, GPT-4o-mini, and Claude Code, the signal achieved 100% recusal when present versus 100% task completion in a no-signal control, though the signal behaved cooperatively rather than absolutely—an explicit operator-authorization framing caused the most capable model to proceed rather than defer. The work establishes an empirical baseline for cooperative governance of autonomous agents operating on real infrastructure and releases the standard, adapters, and experiment harness publicly.