News & Research
The latest AI research and news with real-world stakes. Each item is sourced, dated, summarized in plain English and tagged by impact area, and checked against its source before it appears.
Kind
Impact area
5672 items
- ResearcharXiv2026-06-11QC
Cascade Classification of Dermoscopic Images of Skin Neoplasms with Controllable Sensitivity and External Clinical Validation · Elena S. Kozachok, Sergey S. Seregin, Aleksandr V. Kozachok et al.
This study compares four deep learning architectures (ViT-B/16, Swin-S, ConvNeXt-S, EfficientNetV2-S) across three classification schemes—binary, single-stage four-class, and a two-stage cascade—for identifying malignant skin neoplasms in dermoscopic images. Models trained on open ISIC Archive data achieved strong internal performance (ROC-AUC 0.952–0.966) but showed a marked generalization gap on independent Russian clinical datasets, with AUC dropping to 0.797–0.893, sensitivity falling to 0.53–0.67, and calibration error (ECE) rising from 0.02 to 0.27–0.39. The cascade approach improved macro F1 over single-stage classification for most architectures by recovering malignant lesions misassigned to the dominant benign class, and its tunable triage threshold enables sensitivity control not available in standard argmax classification. The authors conclude that external clinical validation and recalibration are mandatory before deployment in real clinical settings.
- ResearcharXiv2026-06-11EP
MÖVE: A Holistic LLM Benchmark for the German Public Sector · Camilla Dalerci, Thilo Michael, Robin Schaefer et al.
MÖVE is a holistic benchmark designed to evaluate 39 large language models (LLMs) specifically for the German public sector, addressing the gap left by existing benchmarks that are predominantly English-centric, US-focused, and limited to task performance. It assesses models across both performance criteria (summarization, question answering, topic extraction) and governance criteria (hallucination tendencies, energy consumption, provider transparency, and alignment with German constitutional values and political party positions), using ten German-language datasets. Key findings show that no single model dominates across all criteria, top performers vary by task, and model size alone is a poor predictor of quality. This matters for public administrations seeking principled, evidence-based model selection rather than ad hoc choices when deploying LLMs in government contexts.
- ResearcharXiv2026-06-11P
"Is This Not Enough?": Asymmetries in Institutional Accountability and Collective Sensemaking in the Case of Canada's Algorithmic Visa Triage System · Dipto Das, Matthew Tamura, Syed Ishtiaque Ahmed et al.
This paper investigates Canada's Immigration, Refugees and Citizenship Canada (IRCC) algorithmic triage system for temporary resident visa (TRV) applications, analyzing both the official Algorithmic Impact Assessment (AIA) and Reddit discussions among applicants. Using the ADMAPS framework alongside mixed methods, the study finds three key asymmetries between institutional accountability structures and applicant experiences: epistemic asymmetry in access to decision logic, jurisdictional asymmetry shaped by geopolitical positioning, and temporal-relational asymmetry in how uncertainty and waiting are experienced. Applicants engage in collective sensemaking on peer platforms to interpret opaque algorithmic decisions, filling gaps that formal disclosure frameworks do not address. The findings highlight how algorithmic governance in transnational migration produces structured inequities not captured by institutional transparency mechanisms, calling for a shift toward examining the uneven distribution of lived experiences with public-sector AI.
- ResearcharXiv2026-06-11QP
No Hidden Prompts Needed! You Can Game AI Peer Review with Presentation-Only Revisions · Xu Yang, Zhizhou Sha, Junbo Li et al.
This paper investigates whether AI peer-review systems can be manipulated without any hidden prompts, injected instructions, or changes to scientific content. The authors introduce 'adversarial repackaging,' a closed-loop attack that modifies only presentation-level elements—such as the abstract, framing, related work, and narrative structure—while leaving all methods, results, and figures unchanged. Across three mainstream AI reviewers, this approach achieves a 75.1% attack success rate and a mean score gain of +1.21/10, revealing that AI reviewers can mistake the appearance of addressing a limitation for actually resolving it. The findings show that paper presentation itself has become an optimization surface, posing a serious policy-relevant risk to the integrity of AI-assisted peer review.
- ResearcharXiv2026-06-11WP
Fault Lines: Navigating Ethics and Responsible AI Where National Policy Meets Local Practice in Public Sector Transformation · Sitong Lyu, Shabnam Taghiyeva, Mohit Kukadia et al.
This paper investigates how responsible AI policy is translated into practice at the interface between UK central government and local authorities, using Special Educational Needs and Disabilities (SEND) as a high-stakes case study. Through thematic analysis of 17 semi-structured interviews with policymakers, practitioners, and third-sector professionals, the authors identify five key challenges: shadow AI usage and data privacy risks, market-government asymmetry in AI provision, insufficient workforce readiness, a lack of standardised definitions and measurements, and gaps in human accountability. The SEND context sharpens these tensions because high-stakes decisions affecting vulnerable children and families intensify concerns around fairness, accountability, and human oversight. The authors argue that responsible public sector AI requires both national policy adjustments and structural reforms to institutional capacity, values, and governance at the local level.
- ResearcharXiv2026-06-11EQ
SkillChain: Closing the Loop on Skill Evolution for Image-Based E-Commerce AI Assistants · Yimin Hu, Mengtao Xu, Hao Guo et al.
SkillChain is a production system for automating the lifecycle of 'Skills'—per-intent behavioral specifications—for image-based AI assistants on e-commerce platforms. Because a single uploaded image can trigger very different user intents (product search, style recommendation, visual encyclopedia, tool calls), a generic LLM conflates these modes and fails domain quality standards. SkillChain closes the feedback loop through three automated stages—Skill Creator, Route Optimizer, and Body Refiner—using dual-path LLM-Judge evaluation to iteratively refine each Skill. Deployed at production scale, the system substantially improves aggregate response quality (with the strongest gains in structural compliance and content quality) and, in a one-week online A/B experiment, yields significant gains in user engagement, content consumption, and long-term retention.
- ResearcharXiv2026-06-11EQ
Iterating Toward Better Search: A Two-Agent Simulation Framework for Evaluating Agentic Search Architectures in E-Commerce · Jetlir Duraj, Jayanth Yetukuri, Shuang Zhou et al.
This paper introduces a two-agent simulation framework for evaluating AI-powered conversational shopping assistants in e-commerce, pairing a configurable 'buyer' agent with interchangeable 'responder' architectures connected to a real search API. Across 2,011 simulated conversations and 14 persona types, the authors find that rolling-window memory outperforms intent-extraction memory on quality metrics while being 35% faster, and that targeted fixes after failure analysis reduced failure and near-failure rates by 62%. Swapping the LLM backbone from Gemini 2.5 to Llama 3.3 70B costs 0.16–0.45 quality points, and the paper also documents systematic disagreement between frontier LLM judges (Gemini vs. Claude) in how they score outcomes. The framework enables rapid, controlled iteration on agentic search architectures, which is directly relevant to enterprises building and evaluating AI shopping tools.
- ResearcharXiv2026-06-11QP
SafeLLM: Extraction as a Hallucination-Resistant Alternative to Rewriting in Safety-Critical Settings · Julia Ive, Felix Jozsa, Evridiki Georgaki et al.
This paper evaluates extraction-based approaches as alternatives to free-form rewriting in retrieval-augmented generation (RAG) systems used to query safety-critical organisational documents such as NHS acute care, oncology, and NICE guidelines. The authors compare multiple prompting strategies—including line-number-based source selection, safety-annotated sentence extraction, and multi-stage filtering pipelines—finding that line-number selection achieves the strongest overall performance, with term recall up to 95% and close alignment to source text. Safety-oriented strategies improve precision but introduce systematic omissions, and performance varies with document structure. The findings matter for healthcare and compliance settings where hallucinations in AI-generated responses to policy or procedural queries can pose direct safety risks.
- ResearcharXiv2026-06-11WQ
(Human) Attention Is (Still) All You Need: Human oversight makes AI-assisted social science reliable · Chen Zhu, Xiaolu Wang, Weilong Zhang
This paper investigates how structuring human oversight of large language models (LLMs) affects the reliability of AI-assisted social science research. The authors introduce Human-in-the-Loop Economic Research (HLER), an architecture that requires LLMs to reason but not execute data work, enforces deterministic data handling, and installs three human decision gates in the workflow. In a pre-specified factorial experiment with 280 research runs, an unconstrained multi-agent baseline produced critical failures in 72% of runs, while HLER reduced that rate to 16% (Fisher's exact test p<0.001). The findings suggest that reliable AI-assisted research depends on deliberate allocation of cognitive labour between humans and machines, not just model capability alone.
- ResearcharXiv2026-06-11QC
Acquisition state behaves as a structured, measurable variable governing lung-nodule AI: kernel-driven measurement instability and noise-driven detection fragility, invisible to DICOM metadata · Daniel Soliman
This paper demonstrates that CT acquisition parameters—specifically reconstruction kernel and noise level—act as structured, measurable variables that systematically alter the performance of a lung-nodule AI detector (a LUNA16-trained MONAI RetinaNet), in ways invisible to standard DICOM metadata. Kernel differences alone flipped Fleischner size categories in 5.2% of real paired nodules without changing detection confidence, while noise degraded detection sensitivity (especially for sub-6mm nodules) without affecting measurement—two distinct failure modes on dissociated axes. A 4-feature pixel fingerprint recovered reconstruction identity with AUC ~0.95–0.995 across vendors and phantoms, outperforming the ConvolutionKernel DICOM tag entirely. The authors argue this 'acquisition-aware' input-side validation is a currently missing but necessary layer for the acceptance-testing and drift-monitoring requirements now entering imaging-AI accreditation frameworks such as the 2026 ACR-SIIM Practice Parameter and ACR Assess-AI registry.
- ResearcharXiv2026-06-11QP
The Containment Gap: How Deployed Agentic AI Frameworks Fail Public-Facing Safety Requirements · Md Jafrin Hossain, Mohammad Arif Hossain, Weiqi Liu et al.
This paper audits three major agentic AI frameworks—LangChain, AutoGPT, and OpenAI Agents SDK—against six structural safety principles and finds that none of them provide native architectural containment guarantees. In a simulated government benefits agent built on LangChain, a single memory-poisoning attack caused persistent targeted corruption across all tested configurations, driving the wrongful denial rate for targeted applicants to 88.9%, while preserving overall accuracy and thereby evading standard monitoring. The authors introduce two lightweight mitigations—a memory integrity validator and a policy gate—that eliminate both attack vectors with under 0.2ms overhead per call. The findings suggest current agentic frameworks are not secure-by-default for high-stakes public-facing deployments such as government services, healthcare, or financial advising.
- ResearchJournal of Business Economics2026-06-11WEP
The AI workforce and firm maturity: old firms, new tech · Pattanaporn Chatjuthamard, Pornsit Jiraporn, Pandej Chintrakarn et al.
This paper examines how firm age shapes AI workforce adoption using a novel dataset combining resume and job posting data for U.S. firms. The authors find that a one standard deviation increase in firm age reduces the share of AI workers by 5.2%, with entrenched practices and resistance to change as likely barriers. R&D investment, infrastructure upgrades, and board composition—particularly the presence of female and minority directors—are shown to influence AI talent integration, while firms that successfully adopt AI see gains in market valuation and operational efficiency. The findings offer practical guidance for managers and policymakers helping mature organizations navigate technological transformation.
- ResearchJournal of Education & Social Policy2026-06-11WQCP
Artificial Intelligence in Education: Ensuring Equity and Responsible Student Use Through District Policy · Rashad Bigham, Gbolahan Solomon Osho
This paper examines how K–12 school districts are struggling to keep pace with the rapid adoption of AI tools—such as personalized learning systems and generative content platforms—leaving gaps in equitable access, academic integrity, and student data privacy. The authors identify three major risk domains and compare district-level, statewide, and public–private partnership policy models, ultimately recommending a hybrid AI Equity and Ethics Framework that pairs statewide standards with local flexibility. The proposed framework includes AI literacy curricula, educator ethics certification, transparency audits, and public accountability reporting to ensure AI adoption benefits all students responsibly.
- ResearcharXiv (Cornell University)2026-06-11QCP
Exploring Systems-Thinking Approaches to Loss of Control Risk · Aurelio Carlucci, Sean P. Fillingham, James Walpole et al.
This paper applies established systems-safety engineering methods—STECA, STPA, and FRAM—to the problem of losing control over agentic AI systems deployed internally at frontier AI labs for coding and research tasks. The authors find that model-level evaluations alone can miss critical risks: governance responsibilities may be externally unverifiable, monitoring delays can render control actions ineffective, and routine operational variability can gradually erode safeguard calibration and independence. They argue that frontier-AI risk management must combine model-focused evaluations with systems-level hazard analysis and ongoing operational assurance to verify that controls remain effective over time.
- ResearchMultidisciplinary Reviews2026-06-11EQC
Impact of industry 4.0 on the development of sustainable building projects under LEED certification · Galo Anibal Espinosa Chávez, Abel Remache
This systematic review examines how Industry 4.0 technologies—including BIM, IoT, digital twins, AI/ML, blockchain, and 3D printing—can accelerate compliance with LEED v4.1 BD+C sustainability certification in the construction sector. Analyzing 88 high-quality studies from a Scopus search (2020–2025), the authors find that BIM, IoT, and modular prefabrication show the highest technological maturity (TRL 8–9) with measurable impacts on energy use, materials, and indoor environmental quality, while digital twins and AI/ML operate at intermediate maturity (TRL 7). The study proposes a Technology×LEED matrix to prioritize interventions using verifiable performance metrics such as kWh/m²·year, percent waste reduction, and recycled content. These findings offer a structured, replicable framework for decision-making in sustainable construction and LEED certification attainment.
- ResearchCognizance Journal of Multidisciplinary Studies2026-06-11WEQ
Utilization of Artificial Intelligence Technology among Accounting Firms in Isabela, Cagayan Valley: Towards Operational Efficiency · Rhodilet Batarao-Valdez
This study examined how accounting firms in Isabela, Cagayan Valley use AI to improve operational efficiency, finding that AI is currently applied mainly to repetitive and routine tasks rather than higher-level work. Perceived benefits include greater efficiency, accuracy, compliance, and decision-making, but overall adoption remains in early stages due to workforce skill gaps and inadequate infrastructure. The research recommends targeted training, investment, and institutional support to advance AI integration in the accounting sector.
- ResearcharXiv2026-06-10QP
Prefill Awareness in Large Language Models · Andy Wang, Parv Mahajan, David Demitri Africa et al.
This paper investigates whether large language models can detect when their prior assistant-side outputs have been inserted or edited—a capability the authors call 'prefill awareness.' Using a binary preference benchmark across three prefill mechanisms, the researchers find that frontier models show substantial prefill awareness: for example, Claude Opus 4.5 detects opposing prefills in 9–35% of cases with a 0% false positive rate, and models often revert to baseline behavior without explicitly flagging the manipulation. Controlled ablations reveal that stylistic mismatch primarily drives whether a model labels a prefill as foreign, while preference mismatch primarily drives whether it reverts to baseline answers. Because safety evaluations, alignment studies, and AI control protocols commonly rely on prefilling, these findings suggest prefill awareness is already a meaningful confound for such methods, and the authors recommend that model developers actively track this capability.
- ResearcharXiv2026-06-10QP
Rethinking Psychometric Evaluation of LLMs: When and Why Self-Reports Predict Behavior · Rafal Kocielnik, Pengrui Han, Peiyang Song et al.
This paper investigates when and why psychometric self-reports (SR) can reliably predict the actual behavior of large language models (LLMs), a key concern for safe deployment. The authors compare broad personality traits (Big 5) against the Theory of Planned Behavior (TPB), which targets intentions to specific behaviors, finding that SR-behavior coherence exists but is selective: TPB reaches human-level coherence within a shared conversation while Big 5 does not, and cross-conversation coherence survives only for behaviors anchored in training (e.g., implicit bias) but collapses for context-driven behaviors like sycophancy. Persona prompting increases self-report consistency but does not align behavior. The findings suggest that coarse personality frameworks are inadequate tools for predicting LLM deployment behavior, and that more task- and behavior-specific instruments are needed.
- ResearcharXiv2026-06-10QP
Does AI Reviewer See the Full Picture? Attacking and Defending Multimodal Peer Review · Xinyu Zhao, Rana Muhammad Shahroz Khan, Zhen Xu et al.
This paper introduces PaperGuard, the first benchmark for evaluating and defending AI-assisted peer review against adversarial attacks that exploit both text and figures in scientific papers. The authors demonstrate that current large language model (LLM) and multimodal LLM (MLLM) reviewers are broadly vulnerable to domain-specific attacks—such as prompt injections and image perturbations—designed to inflate review scores rather than bypass general safety filters. To address this, PaperGuard provides a multimodal peer-review dataset, a suite of black-box and white-box attacks targeting text and figures, and a chunk-based embedding defense that localizes and neutralizes harmful instructions. The work establishes foundational protocols for building more trustworthy AI-assisted scholarly review systems.
- ResearcharXiv2026-06-10EQ
SMSR: Certified Defence Against Runtime Memory Poisoning in Persistent LLM Agent Systems · Tarun Sharma
This paper identifies a novel attack called Multi-Session Memory Poisoning (MSMP), where an adversary using only normal interaction channels can inject crafted memories into persistent RAG-based LLM agents, steering future users' responses without altering model weights or code. The authors introduce Signed Memory with Smoothed Retrieval (SMSR), the first defence offering a certified robustness bound against this threat, combining HMAC-SHA256 provenance signing at write time with randomised memory ablation and majority voting at query time. Across 15 enterprise scenarios with 3,150 trials, Component 1 reduces attack success from 93–100% to 0% for unsigned injection variants, while Component 2 holds authenticated adversary success to 8.0% (95% CI [5.8, 10.9], n=450); in an end-to-end live-agent test, SMSR cuts attack success from 65.3% to 5.3%. The work matters for enterprise deployments of persistent AI agents, where memory integrity and certified safety guarantees are critical for trust and reliability.
- ResearcharXiv2026-06-10EQ
Deployment-Centered Evaluation: Predicting Query-Level Rejection Risk in a Clinical LLM System · Alyssa Unell, Miguel Fuentes, Brenna Li et al.
This paper presents a deployment-centered evaluation framework for a large language model system embedded in electronic health records at an academic medical center. The authors train a pre-response classifier that predicts, before a response is generated, whether a user (clinician) will reject the LLM's output, achieving an AUROC of 0.719 over a prospective 4.5-month evaluation period. A key finding is that incorporating deployment-specific context—such as provider type, department, and which language model generated the response—improves rejection risk prediction beyond using query content alone. The work demonstrates the feasibility of targeted guardrails and abstention strategies to improve real-world clinical LLM utility, addressing blind spots left by static benchmarks that measure correctness rather than user acceptance.
- ResearcharXiv2026-06-10QP
"Did you lie?" Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms · Alan Cooney, David Africa, Geoffrey Irving
This paper evaluates four lie-detection methods for large language models—a chain-of-thought judge, a logprob classifier, and two activation probes including the new Did-You-Lie (DYL) method—across 31 open-weight models ranging from 2B to 1T parameters and a set of 13 reasoning 'model organisms' whose hidden beliefs are verified via chain-of-thought. The authors find that while all four detectors improve with model scale on prompted-lying tasks, every activation- and logprob-based detector degrades sharply when tested on trained model organisms, with only the chain-of-thought judge achieving strong performance (0.82 balanced accuracy). The study concludes that current lie detectors cannot yet support high-confidence claims about model beliefs, highlighting a critical gap for AI auditing and monitoring applications. The authors release datasets, model organisms, and trained detectors to support further research.
- ResearcharXiv2026-06-10CP
Which Models Are Our Models Built On? Auditing Invisible Dependencies in Modern LLMs · Sanjay Adhikesaven, Haoxiang Sun, Sewon Min
This paper introduces ModSleuth, an agentic system that automatically reconstructs the hidden dependency graphs of large language models (LLMs) by tracing how models rely on other models for data generation, filtering, output judging, and development decisions. Applied to four LLM releases, ModSleuth recovers 1,060 source-verified dependencies and surfaces issues such as multi-hop license obligations, train-evaluation coupling, and discrepancies between released and training-time artifacts. The work highlights that modern LLM development ecosystems are too complex and recursively deep for humans to trace manually, making automated auditing essential. This matters for policy and certification because undisclosed or poorly documented model dependencies can create compliance risks and undermine transparency in AI development.
- ResearcharXiv2026-06-10EQ
Atlas H&E-TME: Scalable AI-Based Tissue Profiling at Expert Pathologist-Level Accuracy · Kai Standvoss, Miriam Hägele, Rosemarie Krupar et al.
Atlas H&E-TME is an AI system built on pathology foundation models that analyzes hematoxylin and eosin whole-slide images to predict tissue quality, tissue region, and cell type labels across multiple cancer types, generating over 4,500 quantitative readouts per slide at cell-level resolution. The system was validated using a dual framework: an IHC-informed multi-pathologist consensus protocol for depth (which improved inter-rater agreement over H&E-only annotation) and benchmarking against more than 200,000 high-confidence pathologist annotations across 1,500+ cases spanning eight cancer types and 25+ sources. Against the IHC-informed consensus, Atlas H&E-TME matches or exceeds pathologist H&E-only performance and generalizes robustly across diverse morphological and technical conditions. This matters because it transforms the most ubiquitous data in pathology into a scalable, quantitative tool for tissue-based biomarker discovery in translational and clinical research.
- ResearcharXiv2026-06-10EP
A Five-Plane Reference Architecture for Runtime Governance of Production AI Agents · Krti Tallam
This paper presents a reference architecture for governing AI agents running in production enterprise environments, where traditional security controls designed for data boundaries fail to address risks that emerge from sequences of AI-initiated actions across tools, connectors, and systems of record. The architecture is built around a five-plane decomposition—a reasoning plane for adjudicating intent plus four enforcement planes covering network, identity, endpoint, and data—combined with composite principals, capability attenuation through delegation chains, and a tamper-evident audit substrate. The authors define six interruption primitives, four correctness invariants, and demonstrate that the architecture forecloses seven identified production-agent threats across five workflows; a reference implementation shows adjudication running in single-digit microseconds with attenuation correctness and evidence reconstructability holding on every trial. The work matters because it provides enterprises with a concrete, measurable governance framework for delegated AI agent action—distinct from model behavior—addressing a gap current policy engines cannot fill.