News & Research
The latest AI research and news with real-world stakes. Each item is sourced, dated, summarized in plain English and tagged by impact area, and checked against its source before it appears.
Kind
Impact area
5732 items
- ResearcharXiv2026-06-08EQ
AI Scientists Are Only as Good as Their Evidence: A Stratified Ablation of Proprietary Data and Reasoning Skills in Drug-Asset Valuation · Yinan Wang
This paper tests whether AI agent capability in drug-asset valuation is primarily limited by reasoning quality or by the evidence the agent can access. Using a controlled three-arm ablation, the authors find that adding reasoning scaffolds (structured tools, valuation playbooks, verifiers) improves calibration and audit discipline, but fails to overcome a factual ceiling imposed by evidence gaps—non-proprietary agents recover only 25–38% of a curated competitive gold record. Adding a proprietary corpus (Noah AI) raises gold-record recovery to 96% and drives 'completeness-aware decision utility' to 7.43 versus 1.76–2.57 for web-only or scaffold-enhanced agents. The key finding is that proprietary evidence sets the upper bound of what an AI scientist can know and therefore decide, which has direct implications for how enterprises deploy AI in high-stakes scientific decisions.
- ResearcharXiv2026-06-08Q
Deterministic Integrity Gates for LLM-Assisted Clinical Manuscript Preparation: An Auditable Biomedical Informatics Architecture · Yoojin Nam, Jinhoon Jeong, Namkug Kim
This paper presents MedSci Skills, an open-source architecture for verifying AI-generated clinical manuscripts by pairing LLM-based generation with deterministic integrity gates. The system uses 43 modular skills and a 21-detector deterministic tier to catch fabricated citations, numerical drift from source tables, and unmet reporting-guideline items (STARD, PRISMA, STROBE) before a manuscript advances to the next stage. In a seeded-defect test, the deterministic gates detected all 27 injected defects with no false positives, while a single-prompt LLM reviewer caught only 11, missing defects hidden in code, bibliography, and style. The architecture produces an auditable, re-executable verification trail intended to support human oversight of LLM-assisted manuscript preparation.
- ResearcharXiv2026-06-08Q
LLM-Orchestrated Conformance Checking in Stroke Care Without Computer-Interpretable Guidelines · Giorgio Leonardi, Stefania Montani, Manuel Striani et al.
This paper presents a modular LLM-based framework that performs conformance checking in healthcare — assessing whether patient care pathways follow clinical guidelines — without requiring formal Computer-Interpretable Guidelines (CIGs). Applied to stroke care at Alessandria Hospital's neurological ward, the system automatically extracts patient traces from clinical discharge letters, derives normative rules from unstructured guideline text, translates them into executable scripts, and computes a Trace Conformance Indicator. Evaluated against 50 rules and hundreds of patient traces, the framework found that more than 86% of traces were conformant, demonstrating both the feasibility of LLM-orchestrated compliance analysis and a high level of adherence to stroke care guidelines at that institution. The approach is significant because it removes a key barrier to real-world clinical guideline compliance monitoring by eliminating the need for pre-existing formal guideline representations.
- ResearcharXiv2026-06-08CP
AI Assurance in UK Defence: Challenges in Operationalising JSP 936 · Callum Cockburn, Sam Farrow
This report analyzes the practical difficulties of implementing JSP 936 Part 1, the UK Ministry of Defence's AI assurance directive, through a structured interpretive review. It identifies eight thematic challenge areas—including adequacy of evidence and argument, human interaction management, operational environment definition, systems-of-systems integration, AI performance assessment, safety and security analysis, ethicality measurement, and managing AI's inherent complexity. The authors argue that while JSP 936 provides a useful governance basis, significant unresolved technical, organisational, and assurance questions remain due to the socio-technical nature of AI-enabled systems and uncertainty in real-world deployment. The report calls for further methods, guidance, and organisational capability to support ambitious, safe, and responsible AI adoption across UK Defence.
- ResearcharXiv2026-06-08QP
Pretrained, Frozen, Still Leaking: Auditing Cross-Encoder Attribute Transfer in EEG Foundation Models · Jianwei Tai
This paper audits privacy risks in EEG foundation models (BIOT, LaBraM, and EEGPT) by evaluating their frozen embeddings across four endpoints simultaneously: raw reconstruction, membership inference, identity linkage, and differential privacy on downstream heads. The key finding is that a single attribute decoder trained on one encoder can transfer to other encoders via a linear bridge, leaking spectral attributes even when individual single-endpoint audits appear to pass, with subject-disjoint confidence interval lower bounds of at least 0.081 across all six cross-encoder directions. The authors introduce an audit-endpoint disagreement score (AEDS), which is positive across all tested datasets (p<0.001), while standard defenses including DP-SGD and LiRA membership inference (AUC 0.50–0.70) fail to close the attribute leakage channel. The framework provides a joint release decision rule that can block model releases based on cross-encoder attribute leakage rather than relying on scattered single-endpoint defenses.
- ResearcharXiv2026-06-08QP
Culturally-Adapted Red-Teaming Across East and Southeast Asian Contexts: A Methodological and Comparative Analysis · Hyeji Choi, Yongtaek Lim, Minwoo Kim
This paper investigates a critical gap in multilingual safety testing of large language models (LLMs): the common practice of directly translating English safety benchmarks into other languages fails to capture culturally embedded threats, social norms, and legal frameworks. The authors built paired direct-translation and culturally-adapted datasets for Korean, Japanese, Thai, and Khmer, finding that culturally-adapted prompts produced higher Attack Success Rates across all 16 language-model combinations (mean +9.3 percentage points), while direct translation underestimated risk in 44 of 48 category-language combinations. Cultural Realism scores for direct-translated prompts averaged only 0.17 out of 3.0, versus up to 2.51 for culturally-adapted prompts, confirming that translation-only approaches produce inputs that diverge systematically from real-world multicultural settings. The findings make a strong methodological case that valid LLM safety evaluation requires culture-specific benchmark adaptation, not merely linguistic conversion.
- ResearcharXiv2026-06-08WE
Autonomous Incident Resolution at Hyperscale: An Agentic AI Architecture for Network Operations · Arun Malik
This paper presents a deployed multi-agent AI architecture designed to autonomously detect, diagnose, and remediate network incidents in large-scale cloud infrastructure without human intervention. The system uses hierarchical agent decomposition, runbook-derived knowledge, and progressive autonomy with layered safety controls including authorization checks and rollback mechanisms. Deployed in production at a major cloud provider, it achieves autonomous resolution rates exceeding 90% for common incident categories. The work demonstrates that agentic AI can meaningfully reduce reliance on human-driven incident response at hyperscale while preserving operational safety guarantees.
- ResearcharXiv2026-06-08EQ
Context Rot in AI-Assisted Software Development: Repurposing Documentation Consistency for AI Configuration Artifacts · Christoph Treude, Sebastian Baltes
This paper introduces 'context rot,' the phenomenon where AI coding assistant configuration files (such as CLAUDE.md, AGENTS.md, and .cursorrules) become stale as software evolves, causing the persistent context guiding AI tool behavior to diverge from the actual codebase. The authors argue that decades of software documentation consistency research provides an immediate toolbox for detecting this problem, and they present a research roadmap connecting existing approaches to this new setting. As preliminary evidence, applying an existing README/wiki consistency checker to a statistically representative sample of 356 repositories found stale code element references in 23.0% of repositories, demonstrating that traditional documentation consistency tools can already surface context rot. The findings matter because undetected context rot may silently degrade the quality of AI-assisted code generation across development teams.
- ResearcharXiv2026-06-08QP
Context-Fractured Decomposition Attacks on Tool-Using LLM Agents: Exploiting Artifact Provenance Gaps · Xiaofeng Lin, Yukai Yang, Daniel Guo et al.
This paper introduces Context-Fractured Decomposition (CFD), a family of jailbreak attacks targeting tool-using LLM agents that operate across multiple steps and artifact states. Unlike existing multi-turn jailbreaks such as Crescendo and Tree of Attacks—which assume a single contiguous conversation—CFD exploits 'provenance gaps,' where harmful behavior is elicited later in a pipeline by combining individually innocuous tool actions and artifacts from earlier interactions. The authors demonstrate that CFD improves jailbreak success rates by up to 28.3 percentage points over state-of-the-art baselines, even against strong single-turn judges, and propose provenance lineage tagging as a mitigation direction. The findings reveal a critical security gap in real-world agent deployments where enforcement is fragmented across tools, modules, and time.
- ResearcharXiv2026-06-08Q
Decoy-Calibrated Failure Audits for Language Models · Vyzantinos Repantis, Ameya Gawde, Harshvardhan Singh
This paper introduces Janus, a statistical auditing procedure for identifying where language models fail and whether those failure patterns are real or artifacts of selection bias. Janus scores candidate error explanations (called descriptors) by their error-rate lift and then compares them against fake 'decoy' descriptors with the same frequencies but random assignments, confirming only those that beat the decoy baseline and replicate on held-out data. In controlled experiments on multi-table lookup tasks and two public benchmarks (MuSiQue and LongBench v2), Janus avoids false positives that uncalibrated methods report — for example, reducing 20 reported descriptors to zero confirmed findings on LongBench v2. The method matters for quality assurance because it provides a principled way to separate proposing failure explanations from credibly reporting them, reducing the risk that auditors mistake noise for real model weaknesses.
- ResearcharXiv2026-06-08Q
TRIAGE: Dialectical Reasoning for Explainable Risk Prediction on Irregularly Sampled Medical Time Series with LLMs · Hyeongwon Jang, Gyouk Chu, Changhun Kim et al.
TRIAGE is a framework that trains large language models to perform dialectical reasoning—generating rationales for competing clinical outcomes—when predicting patient risk from irregularly sampled electronic health record time series. The paper shows that existing LLM approaches collapse graded risk into overconfident binary predictions, and that TRIAGE mitigates this by producing continuous, calibrated risk scores alongside explicit clinical reasoning. Evaluated on three benchmarks, TRIAGE achieves an average 3.3% AUPRC improvement and reduces calibration error by 81% compared to competitive baselines, while its rationales outperform post-hoc explanations by 20% in clinical reasoning quality. These results matter for clinical quality assurance because better-calibrated, interpretable early warning systems give clinicians verifiable reasoning to support patient triage decisions.
- ResearcharXiv2026-06-08EP
Document-Authored Control-Signal Impersonation: A Low-Cost Indirect Prompt Attack on RAG Safety Boundaries · Jianguo Zhu
This paper studies a security vulnerability in retrieval-augmented generation (RAG) systems where attacker-controlled document text can impersonate trusted metadata, provenance, or policy signals — a pattern the authors call Document-Authored Control-Signal Impersonation (DACSI). Unlike traditional prompt injection that issues direct override commands, DACSI exploits the fact that RAG pipelines collapse trusted and untrusted text into a single natural-language channel, allowing document-authored labels to be misread as authorized control signals. The authors evaluate DACSI across six language models (including DeepSeek V4 Pro, Qwen3.5-397B, GPT-5.5, Gemini 3.1 Pro, and GLM-4.7), finding varying levels of susceptibility across model regimes and identifying source/channel separation as a key mitigation direction. This matters for enterprise and policy contexts because RAG systems are widely deployed, and this low-cost attack vector bypasses conventional safety boundaries without any explicit instruction to override policy.
- ResearcharXiv2026-06-08WQ
CARE: A Conformal Safety Layer for Medical Summarization · Suhana Bedi, Bridget Lin, Anson Y. Zhou et al.
CARE (Conformal Assessment for Risk Evaluation) is a post-hoc, model-agnostic safety layer that applies conformal risk control to flag hallucinations and omissions in LLM-generated medical summaries without requiring model retraining. It provides finite-sample, distribution-free guarantees through two controllers—one bounding the probability of unflagged hallucinated sentences and one bounding the expected fraction of important omissions missed—by jointly calibrating over a two-dimensional threshold space. Tested across five medical summarization tasks, CARE satisfies its target risk bound (α=0.15) with 95% confidence across 100 calibration/test resplits using roughly 100 labeled documents per domain, and surfaces up to 5× fewer sentences than alternative calibrated baselines. A preliminary clinician study of 75 document reviews found that calibrated flags improved omission detection by 28.6 percentage points on average, demonstrating a tunable mechanism for balancing residual risk against clinician review burden.
- ResearcharXiv2026-06-08QP
From Statute to Control Flow: Span-Grounded Deontic Trees for Defeasible Scope Parsing · Jian Chen, Siyuan Li, Chucheng Wan et al.
This paper introduces NormBench, a benchmark of 2,290 legal and policy provisions in Chinese, English, and cross-lingual settings, designed to measure whether AI models can correctly parse the nested exception structures in statutes and regulations. The authors identify a failure mode called Silent Scope Omission (SSO), where a model applies a general rule but silently drops nested exceptions, producing outputs that appear compliant but break on edge cases. They propose Span-Grounded Deontic Trees (SG-DT), a structured intermediate representation that anchors every logical branch to source text spans and requires explicit exclusion guards, enabling deterministic compilation and audit. Evaluations of frontier LLMs reveal two recurring problems—Recursion Decay (performance drops as exception depth increases) and an Auditability Trap (models retrieve relevant spans but fail to assemble correct control flow)—with SG-DT improving whole-tree fidelity and exception recovery especially on SSO-prone cases.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-06-08EQCP
Video - Artificial Intelligence Systems in Accounting and Auditing: A Bibliometric and Exploratory Analysis · Ioana Florina Coita, Laura Filip, Marius Vlad Pop et al.
This study analyzes 729 peer-reviewed articles on AI in accounting and auditing using bibliometric methods and evaluates ten AI-based solutions against a structured framework covering AI subset, functional category, entity size, audit relevance, and architectural transparency. Results show the field converges on supervised and deep-learning approaches, and that AI tools cluster into three functional categories—process automation, analytics/business intelligence, and predictive/audit-oriented systems—with adoption patterns differing by company size: SMEs benefit most from process automation and OCR, while large entities gain more from full-population analytics and ensemble-based anomaly detection. The paper also addresses trustworthiness limits of probabilistic AI systems and discusses implications for audit assurance and emerging regulatory frameworks such as the EU AI Act and ISO/IEC 42001. These findings matter for enterprise adoption decisions, quality assurance in financial workflows, and the development of AI governance and certification standards.
- ResearchSustainability2026-06-08WE
Artificial Intelligence and Export Performance in Small and Micro-Enterprises: The Roles of Internal Capability and External Tools · Mengyang Gu, Chuyue Jin
This study examines how internal AI capability and external AI tool utilization jointly affect export performance in 475 small and micro-enterprises at Yiwu International Trade City. Both types of AI resources independently improve export performance, but their interaction produces diminishing marginal returns, meaning that investing heavily in both simultaneously reduces each resource's incremental contribution under resource-constrained conditions. The findings offer practical guidance for small firms on how to allocate AI-related investments efficiently to maximize international competitiveness.
- ResearcharXiv (Cornell University)2026-06-08WP
The Jagged Global Economy: Frontier AI Unevenly Exposes National Economies · Arul Murugan, Tomás Aguirre, Abhishek Nagaraj et al.
This paper introduces a national AI exposure metric applied across 141 countries, combining occupation-level AI exposure scores with international employment data. It finds that high-income countries face substantially greater direct labor-market exposure to frontier AI than low-income countries, and that women are more exposed than men in 91% of countries due to their concentration in white-collar and sales roles. The study also uncovers an indirect exposure mechanism via remittance dependencies—countries like Tajikistan face above-average effective exposure because their GDP relies heavily on remittances from highly-exposed economies. The findings demonstrate that policy responses calibrated to U.S. or European labor markets will not generalize globally, underscoring the need for country-specific workforce and policy strategies.
- ResearchTelecom2026-06-08WEQC
HuntGPT: Integrating Machine Learning-Based Anomaly Detection and Explainable AI with Large Language Models (LLMs) · Tarek Ali, Panos Kostakos, Saeid Sheikhi
HuntGPT is a specialized intrusion detection dashboard that integrates a Random Forest classifier trained on the KDD99 dataset with Explainable AI frameworks (SHAP and LIME) and a GPT-3.5 Turbo conversational agent to make network anomaly detection more understandable and actionable. The system aims to address long-standing barriers to ML adoption in cybersecurity—namely false positives, opaque model outputs, and limited XAI acceptance—by delivering threat explanations in plain language through an interactive interface. The prototype was evaluated for technical accuracy using Certified Information Security Manager (CISM) Practice Exams and assessed for response readability across six metrics, with results suggesting that LLM-backed conversational agents integrated with XAI can generate explainable, actionable outputs for intrusion detection. This matters because it offers a path to increasing analyst trust and operational efficiency in threat hunting workflows.
- ResearcharXiv (Cornell University)2026-06-08QCP
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting · Avijit Ghosh, Anka Reuel, Jenny Chim et al.
This paper introduces EvalCards, a structured reporting layer designed to make AI evaluation results interpretable and comparable across sources like leaderboards, model cards, and benchmark papers. The authors derive a reporting schema from 52 papers and 10 stakeholder interviews, implement four interpretive signals (reproducibility, documentation completeness, provenance and risk, and score comparability), and deploy the system across 5,816 models, 635 benchmarks, and 101,843 results. The work reveals systematic gaps in current AI evaluation reporting and provides different reader modes for research and non-research audiences. This matters because inconsistent evaluation reporting makes it difficult for stakeholders to compare AI systems, verify claims, or assess risks—problems directly relevant to quality assurance, certification, and policy decisions.
- ResearchJMIR AI2026-06-08QCP
AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation · Kate H Bentley, Luca Belli, Adam M. Chekroud et al.
This study validates VERA-MH, an open-source automated benchmark for assessing whether AI chatbots respond safely to users at risk of suicide. Researchers simulated conversations spanning a wide range of suicide risk levels and had licensed mental health clinicians independently rate chatbot behaviors, then compared those ratings to an LLM-based evaluator using the same rubric. The LLM judge achieved strong alignment with clinical consensus (interrater reliability of 0.81), comparable to inter-clinician agreement (0.77), supporting VERA-MH as a reliable automated safety evaluation tool. This matters because millions of people use AI chatbots for psychological support, and a validated benchmark is critical for ensuring these tools do not harm vulnerable users.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-06-08EQCP
Video - Artificial Intelligence Systems in Accounting and Auditing: A Bibliometric and Exploratory Analysis · Ioana Florina Coita, Laura Filip, Marius Vlad Pop et al.
This paper examines AI's role in accounting and auditing through a two-stage approach combining bibliometric analysis of 729 peer-reviewed articles with an exploratory evaluation of ten AI-based solutions. The bibliometric findings show the field converges around supervised and deep-learning methods, while the applied evaluation reveals that AI tools cluster into process automation, analytics/business intelligence, and predictive/audit-oriented systems—with SMEs benefiting most from process automation and OCR, and large entities gaining more from full-population analytics and anomaly detection. The study raises important concerns about vendor black-box opacity and the trustworthiness limits of probabilistic AI, with implications for audit assurance and emerging regulatory frameworks like the EU AI Act and ISO/IEC 42001. These findings matter for enterprise adoption decisions, quality assurance in auditing, and the policy landscape governing AI in financial workflows.
- ResearchInternational Journal of Finance & Economics2026-06-08EP
How Does <scp>AI</scp> Empower Corporate <scp>ESG</scp> Practices? Mechanisms Based on Information Processing Theory · Qin Zhu, Shanshan Jiang, Anna Min Du
Using data from 2,630 Chinese A-share listed companies from 2010 to 2022, this study finds that AI adoption significantly improves corporate ESG performance. Drawing on information processing theory, the research identifies three mechanisms: AI boosts green innovation, enhances price markup capabilities, and reduces agency costs. These effects are strongest for technology-intensive firms, non-highly polluting firms, and firms in competitive industries, and are amplified by a better external capital market financing environment. The findings offer practical guidance for enterprises pursuing sustainable development through 'AI+ESG' strategies and for regulators and capital markets designing supportive governance frameworks.
- ResearchPubMed2026-06-08EQCP
Artificial Intelligence Governance in Health Systems: Systematic Review of Frameworks and Integrative Model Proposal. · Hassane Alami, Renata Pozelli Sabio, Elsury Johanna Pérez et al.
This systematic review synthesized 19 AI governance frameworks for health systems drawn from over 10,000 records across 8 academic databases, identifying six critical governance processes—including data governance, risk assessment, validation, and monitoring—as well as four relational mechanisms such as ethical principles, education, and standards. The authors propose an integrative AI governance model operating across local, national, and international levels that explicitly models interactions between governance dimensions. The findings highlight that most existing frameworks are recent, concentrated in North America, and lack primary study grounding, underscoring the need for more rigorous and globally representative governance approaches. This work is directly relevant to health policy, AI certification standards, and responsible AI integration in health enterprises.
- ResearcharXiv2026-06-07EQ
Building Customer Support AI Agents at 100M-User Scale: An Evaluation-Driven Framework · Aman Gupta, Kevin Rossell, Edesio Alcobaça et al.
This paper presents a unified framework for building and deploying customer support AI agents at Nubank, a fintech company with over 100 million users. The framework integrates context engineering, human-in-the-loop prompt iteration, LLM-based evaluation with inter-rater agreement measurement, and offline-to-online validation to accelerate development while ensuring production quality. Across five deployed use cases—card delivery, debt management, credit-limit support, card management, and product explanation—the approach delivers measurable gains, including a 37 percentage-point improvement in AI transactional Net Promoter Score and a 29 percentage-point gain in self-service rate in card delivery, with AI satisfaction approaching that of expert human agents. The central finding is that evaluation-pipeline quality directly drives iteration velocity and reliably predicts real-world outcomes.
- ResearcharXiv2026-06-07WQ
A Classroom Study of LLM-Generated Feedback Intervention in Introductory Programming · Hasnain Heickal, Andrew Lan
This classroom study deployed AI-generated feedback in a randomized protocol across an introductory Python programming course, collecting 6,693 submissions from 215 students across 17 labs (the ProgFeed dataset). Students received one of three feedback conditions—natural language hints, AI-generated failing test cases, or no AI feedback—allowing comparison of modalities on completion rates, convergence speed, and submission behavior. The study finds that natural language feedback is significantly associated with higher completion rates and faster convergence to correct solutions, while test case feedback shows heterogeneous effects depending on feedback validity. The results highlight that the form and quality of AI-generated feedback, not merely its presence, are critical to pedagogical outcomes in programming education.