News & Research
The latest AI research and news with real-world stakes. Each item is sourced, dated, summarized in plain English and tagged by impact area, and checked against its source before it appears.
Kind
Impact area
5672 items
- ResearcharXiv2026-06-18QC
Confidence Calibration for Multimodal LLMs: An Empirical Study through Medical VQA · Yuetian Du, Yucheng Wang, Ming Kong et al.
This paper investigates a key reliability problem in medical AI: multimodal large language models (MLLMs) often express confidence levels that don't match their actual accuracy, which could contribute to misdiagnosis or ignored correct advice. The authors propose a method combining Multi-Strategy Fusion-Based Interrogation (MS-FBI) with auxiliary expert LLM assessment to better calibrate model confidence in Medical Visual Question Answering tasks. Experiments across three Medical VQA datasets show the approach reduces Expected Calibration Error (ECE) by an average of 40%, making AI-assisted diagnosis more trustworthy. The findings underscore the need for domain-specific confidence calibration before deploying MLLMs in healthcare settings.
- ResearcharXiv2026-06-18EQ
Speeding up the annotation process in semantic segmentation industrial applications · Marta Fernandez-Moreno, Margarita Guerrero, Rosalia Rementeria et al.
This paper investigates how unsupervised computer vision algorithms can accelerate the pixel-level annotation process for semantic segmentation tasks in industrial materials science, specifically microstructure characterization of steel. The authors demonstrate that using unsupervised algorithms as a pre-annotation step reduces labeling time from 170 hours to 37 hours—approximately a 78% reduction—compared to annotating from scratch. They also create and release the largest public steel microstructure segmentation dataset to date (high-resolution images up to 1280x959 pixels, MIT License with permanent DOI), and provide a deep learning model trained on it, validated by field experts, and deployed in an industrial setting. These findings matter because annotation bottlenecks are a major barrier to deploying machine learning in industrial quality-assurance workflows, and quantifying the speedup from unsupervised pre-annotation is a novel contribution.
- ResearcharXiv2026-06-18CP
Measuring Biological Capabilities and Risks of AI Agents · Patricia Paskov, Jeffrey Lee, Kyle Brady et al.
This paper examines how to generate and interpret credible evidence about the biological capabilities and risks of agentic AI systems—AI that can autonomously or collaboratively perform multi-step scientific tasks. The authors synthesize current evidence on AI-enabled biological risks and introduce 'biological agentic evaluations' as a tool for assessing these systems, while highlighting that evaluation design choices around defining, designing, running, scoring, and documenting evaluations materially shape what results imply about risk. The work is aimed at helping policymakers interpret biological evaluation outputs with appropriate caution, guiding funders toward high-leverage investments in AI-biology evaluation research, and supporting biosecurity practitioners assessing emerging AI systems.
- ResearcharXiv2026-06-18QP
Open Weight AI Models Require Proportional Evaluation Approaches · Patricia Paskov, Christopher Rodriguez, Sunishchal Dev et al.
This paper argues that open-weight AI models (OWMs)—those released with publicly available weights—present distinct risk factors that existing evaluation frameworks, designed primarily for closed-weight models, do not adequately address. The authors propose four 'proportional evaluation' approaches specific to OWMs: evaluating without system-level safeguards (PE1), assessing robustness to modifications that undo model-level safeguards (PE2), testing selective capability amplification (PE3), and proxying worst-case misuse (PE4). A systematic review of 37 families of OWMs released between 2025 and April 2026 finds that only one fulfills all four criteria and most fulfill none. The paper is directed at policymakers, funders, and researchers, calling for stronger governance and evaluation standards as OWMs approach the performance levels of leading closed-weight models.
- ResearcharXiv2026-06-18QP
FinRED: An Expert-Guided Benchmark Generation and Evaluation Framework for Financial LLM Red-Teaming · Chaeyun Kim, Daeyoung Park, Junghwan Kim et al.
FinRED is an expert-guided red-teaming framework designed to evaluate the safety of large language models (LLMs) deployed in financial services, addressing gaps left by general-purpose safety benchmarks. It introduces a two-level taxonomy grounded in global standards such as FATF and EU DORA, mapping threats ranging from regulatory evasion to complex fraud, and uses a scalable pipeline that converts real financial documents into context-rich adversarial prompts. An expert-validated, finance-specific evaluation rubric reduces critical false negatives from 28 to 12 compared to generic rubrics, and the framework has been deployed in South Korea's Financial Security Institute regulatory sandbox for real-world generative AI security evaluation. FinRED matters because it provides regulators and financial institutions with a rigorous, standards-aligned tool for identifying compliance and fraud risks unique to financial LLMs.
- ResearcharXiv2026-06-18QP
REDACT: A Systematically Controlled Multilingual Benchmark for Personal Information Detection · Guneesh Vats, Anubha Agrawal, Shikha Singhal et al.
REDACT is a new multilingual benchmark for evaluating personally identifiable information (PII) detection systems, containing 13,427 records, 324,078 entity annotations, 51 entity types, and coverage of 25 languages across 9 scripts. The benchmark uses a systematically controlled design with nine generation axes and GDPR-aligned sensitivity tiers to enable fine-grained evaluation beyond simple aggregate metrics. Testing five detectors—including Presidio, GLiNER, OpenAI Privacy Filter, GPT-4.1, and Claude Sonnet 4.6—reveals that aggregate F1 scores hide architecture-dependent failure patterns: rule-based systems perform poorly on high-sensitivity data (recall of 0.07) while LLM-based detectors are more robust on those same categories. This work matters for quality assurance and policy compliance because it exposes where automated PII detectors are most likely to fail on the data that carries the highest regulatory risk under frameworks like GDPR.
- ResearcharXiv2026-06-18P
Challenges to Grassroots Organization Engagement with AI Policy · Carter Buckner, Jennifer Mickel, Nandhini Swaminathan et al.
This paper examines the barriers that grassroots and marginalized community organizations face when trying to meaningfully engage with AI policymaking. Through a case study of participatory design (PD) efforts focused on queer communities in the US, the authors describe their interactions with US policy bodies and the process of co-developing AI policy positions. They identify structural challenges—such as limited networks and lobbying power—that hinder effective participation, and offer actionable recommendations for both policymakers and community organizers seeking more inclusive AI governance.
- ResearcharXiv2026-06-18EQ
Human-on-the-Loop Orchestration for AI-Assisted Legal Discovery · Anushree Sinha, Srivaths Ranganathan, Abhishek Dharmaratnakar et al.
This paper addresses risks of deploying autonomous LLM agents in legal e-discovery, where multi-step reasoning errors can silently compound—a phenomenon the authors call 'trajectory collapse'—and potentially constitute legal malpractice. The authors propose a taxonomy of agentic failures in legal information retrieval, a four-layer verification architecture (covering planning, reasoning, execution, and uncertainty quantification), and a Human-on-the-Loop (HOTL) escalation framework. A simulation study on a synthetic e-discovery corpus finds that calibrated uncertainty thresholds can reduce privilege-waiver risk by up to 61% compared to fully autonomous deployment, while routing fewer than one quarter of documents to attorney review. These findings matter for legal enterprises adopting AI workflows, highlighting that structured human oversight can substantially mitigate compliance and malpractice risk.
- ResearcharXiv2026-06-18EQ
AgentFinVQA: A Deployable Multi-Agent Pipeline for Auditable Financial Chart QA · Aravind Narayanan, Shaina Raza
AgentFinVQA is a multi-agent pipeline for financial chart question answering that combines auditability and on-premise deployability. The system decomposes queries into planning, OCR, legend grounding, visual inspection, and verification steps, recording each in a traceable Model Evaluation Packet (MEP). On the FinMME benchmark, it achieves a +7.68 percentage-point improvement over a zero-shot baseline using a proprietary backbone (71.24% vs. 63.56%) and +4.84 pp with a locally served open-weights model, while the verifier's confidence signal helps route uncertain answers for human review (68.2% vs. 55.6% exact accuracy on confirmed vs. revised answers). This matters for regulated financial institutions that need both trustworthy, auditable AI outputs and the ability to keep client data on-premise.
- ResearcharXiv2026-06-18Q
Benchmarking Agentic Review Systems · Dang Nguyen, Wanqing Hao, Yanai Elazar et al.
This paper benchmarks agentic AI peer-review systems—OpenAIReview, coarse, Reviewer3, and a zero-shot baseline—across six large language models on real academic papers from ICLR and NeurIPS. The best configuration (OpenAIReview + GPT-5.5) achieves 83.0% pairwise accuracy in tracking paper quality against external signals like citations and acceptance decisions, and detects 71.6% of injected errors in a perturbation benchmark. A public deployment study finds user votes skew positive at 1.44 to 1, though false positives and minor nitpicks are the most common complaints. The findings suggest AI review systems can already align with human quality judgments and catch meaningful errors, but meaningful gaps remain before they can fully support or replace human peer review.
- ResearcharXiv2026-06-18EQ
What sentiment analysis can't see: Measuring whether customers were helped, and what went wrong, across 70,000 support conversations · Jason Potteiger
This study tests whether large language models can extract richer insight from customer support conversations than standard sentiment analysis, which measures tone rather than outcomes. Using GPT on over 70,000 support conversations from an online fundraising platform, the researchers estimated customer satisfaction and flagged reported problems, then validated these readings against actual customer ratings. The LLM-based satisfaction estimate correlated with ratings at 0.47 versus 0.36 for sentiment, produced fewer false alarms on unhappy customers, and revealed that tone and satisfaction disagree in 44% of conversations—including a large 'tolerated friction' segment of customers who are satisfied yet still reporting fixable problems that sentiment dashboards never surface. The findings suggest LLM annotation can ground business metrics in customer outcomes and problem causes rather than mere tonality, offering enterprises a more actionable view of support quality.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-06-18EQCP
Operationalizing Secure-by-Design AI Through Deterministic Runtime Governance · John Willis
This paper argues that current AI governance frameworks—including NIST AI RMF, NIST CSF 2.0, NIST SP 800-53, CISA Secure-by-Design, MITRE ATLAS, and OWASP Agentic AI guidance—focus on lifecycle and model-level oversight but lack concrete mechanisms for governing autonomous AI actions at the moment of execution. The authors propose a 'runtime governance' architecture that deterministically evaluates whether a specific autonomous action is admissible before it is carried out, enforcing authority validation, structural refusal, and governance continuity at the commit boundary. The paper contends that as AI systems become more autonomous and operationally consequential, execution-time governance becomes a necessary complementary layer to existing assurance frameworks, enabling accountability, auditability, and resilience. This work is directly relevant to how organizations and policymakers can operationalize secure and trustworthy AI deployments in practice.
- ResearcharXiv2026-06-18P
The Role of AI in Policy and Regulation · Vidisha Shekhawat, Pranjal Khare, Kiet Hoang Le
This chapter examines the evolving relationship between artificial intelligence and regulatory frameworks, covering the legal limitations and obstacles AI technology presents for governance. It surveys the global regulatory landscape—including the EU AI Act and segmented U.S. approaches—while exploring algorithmic accountability and AI's potential role as a policymaking tool. The chapter identifies key regulatory gaps such as questions of ineligibility and liability, reviews sector-specific frameworks, and proposes future pathways to address the emergent nature of AI through new regulatory approaches. This work matters for policy because it synthesizes foundational governance challenges and offers direction for developing coherent, adaptive legal structures around AI.
- ResearcharXiv (Cornell University)2026-06-18QCP
Generative Responsible AI Data Evaluation Schema (GRAIDES) for AI Assurance in Local Government · Ethan Knights, Christopher Conlan, Temilorun Gbolahan et al.
This paper introduces GRAIDES, an open-source data schema designed to centralize and standardize AI evaluation data across vendors for local government use. The framework addresses the problem of fragmented, inconsistently structured evaluation data by providing blueprints for code, architecture, and statistical evaluation. Case study results from Westminster City Council's AI catalogue demonstrate how the schema measures human-model alignment and detects systematic evaluator disagreement. By framing AI assurance as a data modeling problem, GRAIDES offers a practical path toward more consistent, reproducible benchmarking and oversight of generative AI systems in public sector organizations.
- ResearchThe Journal of Law Medicine & Ethics2026-06-18QCP
Prohibited AI Practices in Healthcare under the European Artificial Intelligence Act · Hannah van Kolfschooten
This paper examines how the European AI Act's Article 5 prohibitions on 'unacceptable risk' AI practices apply specifically to healthcare and public health settings. The authors argue that health-related AI tools—such as emotion recognition systems, biometric categorization, and technologies targeting vulnerable populations—may fall within these outright bans, and that existing medical and safety exceptions risk undermining the vulnerability protections the prohibitions are designed to uphold. Drawing on the European Commission's 2025 interpretive Guidelines, the paper provides a legal framework for interpreting these prohibitions and assesses their implications for ethical AI deployment in healthcare both within and beyond the EU.
- ResearchSoftware2026-06-18EQ
Iterative Audit Convergence in LLM-Managed Multi-Agent Systems: A Case Study in Prompt-Engineering Quality Assurance · Elias Calboreanu
This paper presents a case study of using AI agents to iteratively audit prompt specifications in a multi-agent LLM production system called AEGIS, which has over 7,000 lines of specification across interdependent files. Across nine audit rounds, the system surfaced 51 consistency defects, demonstrating non-monotonic convergence as edits cascaded and audit scope expanded. Partial replications using four frontier LLM vendors showed the multi-vendor approach detected all seeded defects, and inter-rater reliability was moderate-to-good (Cohen's κ = 0.80 on category). The findings matter because they establish a structured, reproducible methodology for quality assurance of complex AI system specifications, an area that has lacked formal inspection rigor.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-06-18EQCP
Operationalizing Secure-by-Design AI Through Deterministic Runtime Governance · John Willis
This paper argues that current AI governance frameworks focus on model oversight and lifecycle management but lack guidance on how to enforce governance at the moment an AI system actually executes an action. The authors introduce 'runtime governance' as a deterministic architectural layer that evaluates whether autonomous AI actions are admissible and legitimate immediately before execution, rather than relying solely on post-hoc assurance. The proposed architecture addresses accountability, auditability, and operational control by complementing existing frameworks such as the NIST AI RMF, NIST CSF 2.0, NIST SP 800-53, CISA Secure-by-Design, MITRE ATLAS, and OWASP Agentic AI guidance. The central claim is that as AI systems become more autonomous and consequential, execution-time governance is a necessary architectural component of Secure-by-Design AI.
- ResearchAcademic journal of management and social sciences2026-06-18WP
The Impact of Artificial Intelligence Development on Firms’ Educational Composition of Labor · Yanxing Shen
Using data from Chinese A-share listed firms (2014–2024), this study finds that AI development significantly shifts firms' labor demand away from low-educated workers toward high-educated workers, with each one-unit increase in AI development associated with a 0.007-unit decrease in the low-education labor share and a 0.006-unit increase in the high-education share. Technological innovation capability mediates this relationship, and the effects are stronger in developed regions and in high-technology industries (approximately 2.5 times larger than in non-high-technology industries). The findings provide empirical support for skill-biased technological change theory in the AI era and offer evidence for governments designing differentiated talent and labor market policies.
- ResearchDigital Humanities and Society Studies2026-06-18WP
Artificial Intelligence, Structural Transformation, and the Rethinking of Labour-Intensive Growth in India · Lijanshi Singh
This study analyzes how AI and automation are reshaping labor market prospects in India, finding that structural transformation remains incomplete—employment growth is concentrated in low-productivity construction and informal services rather than manufacturing. Using data from India's Periodic Labour Force Survey, the India Employment Report (2024), World Bank, OECD, and IMF sources, the authors identify a 'technological dualism' in which AI adoption is confined largely to formal, urban, high-skill environments while most workers remain in low-exposure occupations bypassed by productivity gains. The paper concludes that labour-intensive growth remains relevant but viable only if paired with technological upgrading, broad-based skill development, and supportive institutional frameworks. The findings are significant for workforce and policy debates about how developing economies can pursue inclusive growth amid rapid AI adoption.
- ResearcharXiv2026-06-17EQ
A Layered Security Framework Against Prompt Injection in RAG-Based Chatbots · Gulshan Saleem, Nisar Ahmed, Muhammad Imran Zaman et al.
This paper presents a three-layer security framework designed to defend retrieval-augmented generation (RAG) chatbots against prompt injection attacks, which OWASP ranks as the top vulnerability in LLM deployments. The framework intercepts both direct and indirect injection across the full inference pipeline: Layer 1 screens user inputs with rule-based and fine-tuned semantic classifiers, Layer 2 enforces a provenance-based instruction hierarchy to prevent poisoned retrieved documents from overriding operator policy, and Layer 3 audits model outputs before delivery. Evaluated on 5,080 samples across GPT-4o, Llama 3, and Mistral 7B, the framework reduces Attack Success Rate from 71.4% to 11.3%, outperforming the best single-layer baseline by 27.3 percentage points and a published guardrail system by 23.8 percentage points, with a 4.8% false positive rate and 61.2 ms median latency overhead. The model-agnostic middleware approach offers a practical path for enterprises deploying RAG-based chatbots to substantially reduce their exposure to prompt injection without modifying the underlying LLM.
- ResearcharXiv2026-06-17EQ
Denoising Implicit Feedback for Cold-start Recommendation · Gaode Chen, Shicheng Wang, Shikun Li et al.
This paper addresses the problem of noisy implicit feedback (e.g., clickbait, position bias) in recommender systems, specifically in cold-start scenarios where new items have limited interaction history. The authors propose DIF, a model-agnostic denoising method that infers pseudo-labels for cold items by leveraging content-similar warm items, models confidence based on content similarity, and estimates label uncertainty using relative entropy and cold-start status to correct noisy labels at the sample level. Experiments on real-world datasets support DIF's effectiveness, and the method has been deployed on the Kuaishou short video platform at billion-user scale, yielding significant improvements in commercial metrics within cold-start scenarios. The work is relevant to enterprise recommendation systems that must handle continuous item influx and noisy user signals at scale.
- ResearcharXiv2026-06-17Q
Before the Labels: How Dataset Construction Shapes Suicidality Detection in Clinical Text · Priyanshi Garg, Ishita Rao, Jieqiong Ding et al.
This paper critically examines how EHR-based datasets used to train clinical NLP models for suicidality detection embed structural assumptions that are often mistaken for neutral ground truth. Using the ScAN dataset built on MIMIC-III clinical notes as a case study, the authors show that choices like ICD-based cohort selection, single-annotator labeling, and hospital-stay-level aggregation shape what 'suicidality' means in the resulting labels. A linguistic analysis further reveals that identical labels can cover heterogeneous clinical framings that differ in temporality, negation, and uncertainty. The paper argues that clinical NLP researchers must scrutinize these embedded assumptions before treating dataset labels as reliable ground truth for detecting suicidal behaviors.
- ResearcharXiv2026-06-17EQ
Configurable Clinical Information Extraction with Agentic RAG: What Works, What Breaks, and Why · Osman Alperen Çinar-Koraş, Marie Bauer, Sameh Khattab et al.
This paper presents ACIE (Agentic Clinical Information Extraction), an on-premise agentic retrieval-augmented generation (RAG) pipeline deployed at University Medicine Essen to extract structured clinical information from heterogeneous patient records. Standard RAG approaches fail on this data due to temporal reasoning challenges, cross-document dependencies, and missing metadata, so ACIE reasons over complete patient contexts and grounds every answer in source passages for clinician verification. Evaluated across 7,326 clinician judgments in a retrospective lymphoma registry study, the system achieved a 96.5% acceptance rate, with per-type acceptance ranging from 80% to 99%. The results demonstrate that agentic RAG can meaningfully close the metadata gap in clinical AI pipelines while maintaining clinician oversight.
- ResearcharXiv2026-06-17EQ
IHBench: Evaluating Post-Interruption Recovery in Voice Agents with Structured Workflows · Ahmad Salimi, Wentao Ma, Yuzhi Tang et al.
IHBench is a new benchmark designed to evaluate how well voice agents recover after user interruptions in structured, multi-step workflows such as customer service and healthcare scheduling. The benchmark covers 10 enterprise domains, six interruption types, and scores agents on both task fulfillment and recovery quality across 27 audio-language model configurations from OpenAI, Google, and open-weight sources. Key findings show that closed-weight models are substantially more robust—degrading roughly 3.3x more slowly as conversations lengthen and showing no audio-versus-text modality gap—while open-weight models underperform on all three measured dimensions. A human validation study and cross-benchmark analysis confirm that post-interruption recovery is a meaningfully distinct capability not captured by existing benchmarks.
- ResearcharXiv2026-06-17Q
Reliability without Validity: A Systematic, Large-Scale Evaluation of LLM-as-a-Judge Models Across Agreement, Consistency, and Bias · Justin D. Norman, Michael U. Rivera, D. Alex Hughes
This paper conducts the largest known systematic evaluation of LLM-as-a-Judge models to date, testing 21 judges from nine providers across three benchmarks (MT-Bench, JudgeBench, RewardBench) with roughly 541,000 individual judgments. It finds that exact-match agreement universally overstates judge quality relative to Cohen's kappa (by 33–41 percentage points on MT-Bench), that judge rankings shift by up to 14 positions across benchmarks, and that high test-retest reliability can coexist with severe position bias in production-deployed judges. The authors distill these findings into a Minimum Viable Validation Protocol, offering concrete guidance for more rigorous LLM evaluation practices.