News & Research
The latest AI research and news with real-world stakes. Each item is sourced, dated, summarized in plain English and tagged by impact area, and checked against its source before it appears.
Kind
Impact area
5672 items
- ResearcharXiv2026-06-16WQ
Agentic AI Enhances Physician Trust in Clinical Decision Making · Zhiling Yan, Zhe Fang, David J King et al.
This study examines whether agentic AI—which autonomously invokes external tools and makes its intermediate reasoning steps transparent—earns greater physician trust than non-agentic AI in clinical decision-making. Three physicians evaluated 315 multimodal clinical cases, finding significantly higher cognitive (process-oriented) and behavioral (outcome-oriented) trust for the agentic model (P < 0.001), with physicians preferring agentic reasoning in 89.57% of treatment planning cases. However, the study also identifies measurable over-reliance on incorrect agentic outputs, showing that transparency in decision logic alone is insufficient and that rigorous clinician oversight remains essential.
- ResearcharXiv2026-06-16EQ
Fine-tuning LLMs for Passive Depression Severity Estimation from AI Mental Health Dialogue · Olivier Tieleman, Ziyi Zhu, Ting Su et al.
This paper fine-tunes a large language model (Qwen3.5-27B) to predict PHQ-9 depression severity scores directly from transcripts of user conversations with an AI mental health application, requiring no additional clinical data. Using a dataset of 6,283 users built by augmenting 3,111 ground-truth labels with pseudolabels, the best model achieves a Pearson correlation of 0.80 and AUC of 0.91 at the clinically relevant PHQ-9 ≥ 10 threshold, with AUC above 0.87 across every severity level tested. The work demonstrates that passive, continuous depression monitoring is feasible from routine AI-generated conversation text alone, potentially reducing reliance on self-report measures that suffer from low completion rates and response bias.
- ResearcharXiv2026-06-16QP
Legal Reasoning Is Not Lawyering: Rethinking Legal Benchmarks for Pro Se Access to Justice · Andrew Lou, David Shin
This paper critiques the assumption that large language models (LLMs) can improve access to justice for pro se litigants—people without legal representation—by arguing that current legal AI benchmarks measure only an upper bound of model performance using expert-preprocessed inputs. The authors contend that pro se users submit noisy, incomplete, or informally worded prompts that mirror known LLM failure conditions such as hallucination, long-context sensitivity, and typographical perturbations, yet no benchmark currently measures robustness under these realistic conditions. Using a perturbation experiment on the LEXam legal benchmark, the paper illustrates the performance gap between expert-curated and pro se-like inputs. The authors call for new legal benchmarks specifically designed to test robustness under pro se conditions so that access-to-justice claims about legal AI can be empirically validated rather than assumed.
- ResearcharXiv2026-06-16WQ
ASTRA: A Scalable Next-Generation ATCO Training Simulator with Autonomous Simpilots · Ethan Chew, Enjia Wu, Iruss Eng et al.
ASTRA is an end-to-end AI simulator designed to automate the 'simpilot' role in Air Traffic Control Operator (ATCO) training, replacing specialized human trainers who role-play pilots and controllers in simulated airspace. The system combines a locally fine-tuned Automatic Speech Recognition pipeline—reducing Word Error Rate from up to 107.80% to 23.45% on Singaporean-accented aviation speech—with an AI-assisted performance evaluation framework that scores trainee radiotelephony communications on accuracy (91.7%), brevity (88.2%), and completeness (86.9%). Built on open-source tools including DSPy and Unsloth, ASTRA enables scalable, standardized ATCO assessment while reducing the workload on human instructors. The work directly addresses training capacity constraints in aviation by demonstrating that locally adapted AI can outperform Western-centric off-the-shelf speech models in operational contexts.
- ResearcharXiv2026-06-16WP
FairTutor: Equity-Aware Pedagogical LLM Routing for Budget-Constrained AI Tutoring · Qingyang Xu
FairTutor is a multi-agent AI tutoring framework designed to close the quality gap between students with access to premium AI models and those limited to free or low-cost services. It combines query analysis, pedagogical planning, low-cost model generation, evaluator-guided critique and revision, and selective escalation to premium models. Empirical evaluations show FairTutor achieves 97.1% of premium pedagogical quality (measured via floor-adjusted Likert scale) while reducing serving cost by 71.6%, as assessed on TutorAccessEval, a benchmark covering math, reading, writing, science, and language learning. The work directly addresses AI-driven education inequity and offers a tunable cost–quality Pareto frontier adaptable to diverse student populations.
- ResearcharXiv2026-06-16QP
The Slop Paradox: How Synthetic Standardization Erodes Clinical Uncertainty and Cross-Modal Alignment in AI-Rewritten Radiology Reports · Samar Ansari
This paper investigates how AI-assisted rewriting of radiology reports degrades clinical information, using 450 chest X-ray reports from the Indiana University dataset processed through three LLM rewriting tasks: EHR summarization, standardized rewriting, and teaching case preparation. The central finding is a 'slop paradox': EHR summarization causes the most entity erosion (51.4% of clinical entities, 43.7% of hedging language lost) but barely affects image-text alignment (2.5% drop), while standardized and teaching-case rewrites preserve more entities but cause 14.9–16.5% alignment drops—six to seven times larger. Contrary to expectations, rare pathologies were not preferentially degraded; the dominant driver of degradation is the rewriting task type, not the clinical content. These findings have direct implications for the governance of AI-assisted clinical documentation and the construction of multimodal medical AI training datasets.
- ResearcharXiv2026-06-16WQ
Toward Accessible Psychotherapy Training Using AI-Driven Interactive Patient Avatars · Pascal Riachi, Sofie Kamber, Stella Brogna et al.
This paper presents an AI-driven training system for psychotherapists learning Acceptance and Commitment Therapy (ACT), using large language models to simulate realistic virtual patients derived from real therapy sessions and providing automated turn-by-turn feedback on therapist responses based on established ACT fidelity criteria. Expert evaluation with practicing psychologists confirmed high realism in patient behavior, and quantitative testing across 49 therapy transcripts found GPT-4o-mini achieved the lowest mean absolute error (MAE = 6.12) in replicating human supervisor fidelity ratings. The system is designed to complement rather than replace supervision, offering a scalable, low-risk environment for deliberate practice and immediate feedback. This matters for workforce development by expanding access to standardized psychotherapy training that is otherwise constrained by ethical, logistical, and resource barriers.
- ResearcharXiv2026-06-16QC
Vision-language models for chest radiography do not always need the image · Mahshad Lotfinia, Sebastian Ziegelmayer, Lisa Adams et al.
This paper challenges the assumption that high benchmark accuracy on chest radiograph tasks means medical vision-language models (VLMs) actually use the image. Using a causal audit involving image occlusion, irrelevant-region occlusion, and patient-image swapping, the authors show that a text-only model with no image access comes within 5.7 accuracy points of the best multimodal system, and a 119-billion-parameter multimodal model is statistically indistinguishable from a 7-billion text-only baseline. The audit categorizes nine systems: three that ignore the image entirely, one that is unstable, and five that use it selectively, with findings replicated across a second dataset, resolution, and prompt phrasing. The authors conclude that grounding audits—not accuracy—should gate clinical deployment, as reported confidence only flags ungrounded answers when a model genuinely uses the image.
- ResearcharXiv2026-06-16WP
Mapping the Artificial Intelligence Divide in Africa: Infrastructure, Accessibility and Capacity · Abayomi O. Agbeyangi, Jose M. Lukose
This paper empirically maps the 'AI divide' in Africa across three dimensions: physical infrastructure, accessibility, and human capacity. Key findings include only 38% internet penetration, less than 1% of global data centres located in Africa, high data costs relative to income, gender-based digital divides, and a lack of NLP models supporting African languages. Despite these barriers, the paper identifies positive grassroots trends such as local startups and university-led AI initiatives. Based on these findings, the authors offer concrete policy recommendations to foster a more equitable and comprehensive AI ecosystem across the continent.
- ResearcharXiv2026-06-16EQ
FacProcessTwin: An LLM-Based System for Process Twin Development · Yash Pulse, Yong-Bin Kang, Abhik Banerjee et al.
FacProcessTwin is an LLM-based system that automates the development of process digital twins for manufacturing facilities by extracting process models from plant documentation and natural-language operator input, then binding those models to live operational data. In a real-world case study with an Australian food manufacturer covering 16 production process flows, the system achieved a mean F1 score of 95.2% for process model accuracy and reduced twin development time to roughly one-sixth of the manual effort. A human-in-the-loop governance layer ensures safety-critical data bindings remain correct: at ambiguous points where a baseline approach mis-binds 75% of the time, FacProcessTwin defers to the operator and achieves zero mis-bindings. The work demonstrates that LLMs can substantially lower the cost and time of deploying process twins while maintaining the accuracy and safety oversight required in manufacturing environments.
- ResearcharXiv2026-06-16Q
Understanding LLMs in Title-Abstract Screening: From Disagreements to Recommendations · Mika Mäntylä, Patricia Matsubara, Katia Romero Felizardo et al.
This study investigates why large language models (LLMs) fail at title-abstract screening in systematic reviews (SRs) for software engineering, going beyond simple accuracy metrics to qualitatively analyze disagreements between LLMs and human researchers across six SRs and over 1,000 papers. Screening was performed independently by humans and LLMs in zero-shot mode, yielding Kappa agreement values ranging from 0.52 to 0.77, and qualitative analysis identified recurring failure causes including boundary ambiguity in key terms, keyword overemphasization, and incorrect topic inference. Based on these findings, the authors propose actionable recommendations such as validating semantic understanding before deployment, running multiple LLMs, and focusing validation on borderline cases. The work highlights that community-level normative guidelines for LLM use in systematic reviews are still needed, with implications for research quality assurance and evidence synthesis workflows.
- ResearcharXiv2026-06-16EQ
Scaling Enterprise Agent Routing: Degradation, Diagnosis, and Recovery · Kellen Gillespie, Robyn Perry
This paper investigates how routing accuracy degrades in a deployed enterprise AI assistant as the number of specialized agents and tools grows. Testing three frontier language models on a catalog of 110 agents and 584 tools, the authors find that routing F1 on under-specified requests drops 16–23 percentage points as the catalog scales. They decompose this degradation into a retrieval gap and a confusion gap, and show that embedding-based shortlisting recovers 10–11 percentage points of F1 at full scale. A production annotation study with 1,435 human-labeled utterances confirms real-world recovery of 10–17 percentage points, validating the approach on live traffic.
- ResearcharXiv2026-06-16EQ
LLM-as-Judge in Education: A Curriculum-Grounded Marking Pipeline · Xiwei Xu, Chen Wang, Jacky Jiang et al.
This paper presents a curriculum-grounded LLM-as-Judge pipeline for automated marking of student responses in high-stakes exam preparation contexts. The pipeline grounds LLM outputs in authorised curriculum artefacts—such as syllabus verbs, performance band descriptors, glossary definitions, and marking-guideline principles—to generate question-specific rubrics and evaluate student answers. Preliminary evaluation shows the pipeline produces marking outcomes comparable to human tutors, with justifications more traceable to official curriculum standards. The system has been integrated into an online study platform, with early deployment data offering initial insights into operational usage and manual overrides.
- ResearcharXiv2026-06-16EQ
Simulated Customers Never Walk Away: Decision Fidelity of LLM User Simulators Measured Against Real Purchase Outcomes · Liang Chen
This paper investigates whether large language model (LLM) user simulators accurately replicate the decision-making behavior of real customers in high-stakes conversational settings. Using 2,790 production conversations between an LLM sales agent and real customers—including 793 with verified payment outcomes—the authors find a systematic 'disengagement deficit': simulators closely reproduce the behavior of actual buyers but significantly inflate non-buyers toward purchase-oriented engagement, halving expressed resistance (25.1% to 13.5%) and nearly doubling deliberation (21.9% to 40.1%). This bias persists across model families (e.g., DeepSeek: d=0.41, p=0.008) and is not resolved by simply instructing simulators that they may disengage. The findings matter because AI sales and persuasion agents trained or evaluated against such simulators will systematically overestimate funnel progress precisely among the customers most likely to walk away.
- ResearcharXiv2026-06-16QP
The Consistency Dilemma in LLMs: Generator-Evaluator Agreement and Vulnerability to Mistakes · Marina Mancoridis, Zoë Hitzig
This paper introduces 'generator-evaluator self-consistency,' a measure of whether large language models apply concepts the same way when generating outputs as when evaluating those outputs. Testing 10 frontier models across 491 concepts, the authors find substantial variation in this self-consistency metric. Critically, in a clinical setting using physician-validated mistakes (Proniakin et al., 2025), models with higher self-consistency are paradoxically more vulnerable to mistakes—revealing a 'consistency dilemma' where being operationally consistent does not mean being safe to deploy. This finding has significant implications for agentic AI pipelines that rely on models self-evaluating their outputs without external verification.
- ResearcharXiv2026-06-16QC
AIPatient Arena: EHR-grounded evaluation of large language models in end-to-end clinical consultation workflows · Jiahui Niu, Huizi Yu, Wenkong Wang et al.
AIPatient Arena is an EHR-grounded evaluation framework that assesses large language models (LLMs) across eight dimensions of clinical competence in multi-turn physician-patient consultation workflows. Applied to a primary cohort of 437 patients and two validation cohorts, the framework found that LLMs performed well on interview questioning skills, ethical conduct, and clarity of explanation, but showed persistent weaknesses in handling ambiguous responses, information coverage, and diagnostic accuracy and reasoning. Process-based evaluation revealed recurring failures such as repetitive questioning, omission of past medical history, and inadequate handling of uncertainty. The paper argues that final-answer accuracy alone is insufficient for evaluating clinical readiness, and proposes this framework as a workflow-oriented pre-deployment evaluation tool for medical LLMs.
- ResearcharXiv2026-06-16EQ
PARSE: Provenance-Aware Retrieval Sanitization for Professional Domain LLM Agents · Aaditya Pai
PARSE addresses a critical gap in AI security: existing prompt injection defenses tested on synthetic benchmarks fail to generalize to real enterprise documents such as SEC filings, Federal Register rules, and PubMed abstracts. The authors benchmark 122 tasks across five professional domains and show that paraphrasing—the strongest known synthetic-benchmark defense—produces no statistically significant attack reduction on real documents (p=0.500) while degrading utility from 91.8% to 82.8%. Their proposed system, PARSE, uses provenance-aware, fact-preserving sanitization to classify sentences by injection risk, extract structured facts, and verify preservation, achieving a 38% reduction in attack success rate (from 25.4% to 15.6%) at 86.9% utility with statistical significance (p=0.014). The findings carry a direct practical warning: AI security defenses for enterprise LLM agents must be evaluated on domain-matched real documents rather than synthetic proxies.
- ResearcharXiv2026-06-16QP
AI Transparency: Governance Compliance or Stakeholder Requirements? · Muneera Bano, Didar Zowghi
This paper examines 92 AI transparency statements published by Australian Government agencies under a national AI governance mandate, finding that structural compliance with disclosure requirements does not equate to meaningful transparency for all stakeholders. The authors introduce the Risk–Control–Involvement–Need (RCIN) framework to classify stakeholders by their structural position and transparency needs, revealing that criteria serving high-control stakeholders are consistently met while criteria most critical for high-risk, low-control stakeholders are fewer and less substantively addressed. The authors call this the 'Transparency Illusion'—where compliant artefacts create an appearance of transparency without adequately serving those most exposed to AI-supported decisions. The study reframes transparency as a stakeholder-calibrated validation problem, with direct implications for how AI governance mandates are designed and assessed.
- ResearcharXiv2026-06-16EP
Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems · Xi Chu, Yupeng Hou
This paper investigates how brands compete for recommendations within large language models (LLMs) using skincare products as a test case across GPT-4o-mini, Claude Sonnet, and Gemini 3 Flash. The authors find that well-known brands achieve a 'Conditional Monopoly' — receiving 100% of recommendations when products share identical specifications — but this dominance can be disrupted by a competitor gaining even a marginal rating advantage or by using authority-style marketing language, including fabricated clinical-evidence claims. When multiple brands simultaneously adopt the same generative engine optimization (GEO) strategies, individual payoffs collapse from +0.802 to +0.007 in the authors' payoff proxy, creating a social dilemma. The findings suggest GEO is not merely a security concern but an emerging marketing practice with significant implications for market competition and consumer information integrity.
- ResearcharXiv2026-06-16EQP
Model Validation of Agentic AI Systems: A POMDP-Based Framework for Belief-State, Forecast, and Policy Validation · Matthew Francis Dixon
This paper proposes a structured validation framework for agentic AI systems—autonomous agents that form beliefs, make forecasts, and take actions over time—grounded in Partially Observable Markov Decision Processes (POMDPs). It decomposes autonomous decision-making into distinct components (information, beliefs, forecasts, actions, and utility) so each can be validated independently, and develops a model-risk taxonomy covering state-space, filtering, forecast, policy, utility-specification, and parameter risks. A portfolio-management case study demonstrates the framework, with empirical results indicating that latent-state inference independently contributes to decision quality and that findings are robust across parameter values. The work provides a practical foundation for extending established model risk management concepts to agentic AI, with direct relevance to governance, monitoring, and validation of autonomous systems.
- ResearchOpen Repository and Bibliography (University of Luxembourg)2026-06-16WP
Artificial Intelligence, Skills, and Labor Mobility: Understanding the Transformation of Work · David Marguerit
This PhD dissertation examines how AI reshapes labor markets, skills, and education across the U.S. and Europe. It finds that augmentation AI (which enhances worker output) creates new work and raises wages primarily for high-skilled workers, while automation AI increases employment but depresses wages and harms low-skilled workers. In Europe, AI exposure shifts employer demand toward AI, Data, and Prediction skills while reducing demand for Social skills. AI-driven labor-market signals also propagate upstream to higher education, affecting student enrollment choices and program openings or closures depending on whether AI automates or augments the relevant field.
- ResearcharXiv (Cornell University)2026-06-16QC
Analytics for Quality Assurance for Item Pools (AQuAP): Monitoring and Maintaining Item Bank Health in AI-Driven Assessment Systems · Alina von Davier, Xiaowan Zhang, Yigal Attali et al.
This paper introduces AQuAP (Analytics for Quality Assurance for Item Pools), a dashboard environment designed to monitor item quality and item bank health in AI-driven, high-stakes educational assessments. AQuAP supports the Item Factory framework for automated and human-supported test development, translating psychometric concepts—such as Effective Bank Size, maximum exposure, and rarely-administered fraction—into operational quality assurance tools. The system is illustrated using the Duolingo English Test and aims to ensure item bank security, diversity, and efficiency in high-volume testing programs. This work matters because it shows how operational analytics can sustain the integrity and health of AI-generated item pools at scale.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-06-16QCP
PriyankaPSurve/CEDAR42001: CEDAR-42001: From ISO/IEC 42001 Conformity to Architecture-Aware, Audit-Visible Assurance Posture for AI Cyber-Physical System · PriyankaPSurve
CEDAR-42001 is a two-stage method that converts ISO/IEC 42001 audit evidence for AI-enabled cyber-physical systems into an architecture-aware assurance posture, going beyond simple conformity assessment. The method adds architectural attribution, maturity profiling, risk-proportionate targets, and action recommendations to each audit row, revealing that while 89.9% of audit rows were conforming, only 34.3% reached a baseline High-Assurance category. Applied retrospectively to the 2023 Cruise robotaxi incident, the method mapped documented concerns across governance, perception, decision-making, and human oversight to layer-specific actions. This matters because it helps organizations identify where audit evidence warrants deeper technical assurance, organizational improvement, or remediation beyond what conformity certification alone reveals.
- ResearchBusiness Strategy and the Environment2026-06-16EP
Agentic AI and Circular Procurement Performance: An Empirical Study · Surajit Bag, Susmi Routray, Andrea Chiarini et al.
This empirical study examines how adopting agentic AI in industrial purchasing affects circular procurement performance, drawing on resource-orchestration theory and data from a developing nation analyzed via covariance-based SEM and Process analytical methods. The findings show firms benefit most from agentic AI when three conditions are met: structured resource scanning and evaluation processes, integrated cross-functional procurement systems, and established supplier collaboration routines. The study is the first to theorize the relationship between agentic AI adoption and circular procurement performance, offering practical guidance for purchasing and supply managers on formulating policies and standard operating procedures.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-06-16ECP
Operationalizing NIST AI RMF 1.0 for Federal Training and Academic AI Deployers · Ruchir Bakshi
This paper operationalizes the NIST AI Risk Management Framework (AI RMF 1.0) specifically for federal instructional-design, training, and academic units that deploy AI tools such as adaptive learning systems, AI tutoring, and AI-text detection—organizations not explicitly addressed by the existing framework. The authors develop a use-case AI RMF Profile that applies a deployer lens to all 72 Playbook subcategories, retaining 71 as in-scope and providing applicability analysis alongside blank current-state and target-state fields for adopters to complete. A reproducible build script ensures the Profile stays aligned with the machine-readable NIST Playbook source, eliminating transcription drift. The work matters because it lowers the operationalization burden for a broad class of public-sector AI deployers seeking structured, voluntary risk management guidance.