News & Research
The latest AI research and news with real-world stakes. Each item is sourced, dated, summarized in plain English and tagged by impact area, and checked against its source before it appears.
Kind
Impact area
5672 items
- ResearcharXiv2026-06-17E
Data Intelligence Agents: Interpreting, Modeling, and Querying Enterprise Data via Autonomous Coding Agents · Anoushka Vyas, Aarushi Dhanuka, Sina Khoshfetrat Pakazad et al.
Data Intelligence Agents (DIA) is a production-deployed system of three autonomous coding agents—Data Interpreter, Schema Creator, and Query Generator—designed to reduce the repeated, lossy handoffs between data owners, engineers, and analysts in enterprise data workflows. Rather than generating plain text, the agents produce, execute, validate, and repair concrete artifacts (e.g., SQL queries and schemas), drawing on a shared memory for experience reuse while surfacing outputs for expert review. Evaluated across seven SQL benchmarks spanning four task categories and four dialects, the Query Generator matches or surpasses the best published results on all seven, demonstrating strong generalization across enterprise data intelligence tasks.
- ResearcharXiv2026-06-17EQ
Correct Yourself, Keep My Trust: How Self-Correction and Social Connection Shape Credibility in Social Chatbots · Biswadeep Sen, Yi-Chieh Lee
This study (N=120 between-subjects experiment) tests three error-correction strategies for social chatbots—webpage retraction, self-correction by the same chatbot, and correction by an expert chatbot—finding that all three corrected misinformation equally well, but only self-correction preserved the chatbot's trustworthiness and perceived expertise. Additionally, users' social connection with the chatbot (measured via social attraction and self-disclosure) predicted greater belief change, but only when the chatbot corrected itself; outsourcing corrections to external sources eliminated this effect entirely. The findings suggest that social chatbots should self-correct errors rather than defer to external sources, and that building social connection with users is a functional mechanism that amplifies correction effectiveness, not merely a design feature. This has direct implications for designing AI systems that maintain long-term credibility while managing the risks of inaccurate outputs.
- ResearcharXiv2026-06-17EQ
Language Models as Interfaces, Not Oracles: A Hybrid LLM-ML System for Pediatric Appendicitis · Soheyl Bateni, Maryam Abdolali
This paper introduces ClaMPAPP, a hybrid AI system for diagnosing pediatric appendicitis that uses a large language model (LLM) solely to extract structured clinical features from free-text notes, then feeds those features into an XGBoost classifier for final risk prediction. Evaluated on two independent German pediatric hospital cohorts, ClaMPAPP outperformed end-to-end LLM diagnostic approaches in overall diagnostic performance and minimized missed appendicitis cases — the critical safety concern in acute triage. The study also shows that end-to-end LLMs suffer unstable sensitivity-specificity trade-offs and degrade when narrative text is reordered, weaknesses avoided by the hybrid design. The findings support separating natural-language usability from predictive inference to create more auditable, robust clinical decision support tools.
- ResearcharXiv2026-06-17EQ
JustDiag!: A Diagnostic Justification Engine for Accountable Root Cause Analysis · Tingzhu Bi, Xinrui Jiang, Xun Zhang et al.
JustDiag is a diagnostic justification engine that augments large language model-based root cause analysis (RCA) by maintaining an explicit process state tracking evidence, competing hypotheses, conflicts, and unresolved uncertainties. Evaluated on 66 real-world incidents using a two-layer protocol that separately scores final-answer quality and process quality, JustDiag outperformed a matched control on both outcome and process scores, while accepting slightly lower terminal completion due to more calibrated non-closure. The findings suggest that accountable RCA in high-stakes operations requires explicit diagnostic justification artifacts and process-aware evaluation, not merely fluent final answers.
- ResearcharXiv2026-06-17EQ
TRAP: Benchmark for Task-completion and Resistance to Active Privacy-extraction · Moon Ye-Bin, Nam Hyeon-Woo, Baek Seong-Eun et al.
TRAP introduces a benchmark that tests AI agents on two competing obligations: using private information (like passport numbers) to complete tasks accurately, while never revealing that information in natural-language responses. Evaluating 22 frontier and open-source models, the authors find that all model families exhibit non-trivial privacy leakage and that stronger instruction-following ability correlates with higher leakage rates. The paper proves mathematically that no soft-constraint (prompt-based) defense can simultaneously achieve high task accuracy and zero leakage for softmax-based models. As a remedy, the authors propose structural private field isolation—replacing sensitive fields with hash keys before they reach the model—which largely eliminates leakage without sacrificing task accuracy.
- ResearcharXiv2026-06-17EQ
Decoupling Search from Reasoning: A Vendor-Agnostic Grounding Architecture for LLM Agents · Emmanuel Aboah Boateng, Kyle MacDonald, Amardeep Kumar et al.
This paper introduces Decoupled Search Grounding (DSG), an architecture that separates web search and retrieval from the reasoning model in LLM-based agents, making grounding an explicit, controllable interface rather than a bundled model feature. Tested across five frontier models on SimpleQA, FreshQA, and HotpotQA benchmarks, DSG nearly matches native search accuracy on factual queries (86.1% vs. 87.7% on SimpleQA) while reducing search costs by 91% and latency by 68% through semantic caching. In a production e-commerce query-understanding workload, DSG matches or slightly exceeds native-search accuracy while cutting search cost by over 98%. The work is relevant to enterprises operating large-scale agentic AI systems, showing that grounding can be made vendor-agnostic, auditable, and significantly cheaper without sacrificing answer quality.
- ResearcharXiv2026-06-17WQ
Improving Medical Communication using Rubric-Guided Counterfactual Recommendations · Adrian Cosma, Nicoleta-Nina Basoc, Andrei Niculae et al.
This paper presents a language-model-guided pipeline that generates counterfactual recommendations for improving doctor-patient text communication in telemedicine settings. The system identifies interpretable communication features—such as tone, personalization, actionability, and completeness—and suggests minimal, ordinal changes predicted to increase positive patient feedback, without altering medical content. Evaluated across real interactions, the recommendations yield a mean +6.41% gain in predicted positive feedback probability under independent auditor models, with non-negative gains for 93.31% of recommendations. The work suggests that small, interpretable communication adjustments can meaningfully improve perceived communication quality while preserving physician control over medical reasoning.
- ResearcharXiv2026-06-17EQ
Scaling Learning-based AEB with Massive Unlabeled Data · Xiangyu Wang, Yang Zhan, Mengxiang Hao et al.
This paper presents a semi-supervised learning framework called MF-SSL (meta-feedback semi-supervised learning) for scaling automatic emergency braking (AEB) systems using massive amounts of unlabeled fleet driving data. The method uses a teacher model to generate pseudo-labels for unlabeled data, updated via a small labeled anchor set, and includes two stabilization mechanisms—Noise-Aware Decoupling and kinematics-gated pseudo-labeling—to suppress errors that arise from anchor ambiguity and labeled-unlabeled data mismatch. Trained on up to 1 billion data windows, the resulting student model was deployed to hundreds of thousands of vehicles and validated over 10^9 km of driving, achieving a positive-to-false activation ratio exceeding 100:1 and a 35% improvement in accident-free driving mileage compared to a production rule-only baseline. The work demonstrates that semi-supervised learning can substantially improve real-world automotive safety systems at scale while maintaining driver comfort.
- ResearcharXiv2026-06-17WE
Improving Human-Robot Teamwork in Urban Search and Rescue Through Episodic Memory of Prior Collaboration · Taewoon Kim, Emma van Zoelen, Mark Neerincx
This paper investigates whether robots in Urban Search and Rescue (USAR) scenarios can leverage recorded past collaboration patterns—stored as knowledge-graph episodic memories—to become better teammates from the very start of a new interaction. Using the MATRX simulation environment, the researchers had human participants externalize their teamwork strategies via a chat and reflection interface, then applied graph representation learning to automatically select the most effective prior collaboration pattern to initialize the robot before a new episode. Across 20 participants and 160 round-level observations, initializing the robot with a single automatically selected prior collaboration pattern increased rescue success from 25.7% to 41.3% and reduced average task time by 283 seconds, with the strongest gains appearing early in the interaction. These findings suggest that episodic memory of prior collaboration can meaningfully accelerate human-robot team performance, offering a practical path toward robots that adapt to human partners more quickly and effectively.
- ResearcharXiv2026-06-17EQ
GateMem: Benchmarking Memory Governance in Multi-Principal Shared-Memory Agents · Zhe Ren, Yibo Yang, Yimeng Chen et al.
GateMem is a new benchmark designed to evaluate how well LLM-based memory agents handle shared, multi-user environments such as hospitals, workplaces, campuses, and households. Unlike prior benchmarks that assume a single user, GateMem tests three dimensions simultaneously: utility for legitimate long-horizon requests, access control across different roles and authorization boundaries, and reliable deletion ('active forgetting') of information after explicit requests. Tested across diverse baselines and backbone models, no existing method achieves strong performance on all three dimensions—long-context prompting offers the best governance scores but at high token cost, while retrieval-based methods reduce cost yet still leak unauthorized or deleted information. The findings indicate that current memory agents are not yet reliable enough for shared institutional deployment.
- ResearcharXiv2026-06-17E
ProfiLLM: Utility-Aligned Agentic User Profiling for Industrial Ride-Hailing Dispatch · Tengfei Lyu, Zirui Yuan, Xu Liu et al.
ProfiLLM is an agentic LLM pipeline that generates behavioral user profiles for ride-hailing dispatch at industrial scale, specifically targeting DiDi's production platform. It addresses three core challenges: log data exceeding LLM context windows, sparse data for long-tail users, and profiles that are fluent but not predictively useful. The system uses tool-augmented knowledge mining and utility-aligned profile exploration with DPO fine-tuning to produce profiles that measurably improve downstream prediction and business outcomes. Deployed in a 14-day online A/B test, ProfiLLM achieved +0.47% GMV, +0.33% Completion Rate, and -0.82% Cancel-Before-Accept rate, demonstrating that LLMs can serve as practical semantic feature extractors in latency-sensitive production matching systems.
- ResearcharXiv2026-06-17EQ
EARS: Explanatory Abstention for Reliable Sub-Agent Modeling in Large-scale Multi-Agent Systems · Shuang Xie, Yunan Lu, Han Li et al.
EARS (Explanatory Abstention for Reliable Sub-Agent Modeling) is a framework designed to improve the reliability of large-scale multi-agent systems (MAS) used in enterprise settings, where a coordinator delegates tasks to specialized sub-agents. The core problem is that sub-agents built on smaller fine-tuned models tend to over-answer ambiguous or misrouted requests, producing hallucinated outputs rather than useful feedback. EARS reframes abstention as a structured inter-agent communication protocol: sub-agents are fine-tuned to detect failure conditions and return actionable rationales to the coordinator for clarification, rerouting, or fallback. Evaluated in a production e-commerce business intelligence assistant, EARS improved the overall response pass rate from 68.5% to 78.9%, demonstrating measurable reliability gains in enterprise AI deployments.
- ResearcharXiv2026-06-17WP
Gender Bias in LLM Hiring Decisions: Evidence from a Japanese Context and Evaluation of Mitigation Strategies · Serena A. Hoffstedde, Machiko Hirota, Akshara Nadayanur Sathis Kanna et al.
This study investigates gender bias in LLM-assisted hiring decisions within a Japanese corporate context, using 60 rirekisho-format resumes, 12 name pairs, and five leading LLMs across 43,200 API calls. The results confirm a significant pro-female bias across all five models, replicating patterns previously found in Western research and extending them to a non-Western setting. A prompt-level gender-neutrality instruction failed to meaningfully reduce the bias, but removing candidate names from prompts nearly eliminated the female effect, identifying the name as the primary channel of gender information. The study also surfaces a practical deployment challenge: a 42% refusal rate when combining name anonymization with GPT-4o's content safety filter, complicating the use of privacy filters in real hiring pipelines.
- ResearcharXiv2026-06-17WQ
AI-Driven Assessment of Human Tutors: Linking Training Performance to Real-Life Practice · Danielle R. Thomas, Marie Cynthia Abijuru Kamikazi, Clara Brandt et al.
This paper presents an AI-driven system using Gemini-2.5-pro to assess both training performance and real-life tutoring sessions for human math tutors, bridging the gap between simulation-based training and authentic classroom practice. Across 86 tutors and 405 session-to-lesson pairs, training performance significantly predicted real-life tutoring quality with an effect size of 0.25 SD, with open-response scores being more predictive than multiple-choice. Tutors showed a 7.4% average learning gain during training and demonstrated measurable improvements in both encountering and executing pedagogical opportunities in real sessions over time. The system contributes open datasets, AI prompts, and scoring rubrics to support reproducibility in AI-driven tutor evaluation.
- ResearchInternational Journal of Productivity and Performance Management2026-06-17EQP
Green AI adoption for circular and agile supply chains: ESG-driven pathways in Vietnam's emerging technological economy · Bang Nguyen‐Viet, Luu Chi-Luong, Ngan Nguyen-Khanh
This study surveys 780 Vietnamese medium- and large-sized enterprises to examine how green AI adoption drives sustainable business performance. Using structural equation modeling, the findings show that green AI positively enhances green circular capacity, supply chain agility, and green innovation, all of which improve business performance. Green digital orientation amplifies AI's effect on supply chain agility, and ESG compliance strengthens the link between green innovation and performance. The results demonstrate measurable pathways through which AI and sustainability capabilities translate into accountable outcomes in emerging markets.
- ResearchSustainability2026-06-17EP
The Impact of Government Green Procurement on Corporate Carbon Emission Reduction: A Dual Mediation Perspective of Artificial Intelligence and Green Finance · Z K Zhang, Jianmin Wu
This study examines how government green procurement policies affect corporate carbon emission reduction among Chinese A-share listed companies from 2020 to 2024. Using two-way fixed effects models and Bootstrap methods, the authors find that green public procurement significantly improves firms' carbon reduction performance, with AI adoption and government green subsidies acting as mediating mechanisms that amplify this effect. The impact is strongest for state-owned enterprises, high-tech firms, and companies in regions with more advanced digital economies. The findings offer actionable guidance for policymakers seeking to align procurement policy with digital technology and green finance to accelerate decarbonization.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-06-17WEQCP
Competency Based Education in Information Technology: Designing AI-Driven, Outcome-Focused Learning Pathways · Salmon Oliech Owidi, Kelvin Kabeti Omieno
This systematic review of 124 studies explores how AI-driven Competency-Based Education (CBE) frameworks can replace time-based IT curricula with skill-mastery pathways. The paper finds that adaptive learning improves outcomes by 0.35–0.65 standard deviations, explainable AI boosts complex problem-solving with effect sizes up to 0.58, and automated assessment cuts feedback latency from days to seconds while raising student satisfaction 25–30%. The authors propose a framework integrating explainable AI and PEARL principles, and recommend aligning competency taxonomies with CC2020 and industry certifications, piloting AI assessments in low-stakes settings, and funding longitudinal research on graduate employment outcomes.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-06-17WEP
Human Capital Ratio (HCR): A Fiscal Framework for the Age of Artificial Intelligence · Nandeep Nagarkar
This paper proposes the Human Capital Ratio (HCR), a fiscal policy framework designed to counteract AI-driven labor displacement in the enterprise economy. Modeled on healthcare's Medical Loss Ratio and 401(k) non-discrimination testing, HCR creates a dual-trigger levy system that activates when firms' AI spending outpaces human labor spending or when compensation becomes skewed toward highly paid employees. Levy proceeds fund a cost-of-living-indexed social minimum income for displaced workers, with eligibility tied to trigger events rather than documented terminations, and surplus funds directed to workforce retraining. The framework also mandates forward-looking efficiency disclosures at the point of AI capital commitment to close avoidance vectors that the spending ratio alone cannot address.
- ResearchInclusive Society and Sustainability Studies2026-06-17WEP
Environmental Competitiveness and Green Entrepreneurship: The Mediating Role of Generative Artificial Intelligence in Pakistan’s Manufacturing Sector · Muhammad Ahsan Iqbal, Asifa Muhammad Sabir, Hufza Iqbal
This study examines how environmental competitiveness drives green entrepreneurship in Pakistan's manufacturing sector, with generative AI (Gen-AI) adoption acting as a partial mediator. Using PLS-SEM with survey data from 250 employees, the authors find that environmental competitiveness significantly predicts both Gen-AI adoption (β=0.257) and green entrepreneurship (β=0.178), while Gen-AI adoption in turn predicts green entrepreneurship (β=0.216) with a confirmed indirect mediation effect (β=0.055). The findings suggest that AI adoption can serve as a strategic mechanism linking competitive environmental pressures to sustainable business outcomes, with practical implications for managers and policymakers in emerging economies.
- ResearchAdvances in computational intelligence and robotics book series2026-06-17P
Policy, Regulation, and the Public Good of Artificial Intelligence · Robert Ipiin Gnankob, Jayanta Kumar Mohapatra, Samuel Etse Dzakpasu
This chapter develops a 'Policy–Regulation–Governance Compass' framework grounded in public value theory, responsible innovation, and stakeholder governance to guide ethical AI development. Through comparative analysis of global regulatory approaches—including the EU AI Act, U.S. AI Bill of Rights, and emerging African and Asian strategies—it identifies asymmetries in regulatory capacity and persistent challenges such as regulatory lag, power concentration, and global inequality. The authors argue that governance functions as an ethical enabler rather than a constraint on innovation, and propose actionable pathways emphasizing adaptive frameworks, institutional capacity, inclusivity, and international cooperation. The work matters because it offers policymakers a structured lens for aligning AI development with equitable and sustainable public good outcomes.
- ResearchSustainability2026-06-17EQP
Artificial Intelligence as a Strategic Driver of Environmental Sustainability: Unpacking the Mediating Role of Green Governance in GCC Industrial Firms · Ruaa Binsaddig, Amina Toumi, Reem Khamis et al.
This study examines how AI adoption drives environmental sustainability among 75 publicly listed industrial firms across six GCC countries from 2018 to 2025, using fixed-effects and bootstrapped mediation analyses. Results show AI adoption is positively and significantly associated with environmental sustainability, with green governance (board-level ESG structures) partially mediating this relationship. The findings suggest AI's environmental benefits are most fully realized when embedded within sound corporate governance frameworks, offering actionable guidance for policymakers and managers in resource-intensive industries pursuing digital transformation.
- ResearchFuture Internet2026-06-17EQCP
AI-Augmented Compliance Auditing for Cloud Systems: A Hybrid ML–LLM Approach · Moise Iradukunda Ingabire, Jema David Ndibwile
This paper presents a hybrid AI compliance auditing system combining XGBoost multi-label classification and GPT-4o-mini large language model analysis, evaluated against Rwanda's National Cyber Security Authority standards (169 controls across 14 families). The system achieves 85.1% F1 on real-world logs with a 6.4% false-positive rate, and 92.8% macro detection across adversarial MITRE ATT\_CK scenarios with 0.0% false-positive rate on compliant logs. A key contribution is identifying and correcting an 86.3% data-leakage flaw that had artificially inflated prior results to 99.99%, improving result credibility. The system runs on $50/month cloud infrastructure and generates audit reports in 2–5 seconds, demonstrating that effectiveness-based compliance auditing is viable without enterprise-grade resources.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-06-17WQCP
Competency Based Education in Information Technology: Designing AI-Driven, Outcome-Focused Learning Pathways · Salmon Oliech Owidi, Kelvin Kabeti Omieno
This systematic review of 124 studies proposes a framework for integrating AI-driven Competency-Based Education (CBE) into IT programs, combining explainable AI tools with adaptive learning to create transparent, personalized learning pathways. The paper reports that adaptive learning improves outcomes by 0.35 to 0.65 standard deviations, automated assessment reduces feedback latency from days to seconds, and increases student satisfaction by 25 to 30 percent, while cutting per-learner assessment costs from $45–$60 (human) to $12–$18 (AI). Key recommendations include aligning competency taxonomies with CC2020 and industry certifications, piloting AI assessment in low-stakes contexts, and investing in faculty AI literacy to build workforce-ready graduates. The findings are directly relevant to educators, certifying bodies, and policymakers seeking to modernize IT credentialing and improve graduate employment outcomes.
- ResearchInternational Journal of Computational and Experimental Science and Engineering2026-06-17EQCP
An Adaptive Governance-Centric MLOps Framework for Risk-Tiered Control and Continuous Assurance of Responsible AI in High-Stakes Domains · Sunilkumar Reddy Eraganeni
This paper proposes an Adaptive Governance-Centric MLOps Framework that embeds responsible AI principles—including explainability, fairness checking, and audit logging—directly into the machine learning lifecycle for high-stakes domains like finance and healthcare. Validated on the German Credit Dataset using six ML and deep learning models, the framework shows that governance enforcement (regulatory compliance, bias detection, drift monitoring) can be integrated without sacrificing predictive performance, with XGBoost reaching 94.57% accuracy and the MLP model meeting strict fairness thresholds. The study demonstrates that continuous compliance monitoring, including automated retraining triggers upon distributional shift, is achievable within a unified MLOps architecture. These findings are directly relevant to organizations and regulators seeking to operationalize requirements under the EU AI Act and GDPR in production AI systems.
- ResearcharXiv2026-06-17EQC
Pedagogical Audit Framework: The Multimodal Video Evaluator (MVE) Knowledgebase Version 1.1 · Koichi SATO
This paper presents the Multimodal Video Evaluator (MVE) Version 1.1, an AI-assisted system designed to automate pedagogical auditing of educational video content in higher education. The system uses a knowledgebase grounded in the Cognitive Theory of Multimedia Learning, video engagement research, and visual scaffolding frameworks to evaluate multimodal attributes like visual transitions, text density, audio-visual synchronization, and instructional pacing. Version 1.1 refines the prior architecture by eliminating semantic inconsistencies in its rule set to improve deterministic reasoning and reduce LLM processing conflicts. The system addresses the scaling challenge of auditing large institutional media libraries, moving beyond manual low-percentage sampling to programmatic identification of cognitive bottlenecks and prediction of student engagement patterns.