News & Research
The latest AI research and news with real-world stakes. Each item is sourced, dated, summarized in plain English and tagged by impact area, and checked against its source before it appears.
Kind
Impact area
5672 items
- ResearcharXiv2026-06-19ECP
The AI Evaluability Gap: The Missing Layer for Managing Risk and Sustaining Value · Vishal Srivastava, Tanmay Sah
This paper identifies what it calls the 'AI Evaluability Gap': the condition in which organizations lack sufficient evidence to make high-confidence governance decisions about AI systems, covering both risk management and value creation. The authors argue that current AI governance frameworks focus on system properties—such as safety, fairness, and compliance—while neglecting the evidentiary foundations needed to justify decisions about those properties. To address this, they introduce 'Evaluability' as a formal capability for AI systems to generate and renew evidence over time, characterized by six properties: observability, attributability, intervenability, verifiability, calibration, and temporal validity. The framework also distinguishes Operational Certification (structural evidence for deployment) from Investment Certification (causal evidence for continued resource allocation), positioning evidence sufficiency as a prerequisite for sound AI governance.
- ResearchIntegration of Education2026-06-19WQCP
The Role of Micro-Credentials in the Future Digitalized Artificial Intelligence-Driven Education · Wadim Striełkowski, Akima Orozalieva, Larisa Gorina et al.
This bibliometric study analyzes 664 publications from the Scopus database to assess the role of micro-credentials in AI-driven higher education, finding a dramatic rise in research output from one publication in 1992 to 162 in 2025. Using VOSviewer network analysis, the authors identify key thematic clusters centered on employability, digital transformation, and lifelong learning, demonstrating that micro-credentials can address skills shortages and align higher education with labor market demands. The findings suggest micro-credentials offer flexible, targeted learning pathways that enhance learner employability in AI-enabled education systems, with implications for educators, policymakers, and institutional stakeholders building adaptive digital education frameworks.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-06-19QCP
Human in the Log: Public Evidence Chains for Public-Sector AI Oversight · Anton Sokolov
This article introduces the concept of the 'human in the log'—a record-centered framework for evaluating whether human oversight of public-sector AI systems is genuine and reconstructable. Using Colombia's Constitutional Court decision T-323/24 as an anchor case, the authors argue that assurances of human oversight are unverifiable without structured records capturing the sequence, transparency, human capacity, and repair mechanisms involved in AI-assisted decisions. The paper proposes a seven-part oversight record architecture and synthesizes legal, regulatory, audit, and procurement evidence across 315 documented cases to identify the evidence functions needed to make human control inspectable by the public and policymakers.
- ResearchJournal of Business Economics and Management2026-06-19WEP
A comparative study of the relationships between AI use, employment, economic performance, and sustainability in the EU countries · Anca Antoaneta Vărzaru, Claudiu George Bocean
This study examines how enterprise-level AI adoption relates to employment, economic performance, and sustainability across EU countries using factor analysis, general linear models, and cluster analysis. Results show consistent positive links between AI adoption and higher GDP per capita and a larger share of science and technology professionals, while relationships with overall employment and sustainability indicators are weaker but present. Cluster analysis reveals diverse national profiles shaped by differences in digital readiness, human capital, and institutional factors. The findings offer policymakers an empirical basis for understanding how AI diffusion may support inclusive growth and sustainability goals across the EU.
- ResearcharXiv (Cornell University)2026-06-19QCP
CEDAR-42001: From ISO/IEC 42001 Conformity to Architecture-Aware, Audit-Visible Assurance Posture for AI Cyber-Physical Systems · Priyanka Prakash Surve, Asaf Shabtai, Yuval Elovici
CEDAR-42001 is a two-stage method that transforms ISO/IEC 42001 audit evidence for AI-enabled cyber-physical systems into an architecture-aware assurance posture, going beyond simple conformity determinations. The method adds layer attribution, maturity profiling, risk-proportionate targets, and action recommendations to each audit row, revealing that while 89.9% of audit rows were conforming in a synthetic autonomous-fleet evaluation, only 34.3% of those conforming rows reached the baseline High-assurance category. A retrospective analysis of the 2023 Cruise robotaxi incident demonstrates how the method surfaces governance, perception, decision-making, and human oversight gaps and maps them to specific corrective actions. The work matters because it shows that ISO/IEC 42001 conformity alone is insufficient to characterize actual assurance levels in high-stakes AI systems, and provides a structured path toward deeper technical and organizational improvement.
- ResearcharXiv (Cornell University)2026-06-19EQP
A Multi-Agent Audit Framework for High-Stakes Reasoning: Evaluation and Interpretability in Clinical Mental Health Screening · Jingchen Ye, Yanpei Yu, Luyao Zhang
This paper presents a Multi-Agent Audit Framework for clinical mental health screening that decomposes reasoning into specialized agents—including perception, retrieval-augmented generation, chain-of-thought inference, and an audit verification stage—to improve transparency and accuracy over single-model baselines. Evaluated on the DAIC-WOZ dataset, the multi-agent pipeline reduces Mean Absolute Error for PHQ-8 depression severity prediction from 5.35 to 5.02 compared to single-agent approaches. The framework exposes cross-agent validation traces to mitigate hallucination and reasoning drift, providing interpretable diagnostic rationales suited for high-stakes AI-assisted decision support in clinical settings.
- ResearchScientific Reports2026-06-19QCP
An interpretable attention-based TabTransformer framework with feature fusion for green architecture classification · Yiyuan Zhao, Chang Li, Li Xintong et al.
This paper presents EcoArch-TabFusionNet, a deep learning framework that automatically classifies buildings as green or non-green architecture using a TabTransformer model that fuses architectural, energy, and contextual features. The model achieves 97.5% accuracy on this classification task, outperforming four tabular transformer baselines, while incorporating explainable AI tools (SHAP, LIME, and attention heatmaps) to make its decision logic transparent. By enabling scalable, data-driven assessments of building sustainability, the framework could reduce reliance on costly manual certification processes. This work has direct implications for automating and improving green building certification as well as informing policy around sustainable construction.
- ResearchEngineering and Technology Journal2026-06-19WEQC
Competency Based Education in Information Technology: Designing AI-Driven, Outcome-Focused Learning Pathways · Salmon Oliech Owidi, Kelvin Kabeti Omieno
This systematic review of 124 studies proposes an AI-driven Competency-Based Education (CBE) framework for IT programs that combines explainable AI (Req2XAI) with PEARL design principles to create personalized, transparent learning pathways. The paper finds that adaptive learning improves outcomes by 0.35–0.65 standard deviations, explainable AI yields effect sizes of 0.32–0.58, and automated assessment reduces feedback latency from days to seconds while boosting student satisfaction by 25–30%. AI assessment costs $12–$18 per learner versus $45–$60 for human grading, though upfront development runs $50,000–$200,000, and out-of-the-box AI agreement with human evaluators rises from 67% to 89% after rubric alignment. The framework recommends aligning competency taxonomies with CC2020 and industry certifications, piloting AI in low-stakes assessments first, and investing in faculty AI literacy to support verifiable, job-ready graduate outcomes.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-06-19QCP
Human in the Log: Public Evidence Chains for Public-Sector AI Oversight · Anton Sokolov
This paper introduces the concept of the 'human in the log'—a record-centered framework for evaluating whether human oversight of public-sector AI systems is real and reconstructable rather than merely asserted. Using Colombia's Constitutional Court decision T-323/24 as an anchor case, the authors develop a seven-part oversight record architecture (entry point, input/output, human view, human act, reasons, contest path, and retention/repair) drawn from a synthesis of legal, regulatory, audit, procurement, and standards sources across 315 documented cases. The framework aims to make human control in automated decision systems inspectable and auditable, addressing the gap between official claims of human oversight and the evidentiary record needed to verify them. This matters for public-sector accountability because it offers a practical architecture for governments and auditors to assess whether AI advisory roles are genuinely subject to meaningful human review.
- ResearchSystems and Control Transactions2026-06-19WEQCP
Artificial Intelligence (AI) Usage in an Undergraduate Chemical Engineering Course: Strengths, Pitfalls, and Future Insights · Sourojeet Chakraborty, Stuart R. Gray, Daniela Galatro
This paper examines the integration of generative AI (ChatGPT) into an undergraduate chemical engineering course at Johns Hopkins University, identifying both capabilities and significant limitations. Quantitative benchmarks revealed an 88% AI accuracy rate on standard problems but only 15% on complex mass/energy balance problems, a 73% performance gap, though Effective Prompt Engineering and Chain-of-Thought strategies restored accuracy to 91% in instructor-led trials. A dual-layer quality assurance audit using Turnitin showed a mean AI-probability below 12% in student derivations, suggesting students used AI for brainstorming while retaining technical ownership. The findings offer a scalable framework for higher education institutions to integrate AI as a collaborative tool without compromising engineering judgment or academic integrity.
- ResearcharXiv2026-06-19WEP
Cheap Expertise: Mapping and Challenging Industry Perspectives in the Expert Data Gig Economy · Robert Wolfe, Aayushi Dangol
This paper examines how leading AI data annotation companies and their executives publicly frame the emerging 'expert gig economy,' finding that industry rhetoric positions AI-derived expertise as cheaper and higher-ROI than human expertise, treats human expertise as an extractable resource, and views institutional expertise as needing reform to be absorbed into AI systems. The study analyzes social media feeds and podcast appearances from five major annotation organizations to map this industry vision. The findings raise significant concerns for white-collar professionals whose work and perceived value may be fundamentally restructured, as well as for universities and other institutions that traditionally confer and validate expertise. The authors offer provocations for how society might critically engage with an AI-driven gig economy built around 'cheap expertise.'
- ResearcharXiv2026-06-19WEP
Peer Review Report For: Estimation of Firm Labour Productivity and Sales Growth from Artificial Intelligence in Sub-Saharan African Countries [version 1; peer review: 1 approved with reservations] · Olanrewaju Adewole Adediran
This study examines how AI adoption affects firm-level labour productivity and sales growth across 10 sub-Saharan African countries using World Bank Enterprise data from 2007 to 2024. Applying FGLS, robust OLS, and high-dimensional fixed effects regression, the researchers find that AI adoption has a significant positive relationship with both labour productivity and sales growth, though effects vary by country due to differences in technological readiness and industrial structure. The findings highlight the need for targeted policy interventions—including upskilling initiatives and supportive regulatory frameworks—to maximize AI's economic benefits while protecting workers from adverse impacts.
- Researchnpj Digital Medicine2026-06-19EQP
Young people’s perceptions and recommendations for conversational generative artificial intelligence in youth mental health · Adam Poulsen, Ian B. Hickie, Carla Gorban et al.
This study gathered perspectives from 32 young people in Australia through co-design workshops to understand how conversational generative AI chatbots—specifically a tool called Mia—might be integrated into youth mental health services. Using reflexive thematic analysis, researchers identified four key themes including concerns about humanizing AI without dehumanizing care, transparency about how the system works, context-appropriate use, and personalization within safe boundaries. The findings highlight that young people want meaningful input into the ethics, design, and governance of AI tools in mental health settings. The work has direct implications for how genAI chatbots are developed, implemented, and regulated in youth mental health contexts.
- ResearcharXiv2026-06-18EP
Whose Agent Are You? Multi-Layer Fingerprinting and Attribution of Autonomous Web Agents · Dayeon Kang, Hyejun Jeong, Jade Sheffey et al.
This paper presents a multi-layer fingerprinting framework that identifies and distinguishes AI web agents from humans and traditional crawlers using network-level signals (TLS, HTTP) combined with browser interaction behavior. Testing against six prominent agent frameworks—AutoGen, Browser Use, Claude, Gemini, Operator, and Skyvern—a decision tree classifier achieves 97% attribution accuracy by exploiting structural differences in how agents assemble HTTP requests, establish connections, and execute browser actions. The findings show that existing defenses like robots.txt are insufficient, and that cross-layer fingerprinting offers a robust, evasion-resistant approach to content protection and web security policy enforcement. This work matters for enterprise web security and policy, as it provides a deployable mechanism to enforce access controls against unauthorized AI-driven scraping.
- ResearcharXiv2026-06-18WE
The Token Tax of Epistemic Accuracy: Comparing RAG and Long-Context Architectures for Document-Grounded Generative AI Applications · Austin Hamilton, Ryan Singh, Michael Wise et al.
This paper compares two architectures for document-grounded AI assistants—retrieval-augmented generation (RAG) and long-context prompting—on a benchmark of 972 answers drawn from a manufacturing safety training case study. Long-context prompting, which loads the entire document collection into context, achieved higher correctness (73.1%) than semantic RAG (65.4%), but consumed 26 times more tokens per query. The authors frame this trade-off as a 'token tax of epistemic accuracy,' arguing that broader evidentiary access improves answers but at substantial cost, with direct implications for resource-constrained organizations deciding how to deploy AI in high-stakes, knowledge-intensive workflows.
- ResearcharXiv2026-06-18Q
PeerCheck: Enhancing LLM-Generated Academic Reviews Towards Human-Level Quality · Zeyuan Chen, Ziqing Yang, Yihan Ma et al.
PeerCheck is a framework that investigates differences between human-written and LLM-generated academic peer reviews and explores methods to close the quality gap. The study finds that LLMs and humans focus on different aspects of papers—LLMs prioritize theory while humans emphasize methodology and experiments—highlighting a meaningful misalignment. Using prompt engineering techniques like Chain-of-Thought (CoT) and retrieval-augmented generation (RAG), the authors find that CoT significantly improves LLM review quality, while RAG produces inconsistent results across models and can even reduce quality, a phenomenon they term the 'RAG paradox.' These findings reveal both the promise and limitations of LLM-assisted peer review, with implications for the fairness and effectiveness of academic quality-assurance processes.
- ResearcharXiv2026-06-18EQ
Can LLMs Reason About Brand Ownership? An Empirical Study of Domain Attribution Intelligence · Fathima Mashood, Mohamed Nabeel
This paper presents the first systematic empirical evaluation of large language models (LLMs) for brand intelligence tasks — specifically, determining whether a domain belongs to a known brand or is a squatting/phishing site. Testing four models (Gemini 2.5 Flash, Gemini 3.5 Flash, Claude Sonnet 4.5, and Claude Sonnet 4.6) across four retrieval settings on 36 heavily phished brands, the authors find that LLMs achieve up to 82% precision enumerating brand-owned domains from memory but fail at ownership verification without external tools (macro F1 at most 0.37 in in-context learning mode). Augmenting LLMs with WHOIS lookups raises the macro F1 for binary ownership classification by up to 0.65 points and yields near-perfect precision (≤0.99), sharply reducing false positives that could harm end users and brand reputations. The findings offer concrete guidance for integrating LLMs into brand-protection pipelines with appropriate retrieval augmentation.
- ResearcharXiv2026-06-18QP
StylisticBias: A Few Human Visual Cues Drive Most Social Biases in MLLMs · Shaghayegh Kolli, Timo Cavelius, Nafiseh Nikeghbal et al.
StylisticBias introduces a controlled benchmark to identify which visual cues cause multimodal large language models (MLLMs) to make biased social judgments about people. By generating ~25K photorealistic face images that each vary only one attribute at a time across 500 base identities, the study isolates how specific appearance cues—rather than identity differences—shift model outputs across 25 binary social judgment scenarios. The findings show that age and body type dominate identity-level bias, while fashion style and a small set of roughly 15 attributes account for nearly 80% of total variation, with bias concentrated especially in socioeconomic and style-related judgments. This matters because it pinpoints a tractable set of visual cues that developers and auditors can target to reduce discriminatory behavior in AI systems deployed in high-stakes settings.
- ResearcharXiv2026-06-18QC
Calibration Without Comprehension: Diagnosing the Limits of Fine-Tuning LLMs for Vulnerability Detection in Systems Software · Arastoo Zibaeirad, Marco Vieira
This paper introduces CWE-Trace, a diagnostic framework that tests whether large language models (LLMs) genuinely reason about software vulnerabilities or simply pattern-match on familiar data. Using 834 manually curated Linux kernel samples across 74 vulnerability categories and a strict temporal train/test split, the authors evaluate eight base LLMs and 15 fine-tuned variants and find that fine-tuning adjusts output distributions without improving underlying security reasoning—a phenomenon they call 'calibration without comprehension.' The best model reaches only 52.1% accuracy on binary vulnerability detection (just above chance), and exact vulnerability classification tops out below 1.3% Top-1 accuracy, demonstrating that current LLMs cannot reliably detect or classify vulnerabilities in systems software regardless of fine-tuning strategy. These findings matter for quality assurance and certification contexts where LLM-based vulnerability scanning tools might otherwise be trusted without rigorous empirical validation.
- ResearcharXiv2026-06-18QC
NRT-Bench: Benchmarking Multi-Turn Red-Teaming of LLM Operator Agents in Safety-Critical Control Rooms · Hanwool Lee, Dasol Choi, Bokyeong Kim et al.
NRT-Bench is a benchmark for stress-testing large language model (LLM) agents that act as operators in safety-critical environments, instantiated as a simulated nuclear power plant control room. The benchmark uses multi-turn adversarial attacks across four communication channels, with harm measured objectively by whether any of six critical safety functions (CSFs) is lost—not by LLM-judged text. Testing four frontier models, the authors find that 8.7%–12.1% of adversarial sessions cause a CSF failure, but crucially, vulnerabilities are nearly disjoint across models: roughly one-third of sessions defeat at least one model while none defeat all four, meaning models fail in different ways rather than sharing common weaknesses. Additionally, defensive measures such as guardrail stacks or safety-advisor agents have strongly model-dependent effects, sometimes increasing attack success for models they were intended to protect, highlighting the difficulty of generalizable safety solutions for LLM-based control systems.
- ResearcharXiv2026-06-18WP
From Novelty to Normalisation: Tracking Changing Perceptions of AI in Higher Education, 2024-2026 · Juliana Gerard, Morgan Macleod, Kelly Norwood et al.
This longitudinal study tracked AI perceptions among 1,665 undergraduates, doctoral researchers, teaching staff, and non-teaching staff at Ulster University across three survey waves from 2024 to 2026. Results show students rapidly normalized generative AI use—moving from tentative experimentation to routine engagement—while teaching staff maintained persistent concerns about academic integrity, assessment design, and critical thinking. A widening gap emerged between student practice and institutional policy, with doctoral and non-teaching staff occupying intermediate positions. The findings highlight the need for adaptive institutional policy, AI literacy initiatives, and targeted staff training in higher education.
- ResearcharXiv2026-06-18QP
NAMESAKES: Probing Identity Memorization in Text-to-Image Models · Morris Alper, Vasudha Varadarajan, Moran Yanuka et al.
This paper addresses privacy risks in text-to-image (T2I) models that can generate realistic likenesses of real individuals when prompted with their names. The authors introduce a fully black-box behavioral probe—requiring no reference photos, training data access, or model internals—that distinguishes whether a generated face reflects memorized identity or is fabricated. To support benchmarking, they release the NAMESAKES dataset of over one thousand public figures spanning varying fame levels, along with perturbed name variants. Experiments on state-of-the-art T2I models show the probe substantially predicts identity memorization, with additional insights into differences across model families.
- ResearcharXiv2026-06-18E
Autonomous Event-Driven Multi-Agent Orchestration for Enterprise AI at Scale · Harsh Rao Dhanyamraju, Leonidas Raghav, Aaron Lee
This paper evaluates two multi-agent AI orchestration architectures—DAG Plan and Execute and ReAct—across 208 production-derived enterprise scenarios at three scales (Persona, Department, and Enterprise), finding that scale rather than task complexity is the dominant factor in orchestration performance. At enterprise scale (up to 200 agents), agent discovery noise becomes the primary bottleneck, causing simple tasks to degrade more sharply than complex ones, while ReAct proves more robust due to its incremental failure handling compared to DAG's higher overhead. The authors also introduce a Task Manager component for continuous event-driven operation that reduces high-priority queue latency by 14–75% and improves related-event correctness by over 20 percentage points at enterprise scale. These findings provide actionable architectural guidance for deploying AI agent systems in large-scale enterprise environments.
- ResearcharXiv2026-06-18Q
Tri-Info: Generalizable, Interpretable Failure Prediction for VLA Models via Information Theory · Jinghan Yang, Yunchao Zhang, Wang Yuan et al.
This paper introduces Tri-Info, an information-theoretic framework for detecting when Vision-Language-Action (VLA) robotic models are about to fail. The authors formalize VLA control as a closed-loop information pipeline and derive three signals—action diversity, temporal consistency, and state-action coupling—that distinguish successful from failed robot rollouts. Tested across six VLA models and three benchmark environments, Tri-Info matches the best existing detectors in-domain and generalizes across architectures, environments, and the sim-to-real gap without retraining, achieving 83% accuracy on real-world tasks where prior methods perform at chance. This matters for quality assurance in AI-driven robotics, as it provides both generalizable failure detection and interpretable diagnostics of underlying failure modes without requiring access to model internals.
- ResearcharXiv2026-06-18WP
The Algorithmic-Human Manager: AI, Apps, and Workers in the Indian Gig Economy · Omir Kumar, Krishnan Narayanan
This paper investigates how algorithmic management systems — AI-driven tools that allocate, monitor, and evaluate work — affect blue-collar gig workers in India's ride-sharing and delivery sectors. Drawing on interviews with 16 gig workers and 21 key stakeholders, the study finds that while these AI systems expand job access and improve operational efficiency, they are opaque by design, produce inequitable outcomes, and fail to proportionately reward additional labor. The authors propose a hybrid 'Algorithmic Human Manager' governance framework that combines technological efficiency with human accountability. The findings have implications for policymakers, platform companies, and civil society organizations designing equitable AI governance in India and the broader Global South.