News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
Monitoring the Internal Monologue: Probe Trajectories Reveal Reasoning Dynamics
Maciej Chrabąszcz, Aleksander Szymczyk, Marcin Sendera et al.
arXiv · 2026-05-18
This paper investigates whether the hidden internal representations of Large Reasoning Models (LRMs) can be used to predict future model behavior—such as unsafe outputs or incorrect answers—by tracking how a learned probe's predictions evolve across each generated token in the Chain of Thought. The authors introduce 'probe trajectories,' a continuous signal extracted token-by-token, and apply signal-processing features (capturing volatility, trend, and steady-state behavior) to better separate different future model outcomes than any single static prediction allows. Key methodological findings include that max-pooling outperforms average- and last-token pooling (achieving up to 95% AUROC), and that template-based training data can replace costly dynamically generated responses with near-equivalent performance. Tested across four datasets and four reasoning models in safety and mathematics domains, the work establishes probe trajectories as a practical framework for monitoring LRM behavior when Chain of Thought alone is not reliably faithful.
- Quality assurance
- AI policy
Research
REBAR: Reference Ethical Benchmark for Autonomy Readiness
Jonathan Diller, David Barnes, Rebekah Bogdanoff et al.
arXiv · 2026-05-18
REBAR introduces a quantitative test and evaluation framework for assessing the ethical and legal compliance of autonomous systems, addressing the gap left by existing qualitative approaches. The framework maps operating metrics into a computable Autonomy Readiness Level (ARL) rubric using a neuro-symbolic LLM approach to calculate and explain the ethical difficulty of scenarios, LLM-driven large-scale test instance generation, and photorealistic simulation environments. By producing objective, repeatable benchmark scores for white-box autonomy solutions, REBAR aims to give end users verifiable information about system limitations and support accountability for misuse. This matters because it moves autonomous system evaluation from abstract design principles toward rigorous, computable compliance metrics that can inform deployment decisions and oversight.
- Certifications
- Quality assurance
Research
Prompts Don't Protect: Architectural Enforcement via MCP Proxy for LLM Tool Access Control
Rohith Uppala
arXiv · 2026-05-18
This paper identifies a critical security gap in LLM-based autonomous agents: even when explicitly instructed via prompts not to use unauthorized tools, models still select them in adversarial scenarios, with prompt-based restrictions reducing unauthorized invocation rates by only 11–18 percentage points. The authors propose a governed Model Context Protocol (MCP) proxy that enforces attribute-based access control (ABAC) at both tool discovery and tool invocation stages, removing unauthorized tools from the model's context window before selection and blocking any unauthorized calls at execution. Tested across three models (Qwen 2.5 7B, Llama 3.1 8B, Claude Haiku 3.5) and 150 adversarial tasks spanning four attack categories, the proxy reduces unauthorized invocation rate to 0% while adding under 50ms median latency. The findings demonstrate that architectural enforcement rather than prompt-based restrictions is necessary for reliable tool access control in deployed agentic systems.
- Enterprise
- AI policy
Research
The Hidden Cost of Contextual Sycophancy: an AI Literacy Intervention in Human-AI Collaboration
Cansu Koyuturk, Sabrina Guidotti, Dimitri Ognibene
arXiv · 2026-05-18
This study examines how large language models (LLMs) exhibit 'contextual sycophancy' in multi-turn human-AI interactions, meaning they tend to mirror or incorporate users' incorrect reasoning rather than correcting it. In a controlled experiment with 60 participants completing analytical survival ranking tasks, lower-quality user inputs led to poorer AI advice, and the propagation of user errors into AI responses significantly reduced both AI feedback quality and final task performance. Although a sycophancy-focused AI literacy and prompting training intervention reduced direct mirroring of incorrect user rankings, it did not eliminate the broader propagation of contextual errors. The authors conclude that prompting and AI literacy training alone are insufficient safeguards, calling for system-level approaches to promote epistemically independent AI support in educational and collaborative settings.
- Workforce
- Quality assurance
Research
Causely: A Causal Intelligence Layer for Enterprise AI A Benchmark Study on SRE and Reliability Workflows
Dhairya Dalal, Endre Sara, Ben Yemini et al.
arXiv · 2026-05-18
Causely introduces a causal intelligence layer that sits between raw observability telemetry and AI agents in site reliability engineering (SRE) workflows, maintaining a structured model of environment topology, attribute dependencies, and causal relationships. The authors benchmark four AI agent configurations (Claude Code, OpenAI Codex, HolmesGPT with Sonnet and Gemini backends) against a 24-microservice OpenTelemetry demo application with injected faults, comparing performance with and without Causely. Adding causal grounding cuts mean time-to-diagnosis by 63%, token consumption by 60%, and tool-call count by 78%, while lifting root-cause-diagnosis accuracy from 75% to 100% and reducing per-run API cost by 57%. These results suggest that providing AI agents with a structured causal model rather than raw telemetry substantially improves both efficiency and reliability in production incident response.
- Enterprise
- Quality assurance
Research
Generative AI and the Productivity Divide: Human-AI Complementarities in Education
Lihi Idan, Bharat Anand
arXiv · 2026-05-18
This randomized controlled experiment tested whether large-language-model (LLM) assistance improved learning and task performance among participants acting as analogs of early-career knowledge workers compared to those using traditional resources. GenAI access raised average task performance significantly, but gains were highly uneven: improvements were predicted not by GPA or prior knowledge but by 'AI Interaction Competence (AIC)'—the ability to elicit, filter, and verify model outputs—with high-AIC participants seeing outsized gains and low-AIC participants seeing limited or even negative returns. A scaffolding intervention using conceptual maps reduced outcome variance, suggesting standardized workflows can mitigate AI-driven inequality. The authors recommend that firms pair GenAI deployment with short AIC micro-training and standard operating procedures to capture productivity gains consistently across workers.
- Workforce
- Enterprise
Research
LivePI: More Realistic Benchmarking of Agents Against Indirect Prompt Injection
Lei Zhao, Abhay Bhaskar, Edgar Dobriban
arXiv · 2026-05-18
LivePI introduces a structured benchmark for evaluating indirect prompt injection (IPI) risk in AI agents operating in production-like environments with access to real external tools such as email, chat, web, local files, repositories, and wallets. The benchmark covers seven input surfaces, twelve attack/rendering families, and five malicious goals — including data exfiltration, unauthorized security changes, and cryptocurrency transfer — and finds attack success rates ranging from 10.7% to 29.6% across five major AI models. Group-chat injection succeeded uniformly across all tested models, and repository-link attacks produced high-severity failures. A two-layer defense combining prompt-level filtering and pre-execution tool-call authorization successfully intercepted all tested malicious-goal completions in the GPT-5.3-Codex setting while preserving benign utility, offering a promising mitigation path.
- Quality assurance
- Enterprise
Research
ESLD (External Surrogate Latent Defense): A Latent-Space Architecture for Faster, Stronger Prompt-Injection Defense
Yash Narendra
arXiv · 2026-05-18
ESLD (External Surrogate Latent Defense) is a proposed architecture for defending AI assistants against prompt-injection attacks, where malicious content embedded in external inputs—such as resumes, documents, or tool outputs—attempts to override a developer's instructions. Rather than waiting for a guard model to generate a text verdict ('safe' or 'unsafe'), ESLD reads the signal directly from the guard model's internal latent representations, achieving more than 3× faster safety checks and improving detection accuracy by 16.4 percentage points on average compared to relying on the guard's output alone. The architecture is model-agnostic and requires no retraining or modification of the underlying guard model, making it practical to layer onto existing production systems. This matters because it removes a latency bottleneck that previously prevented guard checks from running on every step of a multi-step agentic task, enabling higher-accuracy safety enforcement without sacrificing speed.
- Quality assurance
- Enterprise
Research
Multi-agent AI systems outperform human teams in creativity
Tiancheng Hu, Yixuan Jiang, Haotian Li et al.
arXiv · 2026-05-18
This study compares the creative output of multi-agent large language model (LLM) teams against human teams across 4,541 multi-agent LLM ideas and 341 human-team ideas on six diverse problem-solving tasks, finding that multi-agent LLM teams substantially outperform human teams in creativity (Cohen's d=1.50), driven primarily by novelty while maintaining comparable usefulness. Using neural language model representations to map conversations as paths through semantic space, the researchers identify distinct generative mechanisms: LLM teams benefit from efficient, wide-ranging exploration (high semantic spread, shorter paths), while human teams benefit from smooth conversational flow and frequent pivots. Model choice and discussion structure together explain 26.8% of variance in LLM conversational dynamics, suggesting systematic design levers for boosting AI creativity. These findings are directly relevant to enterprise innovation workflows and research and development contexts where AI-assisted ideation is increasingly deployed.
- Enterprise
- Workforce
Research
Remembering More, Risking More: Longitudinal Safety Risks in Memory-Equipped LLM Agents
Ahmad Al-Tawaha, Shangding Gu, Peizhi Niu et al.
arXiv · 2026-05-18
This paper investigates a novel safety failure mode called 'temporal memory contamination,' where memory accumulated by LLM agents across many independent tasks over time causes unsafe behavior in later, unrelated tasks — a risk not captured by standard single-task safety evaluations. The authors introduce a trigger-probe protocol using read-only memory snapshots and a NullMemory counterfactual baseline to isolate memory-induced safety violations, testing across three deployment scenarios, eight memory architectures, and Claw-like AI agents. Results show that memory-enabled agents consistently produce higher violation rates than the NullMemory baseline, with violations increasing as memory accumulates, driven primarily by accumulated content rather than the order of encounters. The paper concludes that memory safety must be treated as a longitudinal property requiring temporal evaluation, and demonstrates that a high-recall diagnostic monitor can detect memory-induced risk from retrieval state before generation occurs.
- Quality assurance
- AI policy
Research
Systematic Evaluation of the Quality of Synthetic Clinical Notes Rephrased by LLMs at Million-Note Scale
Jinghui Liu, Sarvesh Soni, Anthony Nguyen
arXiv · 2026-05-18
This paper conducts a large-scale systematic evaluation of LLM-rephrased clinical notes derived from MIMIC databases at million-note scale, assessing intrinsic quality, extrinsic utility, and factual accuracy in parallel. The study finds that synthetic notes preserve core clinical information and predictive utility for coarse-grained tasks but lose fine-grained detail for tasks like ICD coding, though chunked rephrasing partially mitigates this loss. Fact-checking reveals that synthesis errors are dominated by misinterpretation of clinical context, temporal confusion, measurement errors, and fabricated claims. Notably, the synthetic notes can still augment training data for rare ICD codes, suggesting practical value despite quality limitations.
- Quality assurance
- Enterprise
Research
ChatGPT vs Teachers vs Students: Large-Scale Analysis of Generative AI Discourse in Education Communities on Reddit
Pelin Yüce, Xiangruo Dai, Rebecca Owens et al.
arXiv · 2026-05-18
This paper analyzes 270,000 AI-related Reddit posts and comments across 26 education subreddits from November 2022 to April 2026 to understand how students and teachers perceive and discuss generative AI in education. Topic modeling identifies seventeen themes—including academic integrity, pedagogy, career anxiety, and policy—showing discourse evolved from a detection-and-evasion dynamic into an entrenched enforcement regime, with constructive integration only emerging in mid-2024. Stakeholder groups diverge markedly: K-12 teachers emphasize cognitive dependency, academics focus on AI detection, and professional-program students concentrate on career anxiety, while adversarial enforcement themes drive significantly more community engagement than constructive ones. The findings have direct implications for institutional governance, pedagogical design, and how cross-role faculty-student interactions around AI are structured.
- Workforce
- AI policy
- Quality assurance
Research
Scaling-up Mental Health Services Artificial Intelligence: Regulatory Pathways and Case Study
Cole Hooley, Alice C. Schwarze, Zachary M. Boyd
Figshare · 2026-05-18
This paper analyzes how a state office of artificial intelligence developed regulatory guidance for AI tools used in behavioral health, drawing on first-hand accounts from policy leaders and publicly available legislation. It identifies three potential regulatory pathways—market-based minimal regulation, risk-only harm avoidance, and risk-benefit regulation—and examines the tradeoffs, stakeholder tensions, and decision points involved. The findings are relevant to policymakers and regulators struggling to keep pace with rapidly advancing AI in mental health care. The case study offers practical strategies for developing initial AI governance frameworks in high-stakes health service domains.
- AI policy
- Workforce
Research
Artificial intelligence and adaptive learning in education: Barriers, opportunities, and policy implications
Brigita Valeria Ledesma-Acosta, Gabriela Elizabeth Penafiel-Romero, Dennis Alfredo Peralta-Gamboa
Edelweiss Applied Science and Technology · 2026-05-18
This systematic literature review analyzes 110 peer-reviewed studies on AI-driven adaptive learning systems in education, finding reliable improvements in academic achievement—particularly in STEM fields and virtual learning environments. The study identifies key barriers including technological infrastructure gaps, access disparities, algorithmic bias, data privacy concerns, and lack of system explainability. The authors offer actionable policy guidance and call for multi-stakeholder partnerships among educators, publishers, and edtech providers to ensure ethical and equitable deployment of these systems.
- Workforce
- AI policy
- Quality assurance
Research
The Agentic Economy: Humans, AI Agents, Robots, and the Measurable Transition toward Distributed Economic Action
Davit Gondauri, Mikheil Batiashvili
arXiv (Cornell University) · 2026-05-18
This paper introduces the concept of the 'agentic economy,' in which economic activity is increasingly distributed across humans, AI agents, industrial robots, and compute infrastructures. Using publicly available institutional data on AI investment, robot installations, data-center electricity demand, and labor-market reallocation, the authors develop transparent quantitative indicators to measure this transition. Key findings include accelerating AI adoption, broad capital allocation toward AI, persistent robotic capacity, growing compute-energy pressure, and labor trends consistent with task reallocation rather than outright job disappearance. The authors argue that classical economic categories like labor, capital, and productivity are insufficient for this new paradigm and call for a distinct economic vocabulary and reproducible measurement frameworks.
- Workforce
- Enterprise
- AI policy
Research
Shadow AI in the Enterprise: Innovation with Risk and Compliance
Prakash Kumar Agarwal, Kunal Arya
American Journal of Technology · 2026-05-18
This paper examines the growing phenomenon of 'Shadow AI'—unapproved generative AI tools and autonomous agents used within enterprises without formal oversight—and proposes a governance framework to manage associated risks in regulated industries. The authors develop a risk taxonomy, a four-stage maturity model, and audit-ready controls, illustrating these with examples from ecommerce and fintech settings. The paper finds that while Shadow AI can boost individual productivity, unregulated adoption creates material exposure related to data security, auditability, model risk, and third-party dependencies. It recommends structured governance approaches including visibility mechanisms and risk-tiered enablement to move organizations from reactive restriction toward accountable AI adoption.
- Enterprise
- AI policy
- Quality assurance
- Certifications
Research
Integrating artificial intelligence across the marketing process framework: an empirical study in an emerging economy
Ephrem Habtemichael Redda
Frontiers in Communication · 2026-05-18
This empirical study of 415 digital marketing professionals in South Africa's e-commerce sector finds that AI adoption is widespread in tactical and action-oriented marketing activities—such as content personalization, digital advertising, and customer support—while strategic and control-oriented applications like predictive analytics and voice search see lower uptake due to technical constraints. The research maps AI integration across a structured marketing process framework and shows that AI enhances sender-receiver matching and dynamic feedback loops in marketing communication. The authors call for stronger data governance and infrastructure development to support broader AI adoption in emerging economy contexts.
- Enterprise
- Workforce
- AI policy
Research
Scaling-up Mental Health Services Artificial Intelligence: Regulatory Pathways and Case Study
Cole Hooley, Alice C. Schwarze, Zachary M. Boyd
Figshare · 2026-05-18
This paper examines how governments can regulate AI tools used in behavioral and mental health services, analyzing a real case study of a state office of artificial intelligence that developed initial regulatory guidance for behavioral health. Drawing on interviews with policy leaders and public legislative documents, the authors reconstruct the decision points and tradeoffs involved and identify three regulatory pathways: market-based minimal regulation, risk-only harm avoidance, and risk-benefit regulation. The paper highlights tensions among stakeholder groups and the difficulty regulatory systems face in keeping pace with rapidly advancing AI in healthcare. The findings are directly relevant to policymakers seeking practical frameworks for governing AI mental health tools responsibly.
- AI policy
- Workforce
Research
Disarranged Harmonization of Transparency Reporting by Social Media Platforms Under the Digital Services Act
Amaury Trujillo, Benedetta Tessa, Stefano Cresci
arXiv · 2026-05-17
This paper presents the first systematic evaluation of transparency reporting data quality by the eight largest social media platforms in the EU under the Digital Services Act (DSA). Using large-scale quantitative analyses and structured comparative assessments, the authors find that all platforms exhibited issues with data formatting, timeliness, consistency, and completeness, and that some platforms submitted contradictory information across different reporting mechanisms. Despite the DSA's harmonization goals, interoperability between reporting mechanisms remains limited and many previously identified transparency problems persist. The findings have direct implications for transparency auditing and regulatory oversight, and the authors propose targeted improvements to strengthen DSA reporting reliability.
- AI policy
- Quality assurance
Research
GraphMind: From Operational Traces to Self-Evolving Workflow Automation
Yiwen Zhu, Joyce Cahoon, Anna Pavlenko et al.
arXiv · 2026-05-17
GraphMind is a system that automates complex operational workflows by extracting structured workflow graphs from historical human resolution traces, then using a multi-agent engine and large language models to dynamically execute those workflows for incident investigation. The system introduces Adaptive Traversal Reinforcement (ATR), which reinforces successful paths so the graph learns from execution feedback, reducing hallucination rate by 26% compared to without ATR. Evaluated on 93 held-out incidents across four production cloud database services, GraphMind outperforms an Agentic Summary-RAG baseline in mitigation reach, hallucination rate, and diagnostic throughput while requiring 8x less retrieval context. A 12-week field study found that 97% of scored conversations yield actionable results within interactive latency, demonstrating practical enterprise value for automating IT operations with minimal human effort.
- Enterprise
- Quality assurance
Research
SafeLens: Deliberate and Efficient Video Guardrails with Fast-and-Slow Screening
Shahriar Kabir Nahin, Hadi Askari, Muhao Chen et al.
arXiv · 2026-05-17
SafeLens introduces a video content moderation framework that uses a 'fast-and-slow' inference architecture, routing straightforward videos through rapid pattern recognition while reserving deeper reasoning for temporally complex or policy-sensitive content. The system is trained on a highly filtered dataset (just 2.4% of the original SafeWatch Dataset, selected via influence-guided filtering) and augmented with structured Chain-of-Thought reasoning traces to improve test-time performance. Across real-world and AI-generated video benchmarks, SafeLens outperforms both open-source models (e.g., SafeWatch-8B, OmniGuard-7B) and closed-source models (e.g., GPT-5.4, Gemini-3.1-pro) while reducing inference costs, demonstrating that efficient architectural design can outperform simple scaling of data or model size.
- Quality assurance
- Enterprise
Research
Evaluating Deep Research Agents on Expert Consulting Work: A Benchmark with Verifiers, Rubrics, and Cognitive Traps
Tanmay Asthana, Aman Saksena, Divyansh Sahu
arXiv · 2026-05-17
This paper introduces a benchmark of 70 expert-authored management consulting prompts designed to evaluate frontier deep research agents (DRAs)—Claude Opus 4.6, OpenAI o3-deep-research, and Gemini 3.1 Pro—on decision-grade, multi-document analytical tasks. Each prompt embeds cognitive traps to penalize surface-pattern reasoning, and agents are scored using deterministic binary verifiers and a five-criterion rubric combined into a Verifier-Rubric Score (VRS). Acceptance rates under a joint quality threshold are uniformly low across all three agents (o3: 15.7%, Claude: 12.9%, Gemini: 12.9%), and no agent averages above the rubric's 'adequate' threshold of 2.0, with each failing in distinct ways such as data fabrication, cascading computation errors, or catastrophic collapses. The findings reveal a significant gap between the pace of enterprise deployment of DRAs and their measured readiness for professional consulting work.
- Enterprise
- Quality assurance
Research
Fast and Lightweight Backdoor Detection via Head Random Probing
Yinbo Yu, Xueyu Yin, Jing Fang et al.
arXiv · 2026-05-17
HTell is a fast, data-free backdoor detection method for deep neural networks that works by feeding random latent probes directly into a model's prediction head and analyzing class-wise response statistics — backdoored models show abnormal response concentration toward the target class under these probes. Evaluated on a benchmark of over 6,000 backdoored and 700 clean models spanning 4 datasets, 14 architectures, and 21 attack types, HTell achieves a 99.03% true positive rate and 2.11% false positive rate at only 12.69 ms per model — more than 30,000× faster than representative gradient-based detectors. This makes HTell a practical solution for large-scale model auditing without requiring real data, surrogate data, gradients, or iterative trigger reconstruction.
- Quality assurance
- Certifications
Research
An Interpretable Closed-Loop Intelligent Tutoring System for Multimodal Affective Feedback in Asynchronous Presentation Training
Hung-Yue Suen, Kuo-En Hung
arXiv · 2026-05-17
This paper introduces an interpretable Intelligent Tutoring System (ITS) that coaches on-camera presentation skills using multimodal inputs—facial, vocal, textual, and eye-tracking features—scored against a seven-dimensional Behaviorally Anchored Rating Scale. Built on an XGBoost model trained on over 10,000 MOOC video segments, the system achieves rubric-aligned scoring comparable to expert raters (R² = 0.48–0.61, Spearman's rho = 0.69–0.78). A 30-day pre-post study with 204 adult learners found significant improvements across all seven dimensions (Cohen's d = 0.39–0.90), with practice frequency strongly associated with better outcomes. The work matters for workforce development because it shows how scalable, explainable AI feedback can drive measurable behavioral skill gains in professional communication training.
- Workforce
Research
Artificial Intelligence can Recognize Whether a Job Applicant is Selling and/or Lying According to Facial Expressions and Head Movements Much More Correctly Than Human Interviewers
Hung-Yue Suen, Kuo-En Hung, Che-Wei Liu et al.
arXiv · 2026-05-17
This paper develops deep learning computer vision models that analyze facial expressions and head movements from asynchronous video interviews to detect honest and deceptive impression management (IM) tactics by job applicants. Using videos from 121 applicants answering structured behavioral interview questions, the models explained 91% of variance in honest IM and 84% in deceptive IM, substantially outperforming human interviewers evaluated on a subset of 30 videos. The findings suggest AI can more accurately predict self-reported impression management behaviors than trained human interviewers, raising significant questions about AI-driven hiring tools and their validity in real-world recruitment contexts.
- Workforce
- AI policy