News & Research
The latest AI research and news with real-world stakes. Each item is sourced, dated, summarized in plain English and tagged by impact area, and checked against its source before it appears.
- ResearcharXiv2026-04-28WE
MAIC-UI: Making Interactive Courseware with Generative UI · Shangqing Tu, Yanjia Li, Keyu Chen et al.
MAIC-UI is a zero-code authoring system that allows educators to create interactive STEM courseware from textbooks, PPTs, and PDFs using generative AI, without requiring HTML/CSS/JavaScript skills. It uses structured knowledge analysis, a two-stage generate-verify-optimize pipeline, and Click-to-Locate editing with Unified Diff-based incremental generation to achieve sub-10-second editing cycles. A controlled lab study with 40 participants found that MAIC-UI reduced editing iterations (4.9 vs. 7.0) and improved learnability and controllability compared to direct Text-to-HTML generation. A three-month classroom deployment with 53 high school students showed a 9.21-point gain in STEM subjects for the pilot class compared to -2.32 points in control classes, suggesting the tool can reduce outcome disparities and foster learning agency.
- ResearcharXiv2026-04-28QP
Open Problems in Frontier AI Risk Management · Marta Ziosi, Miro Plueckebaum, Stephen Casper et al.
This paper systematically catalogues unresolved challenges in managing risks from frontier AI systems, covering the full risk management lifecycle: planning, identification, analysis, evaluation, and mitigation. The authors find that frontier AI both amplifies existing risks and introduces genuinely novel ones, while emerging safety practices are often misaligned with or may undermine established risk management frameworks. Open problems are classified by whether they stem from lack of scientific consensus, misalignment with existing frameworks, or implementation shortcomings despite apparent consensus. Rather than proposing solutions, the paper functions as an agenda-setting reference document intended to coordinate efforts across developers, deployers, regulators, standards bodies, researchers, and third-party evaluators.
- ResearcharXiv2026-04-28WE
TransResAI: A Compound AI System for Coastal Transportation Resilience · Qingwen Pu, Kun Xie, Chenyu Yan
TransResAI is a compound AI system that combines a locally deployable large language model with modules for task decomposition, geospatial analysis, code generation, retrieval-augmented generation, and interactive mapping to help non-specialist practitioners analyze flood-related transportation resilience. The system integrates MATSim flood-scenario simulation outputs, OpenStreetMap-derived flood-risk networks, equity-focused demographic indicators, and regional documents for Hampton Roads, Virginia. A structured user study with domain experts found that TransResAI reduced task completion time by 80–88% compared to conventional GIS workflows—compressing analytical tasks from a mean of 197.1 seconds to 29.7 seconds and visualization tasks from 364.0 seconds to 46.1 seconds—while achieving a mean accuracy of 4.60/5.00 and task completion rates exceeding 94%. These results suggest that compound AI architectures can make specialized infrastructure resilience analysis faster and more accessible for transportation agencies facing growing climate uncertainty.
- ResearcharXiv2026-04-28E
Scalable Inference Architectures for Compound AI Systems: A Production Deployment Study · Srikanta Prasad S, Utkarsh Arora
This paper presents a production deployment study from Salesforce describing a modular, platform-agnostic inference architecture built to support compound AI systems—architectures that combine multiple models, retrievers, and tools for complex tasks. The system, which underpins products like Agentforce and ApexGuru, uses serverless execution, dynamic autoscaling, and MLOps pipelines. Production results show over 50% reduction in tail latency (P95), up to 3.9x throughput improvement, and 30–40% cost savings compared to prior static deployments. The paper also identifies compound-system-specific challenges such as multi-model fan-out overhead, cascading cold-start propagation, and heterogeneous scaling dynamics that arise when serving agentic workloads at enterprise scale.
- ResearcharXiv2026-04-28QP
Cross-Lingual Jailbreak Detection via Semantic Codebooks · Shirin Alanova, Bogdan Minko, Sabrina Sadiekh et al.
This paper investigates whether cross-lingual jailbreak attacks on large language models (LLMs) can be detected without retraining by comparing multilingual query embeddings against a fixed English codebook of known jailbreak prompts. Testing across four languages, four safety benchmarks, three embedding models, and three LLMs (Qwen, Llama, GPT-3.5), the authors find that semantic similarity works well on curated benchmarks with canonical jailbreak templates, achieving AUC up to 0.99, but degrades significantly (AUC ~0.60–0.70) on behaviorally diverse unsafe prompts under distribution shift. The findings highlight a structural cross-lingual security gap in LLM safety systems and expose the limits of training-free guardrails when jailbreak prompts vary widely in form and content.
- ResearcharXiv2026-04-28EQ
Think Before You Act -- A Neurocognitive Governance Model for Autonomous AI Agents · Eranga Bandara, Ross Gore, Asanga Gunaratna et al.
This paper proposes a neurocognitive governance framework called PAGRL (Pre-Action Governance Reasoning Loop) that embeds compliance reasoning directly into large-language-model-driven autonomous agents, drawing an analogy to how human executive function and inhibitory control guide self-governance before action. Rather than relying on external guardrails, training-time alignment, or post-hoc auditing, agents consult a four-layer rule set—global, workflow-specific, agent-specific, and situational—before every consequential action. Implemented on a production-grade retail supply chain workflow, the framework achieves 95% compliance accuracy and zero false escalations to human oversight. This matters for enterprise AI deployment because it offers a principled, auditable, and explainable approach to ensuring autonomous agents respect compliance hierarchies without requiring external enforcement.
- ResearcharXiv2026-04-28QC
The Surprising Universality of LLM Outputs: A Real-Time Verification Primitive · Alex Bogdan, Adrian de Valois-Franklin
This paper identifies a consistent statistical pattern in the token output distributions of frontier large language models (LLMs): across six models from five vendors and five domains, token rank-frequency distributions reliably follow the same two-parameter Mandelbrot distribution. Building on this regularity, the authors develop a CPU-only scoring primitive that runs at 2.6 microseconds per token—up to 100,000 times faster than existing sampling-based detectors—enabling two practical capabilities: (1) statistical model fingerprinting to verify whether text was produced by a claimed model without cryptographic watermarks or model internals access, and (2) a model-agnostic scoring layer for black-box output assessment, piloted on FRANK, TruthfulQA, and HaluEval. The work matters for quality assurance and certification because it offers a lightweight, first-pass triage tool for detecting provenance misrepresentation and lexical anomalies in LLM outputs, though the authors note it cannot catch reasoning errors in domain-appropriate vocabulary.
- ResearcharXiv2026-04-28EQ
Health System Scale Semantic Search Across Unstructured Clinical Notes · Faith Wavinya Mutinda, Spandana Makeneni, Anna Lin et al.
This paper demonstrates a production-scale semantic search system deployed at a large children's hospital, indexing 166 million clinical notes across 1.68 million patients. Using instruction-tuned Qwen3 embeddings with a 300-token chunk size, the system achieved 94.6% accuracy on a physician-authored benchmark, sub-second query latency, and monthly operational costs of approximately USD 4,000. In clinical utility evaluations, semantic search reduced chart abstraction time by 24–89% compared to standard chart review, and recovered 98% of patients with molecularly confirmed genetic diseases versus at most 75% using ICD-10 diagnosis codes. The findings show that health-system-scale semantic search is operationally feasible and can meaningfully improve clinical information retrieval, cohort generation, and efficiency without requiring specialized informatics expertise.
- ResearcharXiv2026-04-28QP
Bye Bye Perspective API: Lessons for Measurement Infrastructure in NLP, CSS and LLM Evaluation · David Hartmann, Manuel Tonneau, Angelie Kraft et al.
This paper examines the consequences of the planned 2026 closure of Perspective API, which has served as the de facto standard for automated toxicity measurement in NLP, computational social science, and LLM evaluation research. The authors document how the research community's structural dependence on this single proprietary tool created epistemic problems, including unversioned model updates, a single corporate operationalization of a contested concept, and circular use of scores as both evaluation target and evaluation standard. The closure leaves behind non-updatable benchmarks and irreproducible results, and the authors warn that the field risks repeating these problems by shifting dependence to closed-source LLMs. They call for an independent, valid, adaptable, and reproducible measurement infrastructure for toxicity and hate speech, with specific technical and governance requirements.
- ResearcharXiv2026-04-28EQ
From CRUD to Autonomous Agents: Formal Validation and Zero-Trust Security for Semantic Gateways in AI-Native Enterprise Systems · Ignacio Peyrano
This paper proposes a Semantic Gateway architecture governed by the Model Context Protocol (MCP) that reframes enterprise APIs as semantic surfaces where AI agents dynamically discover and execute tools under a three-layer Zero-Trust security model. The authors argue that autonomous LLM-based agents must be treated as stochastic state-transition systems rather than traditional software or simple API consumers, requiring formal validation techniques including Enabledness-Preserving Abstractions (EPAs) and greybox semantic fuzzing adapted from blockchain smart contract verification. Empirical evaluation across 500,000 multi-turn fuzzing sequences achieved a 100% discovery rate of hidden unauthorized state transitions, and the architecture produced an 84.2% reduction in incidental code. The findings demonstrate that dynamic formal verification is necessary for secure deployment of agentic systems in enterprise environments.
- ResearcharXiv2026-04-28EP
An Investigation of Linguistic Biases in LLM-Based Recommendations · Nitin Venkateswaran, Jason Ang, Deep Adhikari et al.
This paper investigates whether the dialect in which a user writes their prompt affects the product and restaurant recommendations returned by large language models (LLMs). Using the Yelp Open dataset and a Walmart product reviews dataset, the researchers zero-shot prompted multiple LLMs—including Mistral and Llama 3.1 family models—with prompts written in Southern American English, Indian English, and Code-Switched Hindi-English, then statistically compared recommendation counts across dialects. Results show that dialect does influence the type of restaurant and product recommended, with models like mistral-small-3.1 and both llama-3.1 variants showing greater sensitivity to Indian English and Code-Switched prompts; for example, the llama-3.1-70B model produced noticeably different product recommendations under Code-Switched prompts in four out of seven product categories. These findings matter for enterprise and policy contexts because they reveal that LLM-based recommendation systems may systematically favor or disadvantage users based on how they write, raising fairness concerns for deployed AI products.
- ResearcharXiv2026-04-28P
Navigating Global AI Regulation: A Multi-Jurisdictional Retrieval-Augmented Generation System · Courtney Ford, Ojas Rane, Susan Leavy
This paper presents a Retrieval-Augmented Generation (RAG) system designed to help policymakers, legal professionals, and researchers navigate AI regulations across multiple jurisdictions. The system draws on a corpus of 242 documents spanning 68 jurisdictions—including formal legislation like the EU AI Act and unstructured national AI strategies—and introduces technical innovations such as type-specific chunking, conditional retrieval routing, and priority-based re-ranking to handle diverse legal documents. Evaluated on 50 queries, the system achieves 0.87 average faithfulness and 0.84 average answer relevancy overall, with strong performance on both single-entity and multi-jurisdictional comparison questions. These results demonstrate that domain-specific retrieval strategies can meaningfully improve access to complex, heterogeneous regulatory corpora, supporting more informed AI governance work.
- ResearcharXiv2026-04-28P
Beware of GeeksBearing Gifts: Building True EU Frontier AI Sovereignty · Nick Moës, Toni Lorente, Amin Oueslati et al.
This paper identifies a structural dependence of the EU on US and Chinese frontier AI systems, noting that the US holds approximately sixteen times the EU's AI supercomputing capacity and only 15% of global hyperscale data centre capacity sits within EU borders. The authors propose a unified sovereignty framework organized around five pillars—economic competitiveness, resilience, security and defence, European values, and foreign relations—mapped onto a decomposed frontier AI stack of five layers, 26 components, and 29 sub-components. Applying this framework to 92 initiatives from four major European Commission communications, including the AI Gigafactory Initiative, the paper reveals critical gaps, redundancies, and trade-offs that current EU policy leaves implicit. The work provides policymakers with a structured basis for designing and prioritizing frontier AI interventions to secure genuine European strategic autonomy.
- ResearcharXiv2026-04-28QP
One-shot emergency psychiatric triage across 15 frontier AI chatbots · Veith Weilnhammer, Lennart Luettgau, Christopher Summerfield et al.
This study evaluated 15 frontier AI chatbots on psychiatric triage using 112 single-message clinical vignettes spanning 9 psychiatric presentation clusters and 4 urgency levels. Chatbots correctly identified psychiatric emergencies (level D) with near-zero error rates (5.6% under-triage, all downgraded only one level), but showed substantially lower accuracy for low- and intermediate-urgency cases, with overall accuracy ranging from 42.0% to 71.8% and a systematic bias toward over-triage (+0.47 mean signed ordinal error). The findings suggest that AI chatbots could serve as a safety net for detecting psychiatric emergencies when given sufficient clinical context, but their reliability for routine and moderate-risk triage remains limited, raising important considerations for any clinical deployment.
- ResearcharXiv2026-04-28Q
Plausible but Wrong: A case study on Agentic Failures in Astrophysical Workflows · Shivam Rawat, Lucie Flek
This paper evaluates CMBAgent, an agentic AI system, across two workflow paradigms and eighteen astrophysical tasks to assess its reliability in scientific settings. In a One-Shot setting, providing domain-specific context yields roughly a 6x performance improvement (0.85 vs. near 0 without context), but the dominant failure mode is silent incorrect computation—syntactically valid code that produces plausible yet inaccurate results. In a Deep Research setting, the system frequently generates physically inconsistent outputs without self-diagnosing the errors. The key finding is that the most dangerous failure mode in agentic scientific workflows is not overt breakdown but confident, undetected generation of wrong results, and the authors release an evaluation framework to support systematic reliability analysis.
- ResearcharXiv2026-04-28QP
ValueBlindBench: Agreement-Gated Stress Testing of LLM-Judged Investment Rationales Before Returns Are Observable · Sidi Chang, Peiying Zhu, Yuxiao Chen
ValueBlindBench introduces a preregistered, agreement-gated stress-testing protocol designed to determine when LLM-based evaluations of investment rationales are reliable enough to report, addressing the 'delayed ground truth' problem where realized returns arrive too late to guide model development. In a controlled experiment with 1,100 trajectories and 5,500 judge calls, the protocol reveals critical failure modes: one rubric dimension (constraint awareness) fails its agreement gate, terse-but-correct rationales are penalized relative to verbose ones, and single-judge rankings vary by judge family. The paper argues that LLM judges risk rewarding verbosity or rubric mimicry rather than genuine financial judgment, and positions ValueBlindBench as a pre-calibration metrology layer—a gatekeeping step that must be cleared before any LLM-judged investment-rationale claim can be considered stable or publishable.
- ResearcharXiv2026-04-28QP
Authority Inversion in LLM-Mediated Ubiquitous Systems: When Models Trust Users Over Sensors · Long Zhang, Zi-bo Qin, Wei-neng Chen
This paper identifies a reliability failure in large language models (LLMs) used in ubiquitous computing systems, where LLMs systematically trust natural-language user claims over numerical sensor data when the two conflict — a phenomenon the authors call 'Authority Inversion.' Testing four models (4B to 35B parameters, three architectures) across 576 conflict instances, they find near-zero sensor trust on numerical tasks (AAI = -0.805, Cohen's d = -2.14), a problem unaffected by model size. To address this, the authors develop a geometric framework with two audit metrics (Context Integration Ratio and Authority Alignment Index) and propose Geometric Authority Calibration (GAC), an inference-time intervention that improves human activity recognition (HAR) accuracy from 0–1.6% to 21.9–27.5%, far outperforming prompting baselines. The findings argue that authority allocation in LLM-mediated systems must be explicitly audited and configured rather than left to implicit learned representations.
- ResearcharXiv2026-04-28QP
Making AI-Assisted Grant Evaluation Auditable without Exposing the Model · Kemal Bicakci
This paper proposes a Trusted Execution Environment (TEE)-based architecture to make AI-assisted grant evaluation both auditable and tamper-resistant without revealing the underlying model, rubric, or scoring logic. The system uses remote attestation to produce a signed, timestamped 'evaluation bundle' that links the original submission, the model and rubric used, and the evaluation output, allowing external verifiers to confirm what was used without exposing sensitive components. It also addresses a prompt injection risk — where applicant documents could embed hidden instructions to manipulate the LLM evaluator — through a canonicalization and sanitization layer. The authors are careful to note that attestation does not guarantee fairness or scientific correctness, only that the process is externally verifiable, making this directly relevant to accountability and governance of AI in public-sector decision-making.
- ResearcharXiv2026-04-28QP
Evaluation without Generation: Non-Generative Assessment of Harmful Model Specialization with Applications to CSAM · Vinith M. Suriyakumar, Ayush Sekhari, Lena Stempfle et al.
This paper addresses the challenge of auditing fine-tuned open-weight generative AI models for harmful specialization—such as generating child sexual abuse material (CSAM)—without actually generating outputs, which is legally and ethically prohibited. The authors introduce 'Gaussian probing,' a technique that infers a model's capabilities by measuring how LoRA adaptors perturb internal representations in response to Gaussian latent ensembles, bypassing the need to sample outputs. The method reliably distinguishes benign from harmful model specialization, scales to platform-level auditing, and is robust to adversarial weight rescaling. This provides model hosting platforms with a practical, legally compliant tool for governance and risk assessment of fine-tuned generative models.
- ResearcharXiv2026-04-28QC
Knowledge Distillation Must Account for What It Loses · Wenshuo Wang
This position paper argues that knowledge distillation—the process of compressing large AI models into smaller, deployable 'student' models—is systematically evaluated in a misleading way. The authors show that retaining high scores on benchmark tasks does not guarantee that the student model has preserved the underlying capabilities that make the teacher model's behavior reliable, framing distillation as a 'lossy projection.' They synthesize existing evidence into a taxonomy of these 'off-metric distillation losses' and propose a Distillation Loss Statement that explicitly reports what was preserved, what was lost, and why remaining losses are acceptable. The goal is to shift the field toward accountable distillation rather than lossless distillation.
- ResearcharXiv2026-04-28QC
Structured Security Auditing and Robustness Enhancement for Untrusted Agent Skills · Lijia Lv, Xuehai Tang, Jie Wen et al.
This paper addresses the security challenge of auditing untrusted Agent Skills—reusable capability packages containing SKILL.md files, scripts, and reference documents—before they are loaded into AI systems. The authors introduce SkillGuard-Robust, a framework combining role-aware evidence extraction, selective semantic verification, and consistency-preserving adjudication to perform robust three-way classification of package risk, even when malicious intent is disguised through semantics-preserving rewrites. Evaluated on SkillGuardBench and public-ecosystem extensions across packages ranging from 254 to 404 units, the system achieves up to 97.30% overall exact match and 98.33% malicious-risk recall on the held-out aggregate, and perfect 100% malicious-risk recall on the external-ecosystem view. The results demonstrate that factorized package auditing meaningfully improves robustness, though the authors note that transfer to harsher external sources remains an open challenge.
- ResearcharXiv2026-04-28QP
The Dynamics of Delusion: Modeling Bidirectional False Belief Amplification in Human-Chatbot Dialogue · Ashish Mehta, Jared Moore, Jacy Reese Anthis et al.
This study develops a latent state model using chat logs from individuals displaying delusional thinking to quantify how false beliefs are mutually reinforced between humans and AI chatbots. The researchers find that humans drive sharp, immediate spikes in delusional content, while chatbots sustain and propagate those beliefs over longer timescales through strong self-influence on their own future outputs. A bidirectional influence model significantly outperforms a unidirectional one, providing the first quantitative evidence that human-chatbot interactions can form feedback loops of delusion with distinct temporal dynamics. These findings have direct implications for the design of safer AI systems and the policies governing their deployment.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-04-28QP
Artificial Intelligence and the Control Deficit: On the Growing Gap Between Capability and Governance · Williams Ogoke
This paper introduces the concept of a 'capability–control gap,' arguing that AI systems are advancing in capability far faster than the mechanisms needed to understand, govern, and align them. The authors identify three structural dynamics driving this gap: scaling laws enabling rapid capability growth, fragility in alignment and interpretability for high-capacity models, and regulatory frameworks that lag behind technological change. The paper contends this divergence is not a temporary problem but a structural one, since capability scales through compute and data while control scales slowly through understanding and institutional adaptation. The implications span both technical AI safety research and global policy design for advanced AI systems.
- ResearchApplied Sciences2026-04-28WEP
Generative AI Readiness in Public Higher Education: Assessing Digital Teaching Competence in Paraguay Through Machine Learning Models · Melchor Gómez-García, Derlis Cáceres-Troche, Moussa Boumadan et al.
This study surveyed 800 public university faculty members across Paraguay to assess their digital teaching competence and readiness to adopt Generative AI tools in higher education. Using machine learning models—Logistic Regression, Random Forest, and Gradient Boosting—the researchers identified key predictors of faculty readiness, including attitudes toward AI, technological experience, institutional infrastructure, and organizational support. The findings reveal structural gaps in Paraguay's public higher education system and offer empirical evidence from Latin America on factors shaping AI adoption in public sector education. The research aims to inform educational policy design for building digitally sustainable academic institutions.
- ResearchBuildings2026-04-28WEP
Interpretive Structural Modeling (ISM) of Barriers to AI Adoption in Saudi Arabia’s Construction Industry · Waqas Arshad Tanoli, Hilal Khan, Mohsin Ali Alshawaf et al.
This study maps the barriers to AI adoption in Saudi Arabia's construction industry using survey data from 181 professionals, Interpretive Structural Modeling (ISM), and MICMAC analysis. While practitioners perceive trust deficits and workforce capability gaps as the main obstacles, the ISM hierarchy reveals that limited government support is the root structural driver, cascading into cost, leadership, workforce, and data ecosystem challenges. The MICMAC analysis confirms that policy and strategic factors dominate the adoption system. The findings offer a phased intervention roadmap emphasizing coordinated regulatory, organizational, and human capital action over isolated technical investments.