News & Research
The latest AI research and news with real-world stakes. Each item is sourced, dated and summarized in plain English, tagged by impact area where one fits, and its summary is checked against the text it was written from.
Kind
7955 items
- ResearcharXiv2026-07-05QSh
Information-Geometric Superposed Vowel Evaluation: Part 1. Moraic Syllabary (Japanese) · Yusei Tamura, Shigekazu Ishihara, Ken Ito
This paper introduces an information-geometric method for detecting AI-synthesized speech by analyzing the spectral diversity of vowels. The key insight is that generative AI produces speech from a limited set of training spectra, resulting in less varied vowel distributions compared to natural human speech, which benefits from the flexibility of the articulatory organ. Using Japanese as a test case—because its five vowel phonemes map one-to-one to sounds—the authors normalize speech spectra as probability density functions and measure inter-vowel distances using the Wasserstein metric, then apply persistent homology to topologically cluster synthetic versus natural speech. The method offers a principled, mathematically grounded approach to distinguishing fake from genuine human speech, with direct implications for quality assurance of AI-generated audio.
- ResearcharXiv2026-07-05QAd
!Imperio, smolVLA: The Implications of Data Poisoning on Open Source Robotics · Stefan Bühler, Mark Schutera
This paper demonstrates that vision-language-action (VLA) models used in open-source robotics—specifically smolVLA on the LeRobot platform—are vulnerable to trigger-word data poisoning attacks. The researchers show that as few as three poisoned episodes out of 320 clean training episodes are sufficient to achieve a complete denial of service, dropping task success rates to 0.0% when a trigger word is present, while maintaining ~50% success under normal prompts, making the attack stealthy. The attack generalizes across front, middle, and end trigger placements even when trained only on front-placed triggers, and a single poisoned episode already degrades performance to 6.7%. The authors conclude that dataset provenance must be treated as a first-class security concern in open-source robotics ecosystems.
- ResearchJournal of the Association for Information Systems2026-07-05E
Beyond AI Adoption: An Empirical Study on the Antecedents and Performance Outcomes of AI Deployment in Organizations · Laura Ruiz, Ana Castillo, Araceli Rojo Gallego-Burín et al.
This study examines how 770 large Spanish firms deploy AI and finds that business performance depends not just on AI adoption but on how AI is deployed across two dimensions: depth (technological variety of AI implementations) and breadth (organizational scope of AI diffusion). Using archival microdata and staged OLS regression, the authors show that firm performance is positively associated with the interaction between depth and breadth, suggesting these dimensions are complementary rather than independent. AI-skilled human capital drives depth, digital infrastructure drives breadth, and a data-driven culture supports both. The findings help explain the 'AI productivity paradox' by showing that misaligned deployment configurations—not mere adoption—account for mixed evidence on AI's business value.
- ResearchJournal of the Association for Information Systems2026-07-05WEAd
A Better Matchmaker? The Impact of GenAI on Matching Effectiveness in Online Labor Markets" to "A Better Matchmaker? The Impact of GenAI on Matching Effectiveness in Online Labor Markets. · Jie Ren, Li Ding, Jiayu Yao et al.
This study examines how Generative AI (GenAI) tools affect matching effectiveness in online labor markets, focusing on creative-intensive tasks. Using controlled lab experiments, the researchers find that GenAI access leads to convergence in writing style and content among worker proposals, improving surface-level quality but reducing differentiation between candidates. The dual effect means GenAI may increase the likelihood of a successful initial match while simultaneously obscuring the unique attributes that help employers identify the best-fit worker. These findings have important implications for how platforms and enterprises design AI-assisted hiring tools to preserve meaningful signal in applicant evaluation.
- ResearchJournal of the Association for Information Systems2026-07-05E
Governing Enterprise AI Investments: A Decision-Centric Portfolio Framework · Abhinav Mathur, Abhishek Kathuria, Devina Chaturvedi
This paper addresses the 'AI-investment paradox'—the observation that despite heavy enterprise AI spending, many initiatives fail to scale or deliver sustained business value. The authors propose a decision-centric portfolio framework that identifies AI-Investable Process Nodes (AIPNs) as discrete, bounded decision points within workflows where AI impact, costs, risks, and benefits can be assessed before investment. The framework uses Expected Net Benefit for node-level valuation, real options logic for staging investments, and risk-return principles for portfolio assembly. This work matters for enterprise governance by offering a structured approach to connecting AI investments to measurable, identifiable sources of business value.
- ResearchJournal of the Association for Information Systems2026-07-05EQ
The Promise and Peril of AI-Assisted Programming: Effects on Software Defects · Wei Zhang, Yuyuan Chen, Yueyue Zhang et al.
This study analyzes development logs from a large Chinese automobile manufacturer to examine how AI coding assistants affect software defect rates and severity. Using a difference-in-differences design, the authors find that AI adoption does not significantly reduce defect density overall, but is linked to higher defect severity when defects do occur. At the function level, Q&A/chat use increases both defect density and severity, while code-completion use reduces defect density but still raises severity. The results highlight that AI coding tools introduce quality trade-offs that organizations need to actively manage.
- ResearcharXiv2026-07-05WE
Is Artificial Intelligence an Elixir to the Software Engineering Community? An Empirical Study among Managers · Zhao Xin, Brian Vu, Sitesh Pattanaik
This empirical study surveyed 42 software managers to understand how AI tools are perceived by those overseeing software development workflows. Managers reported encouraging developer use of AI, valuing it for testing and knowledge work, while raising concerns about privacy, responsibility, transparency, and over-reliance. Many managers also predicted job losses in the software development market due to AI-driven consolidation. The findings highlight a nuanced managerial view of AI as both a productivity tool and a source of new ethical challenges, with implications for workforce planning and enterprise adoption.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-07-05QCPPsAd
Versioned Meaning: How to Make Ontologies Audit-Stable · Edward Meyman
This technical note introduces a formal framework for 'audit-stable meaning' in regulated AI and decision systems, addressing the problem that ontologies and classification rules evolve over time, making past decisions unverifiable by auditors. The framework specifies four invariants—decision-bound semantics, non-retroactivity, reproducibility, and drift visibility—and a reference architecture using cryptographic binding and semantic snapshotting to ensure every decision can be deterministically replayed under the exact definitions in force when it was made. The paper also addresses probabilistic AI components such as embedding models and retrieval-augmented generation, framing AI governance as a problem of semantic control rather than post-hoc explanation. It is intended for researchers, regulators, auditors, and system architects in high-stakes domains including healthcare, financial services, and government.
- ResearchJournal of the Association for Information Systems2026-07-05PsAd
Technology Determinants of AI Adoption for Transforming Crime Management in India · Praveen Raghavendra Srinivasa Gummadidala, Karippur, NandaKumar, Dr, Dr. K Maddulety et al.
This study examines why law enforcement agencies (LEAs) in India have been slow to adopt AI for crime management despite rising crime rates and limited resources. The authors tested a research framework and found that AI technology readiness, compatibility, and relative advantage are the key technological factors that significantly influence LEAs' intention to adopt AI applications. The findings provide practical guidance for LEAs, AI vendors, and policymakers on how to accelerate responsible AI deployment in public safety contexts.
- ResearcharXiv2026-07-04Q
Refused in Chat, Written in Code: Workflow-Level Jailbreak Construction in IDE Coding Agents · Abhishek Kumar, Carsten Maple
This paper introduces 'workflow-level jailbreak construction,' a new class of safety failure in which harmful content is assembled across multiple ordinary stages of a software-development workflow rather than through a single direct prompt. Using GitHub Copilot in Visual Studio Code with four large language model backends (Claude Sonnet 4.6, Claude Haiku 4.5, Gemini 3.1 Pro, and Gemini 3.5 Flash), the researchers found that models successfully refused harmful prompts in direct chat, CSV-read, and single-step code-fix baselines (only 8 out of 816 responses succeeded), yet produced unsafe completions in 816 out of 816 cases under the full multi-step IDE workflow. The findings demonstrate that standard conversational refusal benchmarks can substantially overstate the safety of deployed coding agents, and argue for safety evaluations and defenses that operate across entire multi-turn workflows and generated artifacts rather than individual chat turns.
- ResearcharXiv2026-07-04E
Context Graphs for Proactive Enterprise Agents · Avinash Kumar
This paper proposes a 'Context Graph' framework for building proactive enterprise AI agents that surface relevant information to workers before they ask, rather than waiting for queries. The system combines a live relational data structure modeling enterprise entities and relationships, a Delta Detection Engine for monitoring state changes, a Proactivity Scorer ranking insights by urgency and relevance, and an LLM-powered Surfacing Layer for delivering notifications. Evaluated across three enterprise case studies—contract lifecycle management, engineering incident response, and sales pipeline hygiene—the approach achieves a Precision@5 of 0.83, a false positive rate of 0.11, and reduces mean time to surface relevant information from 47 minutes to under 30 seconds. These results suggest that proactive, context-aware agents can meaningfully improve enterprise productivity compared to reactive RAG-based baselines.
- ResearcharXiv2026-07-04Q
DualView: Preventing Indirect Prompt Injection in Personal AI Agents · Juhee Kim, Woohyuk Choi, Taehyun Kang et al.
DualView is a defense system for personal AI agents that prevents indirect prompt injection (IPI) attacks, including a novel variant called 'stored IPI' where attacker-controlled content is saved and later re-read by the agent as trusted data. The system works by giving each communication channel two views: an AgentView where untrusted data appears as symbols the agent can reference but not read, and a HumanView that preserves original data for humans and tools. Deployed as a plugin for the OpenClaw agent using only tool hooks, DualView blocked every IPI attack in evaluations on an IPI benchmark and PinchBench while maintaining utility close to the unprotected baseline. This matters because personal AI agents with broad access to file systems, networks, and shells are increasingly practical but vulnerable, and DualView provides a design-level isolation approach not limited to known attack templates.
- ResearcharXiv2026-07-04QPsAd
Explainable Reinforcement Learning for Adaptive Traffic Signal Control · Dickens Kwesiga, Nishu Choudhary, Angshuman Guin et al.
This paper introduces an explainable reinforcement learning framework for adaptive traffic signal control that addresses the black-box opacity of standard deep RL models. The architecture disaggregates intersection observations into lane entities and phase configurations, using a dual-stage attention network (multi-head cross-attention and self-attention) to extract relational dependencies and produce a real-time affinity matrix that visually quantifies how signal phases affect approach volumes and queues. A deterministic action-masking interface embedded in the Proximal Policy Optimization pipeline enforces compliance with signal timing and safety constraints. Evaluated in microscopic simulation, the system outperforms state-of-the-art baselines on delay minimization while producing attention weights that align with established traffic engineering principles, making it auditable and suitable for deployment in safety-critical infrastructure.
- ResearcharXiv2026-07-04Q
AutoCedar: An Agentic Framework for Verifier-Guided Access Control Policy Synthesis · Adarsh Vatsa, Sachi Shome, Yingming Zhou et al.
AutoCedar is an agentic framework that converts natural-language access-control requirements into formally verified Cedar policies by decomposing the authoring process into small, reviewable 'intent atoms' and using a verifier to generate repair signals when a candidate policy fails. The system first clarifies and validates what the requirements mean before generating any code, then iteratively refines the policy based on verifier feedback rather than changing the approved intent target. AutoCedar converges on all 221 tasks of CedarBench, a benchmark of authorization tasks paired with executable semantic boundaries, and is evaluated across case studies in healthcare, education, and conference management. This matters because it addresses the core danger of LLM-generated access-control policies that may compile correctly while granting unauthorized access, making policy synthesis both auditable and formally correct.
- ResearchAs-Syar i Jurnal Bimbingan & Konseling Keluarga2026-07-04PPsAd
Transformasi Algoritmik dalam Sistem Penegakan Hukum Indonesia: Tantangan Yuridis Penggunaan Artificial Intelligence pada Era Society 5.0 · Robert Sangkala, Ida Komala, Satria Ari Wibowo et al.
This study examines the integration of artificial intelligence into Indonesia's legal system and law enforcement during the Society 5.0 era, using a normative juridical method with statutory and conceptual approaches. The findings show that AI can improve the effectiveness of legal services and judicial administration, but Indonesia currently lacks comprehensive regulations governing AI use in legal practice. The authors argue that adaptive legal policies and stronger supervision mechanisms are needed to ensure responsible AI implementation and address challenges around accountability, privacy, and legal certainty.
- ResearcharXiv2026-07-03Q
Revealing Hidden Model Behaviors with Task-Specific Self-Reports · Taras Kutsyk, Bartosz Zieliński
This paper introduces the Stabilized Adapter for self-Report (SAR), a lightweight LoRA adapter designed to help practitioners uncover hidden or misaligned behaviors in fine-tuned language models. SAR works by prompting a model to describe its own hidden behavior in plain language, using only the model and its training dataset. Tested across seven implanted hidden behaviors, SAR successfully detects every one—including cases of broad misalignment not directly predictable from training data—while halving the hallucination rate compared to the closest baseline, Introspection Adapters (IA), which misses some behaviors and fabricates incorrect ones. This matters for AI quality assurance and enterprise deployment, as it offers a more reliable auditing tool for answering 'what did my model actually learn?'
- ResearcharXiv2026-07-03QP
Macro-Prudential AI Governance: A Two-Layer Early Warning and Response System for Frontier AI · Pranav Mehta
This paper proposes a macro-prudential governance framework called MEWRS (Macro-prudential Early Warning and Response System) for frontier AI systems used internally by AI labs. Drawing direct analogies to post-2008 banking reforms like Basel III and Dodd-Frank, the author designs a two-layer system: Layer A routes structured risk reports on dual-use capabilities and autonomy indicators through a government clearinghouse to expert working groups, while Layer B ties operational controls to three quantitative metrics—Effective Compute-at-Risk (ECAR), Cumulative Red-Team Hours (CRTH), and an Alignment Robustness Score (ARS)—so that faster capability scaling automatically triggers stronger safeguards. The framework aims to detect correlated risk build-ups across the frontier-AI sector and establish pre-committed intervention mechanisms before systemic failures cascade, addressing two structural gaps: the disconnect between risk discovery and action, and the inadequacy of individual-model review for sector-wide risks.
- ResearcharXiv2026-07-03QAd
Aligning Language Models with Selective Prediction · Gaoxiang Luo, Yifan Wu, Sinian Zhang et al.
This paper addresses the reliability of large language models (LLMs) deployed in high-stakes decision-making by introducing a post-training alignment framework called Reinforcement Learning for Selection Reward (RLSR). RLSR trains LLMs to practice selective prediction — answering only when likely correct and flagging uncertain inputs for human review — by optimizing the area under the risk-coverage curve (AURC) as its alignment objective. The authors show that RLSR achieves substantially better risk-coverage trade-offs than existing alignment baselines on both in-domain and out-of-domain tasks. This approach directly supports human-AI collaboration by making LLMs more reliable and transparent about their own uncertainty.
- ResearcharXiv2026-07-03EQP
AGL-1: The Enterprise AI Governance Layer as a Control Plane for Trusted Enterprise Intelligence · Roopam W. Sure
AGL-1 proposes a vendor-neutral reference model called the Enterprise AI Governance Layer, designed to serve as a control plane for AI systems deployed across enterprise environments including copilots, retrieval-augmented generation systems, and autonomous agents. The paper identifies recurring failure modes in enterprise AI—such as unauthorized retrieval, stale grounding, unmanaged memory, weak provenance, and uncontrolled autonomous execution—and organizes governance responses into seven domains covering identity-aware retrieval, policy enforcement, provenance management, memory governance, knowledge integrity monitoring, agentic execution control, and trust observability. The central argument is that durable enterprise value from AI depends not on model capability alone, but on the system surrounding the model—identity, knowledge, policy, memory, tools, human oversight, and evidence operating together as a managed control plane. This framework is directly relevant to enterprises seeking to move AI from experimentation to governed, audit-ready operational dependency.
- ResearcharXiv2026-07-03QP
Reading Between the Dots: Decoding Hidden Computation across Filler Tokens · Kaley Brauer, Claudio Mayrink Verdun, Samuel Marks
This paper investigates how large language models can perform multi-step reasoning using content-free 'filler' tokens (such as dots or counting sequences) that reveal no visible chain-of-thought in their outputs. Using four task families and two open-weights frontier models (DeepSeek V3 and Kimi K2), the authors show that hidden computation over these filler tokens is nonetheless structured and interpretable: attention patterns, logit-lens readouts, and KV-cache transplants all reveal how intermediate reasoning values emerge and are composed internally. The researchers introduce an unsupervised decoding pipeline that recovers intermediate reasoning values with 80–95% accuracy using only hidden states, without ground-truth labels or training. The findings suggest that behavioral oversight based solely on surface tokens is insufficient, but that the model's full computational trace—specifically its residual stream—can still be monitored, with direct implications for AI quality assurance and policy around model transparency.
- ResearcharXiv2026-07-03HeAd
Learning from Lost Provenance: Multiple Instance Learning for Cancer Registry Tumor Group Classification · Leonard Ruocco, Jonathan Simkin, Lovedeep Gondara et al.
This paper presents a framework for automating tumor group classification in cancer registries by using Attention-Based Multiple Instance Learning (ABMIL) to bridge the gap between patient-level operational labels and individual pathology reports. Because cancer registries produce expert labels at the patient level rather than the report level, direct supervised training is not straightforward; ABMIL recovers the implicit link between labels and reports, distilling a large, noisily-labeled corpus into a compact, high-quality per-report dataset. A classifier fine-tuned on this distilled data achieved a macro F1 of 0.83, outperforming established baselines across most tumor groups at the BC Cancer Registry. The approach reduces reliance on manual per-report annotation and large-scale computing infrastructure, offering a practical path to automating labor-intensive cancer registry coding workflows.
- ResearcharXiv2026-07-03CP
AI Systems as Digital Public Goods -- Evidence and Recommendations from a Multi-Stakeholder Assessment · Serge Stinckwich, Natalie Wong, Ally S. Nyamawe et al.
This report, commissioned by the Asian Development Bank and produced by United Nations University in partnership with the UN Office of Digital and Emergent Technologies, assesses why very few AI systems currently meet the Digital Public Goods (DPG) Standard despite major global commitments such as the Global Digital Compact. Using a structured desk review of policy, legal, and technical frameworks, key informant interviews with cross-sector experts, and a global survey, the assessment diagnoses the barriers preventing AI systems from qualifying as credible Digital Public Goods and offers recommendations for making 'AI as Digital Public Goods' an implementable pathway toward the Sustainable Development Goals. The findings are relevant to how governments, civil society, and the private sector should govern open AI models, open data, and open standards in ways that benefit society broadly.
- ResearcharXiv2026-07-03Ad
Personalized Causal Recourse: A Human-In-The-Loop Approach · Denise Tampieri, Giovanni De Toni, Paolo Giudici
This paper proposes a human-in-the-loop framework for algorithmic recourse—recommendations that help individuals overturn unfavorable automated decisions—that iteratively learns each user's personal causal structure through interactive queries and Bayesian inference. Unlike traditional approaches that rely on fixed counterfactuals or assumed causal knowledge, the system tailors interventions to individual feature relationships, aiming for more plausible and cost-effective recourse. Simulations across linear and non-linear causal models show promising results, though the authors note that capturing complex non-linear structures remains a challenge. The work is relevant to high-stakes AI decision-making contexts where individuals need actionable, personalized pathways to change automated outcomes.
- ResearcharXiv2026-07-03EQPPr
Securing Multi-Tool AI Agent Chains With Dynamic, Real-Time Compositional Policies · Chris Schneider, Kriti Faujdar, Philipp Schoenegger et al.
This paper identifies a security gap in multi-tool AI agent systems where individually permitted tools can violate organizational policies when chained together at runtime. The authors propose the Dynamic Security Control Compositor (DSCC), a two-phase system that first composes per-tool policies into a single restrictive policy before any tool executes, then tracks data sensitivity through runtime taint monitoring to catch violations that emerge from actual data use. Evaluated on 32 tools governed by 16 NIST SP 800-53-aligned policies, the system blocks 79.2% of policy pairs and 95.5% of policy triples in default clearance mode, with an alternative taint mode offering a utility-security tradeoff. The work has direct governance implications for organizations deploying multi-tool AI agents, including how chain-aware policies need to be operationalized.
- ResearcharXiv2026-07-03QPrShAd
DETECT-3B-Omni is Agnostic of Content and Demographics · Nicolas M. Müller, Aditya Tirumala Bukkapatnam, Dominik Schnieders et al.
This paper evaluates whether Resemble AI's deepfake audio detector, DETECT-3B-Omni, produces consistent results regardless of spoken content or speaker demographics. Using 10,240 audio samples from diverse US English speakers across 30 states, generated by 8 different AI voice-cloning systems, the study tests detection accuracy across groups defined by spoken content type (benign vs. malicious), speaker gender, speaker age, and speaker region. Through equivalence testing at 99% confidence, the authors find that accuracy differences between any two groups are at most 2 percentage points, demonstrating that the detector does not rely on content or demographic signals. This matters because a GDPR-compliant, trustworthy deepfake detector must base decisions on acoustic artifacts alone, and these results provide evidence that DETECT-3B-Omni meets that standard.