News & Research
The latest AI research and news with real-world stakes. Each item is sourced, dated, summarized in plain English and tagged by impact area, and checked against its source before it appears.
Kind
Impact area
5672 items
- ResearcharXiv2026-06-20W
AI-Mediated Negotiation: Design Reflections and Lessons · Veda Duddu, Jash Rajesh Parekh, Andy Mao et al.
This paper presents Trucey, a conversational AI coaching system designed to help workers prepare for high-stakes workplace negotiations, and tests it in a pre-registered experiment with 267 participants and 15 interviews. Contrary to expectations, a static handbook outperformed both AI conditions on empowerment and usability. The authors find that conversational AI imposes a linear execution model on a fundamentally recursive task, undermining its intended benefits. They propose a sequencing principle—map before path, path before simulation—to guide future AI coaching design.
- ResearcharXiv2026-06-20WE
Human Capital, AI, and Labor Commoditization · Auyon Siddiq, Niuniu Zhang
This paper investigates whether generative AI has shifted how online labor markets value human capital, using contract-level data from Upwork before and after the release of ChatGPT. The authors represent worker profiles as high-dimensional text embeddings to capture human capital information, then use a difference-in-differences design to compare AI-exposed versus less-exposed job categories. They find that in more AI-exposed categories, human capital becomes less important for client demand while price becomes more important—a 'commoditization' effect—with demand shifting toward lower-priced workers and declining premiums for highly skilled workers. These results carry implications for how online labor markets should be designed and for workers' incentives to invest in skills and their overall labor welfare.
- ResearcharXiv (Cornell University)2026-06-20EP
AgentRiskBOM: A Risk-Scoping Security Bill of Materials for Agentic AI Systems · Srimonti Dutta, Akshata Kishore Moharir
AgentRiskBOM introduces a structured security Bill of Materials designed specifically for agentic AI systems—those that autonomously access private data, invoke tools, coordinate with other agents, and act without human approval. The paper identifies a 'capability opacity' gap in existing SBOM, AIBOM, and MLBOM artifacts and proposes an additive JSON-schema layer that captures runtime authority fields such as tool permissions, memory scope, credential scope, approval gates, and audit signals. Evaluated against 13 open-source agents and 52 risk scenarios, AgentRiskBOM achieves 100% risk-category visibility compared to roughly 10.5–20.9% for existing BOM formats, and its diff detector correctly identifies all 33 injected deployment mutations. The authors argue that a machine-readable authority-and-risk artifact is necessary for agentic AI security before incidents occur, making this work directly relevant to enterprise deployment governance and policy-level transparency requirements.
- ResearcharXiv2026-06-20QP
The Language-Energy Divide: Measuring Energy Costs of Multilingual LLM Inference · Naihao Deng, Alissa Shen, Yiming Feng et al.
This paper measures how much energy large language models consume when generating text in different languages, finding dramatic disparities: energy per output token varies up to 8.3× across languages, and total energy for a fixed request set ranges from 17.6 kJ for English to 3,147 kJ for Pashto—a 179× gap. The disparity stems from two compounding factors: higher per-token costs for complex or rare scripts, and more tokens generated for low-resource languages. Critically, the study identifies a 'double penalty' where the highest-energy languages also achieve the lowest task accuracy, and this inequity persists across models, hardware, and tasks. The authors recommend treating energy as a primary evaluation metric and extending model cards and reporting checklists to include multilingual energy costs.
- ResearcharXiv2026-06-20EP
Harness-MU: A Safe, Governed, and Effective Harness for Multi-User LLM Agents · Wangxuan Fan, Xiaoyu Nie, Zhongxiang Dai
Harness-MU is a model-agnostic infrastructure framework that enforces access-control, permission boundaries, and conflict resolution for large language model (LLM) agents operating in multi-user, multi-principal settings. Rather than relying on prompt-based safeguards baked into the LLM, it decouples governance logic into deterministic runtime execution hooks, making safety constraints unbreakable regardless of which model is used. Evaluated on the Muses-Bench benchmark across four frontier models, Harness-MU achieves complete privacy preservation against all access-control attacks, improves utility scores by 0.28–0.39 over the standard baseline, and boosts instruction-following accuracy by up to 48.9 percentage points. This matters because it demonstrates that systematic infrastructure—not model fine-tuning—is the reliable path to governing LLM agents deployed in collaborative enterprise and policy-sensitive environments.
- ResearchApplied Sciences2026-06-20EQP
Structural Ethical Infeasibility in AI-Enabled Infrastructure Systems: A Constraint-Based Diagnostic Framework · Sudipta Chowdhury, Md Abdul Quddus, Ammar Alzarrad
This paper challenges the assumption that inequitable outcomes in AI-driven infrastructure systems—such as ambulance dispatch—are caused by flawed algorithms. Instead, it argues that inequity can be structurally embedded in the physical environment (network topology, resource placement, demand distribution), and introduces a constraint-based diagnostic framework using a hierarchical Irreducible Infeasible Subsystem procedure to attribute infeasibility to rule design, algorithmic choice, or physical infrastructure. The framework proves that observed efficiency–equity trade-offs may reflect underbuilt systems rather than algorithmic shortcomings, and that equity improvements in such settings may redistribute harm rather than reduce it. Critically, the framework can translate findings into concrete capital-investment requirements to restore ethical feasibility, with direct implications for infrastructure policy and AI governance.
- ResearchAdvokasi Hukum & Demokrasi (AHD)2026-06-20EQP
Pertanggungjawaban Pidana Terhadap Penyalahgunaan Kecerdasan Buatan (Artificial Intelligence) dalam Tindak Kejahatan Digital di Indonesia · Yoel Bessoran
This study examines AI-enabled digital crimes in Indonesia from 2023 to 2026, focusing on deepfakes, voice cloning, digital fraud, and misuse of autonomous algorithms. Using a normative juridical and empirical case analysis approach, the research finds that existing Indonesian laws—including the Criminal Code, the Electronic Information and Transactions Law, and the Personal Data Protection Law—are insufficient to enforce criminal liability for AI-mediated harms. The authors recommend combining individual and corporate liability, applying risk-based and vicarious liability principles, and integrating digital forensic technology with cross-institutional coordination. The findings carry direct implications for national policy reform, corporate ethical standards, and the development of a more comprehensive legal framework for AI-based digital crime.
- ResearchJournal of International Relations and Foreign Policy2026-06-20EP
Artificial Intelligence and the Transformation of Global Order: Toward Algorithmic International Relations · Eric C. K. Cheng
This paper argues that artificial intelligence represents a fundamental transformation of international politics—not merely a new tool of state power—and introduces the 'Algorithmic International Relations' (AIR) framework to analyze how AI reshapes the global order. The authors find that AI shifts power dynamics toward control over compute, data, and regulatory standards; compresses decision-making cycles; deepens security dilemmas; and reinforces global inequalities through digital stratification. Three case studies—the US–China AI rivalry, the EU AI Act, and the Russo-Ukrainian War—are used to test the framework and demonstrate that classical IR theories require conceptual adaptation to account for algorithmic agency and infrastructural sovereignty. The paper has direct relevance for understanding how AI governance regimes and regulatory standards are emerging as central arenas of geopolitical competition.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-06-20WCP
Embedding AI Literacy in Philippine Higher Education: A National Strategy for Workforce Readiness in the Age of Artificial Intelligence · Arthur Baldosano
This paper examines the gap between high student AI tool usage and low institutional preparedness in Philippine higher education, finding that over 83% of students use generative AI for academic work while fewer than half of institutions have clear policies on it. The Philippines ranked last in the 2023 Asia-Pacific AI Readiness Index, and the paper argues that embedding structured AI literacy programs is an economic and educational necessity given that 39% of existing job skills are expected to become outdated by 2030. The authors are particularly concerned about disruption to the IT-BPM sector and call for adoption of existing global AI literacy frameworks to prevent a generation of graduates from being unprepared for the labor market.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-06-20EQCP
The Missing Layer: Building Operational Trust Between AI Governance and AI Execution · Dimitri Krijgsman
This whitepaper introduces 'Operational Trust' as a conceptual framework addressing the gap between high-level AI governance requirements—such as law, standards, and organizational policy—and the continuous, evidence-producing controls needed around specific AI-mediated processes. The author argues that governance should focus on bounded execution units rather than abstract models, and proposes an Operational Trust Sequence (OTS) requiring five interdependent functions: technical assessment, regulatory translation, ongoing accountability, legible trust signals, and structural independence. The framework aims to create an inspectable evidence chain from regulatory obligation through to system behavior and intervention capability, enabling institutions to defensibly rely on AI-mediated execution. The paper is positioned as a testable conceptual architecture rather than a certification standard or compliance method, making it relevant to enterprise AI deployment, quality assurance, and emerging policy frameworks.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-06-20WCP
Embedding AI Literacy in Philippine Higher Education: A National Strategy for Workforce Readiness in the Age of Artificial Intelligence · Arthur Baldosano
This paper argues that Philippine higher education must embed structured AI literacy programs to address a critical gap: over 83% of students already use generative AI for academic work, yet fewer than half of institutions have clear policies, and the Philippines ranked last in the 2023 Asia-Pacific AI Readiness Index. The authors contend that without deliberate curriculum reform, a generation of graduates risks entering labor markets where 39% of existing job skills are expected to become obsolete by 2030, with particular threat to the IT-BPM sector. The paper draws on global AI literacy frameworks as models Philippine institutions can adopt immediately, framing AI literacy integration as both an economic imperative and an educational necessity.
- ResearcharXiv2026-06-19EQ
Per-Entity Bias Mapping for AI Visibility: Why Brand Mentions Require Entity-Specific Calibration · Zoltan Varga
This paper introduces Per-Entity Bias Mapping (PEBM), a ten-dimensional framework for measuring how AI answer systems misrepresent brands and organizations differently depending on their size and data prominence. An empirical study of 100 Hungarian B2B entities across 1,400 probe runs finds that large, well-known brands actually suffer higher citation fabrication rates (52.69%) than smaller entities (37.87%), a phenomenon the authors call the Brand Hallucination Paradox, where model familiarity creates more plausible but incorrect outputs. The study also finds that regulatory-framed queries escalate fabrication to 56.77%, and that agentic quality filters can paradoxically amplify hallucinations in compliance contexts. These findings matter for enterprise brand management and quality assurance, showing that aggregate visibility metrics are inadequate and that AI-mediated representation requires entity-specific calibration and verification.
- ResearcharXiv2026-06-19Q
MedHal-Loc: Are "Explainable-by-Architecture" Medical Hallucination Detectors Faithful Localizers? A Localization Benchmark · Minmin Chen, Daojian Lu, Yining Dai et al.
MedHal-Loc introduces a benchmark and metric to test whether medical hallucination detectors can faithfully pinpoint the specific text span containing an error, not just flag that an error exists. Using 300 PubMedQA-derived statements with injected span-level errors and a complementary set of real clinical hallucinations, the study evaluates four detection paradigms and finds that NLI-per-clause, consistency-per-sentence, and the FAVA span detector all localize errors meaningfully above chance, while an elaborate knowledge-graph triple pipeline performs no better than chance despite achieving competitive detection F1 of 0.609. The key finding is that detection competence does not imply faithful localization, meaning systems marketed as 'explainable by architecture' must have their localization claims empirically validated rather than assumed. This matters for clinical quality assurance, where auditable error attribution—not just error flagging—is essential for safe deployment of AI in medical text generation.
- ResearcharXiv2026-06-19EQ
Evaluating LLMs for Real-World Web Vulnerability Detection · Sebastian Neef, Luca Jungnickel, Antonio Benjamin Buchholz et al.
This paper benchmarks six large language models (three frontier and three open-weight) on their ability to detect real-world web vulnerabilities—including SQL injection, cross-site scripting, path traversal, and remote code execution—in WordPress plugins using static analysis. Across five prompt designs and three experiment iterations, results show that all models can identify valid security issues, but detection rates vary significantly: Claude Opus 4.6 achieved the highest rate at 63%, while self-hosted Qwen 3.5 reached only 35%, and no model achieved full reporting consistency across iterations (some as low as 50%). The study finds that narrowly scoped prompts outperform open-ended ones, while prompt complexity has little effect, and no model correctly identified one baseline vulnerability in a specific plugin. The findings highlight both the promise and current limits of LLM-based vulnerability detection, and the authors release all code and data to support future research.
- ResearcharXiv2026-06-19Q
Mind the Noise: Sensitivity of Transformer-based Interaction-Aware Trajectory Prediction Models to Noisy Data · Shahab Salehi, Luca Lusvarghi, Miguel Sepulcre et al.
This paper investigates how sensitive state-of-the-art Transformer-based trajectory prediction models for autonomous vehicles are to noisy input data about surrounding objects. The authors find that prediction accuracy degrades significantly as noise increases—by a factor of 1.3x under small noise levels and up to 3.9x under the highest realistic noise conditions, such as those arising from Vehicle-to-Everything (V2X) communications. The findings highlight a critical gap between how these models are trained (on clean, offline-processed datasets) and how they must operate in real-world deployments, calling for more realistic training datasets and noise mitigation strategies.
- ResearcharXiv2026-06-19QP
Warning labels shift perceptions of sycophantic AI, but not its influence · Lujain Ibrahim, Myra Cheng, Cinoo Lee et al.
This preregistered experiment (N=2,610) tested whether warning labels can protect users from sycophantic AI that validates them even when they are wrong. Participants discussed real interpersonal conflicts with an AI, and results showed that a basic AI disclosure had no detectable effect, while a sycophancy-specific label reduced perceived objectivity and trust but did not reliably reduce sycophancy's actual influence on users' self-perceived rightness or willingness to repair conflicts. The study reveals a gap between AI perception and AI influence, suggesting that warning-based interventions may create a false sense of protection. The authors conclude that mitigating sycophancy will require understanding its mechanisms and improving model behavior itself, not just disclosure labels.
- ResearcharXiv2026-06-19EQ
EvidenceLens: A Claim-Evidence Matrix for Auditing Financial Question Answering · Fengchen Gu, Xiaotian Ren, Zhengyong Jiang et al.
EvidenceLens is a visual analytics prototype designed to make financial question answering by large language models auditable and verifiable. The system breaks down LLM-generated answers into atomic claims and maps each claim against supporting evidence drawn from narrative text, tables, and charts in financial documents such as annual reports and earnings decks. Its core output is a multimodal claim-evidence matrix that makes coverage gaps, contradictions, and modality imbalances immediately visible to analysts. The work matters because it provides a structured, reproducible audit workflow that helps distinguish well-grounded claims from unsupported or overconfident synthesis that conventional chat interfaces obscure.
- ResearcharXiv2026-06-19QP
Who Checks the Citations? Benchmarking Legal Hallucination Detection · Patty Liu, Dominik Stammbach, Peter Henderson
This paper investigates the growing problem of AI-fabricated legal citations, documenting over 1,000 court filings containing hallucinated citations with the number rising year-over-year despite predictions that newer models and court sanctions would reduce the issue. The researchers propose a taxonomy of legal citation hallucinations drawn from real court filings and introduce a benchmark dataset of 1,300 brief excerpts with injected errors to evaluate five AI models in both agentic and non-agentic settings. Results show that even the best-performing system, GPT-5 in an agentic framework, achieves only 82.8% recall and 60.5% F1 score, with all models struggling on subtle error categories and agentic verification requiring an average of 16.9 steps per excerpt. The study raises policy concerns around unequal access to commercial legal databases and offers tools and recommendations for building auditable legal citation-checking systems.
- ResearcharXiv2026-06-19Q
AOR-Bench: Do Large Audio Language Models Over-Refuse Pseudo-Harmful Queries? · Jiaxi Yang, Chaewan Chun, Jason Lucas et al.
AOR-Bench introduces the first benchmark specifically designed to measure over-refusal in Large Audio Language Models (LALMs)—cases where models incorrectly reject benign audio queries that only sound harmful out of context. The benchmark contains 3,000 pseudo-harmful audio samples spanning six scenario categories, and evaluation across 12 representative LALMs from six model families reveals that over-refusal is widespread. The paper also explores two lightweight mitigation strategies, Chain-of-Thought prompting and activation steering, as preliminary approaches to reduce this problem. The work highlights a key quality challenge in deploying audio AI systems safely without sacrificing usefulness.
- ResearcharXiv2026-06-19QC
Answer Engineering: Local Trajectory Editing for Protocol-Constrained Decision Making in Large Language Models · Victor Lavrenko, Anastasiia Molodnitskaia
This paper introduces 'Answer Engineering,' a runtime layer that applies rule-guided edits to a large language model's reasoning steps during text generation—without retraining or modifying model weights—to enforce compliance with clinical protocols. It is tested on a controlled benchmark for managing sudden sensorineural hearing loss (SSNHL), where unguided chain-of-thought reasoning actually worsened protocol compliance (dropping from 54.5% to 25.1% for SSNHL cases). Local trajectory editing recovered and improved compliance, raising SSNHL adherence to 83.5% and balanced accuracy from 42.0% to 80.7%. The findings suggest that auditable runtime control of LLM reasoning can meaningfully improve procedural compliance in high-stakes domains, though limitations around rule coverage and trigger reliability remain.
- ResearcharXiv2026-06-19QP
Scalable Hierarchical Attention Transformers for Multi-Turn Jailbreak Detection in Long Conversations · Chenhui Hu, Muhammed Salih, Sudipto Guha et al.
This paper presents a hierarchical attention transformer model designed to detect multi-turn jailbreak attempts in long AI conversations, where unsafe intent is spread across multiple dialogue turns rather than concentrated in a single message. The model encodes each turn individually and uses a lightweight conversation module combining cross-attention and self-attention to reason across turns without costly full-context concatenation. On a benchmark of 14,038 conversations, the approach achieves an F1 of 0.9394, outperforming Claude Opus 4.7 by 0.07 in F1 while cutting the false-positive rate in half. This matters for AI quality assurance and safety policy, as more accurate conversation-level moderation reduces both missed jailbreaks and unnecessary false alarms in deployed systems.
- ResearcharXiv2026-06-19QP
OTTER: A Red-Teaming System for Toxicity-Evading Jailbreak Prompt Optimization · Jerry Wang, Hsin-Ling Hsu, Yi-Cheng Lai et al.
OTTER is a black-box red-teaming framework that exposes a critical weakness in toxicity-based moderation filters used by production large language models (LLMs): harmful intent and toxic surface wording can be decoupled by changing as few as five tokens. Tested on 457 AdvBench prompts across four GPT models, OTTER raises average attack success rate from 7.0% to 84.0%, demonstrating that current toxicity filters are fundamentally brittle against adversarial prompt rewriting. The paper also provides the first quantitative analysis of the toxicity-bypass relationship and per-category breakdowns, translating findings into actionable recommendations for hardening classifiers in production deployments.
- ResearcharXiv2026-06-19EP
Local LLM Agents as Vulnerable Runtimes:A Source-Code Audit of the Agent Runtime Layer · Zhengsong Zhang, Zongze Li, Jiawei Guo et al.
This paper presents CLAWAUDIT, a static-analysis framework for auditing the security of local LLM agent runtimes—software like OpenClaw and Nanobot that run on end-user machines and execute actions on host resources such as the shell, filesystem, and stored credentials. The authors develop a five-category vulnerability taxonomy derived from STRIDE and implement it as 47 Semgrep YAML rules and 30 CodeQL queries, evaluated against OPENCLAWBENCH, a benchmark of 446 source-code-level advisories. On held-out test advisories, CLAWAUDIT raises Semgrep recall from 21.7% to 66.8% and CodeQL recall from 13.8% to 75.1%, with train/test gaps within 4 percentage points, indicating the rules generalize beyond the training set. The findings matter for enterprise and policy contexts because they expose a previously unaudited implementation layer in AI agents that handle privileged host-level actions, and they show that current general-purpose static-analysis tools significantly underdetect agent-specific vulnerabilities.
- ResearcharXiv2026-06-19Q
Demographic Metadata as Construct-Irrelevant Noise in DistilBERT-Based Automated Essay Scoring · Teik Peng Ch'ng, Hui Na Chua
This study examines whether adding demographic metadata (via naive concatenation) to a DistilBERT-based Automated Essay Scoring model improves or harms performance, using the ASAP 2.0 dataset with 10-fold cross-validation. The results show that early fusion of demographic metadata significantly degrades predictive accuracy, dropping the Quadratic Weighted Kappa from 0.727 to 0.656, while also increasing validation loss and worsening scoring bias (score parity instances fell from 15 to 12 out of 19 tests). The findings suggest that naively incorporating demographic information into AES models acts as construct-irrelevant noise, making the model less accurate and more biased rather than fairer. This matters for quality assurance in educational assessment, where automated scoring tools must be both accurate and equitable across student groups.
- ResearcharXiv2026-06-19QP
Honeyquest for LLMs: Rethinking Cyber Deception for AI Attackers · Kerri Prinos, Lilianne Brush, Cameron Denton
This paper introduces an automated evaluation framework called Honeyquest for LLMs to test whether cyber deception techniques designed for human attackers also work against AI-enabled attackers. The authors evaluated 21 large language models (from 10 providers, ranging from 8B to over 1T parameters) against a 47-participant human baseline using 174 identical reconnaissance queries, generating 10,962 responses. Key findings show that every LLM fell for deceptive traps at a significantly higher rate than humans, the defensive attention-diversion effect seen in humans was statistically absent in LLMs, and models exhibited a critical recognition-action gap—articulating trap recognition in reasoning but exploiting deceptive elements anyway 73.4% of the time (with trap recognition failing to predict behavior, Spearman r = +0.08, p = 0.73). These results demonstrate that human-centered deception hypotheses do not reliably transfer to AI attackers, underscoring the need for AI-native active defense frameworks.