News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
From Celebrities to Anyone: Characterizing AI Nudification Content, Technology, and Community Dynamics on 4chan
Chi Cui, Yixin Wu, Yang Zhang
arXiv · 2026-06-25
This large-scale empirical study identifies 24,105 synthetic non-consensual sexually explicit AI-generated images and videos ('SNEACI') shared on 4chan, revealing that non-celebrity individuals now account for 55.8% of targets—up from just 4.7% in prior studies—indicating AI nudification has expanded well beyond public figures to harm people in users' personal social circles. Open-source tools dominate production, with the Stable Diffusion family responsible for 42.7% of images and Wan for 66.5% of videos, while shared fine-tuned models and accessible tutorials lower barriers to entry. A small cohort of prolific producers drives the ecosystem, with the most active individual generating 780 items, shaping community engagement, target demographics, and technical knowledge diffusion. The authors argue these findings underscore urgent needs for platform governance interventions, technical safeguards, and protections for affected individuals.
- AI policy
Research
Inherited Circuits, Learned Semantics: How Fine-Tuning Creates Evasion Vulnerabilities Invisible to Standard Evaluation
Ryan Fetterman
arXiv · 2026-06-25
This paper demonstrates that large language models fine-tuned for security classification (specifically PowerShell command detection) can pass standard held-out evaluations while becoming more vulnerable to evasion attacks introduced by the fine-tuning process itself. The authors study Foundation-Sec-8B-Instruct and its base model, finding that fine-tuning concentrates and semantically specializes an inherited late-attention classification circuit from Llama rather than building a new one, creating brittle token-level indicator rules. A three-tier evasion benchmark shows the fine-tuned model fails on behavior-preserving transformations—such as alias substitution, string construction, and case mutation—that the base model handles correctly. The authors propose a pre-deployment monitoring method using a linear probe and indicator-token sign test to identify vulnerable command families, cautioning that task-specific fine-tuning can improve accuracy metrics while silently expanding the real-world evasion surface.
- Quality assurance
- Certifications
Research
Auditing Framing-Sensitive Behavioral Instability in Large Language Models for Mental Health Interactions
Abla Bedoui, Ashley L. Greene, Mohammed Cherkaoui
arXiv · 2026-06-25
This paper investigates how large language models (LLMs) used in mental health support applications respond differently to semantically similar concerns depending on how they are contextually framed. Using controlled matched prompts across multiple instruction-tuned model families, the researchers find that framing systematically alters interpretive response tendencies, with layer-wise probing showing that framing-related information is decodable throughout transformer layers. Activation steering experiments further suggest that framing-associated internal representations can partially influence downstream behavioral outputs. The findings highlight that robustness to contextual framing is an important consideration when evaluating the consistency and trustworthiness of AI systems deployed in mental-health-oriented settings.
- Quality assurance
- Certifications
Research
RedVox: Safety and Fairness Gaps in Speech Models Across Languages
Beatrice Savoldi, Sara Papi, Wafa Aissa et al.
arXiv · 2026-06-25
RedVox introduces a multilingual safety and fairness benchmark for speech-capable AI models, covering English, French, Italian, Spanish, and German using real human voices. The study surveys state-of-the-art model releases and finds that only 8% document any multilingual safety analysis. Evaluating eight models with RedVox, the researchers find that safety vulnerabilities persist even under non-adversarial conditions, worsen in non-English languages, and are amplified when inputs are spoken rather than text-based. The paper also highlights unique privacy and sociotechnical challenges in collecting naturalistic speech data from human participants.
- Quality assurance
- AI policy
Research
A Deterministic Control Plane for LLM Coding Agents
Padmaraj Madatha
arXiv · 2026-06-25
This paper examines how LLM coding agent configuration files (rules files, agent definitions, IDE-specific markdown) are managed across 10,008 public GitHub repositories. The study finds these configurations propagate as undeclared shared components, with 10.1% of tracked paths being SHA-256 exact duplicates across independent repositories and 75.5% of clone pairs crossing organisational boundaries; configurations are rarely revised and almost never declare permission boundaries (<1% vs 33% for CI/CD workflows). To address these gaps, the authors propose Rel(AI)Build, a deterministic control plane that treats agent definitions as a managed supply chain with content addressing, audit logs, tiered permissions, and prompt drift detection. The work highlights significant quality assurance and policy risks in how AI coding agent configurations are currently governed and distributed.
- Quality assurance
- AI policy
Research
SamaVaani: Auditing and Debiasing Multilingual Clinical ASR for Indian Languages
Subham Kumar, Prakrithi Shivaprakash, Abhishek Manoharan et al.
arXiv · 2026-06-25
SamaVaani audits eight state-of-the-art automatic speech recognition (ASR) models—including IndicWhisper, WhisperLargeV3, Sarvam, GoogleS2T, Gemma3n, OmniLingual, Vaani, and Gemini—on real-world psychiatric interview data in Kannada, Hindi, and Indian English, finding substantial performance variability across models and languages, with strong results in Indian English but frequent failures on regional speech. The study identifies systematic gaps tied to speaker role and gender, raising equity concerns for clinical deployment. The authors then fine-tune the two best open-source models (Gemma3n and OmniLingual) using a proposed fairness-aware technique called SamaVaani, which simultaneously improves overall ASR accuracy and reduces demographic performance disparities. These findings matter for healthcare quality assurance and policy around equitable AI deployment in multilingual clinical settings.
- Quality assurance
- AI policy
Research
AIGP: An LLM-Based Framework for Long-Term Value Alignment in E-Commerce Pricing
Chennan Ma, Yanning Zhang, Siqi Hong et al.
arXiv · 2026-06-25
AIGP is an LLM-based pricing framework for large-scale e-commerce that addresses key shortcomings of traditional dynamic pricing: poor interpretability, inability to use unstructured information, and misalignment with long-term business goals. The system combines domain-knowledge-prompted LLMs with a Long-Term Value Estimator trained via offline reinforcement learning, using Direct Preference Optimization to align pricing decisions with objectives like GMV, ROI, and milestone achievement. In large-scale online A/B tests on Tao Factory, AIGP delivered +13.21% in GMV, +7.59% in ROI, and +8.20% in milestone achievement rate over 14 days compared to the production baseline, while also producing interpretable pricing rationales. This demonstrates that LLM-based approaches can meaningfully advance enterprise pricing strategy by bridging short-term decisions with long-term business value.
- Enterprise
Research
Do Safety Guardrails Need to Reason? LeanGuard: A Fast and Light Approach for Robust Moderation
Dongbin Na
arXiv · 2026-06-25
This paper challenges the assumption that safety guardrails for AI systems need chain-of-thought (CoT) reasoning to be effective. The authors train a lightweight 395M-parameter bidirectional encoder (LeanGuard) and a reasoning-based guard on the same data, then show that removing CoT does not hurt moderation accuracy — LeanGuard achieves an average F1 of 82.90 over public benchmarks, matching much larger reasoning-based decoders while using roughly 100x less inference compute. The label-only encoder also proves more robust under training-label noise and maintains better recall at strict false-positive rates, suggesting reasoning guards are not the safer choice either. The findings indicate that current guardrail benchmarks may not be challenging enough to justify the cost of CoT-based moderation, with practical implications for on-device deployments such as embodied robots.
- Quality assurance
Research
Can Large Language Models Reliably Code Qualitative Humanitarian Data? A Benchmark Study Against Human Expert Adjudication
Jerome Marston, Tino Kreutzer, Salomé Garnier et al.
arXiv · 2026-06-25
This benchmark study evaluates 46 large language models (LLMs) against a human Gold Standard for coding qualitative humanitarian data, using 150 synthetic transcripts and inter-rater reliability testing with Krippendorff's alpha. The authors find that multiple LLMs can match experienced human coders on deductive coding tasks, particularly when structured prompts and reasoning-enabled configurations are used, but aggregate reliability metrics alone are insufficient for deployment decisions. Models varied in their ability to recognize indirectly expressed needs, needs outside predefined categories, and protection-relevant concerns such as physical safety and discrimination. The findings indicate LLMs can expand humanitarian analytical capacity but require structured codebooks, tiered human oversight, and — for sensitive data — self-hosted open-weights models to balance scalability with data governance.
- Workforce
- Quality assurance
Research
The Inattentional Gap: Task-Conditioned Language and Vision Models Omit the Safety-Critical Signals They Can Otherwise Report
Kwan Soo Shin, In Seok Kang, Yunkyung Min et al.
arXiv · 2026-06-25
This paper identifies a phenomenon called the 'Inattentional Gap,' where language and vision models conditioned on a specific narrow task suppress reporting of co-present safety-critical signals they are otherwise capable of detecting — a machine analogue of human inattentional blindness. Across radiology and driving text scenarios and chest-radiograph vision tasks, focused task instructions suppressed reporting of off-task hazards by up to 0.92 in report rate, with explicit exclusive instructions abolishing such reporting entirely in radiology. The effect appeared across all tested models, did not diminish with scale, and persisted in reasoning models, meaning benchmark safety scores can look near-perfect while real-world safety hazards go unreported. The authors propose 'reporting-complete evaluation' — scoring what a system fails to report alongside what it is asked to find — and show that routing outputs to an independent open-ended critic can restore omitted findings.
- Quality assurance
- AI policy
Research
Adaptive Evaluation of Out-of-Band Defenses Against Prompt Injection in LLM Agents
Praneeth Narisetty, Shiva Nagendra Babu Kore, Uday Kumar Reddy Kattamanchi et al.
arXiv · 2026-06-25
This paper examines 'out-of-band' defenses against prompt injection attacks on LLM agents—security mechanisms enforced outside the model itself using classical integrity and least-privilege principles, as seen in systems like CaMeL, FIDES, Progent, RTBAS, and FORGE. The authors warn that all such defenses have only been validated on static benchmarks, the same flaw that allowed adaptive attacks to break twelve in-band defenses at over 90% success rates. As an independent test, they ran an adaptive evaluation of Progent on the AgentDojo benchmark using an open-weight model (Qwen2.5-7B), finding that Progent reduced mean attack success roughly sixfold (25.8% to 4.2%) and a hand-crafted adaptive attack did not meaningfully raise it (2.6%). The results are consistent with—but do not conclusively establish—that deterministic out-of-band enforcement is more robust against adaptive attackers than in-band detection, while highlighting that stronger white-box attacks remain untested.
- Quality assurance
- AI policy
Research
Auditing a Robotic System for the AI Act
Laura Lucaj, Felix Bok, Patrick van der Smagt
arXiv · 2026-06-25
This paper presents a framework for auditing robotic AI systems under the EU AI Act, validated on a real-world case of hospital ventilation-cleaning robots. The authors find that existing audit methodologies designed for software-based AI are insufficient for embodied reinforcement learning systems, because compliance evidence from simulation does not guarantee real-world safety. The framework addresses specific challenges such as the Sim2Real gap, policy opacity, and distributed stakeholder responsibility, while translating AI Act obligations—risk management, human oversight, and post-market monitoring—into concrete audit criteria. The work argues for context-sensitive, lifecycle-embedded auditing practices, especially for resource-constrained organizations operating under regulatory uncertainty without harmonized standards.
- Certifications
- AI policy
- Quality assurance
Research
The Impact of the Implementation of the AI Systems in Small and Medium Enterprises in Poland: Scale of Usage, Productivity, and Unperceived Sustainability
Michał Polasik, Marta Czarkowska, Wojciech Śniadkowski et al.
Sustainability · 2026-06-25
This study examines AI adoption among 112 SMEs in Poland's Kuyavian–Pomeranian region, combining survey data with manager interviews to assess organizational, economic, and sustainability impacts. Results show AI most strongly reduces workload and improves time efficiency, especially in service firms with intensive AI use, though benefits come alongside new costs from paid tools, data preparation, and governance. Adoption follows a staged path from experimentation to workflow integration, with barriers shifting from knowledge gaps early on to data quality and security issues at advanced stages. Notably, sustainability considerations such as environmental and ESG impacts remain largely unperceived by SME decision-makers, who instead frame sustainability through resilience and competitiveness.
- Enterprise
- Workforce
- AI policy
Research
When Agents Meet Electric Bus Fleet Operations: Pricing Behavior, Trade-offs, and Policy Implications in an Aggregator Framework
Jônatas Augusto Manzolli, Ali Eslami, Luis Miranda-Moreno et al.
arXiv · 2026-06-24
This paper proposes an agentic aggregator framework for managing electric bus fleet operations, combining optimization-based scheduling with supervisory AI agents that handle disturbance detection, tariff adaptation, and real-time re-optimization across charging and vehicle-to-grid (V2G) activities. A realistic depot case study finds that the framework can maintain feasible schedules and improve use of charging flexibility under various operational disruptions, but also reveals that profit-oriented agent configurations can extract value from the public transport operator at its expense. The authors conclude that deploying agentic aggregators in public-fleet contexts requires transparent coordination modes, auditable tariff-setting, and explicit value-sharing rules to prevent misaligned incentives.
- Enterprise
- AI policy
Research
Instruction Bleed: Cross-Module Interference in Prompt-Composed Agentic Systems
Ching-Yu Lin, Yifan Liu
arXiv · 2026-06-24
This paper formalizes a failure mode called compositional behavioral leakage (CBL), where editing one prompt module in an AI agent silently shifts the behavior of other modules that share the same context window, even without any direct variable or code dependency. The authors probe this on a deployed job-evaluation agent (Claude Sonnet 4.6) across 144 trials using a three-channel perturbation protocol targeting volume, content, and form of non-focal modules; only content-channel perturbations produced a detectable effect (Cohen's d = 0.63 with a bootstrap 95% CI excluding zero), though no individual recommendation flipped. While sub-threshold in standard QA terms, the authors argue this interference compounds silently across thousands of agent decisions, making it a systematic evaluation blind spot. The paper contributes an operational definition, a reusable measurement protocol, and a call for cross-module interference testing as a standard requirement in prompt-composed agent evaluation.
- Quality assurance
- Enterprise
Research
Governing Actions, Not Agents: Institutional Attestation as a Governance Model for Autonomous AI Systems
Jakob Salfeld-Nebgen
arXiv · 2026-06-24
This paper proposes a governance model for autonomous AI agents that perform high-stakes, irreversible actions such as clinical prescribing and production software deployment. Rather than monitoring an agent's internal reasoning, the model requires independently attested evidence at the point of consequential action: execution is conditional on preconditions verified by separate authoritative sources, cryptographically bound to a declared intent, and evaluated by a deterministic policy. Decisions are recorded in a tamper-evident log that supports independent re-verification. The authors present a proof-of-concept implementation illustrated with software deployment and clinical prescribing scenarios, arguing this approach mirrors how human institutions have historically governed powerful autonomous actors.
- AI policy
- Certifications
Research
Real-Time Voice AI Hears but Does Not Listen
Martijn Bartelds, Federico Bianchi, James Zou
arXiv · 2026-06-24
This paper evaluates four leading real-time voice AI systems—OpenAI's GPT Realtime 2, Google's Gemini 3.1 Flash Live, and Alibaba's Qwen3.5 Omni Plus and Omni Flash—on tasks where vocal delivery (emotion, tone, sarcasm) carries meaningful information alongside spoken words. Across three scenarios, all four systems consistently act on verbal content alone: ending calls with distressed callers who deny distress, approving wire transfers authorized in frightened voices, and accepting clearly sarcastic consent. Notably, three of the four systems can correctly identify the emotional cues when directly asked, yet ignore them during decision-making—a disconnect the authors call the 'emotional intelligence gap.' The findings suggest current voice AI should be used cautiously in high-stakes settings where tone and emotional delivery are critical, such as customer service, financial authorization, and consent verification.
- Quality assurance
- Enterprise
Research
Same Evidence, Different Answer: Auditing Order Sensitivity in Multimodal Large Language Models
Akshay Paruchuri, Sanmi Koyejo, Ehsan Adeli
arXiv · 2026-06-24
This paper audits 18 multimodal large language models (MLLMs) for order sensitivity — whether shuffling the sequence of options, evidence chunks, documents, images, or mixed-modality inputs changes model answers despite the underlying evidence remaining identical. Using a framework called Facet-Probe and a Bayesian item-response model to separate noise from bias, the authors find that none of the 18 models are order-invariant, with per-facet flip rates ranging from 24–50%, and even the best model flipping on 13.4% of trials. Prompt-level mitigations are shown to be modality-conditional and do not transfer from text to visual reasoning, suggesting architectural or training-time solutions are needed. The findings are directly relevant to AI evaluation standards, as the authors propose cross-ordering flip rate as a standard reporting metric in line with emerging AI evaluation guidelines.
- Quality assurance
- Certifications
Research
AI translation of literary texts is "fine", but readers still prefer human translations
Yves Ferstler, Adam Podoxin, Ty Brassington et al.
arXiv · 2026-06-24
This study had 15 avid readers compare machine translations (MT) generated by an LLM-based pipeline against recently published human translations (HT) for 15 novels in French, Polish, and Japanese translated into English. Readers found MT 'fine' overall but preferred human translations—especially at the chunk level (522 out of 772 comparisons)—citing ease, clarity, and immersiveness, while MT showed greater within-book quality variation. Notably, readers could not reliably identify which version was human (only 17 of 30 guessed correctly) and tended to prefer whichever version they believed to be human. Automatic metrics, including LLM-as-a-judge approaches, failed to align with reader preferences and systematically favored MT, highlighting a gap between current evaluation tools and actual reader experience.
- Quality assurance
Research
Detect, Unlearn, Restore: Defending Text Summarization Models Against Data Poisoning
Poojitha Thota, Shirin Nilizadeh
arXiv · 2026-06-24
This paper addresses the threat of data poisoning during fine-tuning of large language models used for abstractive text summarization, where adversaries can manipulate small task-specific datasets to produce biased or harmful summaries without triggering standard evaluation metrics. The authors propose a unified post-hoc defense framework that detects poisoned training examples via influence-function analysis and behavioral auditing, then applies gradient-ascent unlearning to remediate identified poisoning. Tested across nine architectures and six benchmark datasets, the framework achieves 85–92% detection precision and restores up to 96% of original model behavior with less than 0.6% ROUGE degradation. These findings matter for quality assurance and enterprise deployment of AI summarization systems, demonstrating that fine-tuning-stage supply chain attacks are both detectable and recoverable without full retraining.
- Quality assurance
- Enterprise
Research
Can Trustless Agents Be Trusted? An Empirical Study of the ERC-8004 Decentralized AI Agent Ecosystem
Xihan Xiong, Zelin Li, Wei Wei et al.
arXiv · 2026-06-24
This paper presents the first empirical study of ERC-8004, a permissionless trust protocol for autonomous AI agent economies deployed on Ethereum, BNB Smart Chain, and Base. The authors find that most agent registrations are inactive placeholders—only 3–15% expose valid service endpoints—and that the reputation registry is deeply flawed: feedback values are incommensurable, rarely grounded in verifiable interactions, and easily manipulated, with 59–91% of reviewers exhibiting coordinated Sybil behavior. After filtering Sybil-flagged feedback, between 16% and 87% of rated agents are left with no valid reputation signal depending on the chain. The findings carry direct implications for the design of trustworthy AI agent markets and offer concrete protocol-design recommendations for future ERC-8004 revisions.
- Enterprise
- AI policy
Research
Financing Artificial Intelligence Infrastructure: Mapping AI Infrastructure Investment and Compute Governance Across Africa
Kai-Hsin Hung, Sumaya Nur Adan, Krupa Suchak et al.
arXiv · 2026-06-24
This paper maps AI infrastructure investment across Africa by systematically analyzing 46 publicly announced projects totalling USD $12.7 billion between 2019 and 2025. Using a value chain framework, the authors find that investment is highly concentrated geographically—clustering in South Africa, Kenya, Nigeria, and Egypt—and structurally dominated by global data center operators, hyperscale technology firms, and development finance institutions. They introduce the concept of 'asymmetrical interdependence' to describe how capital and physical infrastructure account for 73% of total funding while control of the compute layer remains concentrated among a small number of global technology firms. The paper argues that meaningful AI compute governance must account for capital flows, ownership, and control, not just geographic access, because infrastructure presence alone is insufficient for equitable governance capacity.
- AI policy
Research
How Large Language Models Source Brand Reputation Across Languages and Markets
Dmitrij Zatuchin
arXiv · 2026-06-24
This study examines where large language models source their brand-related information by analyzing 167,551 URL-grounded citations across 128 brands, 12 home markets, and 13 languages. Key findings show that 85.7% of citations point to third-party sites rather than brand-owned properties, and the source distribution is highly concentrated, with 80% of citations coming from roughly 18% of domains following a Zipf law (alpha=0.86, R²=0.983). Wikipedia dominates as the most-cited domain in 11 of 12 languages, though market-specific patterns emerge — for example, YouTube is the top source for Polish national brands and HR/careers portals supply twice as many citations as Polish Wikipedia. These findings matter for enterprises seeking to manage AI-driven brand reputation, as they reveal that LLM outputs are shaped primarily by a narrow set of third-party sources rather than brand-controlled content.
- Enterprise
Research
MedGuards: Multi-Agent System for Reliable Medical Error Detection and Correction
Congbo Ma, Hu Wang, Yichun Zhang et al.
arXiv · 2026-06-24
MedGuards is a multi-agent framework designed to improve medical error detection and correction in LLM-generated clinical text, addressing patient safety risks that arise when existing automated and heuristic-based methods fail to generalize across unseen datasets. The system assigns specialized agents to separately detect, localize, and correct errors, with a confidence-guided arbitration mechanism that resolves disagreements using reasoning traces and confidence scores — all without retraining the underlying LLMs. The authors also introduce a new evaluation metric, the Keyword-Prioritized Correction Score (KPCS), which checks whether critical keywords from reference text are correctly reproduced, offering a more thorough assessment than conventional metrics. Experiments on four multilingual clinical note datasets show significant improvements across multiple metrics and models, supporting safer LLM deployment in real-world healthcare.
- Quality assurance
Research
Probabilistic Agents in Deterministic Audits: Evaluating Multi-Agent Systems for Automated Audits Based on the German IT-Grundschutz
Lea Roxanne Muth, Marian Margraf
arXiv · 2026-06-24
This paper implements and evaluates a Multi-Agent System (MAS) combined with Hybrid Retrieval Augmented Generation (HybridRAG) to partially automate German IT-Grundschutz (IT-GS) certification, which is required for NIS-2 Directive compliance. Two novel technical contributions are introduced: a Hypothesis-Verification Loop to reduce hallucinations and a Decoupled Reasoning Pipeline to separate semantic extraction from deterministic protection need inheritance. Evaluated against the BSI's 'RecPlast GmbH' expert-generated case study using Precision, Recall, and F1-scores, the system performs well on semantic tasks like Structural Analysis and Modeling, but struggles in logical reasoning phases such as Protection Needs Assessment and IT-GS Check where the probabilistic nature of LLMs conflicts with the deterministic rigor required by the standard. The findings highlight both the promise and current limits of AI-driven automation for scalable, resource-efficient IT security certification.
- Certifications
- Quality assurance