News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
5608 items
Research
Empirical Grounding Improves the Realism of LLM Agents Simulating Human Behavior During Disruptions
Chen Xia, Zexi Kuang, Yuqing Hu
arXiv · 2026-07-19
This paper tests whether embedding real demographic and behavioral data into large language model (LLM) agents improves how realistically those agents simulate human behavior during disasters. The researchers built an empirically grounded framework using U.S. Census demographic profiles, national time-use survey data, and urban spatial context, then validated it against an independent household survey collected during the July 2024 Philadelphia heatwave. The grounded model dramatically outperformed an ungrounded baseline, raising mean correlation with observed activity profiles from 0.349 to 0.836 under heatwave conditions and capturing 46.4% of observed behavioral response amplitude versus 20.6% for the baseline. The findings matter for disaster and infrastructure-disruption planning, showing that empirical grounding is a concrete method for making LLM-based population simulations more statistically credible while also exposing remaining gaps in modeling human adaptation.
- AI policy
- Enterprise
Research
Abliteration Is Not a Scalpel: Off-Target Effects of Refusal Removal on Decision Disposition Across Model Families
Aleksander Fafuła
arXiv · 2026-07-19
This paper investigates 'abliteration,' a technique that removes refusal behavior from open-weight AI models by deleting a refusal direction from model weights, and finds it produces unintended side effects on decision-making beyond just removing content restrictions. Using 21,600 investment decisions on Warsaw Stock Exchange equities as a refusal-free probe task, the authors show that abliterated versions of two large language model families (Gemma-4-26B and Qwen3-30B) become systematically more optimistic, produce longer justifications, and use fewer uncertainty expressions compared to their base counterparts. Critically, the effect on expressed confidence reverses direction between model families, demonstrating that the same weight surgery can produce opposite behavioral shifts depending on the model. The findings warn that deploying 'uncensored' community-modified models as autonomous agents means deploying a measurably different decision-maker, not simply the base model with safety filters removed.
- Enterprise
- Quality assurance
- AI policy
Research
The Librarian Who Refused to Code: Model-Dependent Identity Enactment in LLM Code Generation
Shayell Aharon Salomon, Noam Israel, Ido Safruti et al.
arXiv · 2026-07-19
This study evaluates how biographical personas in system prompts affect large language model (LLM) code generation across controlled, pre-registered conditions using 480 completions from two frontier models (Claude Opus and GPT-5.5). Results show that persona effects are strongly model-dependent: on Claude Opus, a minimalist engineer persona reduced output length by 30% without improving correctness, while a research-librarian persona triggered in-character refusals in 55 of 60 responses and 12 genuine no-code outputs, dropping mean correctness from 0.92 to 0.67. GPT-5.5 showed neither behavior. The findings suggest that personas function as model-dependent behavioral-policy biases rather than universal quality improvements, with significant implications for how prompt engineering is evaluated and deployed in software development workflows.
- Enterprise
- Quality assurance
Research
Team DACTYL at PAN 2026: Bayesian Data Mixing and Empirical X-risk Minimization for AI-text Detection
Shantanu Thorat
arXiv · 2026-07-19
This paper presents Team DACTYL's approach to detecting AI-generated text at the PAN 2026 competition, tackling the common problem that AI-text classifiers perform well on familiar data but poorly on out-of-distribution (OOD) text. The team addresses this by using BERT-tiny models with Bayesian classification heads to curate a consolidated training set from three datasets, then training DeBERTa-V3-large and ModernBERT-large classifiers via empirical X-risk minimization, with an additional MCGrad model for calibration. The MCGrad model achieves a mean score of 0.974 across five metrics (AUROC, F1, C@1, Brier score, and F0.5u) on the PAN 2026 test set, ranking second on the leaderboard, demonstrating that careful dataset curation can significantly improve OOD generalization. These findings are relevant to quality assurance and policy contexts where reliable, generalizable AI-text detection is needed to authenticate or flag machine-generated content.
- Quality assurance
- AI policy
Research
Auditing Differential Visibility of Political Content on TikTok
Hazem Ibrahim, Tewoflos Girmay
arXiv · 2026-07-19
This study audits TikTok's alleged political content shadow-banning by analyzing over 556,000 hourly view observations across 2,753 videos from 67 accounts spanning pro and anti positions on U.S. immigration, Trump coverage, and Israel/Palestine. While pooled video-hour data appeared to show large, statistically significant reach gaps, account-level analysis—the appropriate independent unit—found no evidence of moderate-to-large suppression on any topic after multiple-comparison correction. The apparent gap was explained by methodological artifacts: pseudoreplication from autocorrelated video-hours treated as independent observations, and confounding factors such as account size and language differences. The study concludes that what has been interpreted as shadow-banning is better explained by audience engagement patterns, and offers guidance on what credible platform visibility audits require.
- AI policy
- Enterprise
Research
Agentic ERP: Multi-Agent Large Language Model Architecture for Autonomous Enterprise Resource Planning
Zhihao Liu, Tianyu Wang, Xi Vincent Wang et al.
arXiv · 2026-07-19
Agentic ERP proposes a multi-agent large language model architecture that transforms Enterprise Resource Planning systems from passive transaction recorders into active decision-makers capable of executing end-to-end business workflows autonomously. The architecture uses role-aligned LLM agents, a graph-based Planner–Executor–Reflector–Responder orchestrator, and a risk-tiered human-in-the-loop harness to handle cross-functional operational decisions that classical rule-based automation cannot manage. Evaluation across scenario-based tasks, a comparison of six orchestration paradigms, and a 365-day simulation shows the proposed system significantly outperforms rule-based baselines—sustaining a simulated year of operation with zero stockouts while the rule-based baseline accumulated hundreds under the same demand conditions. This work provides a reference architecture and evaluation protocol for autonomous ERP operation, with direct implications for how enterprises automate complex operational decision-making with human oversight.
- Enterprise
- Workforce
- Quality assurance
Research
DRNOISE: Benchmarking Deep Research Agents in Misleading Evidence Environments
Jun Nie, Zhiqin Yang, Zhenheng Tang et al.
arXiv · 2026-07-19
DRNOISE introduces a 100-task benchmark designed to test how well deep research agents handle misleading evidence on the open web. Each task includes a gold answer supported by indirect evidence chains, plus one plausible but false document offering a conflicting shortcut answer. Across agents that perform well under clean conditions, adding this single misleading document causes accuracy drops of 66–88 percentage points, with 'verification inertia'—stopping before fully reconciling evidence—identified as the dominant failure mode. The findings highlight that reliable AI-driven research requires active reconciliation of direct claims against record-level evidence, not just retrieval and citation.
- Quality assurance
- Enterprise
Research
Lookahead Branching for Neural Network Verification
Liam Davis, Duo Zhou, Huan Zhang et al.
arXiv · 2026-07-19
This paper investigates lookahead branching strategies for neural network verification using branch-and-bound methods. The authors present a general framework for integrating lookahead into any branch-and-bound verifier, showing that the existing FSB heuristic is a special case of this approach. By also generating additional lemmas during lookahead, the method improves branching decisions and accelerates verification. Experiments with two representative verifiers (Marabou and α-β-CROWN) show consistent speedups and up to 57% more solved verification instances, which matters for ensuring the reliability and safety of neural network systems.
- Quality assurance
- Certifications
Research
Lossless but Not Free: An Empirical Anatomy of Speculative Decoding on Consumer Hardware
Param Chordiya
arXiv · 2026-07-19
This paper presents an empirical study of speculative decoding—a technique that uses a small draft model to propose tokens in bulk, which a larger target model then verifies in one batched pass, preserving the target's output distribution exactly. Tested across five draft/target configurations on a consumer Apple-silicon laptop, the best setup achieves a 1.61× wall-clock speedup, but three of five configurations actually slow things down due to hardware-specific failures such as a quantized Metal backend executing verification serially rather than in parallel. The study rigorously confirms distribution equivalence between standard and speculative decoding via statistical testing, and highlights that the technique's benefits depend critically on genuine batch-parallelism during verification and a real latency gap between draft and target models. The findings are relevant to enterprise and workforce contexts where deploying large language models efficiently on consumer hardware is a practical concern.
- Enterprise
- Workforce
Research
Between Safe Boundaries: Exploiting Temporal Consistency for Jailbreaking Text-To-Video Generation Models
Xingkai Peng, Jun Jiang, Jiayang Liu et al.
arXiv · 2026-07-19
This paper presents BSB, a jailbreak framework targeting text-to-video (T2V) AI generation models by exploiting temporal consistency—the way video frames must flow coherently over time. Instead of adapting text-to-image attack methods, BSB encodes harmful content as a transition between two individually innocuous 'boundary states,' causing unsafe intermediate frames to emerge during video generation. The framework uses Monte Carlo Tree Search in a textual proxy space to efficiently explore attack candidates without excessive video queries, and is evaluated against commercial models including Veo 3.1, Sora 2, Seedance, and Kling v1, achieving an average 18.6% relative gain in attack success rate over the strongest existing baseline. The findings highlight temporal consistency as a critical and underexplored vulnerability in T2V systems, with significant implications for AI safety policy and quality assurance in deployed generative video models.
- Quality assurance
- AI policy
Research
Safety That Does Not Transfer: Cross-Lingual Clinical Correctness Drift in Deployable Medical Language Models
Anthonio Oladimeji Gabriel, Dimeji Olawuyi, Toba Ajayi et al.
arXiv · 2026-07-19
This study investigates whether clinical safety established in English for large language models transfers to Hausa, focusing on locally deployable small models (4–9 billion parameters) used in low-resource health settings in northern Nigeria. Matched English-Hausa question pairs covering malaria, sickle cell disease, and tuberculosis were evaluated across six models and scored against Nigerian national treatment guidelines by two fluent Hausa speakers. Among locally deployable models, mean clinical correctness dropped dramatically from 1.57 in English to -0.03 in Hausa (on a scale where 2 is correct and -1 is actively harmful), while a frontier model remained competent in both languages. The authors conclude that this safety gap is a property of the deployable model tier rather than the language or clinical content itself, raising urgent concerns for AI deployment in multilingual, resource-constrained healthcare settings.
- Quality assurance
- AI policy
- Certifications
Research
A Large-Scale Measurement of AI Bill of Materials Completeness in Hugging Face Models
Md Erfan, Ahmed Ryan, Md Rayhanur Rahman
arXiv (Cornell University) · 2026-07-19
This paper empirically examines the completeness of AI Bills of Materials (AIBOMs) generated from approximately 97,500 Hugging Face model repositories, assessing how well these repositories provide machine-readable documentation on model provenance, licenses, datasets, limitations, and external references. The study finds that while required structural fields are fully represented in generated AIBOMs, AI-specific documentation—including model-card details, responsible-use information, environmental impact, and meaningful descriptions—remains weakly represented or absent. These gaps create transparency and governance risks across the AI supply chain, highlighting the need for improved model-card practices, repository-level traceability, and automated AIBOM validation.
- AI policy
- Quality assurance
- Certifications
- Enterprise
Research
AI_LectureNote: A Retrospective Pilot Study of a Post-ASR Workflow for English-Script Rendering and Semantic Drift in Korean-English Medical Lectures
Kyeongeon Lee, Donghoon Chang, Seungryeol Baek et al.
arXiv · 2026-07-19
This paper presents AI_LectureNote, a post-processing workflow that rewrites automatic speech recognition (ASR) output from Korean-English medical lectures into readable study transcripts by restoring Latin-script medical terms instead of Korean phonetic transliterations. In a retrospective pilot across four lectures and five conditions, the workflow raised English-script rendering rates substantially (e.g., from 0.39 to 0.71 on one ASR path), but improved surface rendering did not guarantee semantic faithfulness—post-processed outputs still showed semantic drift in roughly 34–36 of 282 reference sentences and polarity failures in 11–13 of 101 polarity-cue rows. The study identifies distinct failure patterns between surface accuracy and medical-meaning preservation, arguing these dimensions must be evaluated separately. The findings matter for quality assurance in AI-assisted medical education tools, where surface correctness can mask clinically significant errors in meaning.
- Quality assurance
- Workforce
Research
Robust Summarization of Doctor-Patient Conversations: TalTech Systems for the Beyond Transcription Challenge
Aivo Olev, Tanel Alumäe
arXiv · 2026-07-19
This paper presents TalTech's winning systems for the Beyond Transcription Challenge, which requires generating clinical SOAP notes directly from doctor-patient audio recordings without intermediate transcription. The team adapted Voxtral Mini and Voxtral Small speech language models using LoRA supervised fine-tuning followed by DAPO reinforcement learning, with Open Medical Concept F1 as the reward signal. Their systems ranked first in both lightweight and heavyweight tracks, and an independent LLM-as-a-judge evaluation found they had the lowest hallucination rate among all submissions. The work also shows that fine-tuning on text transcripts can transfer effectively to speech input, improving robustness on out-of-domain real-world recordings.
- Quality assurance
- Workforce
Research
Specifying the Delegated-Autonomy Boundary: Requirements Engineering for Agentic AI
Chetan Arora, Andreas Vogelsang, Abbi Sharma
arXiv (Cornell University) · 2026-07-19
This paper addresses a gap in requirements engineering for agentic AI systems—those that plan, maintain state, and act autonomously in external environments. The authors argue that current practices embed critical autonomy decisions inside prompts and runtime policies rather than treating them as explicit requirements-level commitments. They propose two artifacts: an Agency Justification Record (AJR) to evaluate whether an agent is warranted over simpler alternatives, and an Agentic Delegation Policy (ADP) that formally specifies purpose, authority, information, coordination, assurance, and evolution with a graduated authority model. The framework is illustrated through two contrasting examples—a safety-critical hospital discharge coordination agent and an automated code review agent—highlighting its relevance for safe and effective deployment of agentic AI.
- Enterprise
- Quality assurance
- AI policy
- Certifications
Research
Toward Anthropomorphic Dialogue: A Closed-Loop Framework for Human-Like Chat Generation, Evaluation, and Preference Alignment
Wentao Liu, Siyu Song, Xi Chen et al.
arXiv · 2026-07-19
AnthroDial is a closed-loop framework for generating and evaluating human-like private chat that preserves persona, memory, timing, and multi-turn coherence. The system combines a role-conditioned dialogue runtime, an executable benchmark across ten behavioral dimensions, and a post-training pipeline using supervised fine-tuning and reinforcement learning with a cognitively-informed reward signal. Evaluated across 16 systems and 100 role-conditioned cases per model, the best trained model (Qwen3.6-27B-SFT+RL) achieves 39.00% strict accuracy versus 32.00% for the strongest untrained baseline, and a 9B model improves from 0.00% to 18.37% strict accuracy after training. The results demonstrate that tightly coupling generation, evaluation, and reward shaping around shared behavioral dimensions meaningfully advances anthropomorphic dialogue quality.
- Enterprise
- Quality assurance
Research
PocketPPD: Screening for Postpartum Depression Risk Using Passive Smartphone Sensing
Jia Tang, Helinyi Peng, Akihito Taya et al.
arXiv · 2026-07-19
PocketPPD is a passive smartphone sensing system designed to screen new mothers for postpartum depression (PPD) risk without requiring constant active input from users. Using multi-modal sensor data collected over a four-week feasibility study with 61 postpartum women, a passive sensing-only model achieved an AUC of 0.75, while a combined model integrating passive sensing and self-report features reached an AUC of 0.83. The study identifies morning and late-night routine volatility as top digital biomarkers, moderated by factors like infant developmental stage and employment status. This work supports the feasibility of continuous, low-burden perinatal mental health monitoring as an alternative to traditional screening questionnaires.
- Workforce
- Quality assurance
Research
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions
Yukai Zhou, Feiyang Lu, Xiaokai Mao et al.
arXiv · 2026-07-19
This paper challenges the standard practice of evaluating jailbreak attacks on large language models solely by attack success rate (ASR), arguing that a high ASR does not necessarily translate into useful data for improving model safety. The authors propose a defender-centric framework called A-MESS, which uses Shapley values (AttackSHAP) to attribute how much each jailbreak attack contributes to downstream safety improvements when used as red-teaming training data. Their experiments show that ASR rankings are weakly correlated with actual safety utility, and that directly optimizing compact subsets of attacks via their framework yields stronger safety gains than attacker-centric selection methods. This work has implications for how organizations design safety alignment pipelines and allocate red-teaming resources for LLMs.
- Quality assurance
- AI policy
- Enterprise
Research
SlotGuard: Stop Oversharing Private Local Context in LLM Agent Transcri
Haocheng Xia, Yongjoo Park
arXiv · 2026-07-19
SlotGuard is a privacy-preserving system designed to prevent LLM agents from leaking sensitive local context—such as file paths, emails, and API keys—when agent observations are appended to provider-bound transcripts. The system rewrites structural bindings as typed slots, replaces secrets with format-preserving synthetic values, and links cross-turn references via a session graph, restoring raw values only inside a trusted runtime. On controlled benchmarks, SlotGuard removes all 20,814 annotated sensitive characters and reduces credential leakage to 0.0% across 852 planted values, while maintaining task success rates close to those of raw transcripts and completing rewrites in a median of 14.424 microseconds per turn. This matters for enterprise and policy contexts because it demonstrates a practical, low-overhead method to limit data exposure when deploying LLM agents that interact with local systems.
- Enterprise
- AI policy
- Quality assurance
Research
Fenced Citation-Context Retrieval for Case Law: Temporal Leakage and Degree Control Across Two Jurisdictions
Yao Liu, Tien-Ping Tan, Zhilan Liu
arXiv · 2026-07-19
This paper addresses a methodological flaw in prior case retrieval (PCR) systems that use citation context—text describing how later cases reference earlier ones—as a relevance signal. The authors show that existing evaluations allow 'temporal leakage,' crediting retrievers for citations made after a query case was filed, which inflates reported performance gains. They introduce a zero-training, temporally fenced retriever that restricts citation context to pre-query citations only, and demonstrate on two jurisdictions (U.S. federal CLERC and European ECtHR-PCR datasets) that this fencing reveals up to 14.9% of reported citation-context gains are attributable to future citations not available at query time. The findings establish that citation-context retrieval must be both temporally fenced and degree-controlled for its gains to be meaningfully interpreted, with direct implications for the reliability of legal AI retrieval benchmarks.
- Quality assurance
- AI policy
Research
Teach it to stop, not just to click
Barada Sahu, Shivesh Pandey
arXiv · 2026-07-19
This paper investigates the reliability of benchmarking for large-scale computer-use AI agents (CUAs), specifically a 35B model improved via a verifier-guided repair process across five evaluation environments. Using variance-components analysis, the authors show that single-run evaluations—the current norm in this field—are dominated by data-draw and run-to-run nondeterminism rather than training-seed effects, and that the run-to-run distribution can be bimodal, meaning a single reported result has roughly a 30% chance of reflecting a failure mode. They find that repair works reliably only for tightly constrained actions (e.g., a fixed 'done' token at 0.97 success rate) but degrades for open-ended corrections like spatial clicks or generative field-fills, and that task-level gains appear only when the repaired action is the sole remaining blocker. The paper argues that single-run reporting in agentic AI benchmarks is systematically misleading and releases a library (cua_reliability) to support routine multi-seed evaluation practices.
- Quality assurance
- Certifications
- AI policy
Research
Auto Research for Materials: Auditable AI-Scientist Workflows with Held-Out Transfer
Jingjie Ning, Xiaochuan Li, Shanshan Zhong et al.
arXiv · 2026-07-19
This paper develops an auditable AI-scientist workflow for materials property prediction that rigorously tests whether AI-driven modelling improvements genuinely generalize rather than merely overfitting to the development loop. Across 701 evaluated changes spanning ten Matbench endpoints, the authors freeze selected code and evaluate it once on a held-out dataset never seen during search, finding that nine of ten chosen interventions remain the best single tested change. Key validated gains include 17.4% held-out MAE reduction for band gap, 18.6% for steel strength, and a 26.3% mean improvement when separately discovered feature and model changes are combined. The work establishes an evaluation design for executable AI discoveries that can be audited, reused, and combined beyond the feedback loop, with direct implications for how AI research agents should be assessed in scientific and enterprise settings.
- Quality assurance
- Enterprise
- Certifications
Research
Otap:Structure-Aware Optimal Transport for Evaluating Planning and Execution in Agent Trajectories
Babak Barazandeh, Subhabrata Majumdar, George Michailidis
arXiv · 2026-07-19
This paper introduces OTAP (Optimal Transport for Agentic Planning), a new metric for evaluating the trajectories of large language model agents that interleave planning, tool calls, and intermediate results. Rather than reducing evaluation to a binary success flag or exact reference matching, OTAP frames trajectory evaluation as a distance between the agent's execution graph and a set of valid solution graphs, instantiated via an unbalanced fused Gromov-Wasserstein transport problem over attributed dependency graphs. The metric is provably invariant to dependency-preserving reorderings, handles missing or hallucinated steps, and accommodates variation in plan granularity. On controlled perturbations and three public benchmarks, OTAP better separates valid from invalid trajectories than semantics-only metrics, making it a more principled tool for assessing agent reliability and plan quality.
- Quality assurance
- Enterprise
Research
ALLUDE: A Unified Evaluation System for Configurable Attacks in Differentiable Environments
Mansi Phute, Alexander Greenhalgh, Matthew Hull et al.
arXiv · 2026-07-19
ALLUDE is a unified, open-source evaluation system that assesses adversarial attacks against vision-based object detectors within differentiable rendering environments across diverse, configurable conditions. By sampling from 5,400 configurations spanning multiple scene-object pairs, weather conditions, optimizers, camera trajectories, and detection models, the system reveals that existing attacks (CAMOU, RAUCA, FCA) degrade in success rate across all tested conditions, exposing evaluation gaps in prior work. The framework bridges simulation and differentiable rendering to enable end-to-end optimization and more rigorous benchmarking of adversarial robustness, and runs on both Linux and Windows. This matters for quality assurance of AI vision systems deployed in real-world conditions, where limited evaluation environments can mask significant performance failures.
- Quality assurance
- Certifications
Research
A Systematic Evaluation of Traditional Privacy Policy Analysis Tools Against LLMs
Madhav Aryal, Sudipa Saha, Kaushal Kafle et al.
arXiv (Cornell University) · 2026-07-19
This paper systematically evaluates whether large language models (LLMs) can replace specialized privacy policy analysis tools across three major functionalities: contradiction detection, regulatory compliance analysis, and privacy policy summarization. Testing GPT and Gemini models against six representative tools on ten privacy policies, the authors find that LLMs consistently match or exceed specialized tools, achieving up to 91.4% precision and 70.8% recall for labeling third-party sharing entities compared to the OPP-115 dataset. The findings suggest that off-the-shelf LLMs can broadly perform privacy policy and regulatory compliance analysis without requiring domain-specific training or specialized tools, significantly lowering the barrier to compliance work.
- AI policy
- Quality assurance
- Enterprise