News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
5608 items
Research
Answering Without Referring: How AI Search Rewrites the Web's Economic Bargain
Qiaoni Shi, Kai Zhu, Kai Gu
arXiv · 2026-07-08
This paper examines how AI-powered search (specifically ChatGPT Search) disrupts the traditional web economy in which search engines drive traffic to websites. Using Comscore U.S. desktop clickstream data, the authors find that ChatGPT produces outbound clicks in only 5.2% of conversation sessions, far below Google's referral ratio, and that wider ChatGPT access reduces search use by 9.4%, with the largest losses in informational categories. The clicks that do occur skew toward specialized destinations and away from ad-supported sites. The findings suggest AI search increasingly resolves information needs within the intermediary itself, potentially undermining the referral-based economic bargain that has historically linked search engines, web traffic, and content production.
- Enterprise
- AI policy
- Workforce
Research
User identity conditions moral wrongness ratings in non-reasoning large language models
Willem Fourie, Isabel Ray, Gray Manicom
arXiv · 2026-07-08
This study examines whether implicitly conveyed user identity (e.g., professional role) shifts moral evaluations made by large language models, without explicitly instructing the models to adopt any persona or moral stance. Across 12,000 structured multi-turn interactions with gpt-4.1-mini and gemini-2.5-flash-lite, models were asked to rate wrongness (0–100) on ten common-morality rules from Gert's moral framework. Results show that moral judgments vary with user role in both models—particularly for contestable rule-governed acts—while grave-harm acts like killing show a ceiling effect. The findings raise concerns for AI value alignment by demonstrating unintended contextual conditioning and suggest future research should focus on dynamic moral bounds rather than static principles.
- AI policy
- Quality assurance
Research
Alignment Plausibility: A New Standard for Assuring AI in Healthcare
Gwydion Williams, Sara Zannone, Bilal A Mateen
arXiv · 2026-07-08
This paper argues that large language models used for mental health support are structurally misaligned with patient welfare because they are built on attention-economy incentives that favor engagement over effective care. The authors propose a three-level alignment framework—value specification grounded in clinical norms, training that embeds those values, and deployment oversight analogous to clinical supervision—to address both acute and subtler long-term harms such as dependency, boundary erosion, and amplification of distorted beliefs. From this framework they derive a new regulatory construct called 'alignment plausibility,' modeled on the established concept of biological plausibility, which provides a principled basis for arguing whether an AI system is trustworthy, non-harmful, and likely to produce patient benefit. The construct is intended to guide regulators and developers in assuring AI safety in healthcare settings.
- AI policy
- Certifications
- Quality assurance
Research
Beyond Attack-Success Rate: Action-Graded Severity Scale for Tool-Using AI Agents
Harry Owiredu-Ashley
arXiv · 2026-07-08
This paper argues that current AI agent red-teaming benchmarks rely on a binary attack-success metric that loses critical information about how harmful a compromised agent's actions actually were. The authors introduce a seven-level ordinal severity scale (L0–L6) grading an agent's tool-call trajectory based on reversibility, scope-crossing, and privilege escalation, computed both by a deterministic oracle and a panel of three large language model judges. Applied across four victim models and two defenses on the AgentDojo benchmark suite, the rubric uncovers cases hidden by binary metrics—including a defense reporting zero attack-success rate that still permits an externally visible cross-scope data leak. The LLM judge panel closely reproduces the oracle (Krippendorff's alpha = 0.91) but shares systematic blind spots, notably failing to recognize escalation chains, highlighting both the promise and limits of automated severity assessment for AI agent security.
- Quality assurance
- AI policy
- Certifications
Research
Reason Less, Verify More: Deterministic Gates Recover a Silent Policy-Violation Failure Mode in Tool-Using LLM Agents
Vikas Reddy, Sumanth Reddy Challaram, Abhishek Basu
arXiv · 2026-07-08
This paper identifies a critical failure mode in tool-using LLM agents where the agent silently executes policy-violating actions—such as cancelling a booking or changing passenger counts—without any tool error or self-reported failure. The authors find that 78% of observed failures in a budget agent on the τ²-bench airline domain are these silent wrong-state failures. They propose a lightweight fix: deterministic, read-only pre-execution gates that check proposed tool calls against current state and domain policy before allowing writes. These gates raise full-benchmark success from 29.6% to 42.0% on gpt-4o-mini (+12.4pp) and show consistent lift on a disjoint replication set, offering a bounded but reliable mechanism for preventing a specific class of policy-violating agent behavior.
- Quality assurance
- AI policy
- Enterprise
Research
Reasoning Consistency Scanning: A Framework for Auditing Chain-of-Thought Validity in AI Safety Evaluations
Silvia Santano
arXiv · 2026-07-08
This paper introduces 'reasoning consistency scanning,' a framework for auditing whether the stated chain-of-thought reasoning in AI safety evaluations is logically consistent with the model's final answer. Unlike faithfulness (which requires experimental intervention), consistency can be assessed directly from evaluation transcripts, making it a practical post-hoc auditing tool. The authors formalize a six-subtype taxonomy of inconsistency, build a 60-transcript benchmark, and implement a working scanner that detects inconsistency across four AI models and three evaluation tasks, finding that reasoning inconsistency is present, detectable, and varies systematically. This matters for quality assurance and certification of AI systems, as unreliable reasoning traces undermine trust in safety evaluations used to assess model behavior.
- Quality assurance
- Certifications
- AI policy
Research
Vision Foundation Models in Radiology: A Scoping Review of Data, Methodology, Evaluation and Clinical Translation
Alejandro Vergara-Richart, Xavier Rafael-Palou, Almudena Fuster-Matanzo et al.
arXiv · 2026-07-08
This scoping review examines 67 peer-reviewed studies (2017–2026) on vision foundation models (VFMs) built exclusively for radiological imaging, mapping findings across data scale, architecture, and transferability. Datasets spanned brain MRI, thoracoabdominal CT, and chest X-ray, ranging from under 100,000 to multi-million images, with transformer-based architectures and self-supervised pretraining (masked image modeling, contrastive learning) predominating. Evaluation focused on segmentation and classification, but cross-center, cross-scanner, and modality-shift validation was inconsistently reported, and alignment with FUTURE-AI principles was uneven. The review concludes that clinical translation of radiology-specific VFMs remains constrained by limited data representativeness, heterogeneous benchmarks, incomplete reporting, and insufficient deployment-oriented evaluation—findings directly relevant to quality assurance, certification, and policy frameworks for AI in healthcare.
- Quality assurance
- Certifications
- AI policy
Research
Validate the Dream Before You Trust Its Verdict: Admissibility for World-Model Simulators
Christian Oefinger, Finn Rasmus Schäfer, Korbinian Moller et al.
arXiv · 2026-07-08
This paper argues that World Models (WMs)—AI systems used to simulate and evaluate action policies in robotics and autonomous driving—must themselves be certified before their verdicts can be trusted as safety evidence. The authors propose an 'admissibility ladder' (L0–L4) grounded in established safety-critical simulation practices such as VV&A and SOTIF, defining progressive criteria a WM must satisfy before its closed-loop evaluations count as assurance evidence. Applied to two driving WMs, the framework reveals a key finding: the model scoring higher on visual generation quality (Fréchet Video Distance) ranks lower on action-following fidelity, demonstrating that visual realism does not predict the action-robustness required for reliable policy verdicts. This has direct implications for how AI-generated simulations are used in safety certification and quality assurance processes for autonomous systems.
- Certifications
- Quality assurance
- AI policy
Research
Predicting LLM Safety Before Release by Simulating Deployment
Marcus Williams, Hannah Sheahan, Cameron Raymond et al.
arXiv · 2026-07-08
This paper proposes 'deployment simulation' as a method for predicting how large language models will misbehave once released to the public. By taking de-identified conversation prefixes from prior deployments and regenerating responses with a candidate model, evaluators can estimate the prevalence of unsafe or misaligned behavior before release. Tested across four GPT-5-series deployments, the method outperforms adversarially selected production data baselines and produces estimates much closer to real production traffic than traditional evaluations. The authors also show the approach can be seeded from public chat datasets, enabling external researchers to conduct deployment-grounded safety evaluations without access to private logs.
- Quality assurance
- AI policy
- Certifications
Research
Operational Reframing and Approval-Framed Delegation in Multi-Agent LLM Safety
Lifei Liu, Haoran Yu, Xiaochong Jiang et al.
arXiv · 2026-07-08
This paper investigates why multi-agent LLM systems (planner-executor pipelines) can behave more unsafely than single-model prompting, arguing that the commonly reported 'pipeline effect' conflates three distinct mechanisms: operational reframing of harmful intent, planner refusal or transformation, and approval-framed delegation. Using a five-condition controlled experiment across 30 synthetic harmful scenarios and an external validation set from four agent-safety benchmarks, the authors find that operational reframing is the most consistent risk factor across GPT, Gemini, and DeepSeek models, while Claude shows more resistance. Critically, model safety rankings under direct prompting can mispredict behavior in deployed pipelines—Gemini's compliance rate rose from 8.9% to 38.9% when paired with a Claude planner—demonstrating that aggregate pipeline safety is not a stable architectural property. The findings call for multi-agent safety evaluations to separately report reframing, planner behavior, delegation framing, and model pairing rather than treating pipeline architecture as a single variable.
- AI policy
- Quality assurance
- Certifications
Research
Progressive Crystallization: Turning Agent Exploration into Deterministic, Lower-Cost Workflows in Production
Arun Malik
arXiv · 2026-07-08
Progressive crystallization is a lifecycle framework that converts AI agent behaviors—initially requiring full LLM inference—into cheaper, deterministic workflows once those behaviors are repeatedly validated in production. Applied to a cloud networking AIOps system handling tens of thousands of incidents per month, the approach raised deterministic execution from 0% to 45% over eight months and cut per-incident agent costs by more than 70%, even as incident volume doubled. The framework defines a three-stage taxonomy (fully agent-orchestrated, hybrid, and fully deterministic) with evidence-based promotion and automatic demotion when workflows regress, improving safety through greater reproducibility and auditability. This demonstrates a practical path for enterprises to reduce the ongoing inference costs of AI agents without sacrificing reliability.
- Enterprise
- Workforce
- Quality assurance
Research
Learning social norms enhances compatibility in dynamic human-AI coordination
Yi Yang, Siyuan Liu, Xin Gao et al.
arXiv · 2026-07-08
This paper investigates why AI agents—including large language models—often fail to coordinate smoothly with humans in dynamic, real-world interactions. The researchers collected 3,456 human interactions in a pedestrian-vehicle simulation and identified three principles underlying human social norms: outcome predictability, value alignment, and advantage awareness. Incorporating these principles into an LLM agent produced nearly four times the score of a baseline strategy and outperformed human-human interactions by 43% in closed-loop tests. The findings suggest that explicitly formalizing tacit social norms into quantifiable principles is a promising path toward more natural and effective human-AI coordination in daily life.
- Workforce
- Enterprise
- AI policy
Research
Evaluating LLM Robustness Under Domain-Specific Prompt Perturbations in Public Health Applications
Chuqing Zhao, Haochen Yang
arXiv · 2026-07-08
This paper proposes a domain-specific benchmark to evaluate how robust large language models (LLMs) are when handling realistic, non-clinical user inputs in public health settings. The authors test two perturbation types: misinformation framing (MF), where prompts contain false health claims, and layperson rewriting (LR), where patients describe symptoms in everyday language. Results show MF degrades model accuracy by 7.2 percentage points on average with prediction flip rates of 9–38%, even when claims are labeled as unsupported, while LR causes only 1.4 pp degradation. These findings reveal meaningful deployment risks—models may produce incorrect outputs when users inadvertently introduce misinformation or use informal symptom descriptions—underscoring the need for perturbation-aware robustness evaluation in health AI systems.
- Quality assurance
- AI policy
Research
The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI
Muayad Sayed Ali, Aliaksandra Novik, Anji Boddupally et al.
arXiv · 2026-07-08
This paper investigates how the orchestration layer ('harness') in enterprise agentic AI systems determines token usage and cost, independent of which foundation model is used. In a controlled experiment swapping only the orchestration layer across 22 tasks and six foundation models, the authors find the Writer Agent Harness cuts blended cost per task by 41%, reduces tokens per task by 38%, and cuts median wall-clock time by 44%, with task-completion quality at rough parity. Notably, efficiency gains are model-invariant (every model gets cheaper by 33–61%), while quality improvements correlate strongly with a model's baseline capability (r=0.99). The paper argues that orchestration design is a more decisive lever on AI operating costs than model selection itself, making it a critical consideration for enterprise AI deployment and spend governance.
- Enterprise
- Quality assurance
- AI policy
Research
Governance-Aware Agentic AI for Enterprise Engineering Systems: A Design-Science Reference Architecture and Quantitative Risk-Control Model
Kwan Hong Tan
Zenodo (CERN European Organization for Nuclear Research) · 2026-07-08
This paper presents a governance-aware reference architecture for deploying agentic AI in enterprise engineering systems, addressing the gap between high-level AI governance standards and concrete operational controls. Using design-science methodology, it introduces a six-layer control architecture and three formal constructs—Productivity-Adjusted Residual Risk, Governance Debt, and Human Override Threshold—to make governance principles measurable and enforceable at the system level. Illustrative scenarios across customer service, finance, HR screening, and supply planning demonstrate how the model translates abstract principles into engineering checks. The work argues that trustworthy enterprise AI requires governance built into system architecture rather than applied after deployment.
- Enterprise
- AI policy
- Quality assurance
- Certifications
Research
Developing an Algorithmic Accountability and Data Sovereignty Framework for the governance of Artificial Intelligence platforms in United States education
Oluwatayo Osodein, Oluwatayo Osodein, Ayomide Arowolo Ayodeji et al.
Discover Education · 2026-07-08
This systematic review analyzes 47 documents—including privacy policies, federal and state legislative texts, and case studies—to assess how current U.S. regulations govern AI learning platforms in K-12 education. The study finds that existing laws like FERPA and COPPA are insufficiently adapted to modern AI data processing, that 60% of principals and 25% of teachers are using AI while only 18% of districts have formal usage guidelines, and that schools frequently deploy AI without adequate district oversight. Drawing on data ethics, technological determinism, and surveillance capitalism frameworks, the authors identify risks of algorithmic bias, discriminatory disciplinary actions, and commercial exploitation of student data. The paper calls for immediate policy reforms to establish algorithmic accountability, student data sovereignty, and student-centered protections over data monetization.
- AI policy
- Certifications
- Quality assurance
Research
Governance-Aware Agentic AI for Enterprise Engineering Systems: A Design-Science Reference Architecture and Quantitative Risk-Control Model
Kwan Hong Tan
Zenodo (CERN European Organization for Nuclear Research) · 2026-07-08
This paper presents a reference architecture called the Governance-Aware Agentic AI Control Architecture, designed to help enterprises deploy autonomous AI agents responsibly in engineering systems. Using a design-science methodology, it formalizes governance requirements into six architectural layers and three quantitative constructs—Productivity-Adjusted Residual Risk, Governance Debt, and Human Override Threshold—to translate abstract AI governance standards into measurable engineering controls. An illustrative evaluation across customer service, finance, HR screening, and supply planning demonstrates how the model operationalizes oversight mechanisms such as auditing, human escalation, and reversibility. The work matters because it offers organizations a practical, theoretically grounded blueprint for embedding accountability into agentic AI systems rather than treating governance as post-hoc compliance.
- Enterprise
- Quality assurance
- AI policy
- Certifications
Research
LLMs Silently Correct African American English: Auditing and Mitigating Dialect Bias via Activation Steering
Huan Wu, Ali Emami, Muhammad Furquan Hassan et al.
arXiv · 2026-07-07
This paper audits six large instruction-tuned LLMs (14B–70B parameters) and finds they systematically rewrite African American English (AAE) into Standard American English (SAE), even when context clearly calls for AAE. The authors introduce a bias-auditing framework called conditional Dialect Group Invariance (cDGI) and identify syntactic features—especially negative concord—as universal bias triggers across all models. To counter the bias, they apply activation steering at inference time (no retraining required), reducing dialect bias 5 to 20 times more effectively than prompting alone while preserving SAE fluency. They also release REAL-AAE, the largest real-AAE parallel corpus to date, with 17,479 validated triplets, to support future fairness research.
- AI policy
- Quality assurance
Research
Digital Fragmentation and Generative AI Use Across 103 Million Application Events
Sumer S. Vaid, Ashley V. Whillans
arXiv · 2026-07-07
This large-scale observational study analyzed 103 million application events from 1,017 knowledge workers across eight organizations to understand digital fragmentation—the frequent switching between applications that consumes nearly a tenth of the work year. The researchers found that day-to-day variation within individual employees (44.6%) was the largest driver of fragmentation, exceeding stable individual differences (35.8%) and organizational differences (19.6%), with fragmentation rising over the work week and resetting after weekends and holidays. Notably, generative AI use occurred on more fragmented days, but the period immediately following AI use was characterized by narrower, longer, and more predictable application use, suggesting AI may help structure fragmented work rather than intensify it. These findings have implications for how organizations and workers manage digital work patterns and how AI tools might be leveraged to improve workforce productivity and focus.
- Workforce
- Enterprise
Research
Automated Compliance Mapping in Cloud Security with Domain-Adapted Sentence Transformers
John Bianchi, Luca Petrillo, Fabio Martinelli et al.
arXiv · 2026-07-07
This paper proposes automating the manual process of mapping cloud security controls to technical metrics by fine-tuning domain-adapted Sentence Transformer models. The researchers built a training corpus of 3,499 semantic pairs drawn from five European security standards and expanded it to up to 13,996 samples using back-translation and LLM-based paraphrasing. All five fine-tuned architectures outperformed their zero-shot baselines, with the best model gaining up to 23 nDCG@10 points on the control-to-metric task and one model reaching 0.870 nDCG@10 on cross-standard control association. The findings demonstrate that in-domain training data is a primary performance driver, suggesting this approach could significantly reduce manual compliance mapping effort in cloud security contexts.
- Quality assurance
- Certifications
Research
Large language models create an uneven informational layer over cities
Lin Chen, Guangyuan Weng, Esteban Moro
arXiv · 2026-07-07
This study audits restaurant recommendations from three major large language models across 304 neighborhoods in five U.S. cities, using 320 synthetic user profiles that vary by income, age, sex, and residential status. The researchers find that LLMs both fabricate nonexistent venues and systematically overlook real ones, with fabrication concentrated in neighborhoods with weaker digital and physical footprints, while 47.5% of real establishments are never recommended and 31.9% of these blind spots are shared across all three model families. Recommendation patterns also vary by user profile: higher-income users receive more expensive and less popular venues, while tourists are directed toward costlier but more socially diverse establishments than local residents. Simulating shifts in consumer demand suggests widespread LLM reliance would redirect visits and revenue away from chain and quick-service restaurants toward independent and full-service dining, raising concerns about urban economic inequality.
- Enterprise
- AI policy
Research
Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability
Alicia Parrish, Rajat Shinde, Sanket Badhe et al.
arXiv · 2026-07-07
Pluralis v0.1 is a new multimodal, multilingual benchmark dataset containing 6,448 prompts spanning six Asia-Pacific countries (Bangladesh, India, Korea, Pakistan, Singapore, and Taiwan) and eight languages, designed to evaluate AI safety from a culture-first perspective rather than adapting Western-centric defaults. The benchmark introduces a multimodal evaluation paradigm where text and image inputs are individually innocuous but together can trigger locale-specific legal or cultural violations, and it disentangles universal safety violations from localized cultural appropriateness. Testing Vision-Language Models on a subset of Pluralis reveals recurring locale-specific failure modes—including image misidentifications with downstream harm, missed item-context-locale interactions, and inadequate refusals—that vary systematically across locales and languages in ways that globally averaged metrics conceal. This work matters because it exposes critical blind spots in current AI safety evaluations for global deployments and establishes a foundation for advancing multicultural, multilingual AI alignment research.
- Quality assurance
- AI policy
Research
Open-Ended Scenario Reasoning for Specialist Model Adaptation
Youcheng Zong, Runda Jia, Ranmeng Lin et al.
arXiv · 2026-07-07
This paper introduces ROAM (Reasoning-Driven Open Adaptation for Specialist Models), a framework that uses large language model (LLM) reasoning to adapt frozen industrial process models to new scenarios without retraining or collecting new labeled data. Rather than using LLMs as direct predictors, ROAM fuses LLM-generated scenario judgments with online observations in a low-dimensional latent space, while a risk-constrained mechanism prevents overcorrection when LLM evidence is unreliable. Experiments on a mineral thickening process and the IndPenSim penicillin fermentation dataset show ROAM reduces mean absolute error by over 20% in major shift settings, adding only 839 parameters and under 0.02 ms per-step overhead. The results demonstrate that LLM world knowledge can be converted into a conservative, interpretable adaptation signal for deployed industrial models, reducing the need for costly retraining cycles.
- Enterprise
- Quality assurance
Research
Agents That Teach: Towards Designing Incidental Learning Back into AI-Assisted Software Development
Rohit Mehra, Samdyuti Suri, Prithviraj K Tagadinamani et al.
arXiv · 2026-07-07
This paper examines how AI coding agents, while boosting productivity, undermine the informal, effortful problem-solving through which software developers historically acquired expertise — a phenomenon the authors call 'Knowledge Debt,' a developer-level analogue of Technical Debt. The authors argue that incidental learning will not return on its own and propose six design principles to consciously reintroduce it into developer-agent interactions. They present SHIELD, a multi-agent system that uses the AI coding agent's own reasoning to surface contextual learning moments without disrupting developer workflow. The work aims to make productivity and learning complementary rather than competing outcomes in AI-assisted software development.
- Workforce
- Enterprise
Research
Prompt Coach: An Empirical Evaluation of an Agentic Tutor for Learning Prompt Engineering in Software Development
Rohit Mehra, Kapil Singi, Vikrant Kaulgud et al.
arXiv · 2026-07-07
Prompt Coach (PC) is an agentic AI tutor embedded in developers' IDEs that teaches prompt engineering through Socratic questioning, evaluating prompts across multiple quality dimensions and guiding self-correction grounded in the developer's codebase and target LLM behavior. An empirical study with 15 professional developers found statistically significant improvements in prompt quality after a single 60-minute session, with the largest gains in dimensions developers commonly overlook. Participants also reported strong trust, high adoption readiness, and unanimous agreement that PC improved their prompt-writing skills. The findings suggest that in-flow, agentic tutoring is a promising approach for closing the gap in prompt engineering education among software developers.
- Workforce
- Enterprise