News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
5608 items
Research
The Context Access Divide: Interaction-Level Architecture as a Complementary Dimension of Agentic Inequality
Masahiro Fujita
arXiv · 2026-07-09
This paper introduces the 'Context Access Divide' (CAD) as a new dimension of AI inequality operating at the individual interaction level, complementing Sharp et al.'s (2025) framework of agentic inequality across availability, quality, and quantity. The authors argue that two users with nominally equivalent AI agent access can experience qualitatively different utility depending on whether the system autonomously retrieves relevant context (Dynamic Context Retrieval) or requires users to manually attach documents at each query (Manual Attachment). Using a probabilistic model grounded in the fan effect literature from cognitive psychology, they demonstrate that manual context attachment leads to combinatorial collapse in task-success probability as a user's knowledge corpus grows and tasks become more conjunctive, while dynamic retrieval architectures avoid this collapse. The paper analyzes the technical underpinnings of this divide in Model Context Protocol (MCP) and retrieval-augmented generation (RAG) architectures, and examines implications for knowledge-work stratification and AI platform governance.
- Workforce
- Enterprise
- AI policy
Research
Two Axes of LLM Abstention: Answer Correctness and Question Answerability
Benedikt J. Wagner
arXiv · 2026-07-09
This paper investigates why large language models (LLMs) fail to appropriately refuse both wrong answers and unanswerable or false-premise questions. The authors find that these are two distinct axes: standard confidence scores can detect when a model will answer incorrectly, but are nearly blind to whether a question is actually answerable, and this blind spot does not improve with model scale. A hidden-state linear probe fills the gap for answerability detection (reaching 0.69–0.77 AUROC on false-premise questions), and combining both signals into a two-axis calibrated policy dramatically outperforms single-threshold approaches—achieving 0.75 coverage of correct answers versus 0.31 for a single threshold. These findings matter for quality assurance and policy around LLM deployment, as they show that reliable abstention requires separately certifying correctness and answerability budgets rather than relying on a single confidence score.
- Quality assurance
- AI policy
- Certifications
News
New Fabric Test Material Could Help Strengthen Domestic Supply Chain for Textiles and Clothing
nist.gov · 2026-07-09
NIST News reports that researchers at the National Institute of Standards and Technology have developed a new Research Grade Test Material (RGTM 10279) consisting of five fabric squares made from different fibers, designed to help the textile industry validate and improve methods for identifying and sorting textiles. The material is intended to support AI-enabled sorting technologies, which the agency says have not yet been exhaustively tested for accuracy in fiber identification. NIST is distributing the free test material to labs and manufacturers through July 30, 2026, in exchange for measurement feedback, with the goal of ultimately developing a more robust reference standard that meets real-world industry needs. Researchers note the material could also help verify fabric composition for brands and potentially support quality control across the domestic textile supply chain.
- Certifications
- Enterprise
- Quality assurance
Research
Reverse Engineering Compliance: A Dual-Graph Verification Framework for Auditing Legacy IT Security Concepts
Lea Roxanne Muth, Marian Margraf
arXiv · 2026-07-09
This paper introduces ASSERT, a framework for auditing legacy IT security concept documents by extracting them into formal document graphs and comparing them against a verified reference graph using a five-class graph difference method. The framework exports schema-valid OSCAL artifacts to support auditable, machine-readable compliance evidence, addressing a gap left by prior work that focused on generating new security concepts rather than verifying existing ones. Evaluated on the BSI's RecPlast dataset, ASSERT reveals a trade-off between discovering undocumented infrastructure entities and enforcing a strict schema, making document-infrastructure inconsistencies measurable. This work is directly relevant to organizations facing NIS-2 Directive compliance obligations and to BSI's ongoing Grundschutz++ initiative.
- AI policy
- Certifications
- Quality assurance
- Enterprise
Research
From Legacy Documentation to OSCAL: An MCP-Based Agent Pipeline for Threat-Informed Continuous Compliance in Critical Infrastructure
Lea Roxanne Muth, Marian Margraf
arXiv · 2026-07-09
This paper presents a multi-agent AI pipeline that converts natural-language descriptions of critical infrastructure systems into structured, audit-ready compliance artifacts in the NIST OSCAL format, without requiring active network scanning of sensitive operational technology environments. The pipeline grounds LLM reasoning in authoritative threat-intelligence sources via a Model Context Protocol (MCP) architecture, reducing hallucinated vulnerabilities and attack paths. In a synthetic water utility scenario, the system achieves 0.90 CVE recall and perfect D3FEND recall, producing schema-valid OSCAL System Security Plans and Security Assessment Reports. The key finding is that grounding shifts errors to the asset-extraction phase, making remaining risks visible and suitable for efficient manual review rather than eliminating errors entirely.
- Quality assurance
- Certifications
- AI policy
Research
Psychological Competence as a Missing Dimension in AI Evaluation
Marcos Economides, Paul M. Sacher, Samuel Salzer et al.
arXiv · 2026-07-09
This paper argues that current AI evaluation frameworks are incomplete because they focus on technical metrics like accuracy and robustness while ignoring how AI systems psychologically affect the users they interact with. The authors introduce 'psychological competence' as a new evaluation dimension, defined as an AI system's capacity to support user cognition, emotional interpretation, and behavioral decision-making in contextually appropriate ways. They outline a conceptual framework covering interaction properties such as framing, tone, perceived authority, and uncertainty handling, and propose assessment approaches including scenario-based probes and structured human evaluation. The work has direct implications for model providers, deploying organizations, and regulators seeking to understand the real-world effects of AI systems used as advisors, coaches, tutors, and companions.
- Quality assurance
- Certifications
- AI policy
- Enterprise
Research
MentalHospital: A Virtual Environment for Evaluating Psychiatric Clinical Encounters
Yuming Yang, Xiao Sun, Yuanwei Zou et al.
arXiv · 2026-07-09
MentalHospital is a virtual evaluation environment that benchmarks large language models on complete psychiatric clinical encounters, following the S.O.A.P. workflow across 1,193 de-identified EHR cases spanning all major ICD-11 categories and 76 disorders. The framework uses a dual-track assessment protocol combining objective EHR-derived references with subjective clinical process quality, supported by MentalEval, a set of five domain-specific AI evaluators achieving an average quadratic weighted kappa of 0.944 against expert judgment. Benchmarking results show that even the strongest LLM trails clinicians by 37.28 percentage points in objective psychiatric competence, with mental status assessment identified as a key bottleneck. These findings highlight significant gaps between current AI capabilities and clinical standards in psychiatry, with implications for the quality assurance and certification of AI systems in high-stakes healthcare settings.
- Quality assurance
- Certifications
- Workforce
Research
Prismata: Confining Cross-Site Prompt Injection in Web Agents
Corban Villa, Alp Eren Ozdarendeli, Sijun Tan et al.
arXiv · 2026-07-09
Prismata is a defense system designed to protect autonomous web agents from cross-site prompt injection attacks, where malicious third-party or user-generated content on a webpage hijacks an agent by masquerading as instructions. The system enforces contextual least privilege by dynamically deriving trust labels for page content and using structural confinement — inspired by classical integrity models — to ensure labeling errors only reduce privilege rather than escalate it. Mechanical confinement then redacts suspicious content and restricts agent capabilities accordingly, requiring no developer annotations and thus supporting arbitrary websites. Evaluated against recent published web agent attacks including adaptive variants, Prismata substantially reduces attack success while preserving the agent's ability to complete legitimate tasks.
- Quality assurance
- Enterprise
- AI policy
Research
Gauge dependence and structured-output corruption in sign-branched repetition penalties: measurements across models, inference stacks, and alternative repetition controls
Peter Hollows
arXiv · 2026-07-09
This paper identifies a fundamental flaw in the multiplicative repetition penalty used across major LLM inference engines (HuggingFace, vLLM, llama.cpp, and others): the penalty branches on the sign of raw logits, but because softmax is invariant to constant shifts, a model's logit zero-point is arbitrary and left unconstrained by training. The authors demonstrate two measurable consequences: the penalty is effectively undefined across models (re-centering logits changes 58–96% of greedy tokens at a routine theta=1.3, while alternative penalties change none), and it severely corrupts structured output (dropping valid JSON schema conformance from 97% to 23% across 200 real-world schemas). Experiments spanning five models up to 7B parameters, two code models on HumanEval and JSONSchemaBench, and replication inside vLLM and llama.cpp confirm both effects. The authors show that applying the penalty to normalized log-probabilities instead of raw logits removes both problems, and note that HuggingFace already ships this operator (LogitNormalization) but applies it after the penalty and off by default.
- Quality assurance
- Enterprise
Research
Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring
Jennifer Za, Julija Bainiaksina, Nikita Ostrovsky et al.
arXiv · 2026-07-09
This paper stress-tests chain-of-thought (CoT) monitoring—a safety mechanism that uses visible AI reasoning traces to detect misaligned behavior—against adversarial persuasion attacks. The authors find that in adversarial settings, giving a monitor access to an agent's CoT reasoning actually increases approval of harmful, policy-violating actions by an average of 9.5%, because the scratchpad becomes an additional channel for persuasion. To counter this, they introduce a fact-checking monitoring framework and show that pairing monitors and fact-checkers from different model families (e.g., Claude 3.7 Sonnet with GPT-4.1) reduces approval of policy-violating actions by up to 45%, far outperforming same-family pairings at only 6%. The findings indicate that CoT monitoring alone is insufficient against adversarial persuasion and that model-diverse fact-checking is a more robust mitigation strategy.
- Quality assurance
- AI policy
Research
When LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals
Kaihua Ding
arXiv · 2026-07-09
This paper investigates whether agreement among large language models—either within a single model's repeated outputs or across different models—reliably signals correctness, a key assumption in LLM-as-judge evaluation pipelines widely used in enterprise AI systems. Using 265,000 samples from 53 runners across GPQA Diamond and AIME benchmarks, the authors find that agreement is a positive but weak predictor of correctness (rho 0.20–0.59), and that this relationship is highly regime-dependent. For frontier models, agreement is particularly misleading: the most consistent model showed agreement ≥0.8 on 77% of GPQA cases, yet 48% of those were wrong. The findings caution against treating self-consistency or cross-model agreement as standalone confidence scores in enterprise evaluation pipelines, recommending instead that they be understood as conditional proxies whose usefulness depends on model tier and task saturation.
- Enterprise
- Quality assurance
Research
Who Analyses the Analyser? Self-Validating LLM Hazard Analysis with Constitutional Meta-STPA
Samuel Tetteh, Udip Shrestha, Joshua R. Waite et al.
arXiv · 2026-07-09
This paper addresses a critical blind spot in AI-assisted safety analysis: the LLM tools used to perform hazard analysis are themselves safety-relevant systems that have never been subjected to rigorous safety analysis. The authors introduce Constitutional Meta-STPA, a framework that applies Systems-Theoretic Process Analysis (STPA) to the LLM-assisted safety tool itself, deriving a governance constitution of 21 Tool Principles and 8 Meta-Safety Principles directly from the resulting hazard analysis chain rather than asserting them externally. They formalize a coverage operator over the 29-principle set and report that a frontier model ensemble recovers 18/21 canonical and all 8/8 governance principles from the tool's own design, while a weaker model pair recovers significantly fewer, demonstrating the meta-layer is model-limited rather than constitution-limited. This matters because it provides a self-validating, auditable approach to governing AI tools used in safety-critical certification and quality-assurance contexts, closing a governance gap that existing literature has ignored.
- Quality assurance
- Certifications
- AI policy
Research
A safety-oriented hypothetico-deductive framework for AI-assisted differential diagnosis
Fan Ma, Mauro Giuffrè, Donald Wright et al.
arXiv · 2026-07-09
AegisDx is a safety-oriented AI framework for clinical differential diagnosis that coordinates specialized large language model components with role-specific contracts, verification gates, and evidence-retrieval interfaces to enforce screening for dangerous 'must-not-miss' conditions. Evaluated on case reports from NEJM, JAMA, and Annals of Emergency Medicine, AegisDx achieved Top-3 diagnostic accuracy of up to 85.7% versus 68.6% for a standalone LLM, and captured at least one must-not-miss condition in 78.0% of cases compared to 52.0% for the baseline. In a blinded physician evaluation of 43 real-world emergency department notes, AegisDx improved the physician-rated composite safety score from 4.31 to 4.55 on a 5-point scale (adjusted p = 2.1×10⁻⁴). The findings suggest that structuring diagnostic AI as a safety-oriented reasoning framework, rather than optimizing raw predictive accuracy alone, can provide more transparent and clinically meaningful decision support in acute care settings.
- Quality assurance
- Enterprise
- AI policy
Research
From Prompts to Contracts: Harness Engineering for Auditable Enterprise LLM Agents
Joongho Ahn, Moonsoo Kim
arXiv · 2026-07-09
This paper introduces a 'harness engineering' approach for building auditable enterprise LLM agents, where deterministic behaviors—such as source grounding, entity routing, and output validation—are moved into code, schemas, and manifests rather than relying on prompts alone. The authors instantiate this pattern on a dataset of 25 Korean listed companies and evaluate it across three research questions, showing that the harness reliably enforces behavioral contracts across model substitutions (passing all 270 composition-boundary runs) and catches deliberately injected faults. Critically, prompt-only enforcement fails to block recommendation-language and trace-leakage violations, while a bolt-on external guardrail prevents violations but reduces utility from 120/120 to 88/120—only code-owned enforcement preserves both safety and full utility. The result is a reusable engineering pattern for converting exploratory LLM prototypes into auditable, production-ready enterprise applications with versioned controls and validation artifacts.
- Enterprise
- Quality assurance
- Certifications
Research
Beware What You Autocomplete: Forensic Attribution of Backdoored Code Completions
Anjun Gao, Yueyang Quan, Zhuqing Liu et al.
arXiv · 2026-07-09
This paper introduces CodeTracer, a forensic framework designed to trace malicious code completions produced by backdoored large language models back to the specific fine-tuning data responsible for the unsafe behavior. Operating under realistic post-deployment conditions, CodeTracer extracts a behavioral fingerprint from a compromised output, narrows the search to semantically relevant training samples, and uses LLM-based reasoning to attribute unsafe logic to particular backdoor data. Evaluated across three vulnerability cases and ten backdoor attacks with sixteen competitive baselines, the system achieves high forensic accuracy and low false identification rates. This matters for software quality assurance and enterprise security, as it provides a practical accountability mechanism against stealthy backdoor attacks in AI-assisted coding tools.
- Quality assurance
- Enterprise
- AI policy
Research
Tool-Making and Self-Evolving LLM Agents in Low-Latency Systems
Kalle Kujanpää, Ning Liu, Shahnawaz Alam et al.
arXiv · 2026-07-09
This paper presents an agentic tool-making pipeline for production LLM systems that compiles repeated standard operating procedure (SOP) steps into validated, versioned tools before deployment, rather than regenerating code on every request. Deployed in a Fulfillment Center alarm-triage system diagnosing alarms across a 44-node SOP with heterogeneous metric backends, the approach reduced p50 latency by 42% and cut end-to-end error rates by up to 53% on 1,500 historical alarms. A further architectural simplification enabled by compact structured tool outputs reduced p50 latency by an additional 62% in ablation testing. The results demonstrate that self-evolving agents can make industrial LLM deployments faster, more reliable, and more auditable.
- Enterprise
- Quality assurance
Research
ARTIFICIAL INTELLIGENCE ADOPTION AS A DRIVER OF ORGANIZATIONAL AGILITY IN MANAGEMENT
Srinath T. K., Chandana H. S., Dr. Sagar Manjunath et al.
International Journal of Computer Information Systems and Industrial Management Applications · 2026-07-09
This study investigates how AI adoption (AIA) affects organizational agility (OA) across industries, using PLS-SEM analysis of survey data from managerial and professional employees. Results show AI adoption has a strong positive effect on organizational agility (β = 0.708, p < 0.001), but organizational context can constrain agility even as AI adoption increases (β = −0.378, p = 0.013). The findings suggest that AI functions as a strategic capability rather than a standalone solution, requiring complementary investments in flexible structures, transformational leadership, and innovation-oriented culture. The study draws on the TOE Framework, Resource-Based View, and Dynamic Capabilities Theory to explain how AI enables sustainable competitiveness.
- Enterprise
- Workforce
- AI policy
Research
AI Adoption in S&P 500 Firms
Yang Yu, Martin Fleming, Lucy Hampton et al.
arXiv (Cornell University) · 2026-07-09
This paper tracks AI adoption among S&P 500 firms from 2016 to 2025 using SEC 10-K filings as a credible, legally accountable signal to distinguish genuine deep integration from AI hype. By 2025, 11% of S&P 500 enterprises had AI deeply integrated into their business processes and another 10% were using AI in goods production or service delivery—more than quadrupling from 5% in 2022. Technology firms account for two-thirds of deeply integrated adoption, while non-technology firm adoption is slowly accelerating. Firm profitability follows a 'J-curve' as companies move toward deep adoption, though no measurable differences in capital expenditure or productivity were observed, raising important questions about enterprise AI's broader economic impact.
- Enterprise
- Workforce
- AI policy
Research
Assuring an AI Assistant for IRB Preparation: Replication Reliability, Warrant Stability, and Evidence-Driven Revision in Institutional RAG
Jacob Holster
EdArXiv (OSF Preprints) · 2026-07-09
This paper evaluates IRB Helper, a retrieval-augmented generation (RAG) system designed to help researchers prepare Institutional Review Board (IRB) protocols using institutionally grounded policy documents. The authors conducted a rigorous stress-testing evaluation across 25 repeated runs and three system configurations, measuring retrieval reliability, warrant stability, and fabrication risk across 660 responses. Key findings show high retrieval-score reliability, improving cited-warrant overlap from .449 to .602, and no probable fabrications detected after allowlist regeneration, though systematic failures occurred in decline and referral responses—including false denials of documents present in the corpus. The study establishes replication reliability and warrant stability as formal assurance properties for researcher-built educational AI, and argues that evaluation instruments themselves require the same versioning and audit discipline as the systems they assess.
- Quality assurance
- Certifications
- AI policy
Research
A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents
A. Sayyad, J. Emmons, S. Jones et al.
arXiv · 2026-07-08
This paper evaluates the reliability of large audio-language models (LALMs), specifically models in the Gemini family, as automated judges for scoring full-duplex voice agent conversations directly from stereo audio. Tested against three calibrated human raters on 209 stereo sessions across 8 production dimensions, Gemini 2.5 Flash shows competitive agreement with human raters on the majority of dimensions, with LALM-human Spearman correlations departing from pairwise human-human correlations by at most 0.07 on 5 of 8 dimensions. The study finds that human rating alone costs roughly two orders of magnitude more than the equivalent LALM workload, suggesting strong enterprise and quality-assurance value in deploying LALMs as substitute or supplementary raters where evidence supports it. The authors caution that model swaps within the Gemini family require re-validation on calibration specifically, not just rank-correlation, and identify four areas where deployment requires care.
- Quality assurance
- Enterprise
- Workforce
Research
3100 Opinions on Code Review in an AI World: Building Causal Theory from Practitioner Discourse
Shyam Agarwal, Courtney Miller, Christian Kästner et al.
arXiv · 2026-07-08
This paper examines how AI coding agents are reshaping code review practices by synthesizing 38,709 grey-literature documents—engineering blogs and Reddit threads—into a causal theory built from 3,100 coded items, identifying 26 constructs and 67 relationships. An observational analysis of GitHub activity finds that agent-authored pull requests are reviewed less often, merged faster, and discussed less than human-authored ones, though the direction of these trends shifts under different analysis choices. The central finding is that code review becomes the critical control point determining whether AI coding agents help or harm software quality, and the outcome depends on team expertise and how the review process is structured—not on AI alone. The work matters for software quality assurance and enterprise development teams because it transforms vague claims about AI changing code review into falsifiable propositions and offers a scalable, LLM-assisted grey-literature theory-building method for future research.
- Quality assurance
- Enterprise
- Workforce
Research
When Debiasing Backfires: Counterintuitive Side Effects of Preprocessing-Based Stereotype Mitigation
Yahan Zheng, John Guerrerio, Soroush Vosoughi et al.
arXiv · 2026-07-08
This paper investigates preprocessing-based stereotype mitigation methods in NLP—such as training on debiased corpora by removing or swapping demographic references—and finds that while these approaches reduce stereotyping for targeted groups, they often cause unintended increases in stereotyping or counter-stereotyping for other demographic groups not directly targeted. The authors demonstrate these side effects across encoder-only and decoder-only model families, multiple preprocessing strategies, and varying data scales on Wikipedia, and show that standard benchmarks frequently fail to detect these shifts. Attention-rollout analysis reveals no large changes in attention flow accompanying these side effects, making mechanistic explanations difficult. The findings argue for more transparent, side-effect-aware evaluation and mitigation practices in AI fairness work.
- Quality assurance
- AI policy
Research
Agentic AI and Retrieval-Augmented Models in Straight-Through Underwriting
Robert Richardson, Josh Meyers, Brian Hartman et al.
arXiv · 2026-07-08
This paper investigates how agentic AI architectures—combining large language models, retrieval-augmented generation (RAG), and multi-agent planning—can support actuarial underwriting workflows that require transparency, auditability, and human-in-the-loop governance. The authors develop a synthetic experimental environment for straight-through underwriting of small commercial Business Owner Policies and compare three pipelines: a single-LLM baseline, a naive RAG system, and a multi-agent 'Agentic RAG' pipeline. The agentic system performs best overall, with the largest improvements in multi-step and missing-information scenarios, where structured retrieval and reflection help the model avoid unsupported straight-through decisions. These findings matter for enterprise insurance workflows and quality-assurance processes, demonstrating that structured agentic designs can better align AI decision-making with regulated, auditable standards.
- Enterprise
- Quality assurance
- AI policy
Research
False Confidence: Automated Labels Confound Fairness Audits in Cervical Spine Segmentation
Linus Juni, Aasa Feragen, Aditya Parikh
arXiv · 2026-07-08
This paper presents the first fairness audit of cervical-spine MRI segmentation across sex, age, and race using the CSpineSeg dataset, revealing that the deployed model is demographically fair but that the choice of reference label used to evaluate it is not neutral. Because many segmentation datasets supplement expensive expert-annotated ('gold') labels with cheaper machine-generated ('silver') labels, evaluating against silver labels overestimates performance by approximately 8 Dice points and can flip a fairness verdict—here turning a non-significant age disparity into a significant one. The authors identify a specific distortion mechanism they call 'false confidence,' distinct from previously reported 'false magnitude,' in which silver-label evaluation collapses within-group variance rather than merely inflating group gaps. The paper concludes that reference-label provenance is a first-order confounder in segmentation fairness audits, and that performance and fairness claims should always be reported against expert labels alongside disclosure of reference provenance.
- Quality assurance
- Certifications
- AI policy
Research
Institutional Red-Teaming: Deployment Rules, Not Just Models, Causally Shape Multi-Agent AI Safety
Yujiao Chen
arXiv · 2026-07-08
This paper introduces 'institutional red-teaming,' a methodology for evaluating how deployment rules—not just the AI models themselves—causally shape safety outcomes in multi-agent AI systems. Using IABench-CA, a benchmark spanning 228 contexts, five rule types, and 33,924 simulated games across seven model populations, the authors show that changing a single consequence rule can shift mean fatality rates by 22 to 58 percentage points. A key mechanistic finding is that naming the loss-bearer in a rule ('identity salience') drives targeted elimination of the least-resourced agent from 22% to 81% of games, and even anonymization only delays this targeting as agents re-infer hidden rules from observed outcomes. The authors propose a safety-case workflow for certifying provisional rule regions per deployment context, with explicit residual risks and monitoring obligations—directly relevant to AI policy, certification, and quality assurance.
- AI policy
- Certifications
- Quality assurance