News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI
Muayad Sayed Ali, Aliaksandra Novik, Anji Boddupally et al.
arXiv · 2026-07-08
This paper investigates how the orchestration layer ('harness') in enterprise agentic AI systems determines token usage and cost, independent of which foundation model is used. In a controlled experiment swapping only the orchestration layer across 22 tasks and six foundation models, the authors find the Writer Agent Harness cuts blended cost per task by 41%, reduces tokens per task by 38%, and cuts median wall-clock time by 44%, with task-completion quality at rough parity. Notably, efficiency gains are model-invariant (every model gets cheaper by 33–61%), while quality improvements correlate strongly with a model's baseline capability (r=0.99). The paper argues that orchestration design is a more decisive lever on AI operating costs than model selection itself, making it a critical consideration for enterprise AI deployment and spend governance.
- Enterprise
- Quality assurance
- AI policy
Research
Governance-Aware Agentic AI for Enterprise Engineering Systems: A Design-Science Reference Architecture and Quantitative Risk-Control Model
Kwan Hong Tan
Zenodo (CERN European Organization for Nuclear Research) · 2026-07-08
This paper presents a governance-aware reference architecture for deploying agentic AI in enterprise engineering systems, addressing the gap between high-level AI governance standards and concrete operational controls. Using design-science methodology, it introduces a six-layer control architecture and three formal constructs—Productivity-Adjusted Residual Risk, Governance Debt, and Human Override Threshold—to make governance principles measurable and enforceable at the system level. Illustrative scenarios across customer service, finance, HR screening, and supply planning demonstrate how the model translates abstract principles into engineering checks. The work argues that trustworthy enterprise AI requires governance built into system architecture rather than applied after deployment.
- Enterprise
- AI policy
- Quality assurance
- Certifications
Research
Developing an Algorithmic Accountability and Data Sovereignty Framework for the governance of Artificial Intelligence platforms in United States education
Oluwatayo Osodein, Oluwatayo Osodein, Ayomide Arowolo Ayodeji et al.
Discover Education · 2026-07-08
This systematic review analyzes 47 documents—including privacy policies, federal and state legislative texts, and case studies—to assess how current U.S. regulations govern AI learning platforms in K-12 education. The study finds that existing laws like FERPA and COPPA are insufficiently adapted to modern AI data processing, that 60% of principals and 25% of teachers are using AI while only 18% of districts have formal usage guidelines, and that schools frequently deploy AI without adequate district oversight. Drawing on data ethics, technological determinism, and surveillance capitalism frameworks, the authors identify risks of algorithmic bias, discriminatory disciplinary actions, and commercial exploitation of student data. The paper calls for immediate policy reforms to establish algorithmic accountability, student data sovereignty, and student-centered protections over data monetization.
- AI policy
- Certifications
- Quality assurance
Research
Governance-Aware Agentic AI for Enterprise Engineering Systems: A Design-Science Reference Architecture and Quantitative Risk-Control Model
Kwan Hong Tan
Zenodo (CERN European Organization for Nuclear Research) · 2026-07-08
This paper presents a reference architecture called the Governance-Aware Agentic AI Control Architecture, designed to help enterprises deploy autonomous AI agents responsibly in engineering systems. Using a design-science methodology, it formalizes governance requirements into six architectural layers and three quantitative constructs—Productivity-Adjusted Residual Risk, Governance Debt, and Human Override Threshold—to translate abstract AI governance standards into measurable engineering controls. An illustrative evaluation across customer service, finance, HR screening, and supply planning demonstrates how the model operationalizes oversight mechanisms such as auditing, human escalation, and reversibility. The work matters because it offers organizations a practical, theoretically grounded blueprint for embedding accountability into agentic AI systems rather than treating governance as post-hoc compliance.
- Enterprise
- Quality assurance
- AI policy
- Certifications
Research
LLMs Silently Correct African American English: Auditing and Mitigating Dialect Bias via Activation Steering
Huan Wu, Ali Emami, Muhammad Furquan Hassan et al.
arXiv · 2026-07-07
This paper audits six large instruction-tuned LLMs (14B–70B parameters) and finds they systematically rewrite African American English (AAE) into Standard American English (SAE), even when context clearly calls for AAE. The authors introduce a bias-auditing framework called conditional Dialect Group Invariance (cDGI) and identify syntactic features—especially negative concord—as universal bias triggers across all models. To counter the bias, they apply activation steering at inference time (no retraining required), reducing dialect bias 5 to 20 times more effectively than prompting alone while preserving SAE fluency. They also release REAL-AAE, the largest real-AAE parallel corpus to date, with 17,479 validated triplets, to support future fairness research.
- AI policy
- Quality assurance
Research
Digital Fragmentation and Generative AI Use Across 103 Million Application Events
Sumer S. Vaid, Ashley V. Whillans
arXiv · 2026-07-07
This large-scale observational study analyzed 103 million application events from 1,017 knowledge workers across eight organizations to understand digital fragmentation—the frequent switching between applications that consumes nearly a tenth of the work year. The researchers found that day-to-day variation within individual employees (44.6%) was the largest driver of fragmentation, exceeding stable individual differences (35.8%) and organizational differences (19.6%), with fragmentation rising over the work week and resetting after weekends and holidays. Notably, generative AI use occurred on more fragmented days, but the period immediately following AI use was characterized by narrower, longer, and more predictable application use, suggesting AI may help structure fragmented work rather than intensify it. These findings have implications for how organizations and workers manage digital work patterns and how AI tools might be leveraged to improve workforce productivity and focus.
- Workforce
- Enterprise
Research
Automated Compliance Mapping in Cloud Security with Domain-Adapted Sentence Transformers
John Bianchi, Luca Petrillo, Fabio Martinelli et al.
arXiv · 2026-07-07
This paper proposes automating the manual process of mapping cloud security controls to technical metrics by fine-tuning domain-adapted Sentence Transformer models. The researchers built a training corpus of 3,499 semantic pairs drawn from five European security standards and expanded it to up to 13,996 samples using back-translation and LLM-based paraphrasing. All five fine-tuned architectures outperformed their zero-shot baselines, with the best model gaining up to 23 nDCG@10 points on the control-to-metric task and one model reaching 0.870 nDCG@10 on cross-standard control association. The findings demonstrate that in-domain training data is a primary performance driver, suggesting this approach could significantly reduce manual compliance mapping effort in cloud security contexts.
- Quality assurance
- Certifications
Research
Large language models create an uneven informational layer over cities
Lin Chen, Guangyuan Weng, Esteban Moro
arXiv · 2026-07-07
This study audits restaurant recommendations from three major large language models across 304 neighborhoods in five U.S. cities, using 320 synthetic user profiles that vary by income, age, sex, and residential status. The researchers find that LLMs both fabricate nonexistent venues and systematically overlook real ones, with fabrication concentrated in neighborhoods with weaker digital and physical footprints, while 47.5% of real establishments are never recommended and 31.9% of these blind spots are shared across all three model families. Recommendation patterns also vary by user profile: higher-income users receive more expensive and less popular venues, while tourists are directed toward costlier but more socially diverse establishments than local residents. Simulating shifts in consumer demand suggests widespread LLM reliance would redirect visits and revenue away from chain and quick-service restaurants toward independent and full-service dining, raising concerns about urban economic inequality.
- Enterprise
- AI policy
Research
Pluralis v0.1: Towards a Multicultural, Multimodal, Multilingual Benchmark for AI Risk and Reliability
Alicia Parrish, Rajat Shinde, Sanket Badhe et al.
arXiv · 2026-07-07
Pluralis v0.1 is a new multimodal, multilingual benchmark dataset containing 6,448 prompts spanning six Asia-Pacific countries (Bangladesh, India, Korea, Pakistan, Singapore, and Taiwan) and eight languages, designed to evaluate AI safety from a culture-first perspective rather than adapting Western-centric defaults. The benchmark introduces a multimodal evaluation paradigm where text and image inputs are individually innocuous but together can trigger locale-specific legal or cultural violations, and it disentangles universal safety violations from localized cultural appropriateness. Testing Vision-Language Models on a subset of Pluralis reveals recurring locale-specific failure modes—including image misidentifications with downstream harm, missed item-context-locale interactions, and inadequate refusals—that vary systematically across locales and languages in ways that globally averaged metrics conceal. This work matters because it exposes critical blind spots in current AI safety evaluations for global deployments and establishes a foundation for advancing multicultural, multilingual AI alignment research.
- Quality assurance
- AI policy
Research
Open-Ended Scenario Reasoning for Specialist Model Adaptation
Youcheng Zong, Runda Jia, Ranmeng Lin et al.
arXiv · 2026-07-07
This paper introduces ROAM (Reasoning-Driven Open Adaptation for Specialist Models), a framework that uses large language model (LLM) reasoning to adapt frozen industrial process models to new scenarios without retraining or collecting new labeled data. Rather than using LLMs as direct predictors, ROAM fuses LLM-generated scenario judgments with online observations in a low-dimensional latent space, while a risk-constrained mechanism prevents overcorrection when LLM evidence is unreliable. Experiments on a mineral thickening process and the IndPenSim penicillin fermentation dataset show ROAM reduces mean absolute error by over 20% in major shift settings, adding only 839 parameters and under 0.02 ms per-step overhead. The results demonstrate that LLM world knowledge can be converted into a conservative, interpretable adaptation signal for deployed industrial models, reducing the need for costly retraining cycles.
- Enterprise
- Quality assurance
Research
Agents That Teach: Towards Designing Incidental Learning Back into AI-Assisted Software Development
Rohit Mehra, Samdyuti Suri, Prithviraj K Tagadinamani et al.
arXiv · 2026-07-07
This paper examines how AI coding agents, while boosting productivity, undermine the informal, effortful problem-solving through which software developers historically acquired expertise — a phenomenon the authors call 'Knowledge Debt,' a developer-level analogue of Technical Debt. The authors argue that incidental learning will not return on its own and propose six design principles to consciously reintroduce it into developer-agent interactions. They present SHIELD, a multi-agent system that uses the AI coding agent's own reasoning to surface contextual learning moments without disrupting developer workflow. The work aims to make productivity and learning complementary rather than competing outcomes in AI-assisted software development.
- Workforce
- Enterprise
Research
Prompt Coach: An Empirical Evaluation of an Agentic Tutor for Learning Prompt Engineering in Software Development
Rohit Mehra, Kapil Singi, Vikrant Kaulgud et al.
arXiv · 2026-07-07
Prompt Coach (PC) is an agentic AI tutor embedded in developers' IDEs that teaches prompt engineering through Socratic questioning, evaluating prompts across multiple quality dimensions and guiding self-correction grounded in the developer's codebase and target LLM behavior. An empirical study with 15 professional developers found statistically significant improvements in prompt quality after a single 60-minute session, with the largest gains in dimensions developers commonly overlook. Participants also reported strong trust, high adoption readiness, and unanimous agreement that PC improved their prompt-writing skills. The findings suggest that in-flow, agentic tutoring is a promising approach for closing the gap in prompt engineering education among software developers.
- Workforce
- Enterprise
Research
PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails
Mingyang Song, Luxin Xu, Haoyu Sun et al.
arXiv · 2026-07-07
PolicyShiftGuard addresses a gap in image content moderation: most guardrail systems treat safety as a fixed property of an image, but real deployments require adapting to different or evolving policies across products. The authors introduce PolicyShiftBench, a benchmark of 2,000 policy-discriminative instances across 265 images, each paired with multiple policy-conditioned prompts to test whether models genuinely follow the active policy rather than relying on static image-level priors. They then propose PolicyShiftGuard, a compact model trained with a two-stage recipe (RP-SFT and BP-Adapt) that learns to distinguish blocking from passing policies for the same image; their 7B model achieves 76.9 average F1 and 72.1 average PSS on the benchmark, outperforming existing vision-language models and specialized guardrails. This matters for enterprise content moderation systems where policy boundaries shift across products or over time, and for quality-assurance workflows that must enforce specific, context-dependent safety rules reliably.
- Quality assurance
- Enterprise
Research
Harrison.Rad 1.5 Technical Report: A radiology foundation model that can draft reports from images, priors and clinical context
Suneeta Mall, Vladimir Nekrasov, Ashnil Kumar et al.
arXiv · 2026-07-07
Harrison.Rad 1.5 (HR1.5) is a radiology-specific multimodal large language model designed to draft structured radiology reports by integrating images, clinical history, and prior studies. It is trained through a three-stage pipeline covering domain adaptation, contrastive vision-encoder training on approximately 6 million image-report instances, and visual question-answering fine-tuning. Evaluated across multiple benchmarks—including a simulated FRCR 2B Short Case examination—HR1.5 is the only system reported to meet the simulated FRCR passing standard and achieves the highest accuracy on several closed-format clinical question and reporting tasks. The work is motivated by growing imaging demand outpacing the radiology workforce, positioning AI-assisted report drafting as a direct way to reduce radiologist workload and address reporting backlogs.
- Workforce
- Quality assurance
Research
Beyond Refusal: A Same-Lineage Study of Aligned and Abliterated LLMs for Vulnerability Analysis
Mingchen Li, Meikang Qiu, Zifan Peng et al.
arXiv · 2026-07-07
This paper investigates how removing refusal behavior ('abliteration') from safety-aligned large language models affects their usefulness for software vulnerability analysis tasks such as detection, localization, and patch validation. Using a controlled same-lineage design with Gemma and Qwen model families, the authors show that abliterated models consistently outperform their aligned counterparts on concrete security metrics—for example, abliterated Gemma models achieve patch usability rates of 67.8% versus 29.9% for aligned models, and abliterated Qwen models improve line-level vulnerability localization F1 from 2.08% to 3.91%. The study argues that safety evaluations of LLM-based security tools must go beyond measuring refusal rates to jointly assess response correctness and actionability across real engineering workflows.
- Quality assurance
- Enterprise
Research
The Balkanization of Execution-Security Research for AI Coding Agents: Isolation, Access Control, and Time-of-Check-to-Time-of-Use Vulnerabilities
Mohammadreza Rashidi
arXiv · 2026-07-07
This paper systematizes 39 research papers (2023–2026) on execution-layer security for AI coding agents—covering sandbox isolation, capability and access control, TOCTOU races, Model Context Protocol threats, and related topics—into 17 categories to address the scattered, siloed state of the field. The authors identify five cross-cutting gaps: isolation and capability models are never benchmarked against each other; policy-enforcement studies show denylist failure rates of 69–98% yet isolation papers don't re-evaluate under adversarial conditions; TOCTOU and MCP threats are treated as separate problems despite sharing the same root cause; all enforcement assumes honest policy authors; and benign out-of-scope agent actions occurring at rates up to 17.1% are unaddressed. The survey also confirms four disclosed, patched CVEs directly affecting production agent harnesses. The work matters because it reveals that no existing broader security survey dedicates focused attention to execution security for AI agents, and the proposed research agenda aims to close these systemic gaps.
- Quality assurance
- AI policy
Research
Execution Governance for AI Orchestration and Agentic Systems
Ho Wa Ku
Zenodo (CERN European Organization for Nuclear Research) · 2026-07-07
This whitepaper introduces Execution Governance (EG), a reference framework designed to govern AI agentic and orchestration systems by determining whether AI-mediated actions are authorized before they produce real-world effects. The framework is structured around six conditions—Verified Mandate, Valid Constraints, Live-Context Integrity, Accountability, Reviewability, and Sufficient Verifiable Proof—and provides decision guidance (Proceed, Review, Hold, or Block) for enterprise and operational AI deployments. The paper addresses audit trails, evidence receipts, memory-governed context, and enterprise SaaS integration, positioning EG as a complementary layer to existing AI governance and risk frameworks rather than a formal standard or regulatory requirement. This matters for organizations deploying agentic AI because it offers a practical authorization and accountability structure before consequential actions occur.
- Enterprise
- Quality assurance
- Certifications
- AI policy
Research
Execution Governance for AI Orchestration and Agentic Systems
Ho Wa Ku
Zenodo (CERN European Organization for Nuclear Research) · 2026-07-07
This whitepaper introduces Execution Governance (EG), a reference framework designed to govern AI-mediated actions in agentic and orchestration systems before they produce real-world effects. It proposes six governance conditions—Verified Mandate, Valid Constraints, Live-Context Integrity, Accountability, Reviewability, and Sufficient Verifiable Proof—to determine whether an AI action should proceed, be reviewed, held, or blocked. The framework addresses enterprise integration, audit trails, memory-governed context, and resource-aware governance, positioning EG as a complementary layer alongside existing AI governance and risk frameworks rather than a replacement for compliance or regulatory schemes. This matters for organizations deploying agentic AI systems that interact with tools, APIs, and enterprise infrastructure, where authorization and accountability at the point of execution are critical governance gaps.
- Enterprise
- AI policy
- Certifications
- Quality assurance
Research
Governance, risk, and compliance frameworks for AI Security: A review of emerging standards and challenges
Abimbola Filani, Jochebed Akoto Opoku
Magna Scientia Advanced Research and Reviews · 2026-07-07
This review paper surveys emerging AI governance, risk, and compliance (GRC) frameworks—including ISO/IEC 42001, the NIST AI Risk Management Framework, the EU AI Act, and the OECD AI Principles—and finds that the current ecosystem is fragmented, with overlapping principles but inconsistent enforcement and limited operational guidance on adversarial robustness and lifecycle risk monitoring. Traditional cybersecurity standards like ISO/IEC 27001 and the NIST RMF do not fully address AI-specific risks such as adversarial attacks, model drift, and fairness concerns. The paper identifies best practices for embedding GRC into AI development pipelines and highlights critical gaps in interoperability, audit tooling, and liability regimes. The authors call for unified audit standards, empirical evaluation of governance effectiveness, and integration of AI risk metrics into ESG reporting to support trustworthy AI deployment.
- AI policy
- Enterprise
- Certifications
- Quality assurance
Research
Distributed Artificial Intelligence and Health Governance: A Multidimensional Analysis of the Tensions Between Rules, Ethics and Innovation
Fabio Liberti, Francesco Avolio, Alfonso Laudonia et al.
Zenodo (CERN European Organization for Nuclear Research) · 2026-07-07
This paper examines how Distributed Artificial Intelligence (DAI) technologies—such as Federated Learning and Edge Computing—are reshaping healthcare governance by enabling privacy-preserving, collaborative data analysis without centralized storage. The authors propose a four-axis analytical framework covering Technology, Economy, Law, and Ethics to identify and navigate tensions such as security-versus-cost and innovation-versus-regulation trade-offs in real-world healthcare scenarios. The resulting 'tension matrix' is designed to help policymakers and institutions visualize interdependencies and prioritize governance actions aligned with GDPR, the EU AI Act, and ethical principles like fairness and patient autonomy. The work is relevant to both policy development and enterprise adoption of responsible AI in digital health.
- AI policy
- Enterprise
- Quality assurance
Research
Operator-Blind Secret Mediation for AI Agents: A Formal Model and FHE Construction for Credential Derivation on Untrusted Infrastructure
Shutong Jin, Ruiyi Guo, Ray C. C. Cheung
Mathematics · 2026-07-07
CapSeal is a capability-based credential broker for AI agents that replaces direct secret access (API keys, SSH credentials) with session-bound, non-exportable handles, preventing exfiltration via prompt injection or tool misuse. The core contribution is operator-blind secret mediation: a split-broker architecture where tenant secrets are stored only as fully homomorphic encryption (FHE) ciphertexts on untrusted operator infrastructure, with per-request credential derivations evaluated without decryption. The authors formally prove computational operator blindness from IND-CPA security and implement a TFHE-rs prototype measuring the untrusted-boundary crossing cost at about 9 seconds per request—roughly 17 million times slower than plaintext—while also comparing FHE against TEE- and MPC-based alternatives. This work matters because it provides a rigorous framework for deploying AI agents on hosted infrastructure without exposing tenant secrets to cloud operators.
- Enterprise
- Quality assurance
- AI policy
Research
Impact of Artificial Intelligence (AI) Adoption on Employee Competence and Organizational Performance: The Moderating Role of Digital Leadership
Ruli Haris, Yakup Hermansyah
Jurnal ASIK Jurnal Administrasi Bisnis Ilmu Manajemen & Kependidikan · 2026-07-07
This study of 210 employees at technology and banking companies in Indonesia finds that AI adoption significantly improves both employee competence (β = 0.423) and organizational performance (β = 0.381). Digital leadership moderates the AI adoption–organizational performance relationship (β = 0.267), meaning that stronger digital leadership amplifies performance gains from AI. The findings suggest organizations should invest simultaneously in AI infrastructure and digital leadership development to maximize returns on AI-driven transformations.
- Workforce
- Enterprise
Research
Beyond Accuracy: How Humans Evaluate Legally Correct but Socially Controversial Legal Advice from Machines
Benjamin Minhao Chen, Zhiyu Li
arXiv · 2026-07-06
This preregistered survey experiment with 3,348 adults in mainland China examines how people evaluate legally correct but socially controversial legal advice attributed to either an AI system or a human lawyer, with or without accompanying reasoning. Contrary to expectations of algorithm aversion, attributing advice to an AI has no net effect on perceived reasonableness, but mediation analyses reveal two opposing pathways: AI-attributed advice is seen as more objective (boosting perceived reasonableness) yet less comprehensive and less attentive to special circumstances (reducing it). Providing legal reasoning substantially increases perceived reasonableness regardless of source, primarily by enhancing perceptions of objectivity. The findings suggest public acceptance of AI legal advisors is governed by competing normative expectations around objectivity and contextual sensitivity, with direct implications for the design of AI recommendation systems in high-stakes domains.
- AI policy
- Enterprise
Research
EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems
Kenneth Benavides, Josh Fleischer, Danti Chen
arXiv · 2026-07-06
EvalLoop is a methodology for evaluation-driven iterative improvement of large language model (LLM) systems in business contexts, moving beyond static model benchmarking toward systematic diagnosis and targeted fixing of production failures. The approach organizes evaluation into dimensional metric grouping, failure mode classification, and a structured iteration workflow. Validated through a case study on sales intelligence briefing generation, the methodology revealed that 69% of hallucination failures were prompt-induced interpretation errors invisible to aggregate scoring; a targeted prompt fix raised overall performance from 82.6% to 94.6%, with large gains in diagnosed dimensions (Content Accuracy +16.8pp, Synthesis Power +26.4pp). The framework also demonstrates that a small human review panel (4 models, 16 cases) can confirm dimensional rankings with a 94% reduction in review burden, offering practical guidance for enterprise teams deploying AI systems.
- Enterprise
- Quality assurance
Research
Whose fairness? Structural concentration in AI bias research
Abhash Shrestha, Subigya Gautam, Anu Sapkota et al.
arXiv · 2026-07-06
This paper analyzes 692 publications across five thematic domains of AI bias research, finding that the field is structurally concentrated among a small number of countries, institutions, and authors, with the United States dominating publication output and collaboration networks—especially in the foundational 'general fairness and bias mitigation' domain. Citation influence is highly skewed (median = 9; mean = 93.5), meaning a tiny fraction of publications disproportionately shapes the field's definitions, benchmarks, and debiasing frameworks. Because downstream application areas inherit their standards from this foundational domain, the concentration propagates throughout AI bias research as a whole, raising the concern that mitigation methods validated in narrow contexts may not generalize to all populations and settings. The authors provide an interactive atlas for continuous monitoring of the field's structural composition.
- AI policy
- Quality assurance