News & Research
The latest AI research and news with real-world stakes. Each item is sourced, dated and summarized in plain English, tagged by impact area where one fits, and its summary is checked against the text it was written from.
8151 items
- ResearcharXiv2026-07-02Quality assurance · AI policy
Overview of Risk Assessment and Management for Intelligent Systems under the AI Act and Beyond · Javier Irigoyen, Roberto Daza, Aythami Morales et al.
This paper surveys methodologies for assessing and managing risks in AI systems, motivated by emerging regulatory frameworks such as the EU AI Act. It maps the spectrum of AI-related risks identified in the literature—ranging from technical failures to ethical and social impacts—and reviews general risk assessment frameworks, identifying best practices and gaps that warrant further research. The work is relevant to policymakers and regulators seeking structured approaches to ensure safe and reliable AI deployment.
- ResearcharXiv2026-07-02
Synthetic Contact with AI Reduces Cross-Partisan Animosity · Benjamin Lira, Noah Castelo, Stefano Puntoni et al.
Across five preregistered studies with 3,960 U.S. partisans, this research tests whether brief AI chatbot conversations simulating cross-partisan contact can reduce political animosity. Findings show that such 'synthetic contact' lowers resistance to engagement (partisans were nearly twice as willing to interact with an AI outgroup partner as a human one), corrects factual misperceptions about opposing-party positions by over a standard deviation, and warms affective feelings toward the outgroup. Participants who conversed with an outgroup chatbot about immigration were six percentage points more likely to then choose real cross-partisan dialogue, though warmth effects largely faded within a week. The study establishes AI-mediated contact as a scalable and behaviorally consequential alternative to face-to-face cross-partisan engagement, with information exchange identified as the primary mechanism.
- ResearcharXiv2026-07-02Quality assurance
Behind the Refusal: Determining Guardrail Activation via Behavioral Monitoring · William Hackett, Peter Garraghan
This paper introduces the first black-box methodology for detecting whether an AI system's refusal to respond is caused by an external guardrail or the LLM's own built-in safety alignment, using only observable HTTP, lexical, and timing signals. By monitoring behavioral patterns during interactions, the approach detects guardrail presence with 100% accuracy and distinguishes guardrail blocks from LLM rejections with an average F1 score of 98% on unseen prompts. The work matters because knowing which defense layer is active fundamentally changes which bypass techniques are appropriate, providing critical intelligence for adversarial testing and security research. These findings have direct implications for how organizations assess and validate the robustness of AI safety systems deployed in production.
- ResearcharXiv2026-07-02Quality assurance
Prompt Coverage Adequacy · Florian Tambon, Michael Konstantinou, Cedric Richter et al.
This paper introduces Prompt Coverage Adequacy, a novel testing criterion designed for software developed through large language models (LLMs) and autonomous agents, where prompts rather than code are the primary development artifacts. Instead of measuring how thoroughly tests exercise lines or branches of code, it measures how well a test suite satisfies the requirements expressed in a prompt by leveraging LLM attention mechanisms. Evaluated across two datasets and multiple LLMs, the approach uncovers over 30% more faults than traditional code coverage when used to guide test generation. The findings suggest Prompt Coverage Adequacy could become a foundational metric for quality assurance in the emerging paradigm of LLM-driven software development.
- ResearcharXiv2026-07-02Quality assurance
SPLIT: Cross-Lingual Empathy and Cultural Grounding in English and Ukrainian LLM Responses · Anna Chorna
This paper introduces SPLIT, a 500-prompt benchmark designed to evaluate how well large language models (LLMs) provide emotionally grounded, culturally appropriate responses in English and Ukrainian across five crisis-related categories: Stress, Panic, Loneliness, Internal Displacement, and Tension. The authors evaluate three LLMs—Gemini-2.5-Flash, LLaMA-3.3-70B-Instruct, and DeepSeek-V3—on Empathetic Accuracy, Linguistic Naturalness, and Contextual & Cultural Grounding, finding that Gemini-2.5-Flash and LLaMA-3.3-70B-Instruct degrade in Ukrainian while DeepSeek-V3 remains comparatively stable. The study also finds that human and AI evaluators agree weakly on empathy and naturalness but diverge on cultural grounding, and argues that generating Ukrainian text is not equivalent to providing Ukrainian emotional support. These findings highlight significant gaps in LLM deployment for emotional support in low-to-mid-resource language contexts and call for more culturally tailored benchmarks and human-centered evaluation.
- ResearcharXiv2026-07-02Quality assurance · Privacy & Data Protection · +2
Assessing VLM Reliability for Medical Image Quality Evaluation Under Corruption and Bias · Sofiane Ouaari, Kevin Vorwalder, Nico Pfeifer
This paper benchmarks 16 vision-language models (VLMs) on medical image quality assessment (MIQA) under real-world degradation and contextual bias, using the MediMeta-C dataset zero-shot across seven corruption types, five severity levels, and seven imaging modalities. Results show pixelation caused the largest performance drops (mean -20.58%, up to -34.4% for OCT), while textual metadata such as institutional prestige or equipment age shifted scores by as much as +17.15% and -14.7% respectively, revealing that current VLMs lack objectivity and are sensitive to contextual cues that should not influence image quality judgments. Individual models showed extreme swings (up to +95.62% for InternVL-8B and -37.7% for MedGemma), and same-family models correlated at only 0.67–0.83, indicating inconsistency even within model families. These findings raise concerns about deploying VLMs for clinical MIQA without further safeguards, particularly given the tension between privacy-preserving transformations like pixelation and model reliability, and the risk of demographic or institutional metadata introducing bias into quality assessments.
- ResearcharXiv2026-07-02Quality assurance · Education
AIriskEval-edu: New Dataset for Risk Assessment in AI-mediated K-12 Educational Explanations · Javier Irigoyen, Roberto Daza, Francisco Jurado et al.
This paper presents AIriskEval-edu-db2, a dataset of 1,639 instructional explanations drawn from 170 K-12 science, language arts, and social science questions, designed to train and evaluate LLM-based auditors for pedagogical risk assessment. Each question includes a human-teacher explanation and 11 LLM-generated explanations tied to distinct risk profiles, all assessed against a five-dimension rubric covering factual precision, depth, relevance, age-appropriateness, and ideological bias. A subset of 785 explanations includes structured explainability annotations with risk localization and descriptions, validated by expert teachers. Experiments compare proprietary frontier models against a fine-tuned local Llama 3.1 8B model, testing whether supervised fine-tuning can match stronger models while preserving privacy in educational auditing.
- ResearcharXiv2026-07-02Enterprise · Quality assurance · +1
From Battlefield to Boardroom: Strategic Red Teaming as an Epistemic Governance Instrument in the Age of AI · Jeroen Janssen
This technical report reframes 'strategic red teaming' as a board-level AI governance discipline focused on testing the assumptions underlying AI-enabled strategic decisions before those assumptions become operational risks. The authors propose a six-component model encompassing an assumption register, adversarial mandate, independence criteria, evidence grading, a board-facing decision record, and a follow-up mechanism for unresolved findings. The work is conceptual and design-oriented, explicitly not claiming empirical validation or regulatory endorsement, but offering a candidate governance framework to connect AI strategy, accountability, oversight, and evidence. It matters because it provides organizations with a structured method to make strategic uncertainty about AI systems inspectable at the governance level rather than discovering failures after deployment.
- ResearcharXiv2026-07-02Quality assurance
Has This Checkpoint Been Abliterated? A Two-Signal Audit and Its Failure Map · Gabriel Hurtado
This paper presents a two-signal method for auditing open-weight AI checkpoints to detect whether their refusal mechanisms have been removed (a process called 'abliteration') before deployment. The approach combines a reference-anchored activation refusal-gap with a weight-recovery energy metric derived from the difference between a base model and the candidate checkpoint, achieving an AUROC of 0.95 on a 273-checkpoint registry spanning Qwen, DeepSeek-distilled Qwen, Llama, and Gemma models — significantly outperforming either signal alone. The audit correctly identifies 53 of 57 public abliterations while maintaining a false positive rate of 0.11, and transfers to held-out model families at balanced accuracy 0.89. The authors also map two key failure modes — reference spoofing and adversarial fine-tuning past the threshold — framing the method as effective pre-deployment triage rather than tamper-proof security.
- ResearcharXiv2026-07-02Enterprise · Quality assurance
CLAP: Closed-Loop Training, Evaluation, and Release Control for Domain Agent Post-training · Fangfei Li, Chenyang Zhao, Long Wang et al.
CLAP (Closed-Loop Agent Post-training) is a framework designed to address key challenges in deploying domain-specific AI agents on noisy business data, including unreliable post-training gains, mismatches between offline and live performance, and risks from releasing updated model adapters. The method integrates data validation, reward diagnostics, offline evaluation gates, and application-chain replay into a unified loop that determines whether a fine-tuned adapter is safe to release. Tested on five anonymized manufacturing-scenario batches using QLoRA-style LoRA-SFT, the framework shows modest average improvements in overall score, pass rate, and evidence accuracy, but also reveals that only 3 of 5 batches improve, some regress, and certain training methods (GRPO) introduce high KL risks. The findings argue that domain-agent post-training should be governed by an integrated data-training-evaluation-release pipeline rather than relying on training completion or a single offline metric.
- ResearcharXiv2026-07-02
Decoupling Code Complexity from Newcomer Participation: A Causal Study of AI Coding Agent Adoption in OSS · Weiwei Xu, Xuanning Cui, Hengzhi Ye et al.
This causal study examines whether AI coding agents (e.g., Cursor, Claude Code) crowd out newcomers in open-source software projects. Using difference-in-differences analysis across 1,888 GitHub projects that adopted an agent, the researchers find no significant decline in newcomer inflow, onboarding, or retention after adoption. While adoption does modestly increase code complexity (roughly +11% on a cognitive metric for Python and +3–4% on cyclomatic metrics across languages), this complexity rise does not translate into reduced newcomer participation. The findings suggest that the feared trade-off between AI coding assistance and human newcomer involvement does not materialize in established open-source projects.
- ResearcharXiv2026-07-02Quality assurance · Safety & Harms
Safety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded Verification · Yunhao Feng, Ruixiao Lin, Ming Wen et al.
This paper introduces Vera, an automated end-to-end safety testing framework for large language model (LLM) agents that operate autonomously through external tools. Vera uses a three-stage pipeline: it discovers emerging safety risks from literature, composes executable safety test cases from taxonomies of risks and attack methods, and runs agents in isolated sandboxes where outcomes are verified from observable environment state rather than model self-report. Evaluated on four production agent frameworks, Vera found average attack success rates of 93.9% under multi-channel attacks, exposing substantial safety weaknesses. The authors also release Vera-Bench, a benchmark of 1,600 executable safety cases across 124 risk categories, arguing that modular, scalable testing infrastructure is essential for keeping pace with rapidly evolving agentic AI systems.
- ResearcharXiv2026-07-02Enterprise · Quality assurance
Meta-Benchmarks for Financial-Services LLM Evaluation · Blair Hudson
This paper presents a meta-benchmarking framework designed to evaluate large language models (LLMs) specifically for financial-services work, rather than relying on general-purpose leaderboards. It organises 452 publicly reported benchmarks into 41 O*NET Generalized Work Activities and 38 BIAN banking business domains—covering sales, operations, risk, and support—then applies a multiplicative weighting scheme (discrimination × coverage × recency) to suppress saturated legacy benchmarks and reward those still differentiating top models. A pairwise Elo tournament scaled by these weights produces cross-benchmark-comparable scores at both the work-activity and business-domain level. The framework is demonstrated on a snapshot of 288 models from 25 organisations as of June 2026, and is presented with full methodology to support reproducibility for institutions facing model selection and governance challenges.
- ResearcharXiv2026-07-02Quality assurance
Beyond Pixel Diffs: Benchmarking Image Change Captioning for Web UI Visual Regression Testing · Licheng Zhang, Bach Le, Pengtao Zhao et al.
This paper introduces WUICC (Web UI Image Change Captioning), a new task and benchmark dataset for automatically generating natural-language descriptions of visual changes in user interface screenshots during visual regression testing (VRT). The authors evaluate eleven image difference captioning methods and two zero-shot large language models, finding that current methods struggle with web UI-specific challenges like layout diversity, dense text, and fine-grained changes, but already outperform pixel-level comparison by more selectively suppressing non-meaningful visual noise. The work addresses a gap in quality assurance tooling where developers must manually review large volumes of false positives flagged by semantically blind pixel-diff approaches. By enabling text descriptions of what changed rather than just flagging a difference, the benchmark provides a foundation for reducing manual review burden in software release pipelines.
- ResearcharXiv (Cornell University)2026-07-02Enterprise
AI Writes Faster Than Humans Can Review: A Longitudinal Study of an Enterprise 2x Mandate · H. He, Shyam Agarwal, Yegor Denisov-Blanch et al.
A longitudinal study of 802 developers and 196,212 pull requests at a mid-sized enterprise found that a company-wide AI coding mandate helped per-capita pull request throughput reach 2.09x the pre-mandate baseline by April 2026, among the largest productivity gains reported from a real-world AI coding deployment. A difference-in-differences design links this gain to AI adoption and accumulated usage, with the mandate acting as a catalyst rather than a direct driver. Critically, the surge in code output restructured review workflows: per-reviewer load roughly doubled and automated review overtook human review, raising concerns about whether humans can keep pace with AI-generated code. Merge and revert rates held steady, but the study cannot fully separate quality implications across model generations or causally attribute gains due to non-random adoption.
- ResearchProblems and Perspectives in Management2026-07-02Enterprise
AI adoption and perceived organizational performance in Chinese pharmaceutical sales: Efficiency and motivation pathways · Yunlu Cai, Siti Rohaida Mohamed Zainal
This study of 335 pharmaceutical sales professionals in China finds that AI adoption is positively associated with perceived organizational performance, both directly and through two parallel pathways: operational efficiency and employee motivation. Using PLS-SEM analysis, the results show that AI adoption significantly predicts operational efficiency and employee motivation, which in turn predict performance, confirming partial mediation through both pathways. The findings suggest that AI generates performance value in compliance-sensitive sales environments when integrated into workflows alongside motivational support for employees.
- ResearcharXiv (Cornell University)2026-07-02AI policy
Taxing Artificial Intelligence · Juliette Faivre, Sarah H. Cen
This paper examines whether taxation can serve as an effective policy tool to address harms associated with AI development and deployment, including environmental pressures, labor and creative displacement, and systemic risks. The authors survey a range of potential tax instruments—corporate income taxes, consumption taxes on AI services, and excise taxes tied to specific AI activities—and evaluate their feasibility, measurement challenges, incidence, leakage effects, and innovation costs. The paper argues that AI taxation should go beyond simple Pigouvian correction to also redistribute unevenly borne costs and gains and fund regulatory capacity, but cautions that policy design must be carefully matched to specific harms and objectives.
- ResearchVUNOKLANG MULTIDISCIPLINARY JOURNAL OF SCIENCE AND TECHNOLOGY EDUCATION2026-07-02Enterprise · AI policy
<b>BRIDGING THE DIGITAL TRANSFORMATION GAP IN NIGERIA'S CONSTRUCTION INDUSTRY: CHALLENGES FROM AI ADOPTION</b> · Henry Asuquo Okpo, Numbere Jamaibi Thanks
This systematic review of 20 peer-reviewed studies examines why Nigeria's construction industry has failed to adopt AI and machine learning tools despite their global impact. Using the Technology-Organisation-Environment framework and a PRISMA-compliant methodology, the authors identify five barrier clusters: AI-specific data and infrastructure constraints, human capital deficits, financial limitations, regulatory and governance voids, and organisational inertia. The paper proposes a multi-level intervention framework and offers evidence-based recommendations for policy reform and industry strategy in Nigeria's built environment, with broader relevance for AI adoption in emerging economies.
- ResearchSystems2026-07-02Enterprise
The Impact of Artificial Intelligence Application on Corporate ESG Performance: Evidence from Chinese A-Share Listed Firms · Haixia Feng, Renbo Shi, Qingjin Wang
Using panel data from Chinese A-share listed firms spanning 2009 to 2025, this study finds that AI adoption significantly improves corporate ESG performance, with firms showing higher AI usage achieving better environmental, social, and governance outcomes. The positive effect is strongest among firms with lower technological intensity, non-heavily polluting industries, and regions with stricter environmental regulation. Mediation analysis identifies human capital upgrading and green technological innovation as key mechanisms through which AI drives ESG gains. These findings have direct implications for corporate sustainability strategy and the role of AI-enabled workforce and innovation capabilities in meeting ESG goals.
- ResearchEkoist Journal of Econometrics and Statistics2026-07-02Workforce
The Impact of Artificial Intelligence on Labor Force and Gini Coefficient: An Assessment of some EU/OECD Countries · Hacı Bayram İrhan
This study examines how artificial intelligence adoption affects income inequality (measured by the Gini coefficient) and labor market outcomes across EU/OECD countries from 2010 to 2023. Using multiple analytical methods, the empirical analysis finds strong evidence that AI adoption has positive effects on the labor market in these countries. The paper contributes to filling gaps in the literature on AI-driven workforce transformation and income distribution effects.
- ResearchBig Data & Society2026-07-02AI policy
A question of style? Regulating artificial intelligence in the European Union and the USA · Alison Harcourt, Claudio M. Radaelli, Philipp Trein
This comparative policy study examines how the EU and USA regulate artificial intelligence, finding that the EU's 2024 AI Act offers a comprehensive but potentially rigid framework while the US takes a piecemeal, sector-specific approach that is more enforceable in practice. The authors argue that the EU's top-down architecture makes adaptive regulation difficult and raises concerns about enforceability and human rights protection, whereas the US state-by-state and sector-by-sector model allows for greater learning, imitation, and diffusion through 'laboratory federalism.' The paper uses the concepts of policy style and regulatory style to explain these divergent governance approaches and their trade-offs for AI oversight.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-07-02Quality assurance · AI policy · +1
Assured Automated Peer Review: Proof-Carrying Papers and the Trust Architecture AI Review Requires · Robert Tilley Jr.
This paper proposes a trust architecture for AI-assisted peer review systems, motivated by real-world deployments such as Google's Paper Assistant Tool and the AAAI-26 AI Review Pilot. The authors identify five interacting assurance problems—adversarial submissions, common-cause reviewer error, undisclosed tool use, uncertain reviewer performance, and incomplete accountability—and propose corresponding controls including verification contracts, provenance tracking, reviewer dependence measurement, sequential outcome monitoring, and signed decision receipts. The work is an architectural and evaluation proposal rather than an empirical validation, aiming to establish accountability and trustworthiness standards for automated review in live publication workflows. It matters because it addresses how scientific publishing can maintain integrity as AI tools become embedded in peer review processes.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-07-02Quality assurance · Algorithms & Automated Decisions
Assured Automated Peer Review: Proof-Carrying Papers and the Trust Architecture AI Review Requires · Robert Tilley Jr.
This paper proposes a trust architecture for automated AI-assisted peer review systems, addressing assurance problems such as adversarial submissions, undisclosed tool use, and incomplete accountability. The architecture centers on five controls: verification contracts, layered provenance controls, measurement of reviewer dependence, anytime-valid sequential monitoring, and signed decision receipts. The work is motivated by real deployments like Google's Paper Assistant Tool and the AAAI-26 AI Review Pilot, and aims to provide structured accountability without overclaiming unrestricted verification. The contribution is an architectural and evaluation framework rather than an empirical validation of a deployed system.
- ResearchProblems and Perspectives in Management2026-07-02Enterprise
Entrepreneurial orientation, dynamic capabilities, and startup performance: The amplifying role of AI · Nguyen Kien Quoc, Nguyen Ngoc Long
This study of 315 Vietnamese startup founders and executives finds that entrepreneurial orientation and dynamic capabilities both drive startup performance, largely through the mediating role of innovation capability. AI adoption intensity significantly amplifies all three relationships—between entrepreneurial orientation, innovation capability, and dynamic capabilities and performance—with the strongest amplifying effect observed for dynamic capabilities. The model explains 66.5% of variance in startup performance, highlighting AI as a critical enabler of capability-conversion in resource-constrained, emerging-economy contexts. The findings offer practical guidance for startup managers on building competitive advantage through combined capability development and AI adoption.
- ResearchAdvances in computational intelligence and robotics book series2026-07-02Workforce · Enterprise · +1
Agentic Artificial Intelligence-Powered Automation for Dynamic Labor Markets and Business · Sneh Sharma, Sachin Sharma
This paper presents a conceptual framework for deploying Agentic Artificial Intelligence (AAI) in labor markets and enterprise settings, distinguishing AAI from traditional AI by its autonomy, self-reflection, goal-directed behavior, and situational awareness. The authors argue that AAI can automate key workforce decisions—including recruitment, task allocation, and upskilling—while enabling more decentralized and adaptive business models suited to gig economy structures and real-time operations. The chapter also addresses ethical governance frameworks for AI-driven HR systems, suggesting that responsible deployment is central to the proposed approach. This matters because it outlines both the opportunities and governance challenges of AI-powered workforce automation in rapidly shifting labor ecosystems.