News & Research
The latest AI research and news with real-world stakes. Each item is sourced, dated and summarized in plain English, tagged by impact area where one fits, and its summary is checked against the text it was written from.
Kind
7955 items
- ResearcharXiv2026-07-03Q
Brand-as-Memory: Vision-Language Models Encode Causal, Mechanistically Localizable Credibility Priors for News Sources · Chih-Ting Liao, Xin Cao
This paper investigates how vision-language models (VLMs) encode credibility biases tied to news outlet identities when processing news as images. The authors introduce CueTrust, a benchmark measuring when source identity cues (mastheads, logos, domain names) override article content evidence, quantified via a Source-Override Index across seven VLMs. They find that outlet-identity priors are causally formed at specific model layers (19–21), correlate strongly with professional credibility ratings (rho = 0.88 with Media Bias/Fact Check), and can override content signals by roughly 1.8x — a bias that can be partially reduced (41%) by steering the localized causal direction. This matters for quality assurance of AI systems used in news reading or fact-checking contexts, as VLMs may systematically favor source reputation over actual content evidence in ways that are model- and scale-dependent.
- ResearcharXiv2026-07-03Q
Is Agentic Code Review Helpful? Mining Developers' Feedback to CodeRabbit Reviews in the Wild · Hong Yi Lin, Mingzhao Liang, Kla Tantithamthavorn et al.
This paper presents an empirical study of CodeRabbit, an autonomous AI code review agent, analyzing 31,073 code review–feedback pairs from 10,191 pull requests across 239 GitHub repositories. The results show mixed developer reception: 36.4% of agentic reviews were accepted, 7.3% triggered discussion, and 56.3% were rejected—primarily due to false positives, redundant suggestions, or misalignment with developer intent. Agentic reviews focused more on functional concerns than evolvability, yet these were more likely to be invalid. LLM-based rejection prediction methods achieved up to 76% F1 score, indicating learnable patterns exist that could help improve the effectiveness of AI-driven code review tools.
- ResearcharXiv2026-07-03Ed
Reflective Dialogue or Prompt Refinement? Effects of Tutor Scaffolding on Students' Independent LLM Use for Programming · Jerome Brender, Laila El-Hamamsy, Kim Uittenhove et al.
This study compared two LLM-based tutoring approaches in a graduate-level mobile robotics course: a Socratic-Guidance (SG) tutor that uses dialogic questioning and a Prompt-Refinement (PR) tutor that helps students craft better prompts. Across a 6-week intervention with 66 students followed by a 3-week project phase with 52 students using unconstrained LLMs, SG students achieved higher learning gains in later sessions and were more likely to adopt understanding-driven prompting strategies predictive of higher comprehension. Despite being perceived as less efficient, Socratic guidance appears to build students' capacity to learn independently with LLMs over time, offering important design implications for AI-based tutoring systems in education.
- ResearcharXiv2026-07-03QNs
Agentic and Generative AI for Open-Source Intelligence and Cyber Investigations: Taxonomy, Evaluation, Challenges, and Future Directions · Eduardo Almeida Palmieri, Mohamed Chahine Ghanem, Dipo Dunsin et al.
This survey systematically reviews 74 studies on the use of large language models (LLMs) and agentic AI systems for open-source intelligence (OSINT) and cyber investigations. It establishes agentic AI as a distinct analytical category, organizes the literature through an 11-category taxonomy, and identifies a critical 'hallucination-validation gap'—hallucination is flagged as a major concern in over twenty studies, yet is empirically measured in only one OSINT-specific system under non-reproducible conditions. The survey maps research coverage to the OSINT lifecycle, finding strong support for collection and analysis but limited coverage of verification, reporting, and decision support. It concludes that a human-AI co-pilot model—where LLMs assist collection and triage while human analysts retain responsibility for verification and decision-making—is the most defensible near-term deployment architecture, and proposes a ten-point research agenda covering evaluation, hallucination measurement, adversarial robustness, and governance.
- ResearcharXiv2026-07-03E
Organizational Memory for Agentic Business Process Execution · Lukas Kirchdorfer, Adrian Rebmann, Christian Warmuth et al.
This paper argues that LLM-based agents used to automate business processes need a centralized 'organizational memory' — a shared, governed knowledge layer containing organization-specific procedural knowledge such as policies, process models, and standard operating procedures. Without it, enterprises face knowledge silos, duplicated rules, and inconsistent updates across agents. The authors derive requirements for such a memory, propose an architecture for its curation and consumption, and validate the concept through a proof-of-concept procurement scenario. This matters for enterprises adopting AI agents at scale, as it addresses a critical gap in making multi-agent business process automation reliable and maintainable.
- ResearcharXiv2026-07-03QAd
CONTRA: Red-Teaming Configurations of Personalizable Agents · Jonathan Nöther, Adish Singla, Goran Radanovic
CONTRA is an LLM-assisted tree-search algorithm designed to red-team personalizable AI agents by discovering agent configurations that cause the execution of malicious actions without explicit instruction. Testing against 473 popular skills from a public repository, the study finds that 75.1% of skills have at least one configuration leading to malicious action execution, and CONTRA successfully identifies such a configuration in 39.2% of all tested cases. Most of these dangerous configurations were not flagged by existing security scans, demonstrating that current personalization mechanisms in autonomous agents provide insufficient safety guarantees.
- ResearcharXiv2026-07-03Q
Builder, Defender, Breaker: The Case Against Removing the Human from the AI-Driven Security Lifecycle · Mohamed Chahine Ghanem
This paper argues that full autonomy in AI-driven cybersecurity—where the same generative models build, defend, and test software—is structurally flawed rather than a natural progression. When builder, defender, and breaker roles share the same underlying model distribution, they inherit common blind spots that undermine the independence needed for meaningful verification. The authors contend that removing humans collapses the external reference point for judging machine output, eliminates timely intervention, creates predictable targets for adversaries, and erases accountability. Drawing on evidence from autonomous code generation, adversarial machine learning, software fault tolerance, and all-machine hacking tournaments, they conclude that human involvement is a permanent structural requirement and propose principles for a defensible human-machine division of labor in security.
- ResearcharXiv2026-07-03QC
Detecting Architectural Drift in Safety-Critical Firmware through Runtime Trace Analysis · Domenico Francesco De Angelis, Marco De Luca, Domenico Amalfitano et al.
This paper proposes a methodology for detecting when safety-critical firmware's actual runtime behavior diverges from its original architectural design—a problem called architectural drift—in systems that must comply with ISO 26262 automotive safety standards. The approach captures hardware-assisted execution traces, translates them into component-level message exchanges, and compares these against design-time sequence diagrams using a deterministic differencing algorithm that categorizes discrepancies as confirmed, missing, additional, or inverted. A constrained large language model then generates human-readable reports to aid expert review. Evaluation across 26 test cases shows strong agreement between automatically generated deltas and expert-curated references, with practitioners reporting the tool reduces manual analysis effort and supports safety documentation activities.
- ResearcharXiv2026-07-03Q
Flow-A11y: Flow-Aware Accessibility Testing · Nasr Eddine Fliti, Leisan Kokorina, Florian Tambon et al.
Flow-A11y is an automated accessibility testing system that evaluates web applications during real user interaction flows rather than from static page snapshots. By executing natural-language-described scenarios in a live browser, recording runtime traces, and constructing criterion-specific evidence packets, it can detect dynamic WCAG barriers such as keyboard traps, focus loss, and modal leakage that page-level scanners miss. Evaluated on 19 real public-web scenarios covering 45 dynamic WCAG criteria, Flow-A11y achieves over ten times higher oracle agreement than a generic browser-agent audit and improves fail precision from 23.5% to 41.4%. This work demonstrates a practical path toward automating dynamic WCAG criteria that have traditionally required manual inspection.
- ResearcharXiv2026-07-03Ad
Human-Centric Reflective Architecture for Human-AI Collaborative Decision-Making · Andreas Kouridakis, Dimitrios Patiniotis Spyropoulos, George Vouros
This paper introduces the Human-Centric Reflective Architecture (HCRA), a framework for human-AI collaborative decision-making using Large Language Models and reinforcement learning. It models the collaboration as a stochastic game between an AI agent and a human player, integrating human-calibrated models with RL agents that use linguistic feedback in an iterative, reflective process. The work addresses a key challenge—humans over- or under-relying on AI recommendations and AI systems being poorly calibrated to human expectations—and evaluation results show HCRA improves decision-making effectiveness and recommendation quality.
- ResearcharXiv2026-07-03QP
VISTA: Auditing Semantic Divergence in Vision-Language Models · Junchi Liao, Jiawen Deng, Fuji Ren
VISTA is a black-box auditing framework designed to detect hidden biases and inconsistencies in vision-language models (VLMs) that text-only audits cannot catch. The system couples semantic entropy with distribution-based divergence to identify cases where a model produces unusually uniform or skewed responses when images contain demographic features, corporate logos, or ideological symbols. In a controlled study, the authors confirmed VISTA can detect deliberately implanted concept-conditioned stances in fine-tuned VLMs; auditing six VLMs across 19 topics surfaced 142 high-suspicion cases (1.2%) and revealed that models refuse demographic queries at rates ranging from 0 to 65% across groups. This matters for AI quality assurance and policy because it reveals a class of model behavior—selective visual concept-conditioned divergence—that existing audit methods miss entirely.
- ResearcharXiv2026-07-03Q
VERITAS: Towards a General-Purpose Replication Tool for Scientific Research · Haokun Liu, Filbert Aurelian Tjiaranata, Chenhao Tan
VERITAS is a domain-agnostic AI framework that automates the replication of scientific research by extracting claims from papers and/or code repositories, executing the methodology, resolving issues on the fly, and scoring each claim against experimental evidence. Evaluated on 65 papers spanning computer science, social science, medicine, and astrophysics across two benchmarks (CORE-Bench and ReplicationBench), VERITAS achieves state-of-the-art performance against Claude Code baselines on every metric. The system returns an importance-weighted Replication Score, a severity-rated fix log, and a patched codebase, making independent verification faster and more scalable. This matters because AI tools are accelerating scientific publication faster than peer review can keep up, and manual replication remains slow and expensive.
- ResearcharXiv2026-07-03QHe
Where do LLMs Fall Short in CBT-Guided Affective Reasoning? · Vaishnavi Sinha, Pooja Guttal, Pranay Deep Reddy Katike et al.
This paper investigates why large language models (LLMs) fail to apply Cognitive Behavioral Therapy (CBT) principles effectively in mental health dialogue, even though they can score up to 96% accuracy on CBT licensing exam questions. The researchers built a knowledge-guided framework that decomposes user narratives using Beck's Cognitive Conceptualization structure, grounds clinical concepts in SNOMED CT validated via Natural Language Inference, and selects among three response strategies using a Multiple Chain-of-Thought (MCoT) approach. They introduce a new metric called Protocol Leverage Force (F) to measure how much a given intervention actually shifts a model away from its default behavior, finding across three open-weight LLMs and 14 case studies that even MCoT prompting only shifts behavior by roughly 1.2–1.3%, with all models remaining biased toward Validation & Reflection. The findings demonstrate that possessing CBT knowledge does not translate to effective clinical application, and the proposed metric gives researchers a concrete tool to measure this gap.
- ResearchACM AI Letters2026-07-03QCP
Proxy-Based Evaluation and the Limits of Assurance in AI Governance · Aman Sharma
This policy letter examines how AI governance frameworks rely on proxy-based evaluations—such as benchmarks, audits, and transparency documentation—that can diverge from actual real-world performance. The paper develops explicit mappings between proxies and underlying governance goals, illustrating how formally satisfied evaluation mechanisms can still miss performance degradation, hidden harms, and deployment-time failures. The authors argue that assurance claims must clearly distinguish between proxy evidence and deployment evidence to avoid overstating accountability. This work matters for AI certification and policy by clarifying the structural limits of current oversight mechanisms.
- ResearcharXiv2026-07-03EAd
Artificial Intelligence for Revenue Growth in Developing Economies: Results from a Structured Business Survey in Uganda's Hospitality and Tourism Sector · Venkatesh Andavar, Shankar Raman Rajaraman
A structured survey of 212 hospitality and tourism businesses across five Ugandan cities finds that AI adoption is associated with measurable revenue gains: 73% of AI-adopting firms recorded revenue growth, 64% improved customer retention, and 58% reduced operational costs within one year. Regression analysis confirms a statistically significant positive effect of AI use on annual revenue (p < 0.01), with dynamic pricing and AI-driven marketing showing the strongest influence. The study recommends capacity-building programmes, national AI tourism strategies, and subsidised digital infrastructure to support small and medium-sized enterprises in low-income economies.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-07-03WPEdAd
ARTIFICIAL INTELLIGENCE IN EDUCATION: TRANSFORMING TEACHING, LEARNING, AND EDUCATIONAL ADMINISTRATION · Fredfish, Blessing Godwin, Department of Early Childhood and Special Education University of Uyo, Uyo, Nigeria, Sunday, Ubong Patrick, Department of Early Childhood and Special Education University of Uyo, Uyo, Nigeria, Prof. N. A. Udofi, Department of Early Childhood and Special Education University of Uyo, Uyo, Nigeria.
This paper reviews how artificial intelligence is reshaping education across three domains: teaching methods, student learning experiences, and institutional administration. The qualitative literature review highlights benefits such as adaptive learning systems, automated grading, and predictive analytics, while also identifying risks including algorithmic bias, data privacy issues, digital inequality, and potential teacher displacement. The authors conclude that AI should complement rather than replace educators, and that responsible deployment requires ethical guidelines, policy regulation, and inclusive digital infrastructure. These findings are directly relevant to workforce concerns around educator roles and to policy frameworks governing AI use in schools.
- ResearcharXiv2026-07-03EP
Navigating the Grey Zone: Sustainable AI Governance and Leadership in SMEs · Krisztina Finta, Luca Utassy, Klaudia Gabriella Horváth et al.
This paper examines the tension SMEs in the EU face between adopting Generative AI for competitive reasons and complying with the EU AI Act. Through an integrative literature review, the authors find that resource constraints—including deficits in human capital, financial capacity, and in-house expertise—push SMEs into 'symbolic compliance,' a grey zone of AI-washing even without deliberate deception. The paper introduces the SME-RAIL framework, a three-tier governance model designed to translate Algorithmic Accountability principles into phased, resource-proportionate guidance for smaller enterprises, arguing that genuine compliance requires a shift toward human-centric AI stewardship rather than more compliance tools.
- ResearchInternational Review of Management and Marketing2026-07-03E
The Impact of Artificial Intelligence on Innovation Products in Food Manufacturing in Jordan, Mediated by Supply Chain Resilience and Supply Chain Agility · Sami Mohammad, Aysem Celebi, Abdulmula Mohamed Almahdi Arab
This study examines how AI adoption affects product innovation in Jordanian food manufacturing firms, using data from 490 managers across 11 food manufacturing subsectors analyzed via Structural Equation Modelling. Results show AI adoption has a direct positive impact on both radical and incremental product innovation, and also improves supply chain agility and resilience, which in turn further support innovation outcomes. Supply chain agility and resilience both partially mediate the relationship between AI adoption and product innovation, suggesting multiple pathways through which AI drives innovation. The findings offer empirical evidence from an emerging economy context highlighting AI-driven supply chain capabilities as critical for competitive product innovation in food manufacturing.
- ResearchFrontiers in Sustainable Food Systems2026-07-03EQC
AI capability and organizational performance in halal food supply chains: the role of compliance and health-related quality assurance · Khaliphani Ndlovu, Kejian Wang, Ramdhani Andriansyah Ahmad et al.
This study examines how AI capabilities influence organizational performance in halal food supply chains, focusing on compliance and health-related quality assurance as mediating factors. Using PLS-SEM analysis of survey data from 357 organizations in Iraq, the researchers found that AI capability improves system integration, which in turn enhances halal compliance effectiveness (HCE); HCE then positively relates to health-related quality assurance (HQA), which further connects to organizational performance. Notably, system integration alone did not directly predict organizational performance, suggesting that performance gains depend on the interplay of multiple subsystems rather than integration in isolation. The findings highlight AI's potential to add measurable value to food supply chain compliance and quality assurance processes.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-07-03WPEdAd
ARTIFICIAL INTELLIGENCE IN EDUCATION: TRANSFORMING TEACHING, LEARNING, AND EDUCATIONAL ADMINISTRATION · Fredfish, Blessing Godwin, Department of Early Childhood and Special Education University of Uyo, Uyo, Nigeria, Sunday, Ubong Patrick, Department of Early Childhood and Special Education University of Uyo, Uyo, Nigeria, Prof. N. A. Udofi, Department of Early Childhood and Special Education University of Uyo, Uyo, Nigeria.
This paper reviews how artificial intelligence is reshaping education across three domains: teaching methods, student learning, and institutional administration. It finds benefits in adaptive learning systems, automated grading, and predictive analytics, while raising ethical concerns around data privacy, algorithmic bias, digital inequality, and potential teacher displacement. The authors conclude that AI should complement rather than replace educators, and that responsible adoption requires ethical guidelines, policy regulation, and inclusive digital infrastructure. These findings are directly relevant to workforce impacts on educators, enterprise-level institutional management, and policy frameworks governing AI in education.
- ResearchOpen MIND2026-07-03PPr
REGULATION OF ARTIFICIAL INTELLIGENCE IN INDIA: A LEGAL PERSPECTIVE · Aliza Irshad
This paper examines India's current legal and regulatory approach to governing artificial intelligence, which relies on a sectoral and principle-based framework rather than comprehensive AI-specific legislation like the EU's. It identifies key challenges including privacy, accountability, discrimination, transparency, intellectual property, and cybersecurity, and proposes recommendations for a balanced legal framework that promotes innovation while protecting constitutional rights. The paper is relevant to policymakers and regulators seeking to understand gaps in India's AI governance landscape.
- ResearchJournal of Entrepreneurship & Project Management2026-07-03EP
Artificial Intelligence Adoption Among Small and Medium Enterprises in the United Kingdom: Entrepreneurial Opportunities, Project Management Implications and the Productivity Paradox · Whitney Barrington, James R. Whitmore
This systematic narrative review examines AI adoption trends among UK small and medium-sized enterprises (SMEs) between 2025 and 2026, finding that active adoption rose from 25% to 54% yet only 12% of AI-using firms report AI-attributable revenue increases and just 16% have strategic deployments with defined business purpose. Key barriers include skills gaps cited by over 60% of firms, tool fragmentation, ROI uncertainty, and governance deficits, leaving an estimated £78 billion in economic value unrealised. The authors recommend that Government AI Adoption Hubs embed formal project management frameworks alongside technical assistance, and that SMEs prioritise use-case-led deployment with clear governance, measurable KPIs, and workforce AI literacy investment.
- ResearcharXiv (Cornell University)2026-07-03QPNs
The Foreign Policy AI Evaluation Gap · Charles Pozniak, Jeba Sania
This paper argues that AI systems used in foreign policy and statecraft represent a critical but underserved test case for AI governance research. The authors identify structural properties of foreign policy—such as partial observability, unbounded action spaces, contested ground truth, and multidimensional objectives—that make standard AI evaluation methods inadequate. They review existing evaluation ecosystems and find an asymmetric focus on assessment over access, verification, security, and operationalization, and propose a demand-side evaluation framework that breaks foreign-policy workflows into bounded, evaluable sub-tasks. Given that AI is already being deployed in war and peace contexts with limited public evaluation infrastructure, the authors treat this as an urgent governance priority.
- ResearcharXiv (Cornell University)2026-07-03EQPAd
CAGE-1: Control, Assurance, and Governance Evaluation for Enterprise Agentic AI · Roopam W. Sure
CAGE-1 is an evaluation framework designed to assess whether enterprise AI agents are ready for operational deployment, going beyond simple accuracy metrics. The framework addresses governance concerns that arise when agents autonomously plan, retrieve information, call tools, and update systems—asking questions about authorization, policy enforcement, memory integrity, tool safety, auditability, and human oversight. A key innovation is 'Prebind Assurance,' which evaluates whether an agent's proposed action can be proven controlled before it becomes operationally consequential. This matters because enterprises need structured ways to verify that agentic AI systems can be safely and accountably integrated into business workflows.
- ResearcharXiv2026-07-02Q
Vision Token Manipulation Attacks on Cloud-Edge Inference of Large Vision-Language Models · Zikai Zhang, Rui Hu, Olivera Kotevska et al.
This paper identifies a new security vulnerability in cloud-edge deployments of Large Vision-Language Models (LVLMs), where intermediate vision tokens transmitted between edge devices and cloud servers can be intercepted and manipulated by an adversary. The authors propose a 'VTM-Attack' framework under a black-box man-in-the-middle setting, including four attack strategies and an optimization-based token selection method. Experiments across six state-of-the-art LVLMs and four benchmarks show that manipulating just 10% of vision tokens can reduce model accuracy by up to 88.31%, revealing a critical and previously underexplored attack surface in split-computing inference pipelines. These findings have direct implications for the security and reliability of enterprise and cloud AI deployments.