News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
Quantifying Systemic Risk from Correlated Models: A Methodological Review
Yuanyuan Li, Michael von Gablenz
Zenodo (CERN European Organization for Nuclear Research) · 2026-08-16
This paper examines how deploying multiple AI systems that share training data, architectures, or foundation models can produce correlated errors and synchronized failures that standard single-model evaluations miss. The authors survey methods for measuring behavioral similarity across models, propose a risk-oriented evaluation framework with desirable statistical properties, and identify common drivers of correlated behavior. They argue that effective AI governance requires shifting from isolated model validation to portfolio-level auditing and dependency-aware risk management aligned with emerging regulatory frameworks.
- Quality assurance
- AI policy
Research
AI-Generated Evidence And Judicial Decision-Making In India: Constitutional Limits Of Admissibility, Reliability, Human Oversight, And The Role Of Artificial Intelligence In Judicial Discretion
Dr. Prashant Yadav
Zenodo (CERN European Organization for Nuclear Research) · 2026-08-16
This article examines how AI-generated evidence—including facial-recognition outputs, algorithmic analytics, and synthetic media—is being introduced into Indian courts under the Bharatiya Sakshya Adhiniyam, 2023, which admits such material as electronic records upon certification but sets no reliability or explainability standards. The authors argue that the black-box nature of AI systems, combined with risks of bias, hallucination, and deepfake manipulation, threatens constitutional guarantees under Articles 14 and 21 and the right to a fair trial. The paper proposes a framework of mandatory disclosure, independent expert validation, and a human-in-the-loop requirement to ensure that AI may assist but never replace judicial discretion in Indian adjudication.
- AI policy
- Quality assurance
Research
GALENA: A Governance-Aware LLM Enterprise Navigation Architecture for Autonomous Multi-Agent Workflow Automation with Compliance Enforcement
Narasimha Rao Boinapalli
arXiv · 2026-08-16
GALENA is a multi-agent LLM orchestration framework that embeds regulatory compliance (GDPR, HIPAA, and domain-specific rules) as a formal architectural constraint evaluated before every agent action, rather than as a post-hoc filter. Across three enterprise task categories—Regulated Document Processing, Financial Workflow Automation, and IT Service Management—the system achieves 97.0% task completion accuracy, outperforming the strongest baseline by 18.3%, while reducing governance violation rates by 72.7% to a median of 0.03 with a median latency of 164 ms. The paper argues that treating governance as a first-class invariant, including drift-resilient lifecycle management and role-aware agent routing, is both feasible and necessary for production-grade enterprise automation. These results matter for enterprises seeking to deploy autonomous AI workflows without sacrificing regulatory compliance.
- Enterprise
- AI policy
Research
AI-Generated Evidence And Judicial Decision-Making In India: Constitutional Limits Of Admissibility, Reliability, Human Oversight, And The Role Of Artificial Intelligence In Judicial Discretion
Dr. Prashant Yadav
Zenodo (CERN European Organization for Nuclear Research) · 2026-08-16
This article examines the constitutional and procedural challenges posed by AI-generated evidence—including facial-recognition outputs, algorithmic analytics, and synthetic media—in Indian courts. It argues that while the Bharatiya Sakshya Adhiniyam, 2023 treats such material as admissible electronic records upon certification, the statute lacks reliability standards, explainability requirements, or scrutiny of generative processes, threatening due process under Articles 14 and 21. The authors contend that the black-box nature of AI, combined with risks of bias, hallucination, and deepfake manipulation, makes rigorous human oversight essential, and that AI may assist but cannot replace judicial discretion. The article proposes a framework of mandatory disclosure, independent expert validation, and an explicit human-in-the-loop requirement to safeguard the integrity of Indian adjudication.
- AI policy
- Quality assurance
Research
Nurse educators' experiences and perceptions using generative artificial intelligence: a systematic review
Ani Henttonen, Maria Christidis, Helena Kullenberg et al.
BMC Medical Education · 2026-08-15
This systematic review of 13 studies (3,082 participants) examines nurse educators' experiences with generative AI in teaching, finding a tension between optimism about pedagogical efficiency and concerns over academic integrity, critical thinking erosion, and professional role loss. Educators' confidence and use were shaped by institutional position, organizational policy, and prior GenAI experience. The review concludes that effective integration requires structured training, competency development, and clear governance frameworks. These findings highlight the urgent need for policy and capacity-building support within nursing education institutions.
- Workforce
- AI policy
Research
Insurance as AI Risk Infrastructure: A Generative-Agent Simulation of AI Adoption
Yixuan Yuan, Dedai Wei, Chudong Qian et al.
arXiv (Cornell University) · 2026-08-15
This paper proposes using insurance as a financial risk-transfer mechanism to address enterprise hesitancy around AI adoption, arguing that technical safeguards alone cannot protect firms from residual financial losses caused by AI system failures. The authors build an LLM-driven agent-based social simulation to model how insurance frameworks affect firm-level behavior, finding that insurance reduces financial exposure, accelerates aggregate AI adoption, and improves firm solvency and capital. The work frames insurance as complementary infrastructure to existing AI safety measures, with implications for how businesses and policymakers approach AI integration risk.
- Enterprise
- AI policy
Research
IA generativa e LGPD: o desafio da transparência algorítmica
Marise Miglioli Lorusso
Derecho y cambio social. · 2026-08-15
This article examines the legal and technical obstacles to implementing Brazil's General Data Protection Law (LGPD), specifically the right to review automated decisions (Article 20), in the context of Generative AI systems that operate as 'black boxes.' Drawing on case studies involving Meta, X (Grok), the Smart Sampa surveillance program, AI use in the judiciary, and the platformization of public education, the study finds a deep structural conflict between the legal requirement for transparency and the inherent opacity of Large Language Models. The authors conclude that reactive enforcement models or direct transplantation of Global North regulations are insufficient for Brazil's socioeconomic context, and recommend regulatory sandboxes, Algorithmic Impact Assessments, and mandatory Privacy by Design as more appropriate governance approaches.
- AI policy
- Enterprise
Research
Working with semantic machines: The role of semantic reconciliation in public sector data governance
Jonny Holmström
Government Information Quarterly · 2026-08-15
This qualitative case study of a large Swedish municipality examines an AI interviewing system used in early-stage recruitment, which translated candidates' open-ended responses into competency indicators and narrative summaries. The authors develop a 'semantic reconciliation' process model identifying five recurring forms of organizational work needed to make AI outputs meaningful across a decentralized public sector hierarchy—including mapping vendor categories to local vocabularies and integrating AI reports with human judgment. The study shows that while AI tools can make evaluative information portable, continuous organizational work is required to restore contextual adequacy and preserve human decision authority. The findings offer practical guidance for municipalities seeking to adopt common AI tools without letting standardized outputs displace role-specific knowledge.
- Workforce
- AI policy
Research
Artificial Intelligence in Nursing Education: A Scoping Review of Academic Perspectives
Natasha Hawkins, Anthea Fagan, Yumiko Coffey et al.
Journal of Advanced Nursing · 2026-08-15
This scoping review of 15 studies across 8 countries (2,004 nursing academics) finds an 'adoption paradox' in AI integration in nursing education: while most academics believe AI will revolutionize the field, implementation remains conservative, with two-thirds of applications operating only at an augmentation level and none achieving transformative redefinition. Key barriers include knowledge gaps, institutional policy vacuums, and global access inequities, with academics expressing concern about critical thinking erosion and threats to professional identity. The findings point to a need for faculty development, institutional policy frameworks, and curriculum strategies that balance technological advancement with person-centred nursing values.
- Workforce
- AI policy
Research
The Complementarity Paradox: Human-AI Collaboration, Overreliance, and Skill Decline in Entrepreneurial Contexts
Srinivas Subramanya, Satish Krishnan, Nasreen Azad
Information Systems Frontiers · 2026-08-15
This paper examines how knowledge workers with strong entrepreneurial orientations adopt generative AI tools deeply, which can lead to over-reliance and perceived erosion of professional skills—a pattern the authors call 'hustle-to-handicap.' Drawing on cognitive offloading theory and skill decay literature, the researchers theorize and find support (via survey data from 143 knowledge workers across five countries) for a sequential mechanism: entrepreneurial orientation drives deeper AI collaboration, which fosters AI dependence, which in turn correlates with perceived skill decline. A long-term orientation amplifies this effect because future-focused individuals reframe routine AI use as strategic investment rather than a threat to skill maintenance. The findings contribute to debates on human-AI collaboration by showing how individual mindset and temporal outlook shape whether AI augments or erodes capability.
- Workforce
- Enterprise
Research
Digital Development and Participation in EU Digital Governance: Evidence from Central and Eastern Europe
Xiaoqing Wang
Journal of International Relations and Foreign Policy · 2026-08-15
This article investigates how levels of digital development among Central and Eastern European (CEE) EU member states affect their ability to participate in EU digital governance processes, using the Digital Economy and Society Index 2020 as a baseline. Examining the Digital Markets Act and the Artificial Intelligence Act, it finds that CEE countries remain relatively peripheral during agenda-setting and legislative negotiation, while implementation capacity varies across the region. The study argues that common EU digital rules do not produce equal participation, meaning digital development is both a policy outcome and a structural condition shaping governance influence. This matters for policy because it highlights systemic inequalities in how EU AI and digital regulations are shaped and adopted across member states.
- AI policy
Research
From data to decisions: how neural network input structures propagate to air quality policies
Laura Zecchi, Michele Francesco Arrighini, Claudio Marchesi et al.
npj Clean Air · 2026-08-15
This study examines how different artificial neural network (ANN) architectures used as surrogate models in air-quality planning affect both predictive accuracy and the policy recommendations that emerge from optimization. The researchers tested multiple ANN designs varying in spatial aggregation for the Po Valley basin and found that regionalized ANNs achieve relative mean absolute errors of 2.4–3.3%, outperforming a single basin-wide ANN at 6.4%. At higher cost levels, these architectural differences propagate into divergent optimal emission-reduction policy portfolios, affecting how resources are allocated across sectors like domestic heating, transport, and agriculture. The findings demonstrate that surrogate model design choices meaningfully shape air quality policy outcomes, not just prediction quality.
- AI policy
- Quality assurance
Research
Brain Capital Management: A Firm-Level Theory of Cognitive Capability, Its Three Constraints, and an Agenda for Its Measurement in the Age of AI
Naoki Kadowaki
Zenodo (CERN European Organization for Nuclear Research) · 2026-08-15
This working paper introduces a firm-level theory of 'enterprise brain capital,' defining it as the cognitive capability a company can actually deploy, shaped by three constraints: Belonging (whether cognitive capacity can be expressed organizationally), Base (whether strain consumes available capacity), and Build (how capability accumulates or depreciates over time). The framework develops fourteen falsifiable propositions and a three-tier measurement architecture, while distinguishing genuine capability accumulation from AI-assisted productivity gains—arguing the two do not necessarily coincide. A key finding is a measurement asymmetry in existing human-capital disclosure regimes, which report inputs and costs but largely omit validated measures of workforce cognitive and psychological states. The paper positions Brain Capital Management as a diagnostic and research agenda requiring further empirical validation rather than a finished framework.
- Enterprise
- Workforce
Research
From Hype to Evidence: Evaluating LLM Reliability in Supply Chain Management
Gökhan Cenk, Tobias Engel, Jonathan Kreßel et al.
Journal of the Association for Information Systems · 2026-08-15
This paper investigates whether Large Language Models (LLMs) can reliably support supply chain management (SCM) tasks such as forecasting and automated decision-making. Using repeated forecasting trials with agentic LLM orchestration, CrewAI, and Retrieval Augmented Generation (RAG), the authors find that LLM-generated forecasts do not outperform traditional algorithmic approaches, highlighting a gap between LLM hype and domain-specific performance. The paper calls for a dedicated SCM benchmarking framework covering accuracy, consistency, contextual fit, and cost efficiency, and specifically addresses how small and medium-sized enterprises (SMEs) can build evaluation capabilities for responsible AI adoption in compliance with data sovereignty requirements.
- Enterprise
- Quality assurance
- AI policy
Research
Whose doctor does the AI recommend? An algorithm audit of reputation and demographic signals in large language model-assisted physician choice
Syeda Anshrah Gillani, Mirza Samad Ahmed Baig
arXiv (Cornell University) · 2026-08-14
This paper audits seven large language models acting as physician recommendation assistants, testing whether demographic signals (gender, ethnicity via names) and reputation attributes (ratings, fees) causally influence which doctors the AI recommends. Using 40,068 randomized choice sets, the study finds that reputation signals dominate recommendations—higher ratings and lower fees strongly drive selections—but also detects statistically significant demographic biases favoring female- and minority-signaled names over White-signaled names, effects the models never mention in their own explanations. The findings reveal that LLM self-reported reasoning is an unreliable transparency mechanism, since demographic tilts appeared in fewer than 0.03% of stated reasons, and the authors argue that repeatable behavioral audits are necessary for meaningful AI accountability in high-stakes intermediary roles.
- AI policy
- Quality assurance
Research
Engineering Reliable Coding Agents: Evaluating and Operating the System Around the Model
Stephanie Jarmak
arXiv (Cornell University) · 2026-08-14
This monograph argues that AI coding agents are commonly evaluated as isolated models but deployed as complex systems, and that reliability depends on the broader infrastructure—including execution environments, retrieval, memory, permissions, and observability—not just model capability alone. Drawing on a structured multivocal review of 164 scholarly works, 100 practitioner records, 29 benchmarks, and 17 author-system case records, the authors find that many apparent model failures originate elsewhere in the system, and that improvements at one layer often fail to propagate to end-to-end outcomes. The paper contributes a versioned catalog of 206 reliability records, runnable evaluation protocols, and a dependency-and-repair framework to help practitioners distinguish model capability from infrastructure effects and build systems that recover safely when components fail. This work is directly relevant to enterprise and quality-assurance teams deploying coding agents, providing structured guidance for defensible evaluation and safer operation.
- Enterprise
- Quality assurance
Research
Optimization of organizational processes in healthcare using artificial intelligence and implications for BRICS countries: a systematic review with narrative synthesis
Aleksey V. Cherepov, Laila M. Malikova, M.V. Makarovskaya et al.
The BRICS Health Journal · 2026-08-14
This systematic review of 78 publications examines how AI technologies affect organizational and managerial processes in healthcare, with particular attention to applicability in BRICS countries. The strongest evidence found that predictive analytics for patient flow management reduced waiting times by 18–26%, while machine learning for operating room scheduling cut idle time and improved resource utilization; ambient and generative AI documentation tools were linked to reduced administrative burden and clinician burnout. Evidence for other applications such as revenue cycle management and supply chain optimization was more limited, and major barriers including fragmented infrastructure, interoperability gaps, and governance challenges temper conclusions about scalability and long-term effectiveness.
- Workforce
- Enterprise
Research
AI-mediated relational competence and its limits: Psychedelic-assisted therapy as a stress case and policy signal
Robert McGrath, Everett B. Sackett
Psychedelics · 2026-08-14
This paper examines how AI is entering psychedelic-assisted therapy (PaT) through facilitator training simulators and in-session administrative tools, using PaT as a stress test for a six-part construct of AI-mediated relational competence. The authors find that while empathic-presence and relational-judgment components of that construct largely hold—and may be intensified—in altered-consciousness settings, trustworthiness, transparency, accountability, and equity-awareness components strain significantly given PaT's unique vulnerabilities and documented history of racial and cultural inequity. The paper raises a pointed policy question: without an accreditation infrastructure comparable to medical education, a privately developed AI platform risks becoming the de facto standard for facilitator competency certification by default.
- Certifications
- AI policy
Research
LegacyWorld: Atomicity-Aware Evaluation of GUI Agents for Legacy Workflows
Thilo Reintjes, Sivajeet Chand, Derui Zhu et al.
arXiv (Cornell University) · 2026-08-14
This paper introduces LegacyWorld, a benchmark of 28 Windows GUI workflows designed to evaluate AI agents automating legacy enterprise systems that lack programmable interfaces. The authors assess six multimodal LLM-based computer-use agents using an 'atomicity' criterion—agent runs must either complete the workflow correctly or fail without leaving unintended persistent changes in business or healthcare records. Results show that useful completion, safe failure, and non-atomic side effects represent distinct operational profiles, and that expert-crafted prompts and screen-recording-derived prompts produce meaningfully different outcomes. The authors conclude that workflow capture, state validators, and atomicity-aware acceptance tests should be treated as first-class requirements for deploying AI agents in legacy workflow automation.
- Enterprise
- Quality assurance
Research
Regime-Conditional Verification: Correctness Estimation for Adapting and Monitoring Safety Classifiers
Thiago Sandoval, Ufuk Topcu
arXiv (Cornell University) · 2026-08-14
Regime-Conditional Verification (RCV) is a lightweight wrapper that adapts off-the-shelf safety classifiers for large language models without retraining them, addressing two common failure modes: misalignment with the deployer's desired policy and performance degradation from distribution shift. By estimating correctness probabilities from the classifier's internal representations, RCV selectively corrects likely-wrong predictions and provides a label-free signal for detecting drift. Across three classifiers and two benchmark datasets, RCV improved policy adherence in every combination, catching up to 81% of previously missed unsafe content, and detected every simulated attack campaign in a deployment study. This matters for organizations deploying LLMs at scale who need reliable, maintainable safety filtering without costly retraining cycles.
- Quality assurance
- Enterprise
Research
Causative-Victim Tracing for Integrative Criminal Policy on AI Deepfake Face and Voice Crimes in Indonesia
Muhammad Abdul Azis, Pujiyono
KRTHA BHAYANGKARA · 2026-08-14
This study examines gaps in Indonesia's criminal law framework for addressing AI deepfake crimes involving faces and voices. The authors find that existing regulations—spread across multiple laws on personal data, electronic transactions, the criminal code, and sexual violence—are fragmented and fail to treat deepfakes as a sequence of interconnected crimes involving biometric data, layered actors, and digital evidence. The paper proposes a 'Causative-Victim Tracing' methodology to systematically link the tracing of data sources, technology, and actors with victim mapping, harm assessment, and recovery mechanisms. The work matters for policy by offering an integrative criminal law framework tailored to deepfake-specific harms.
- AI policy
Research
Coupling Coordination between Corporate Digitalization and Green Efficiency under the Data Elements × Action Plan
Shiman Zhou
Journal of innovation and development · 2026-08-14
Using panel data from 2,683 Chinese listed manufacturing firms (2014–2023), this study constructs a composite digitalization index across five technology dimensions and measures green efficiency via a super-efficiency SBM model, then quantifies their dynamic synergy through a coupling coordination degree model. Results show digitalization rose from 0.142 to 0.387 (10.5% annually) and green efficiency from 0.524 to 0.681, with coupling coordination advancing from near-disorder to a primary coordinated stage. Government R&D subsidies, environmental regulation intensity, and firm human capital all significantly drive higher coordination. The findings offer a quantitative framework for policymakers seeking to align digital transformation initiatives with green low-carbon manufacturing goals.
- Enterprise
- AI policy
Research
Mandato: Protocol-Level Enforcement of Digitally Signed Mandates on AI Agent Actions with Cryptographically Chained Audit Trails
Giovanni Racioppi
arXiv (Cornell University) · 2026-08-14
Mandato is a governance proxy system that enforces cryptographically signed authorization mandates on AI agent actions at the protocol level, addressing the gap where current AI tool-calling protocols (like the Model Context Protocol) lack verifiable, auditable authorization infrastructure. The system intercepts every tool call, checks it against a signed mandate specifying permitted tools, parameter constraints, and conditions, blocks non-conforming calls, and records all decisions in a tamper-evident, hash-chained audit log. The paper also maps the mechanism onto key EU regulations including the EU AI Act (Articles 12 and 14), GDPR, NIS2, and eIDAS 2, with a roadmap toward qualified attestation via Qualified Trust Service Providers. This matters because it provides a technical and legally legible accountability layer for AI agents acting on external systems, directly relevant to enterprise governance, regulatory compliance, and policy enforcement.
- AI policy
- Enterprise
Research
A Four-Axis Trustworthiness Benchmark for LLM-as-Judge in Principle-Based Regulation
Dipankar Sarkar
arXiv (Cornell University) · 2026-08-14
This paper introduces Principle-Bench, a benchmark of 168 cryptoasset financial-promotion scenarios mapped to UK FCA regulatory principles, designed to evaluate LLM-based automated judges across four axes: accuracy, paraphrase robustness, adversarial robustness, and calibration. The authors find that no tested method dominates all four axes, and that a 120B-parameter LLM judge drops 47 accuracy points when faced with keyword-stuffed adversarial inputs—a phenomenon they call 'compliance theatre.' The paper argues that any deployment-grade LLM judge used in principle-based financial regulation must report per-principle adversarial deception rates and calibration metrics alongside aggregate accuracy, not just overall performance figures.
- Quality assurance
- AI policy
Research
Determinants of AI Adoption in Banking: Evidence from Tunisia
Rahma Khattab, Sami Bacha
Arab Economic and Business Journal · 2026-08-14
This study examines what drives Tunisian retail banking customers to adopt AI-based banking services, surveying 400 customers and applying regression analysis within a multi-theory framework. Key findings show that perceived usefulness, ease of use, social norms, and trust positively predict AI adoption intent, while perceived risk reduces it; awareness, attitude, and knowledge show no significant effect. Age and education level moderate adoption behavior, underscoring the role of demographic factors. The results provide actionable guidance for banks and policymakers seeking to accelerate digital financial innovation in emerging markets.
- Enterprise
- AI policy