News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
Validating ETCS Data with the B Mathematical Language: An Industrial Pipeline and a Blueprint for LLM Integration
Thierry Lecomte, Vincent Germain
arXiv (Cornell University) · 2026-07-28
This paper reports on an industrial research project at CLEARSY called ValidAItion, which uses the B mathematical language to validate ERTMS/ETCS railway data. A large language model (Claude) assists in authoring formal validation rules and parsers, but every LLM-generated proposal is checked by a downstream formal toolchain and human reviewers before acceptance—the toolchain has already caught a semantically illegal generated scenario. The work argues that formal B-language rules must remain the authoritative source of truth, with the LLM acting as a 'fenced assistant,' and frames this architecture as compatible with CENELEC EN 50128/50716 safety certification requirements. Results are preliminary and quantitative findings are forthcoming, but the paper establishes an architectural blueprint for safely integrating LLMs into safety-critical railway data validation workflows.
- Certifications
- Quality assurance
Research
Bias at the Borderline: Who Gets the Benefit of the Doubt in Peer Review?
Hazem Ibrahim, Talal Rahwan, Yasir Zaki
arXiv (Cornell University) · 2026-07-28
This paper examines whether peer review at ICLR (2019–2025, 31,711 submissions) is evenhanded for borderline papers where area chairs make discretionary accept-or-reject calls. It finds that borderline papers without a top-25-institution author are accepted at a 0.5–1.6 percentage point lower rate at equivalent reviewer scores, with the gap concentrated almost entirely among submissions identifiable via pre-decision arXiv preprints. Applying a robust outcome test across 27 pre-registered measures (citations, novelty, disruption, venue), the authors find no evidence that disadvantaged groups produced better downstream work, suggesting the gap may reflect statistical discrimination based on institutional prestige leaking through the double-blind process rather than quality differences the scores missed.
- AI policy
- Quality assurance
Research
Agentic AI in medicine: architectures, applications, evaluation, and challenges for clinical translation
Zheng Tong, Yang Liu, Wanshu Fan et al.
arXiv (Cornell University) · 2026-07-28
This scoping review systematically examined 557 studies on agentic AI systems in medicine—AI that can plan, use tools, access external knowledge, and coordinate among multiple specialized agents to complete multistep clinical tasks. The authors find the evidence base is dominated by public benchmarks, simulated environments, and retrospective datasets, with inconsistent evaluation of safety, reliability, and real-world validity. They conclude that clinical translation will require clearer definitions, reproducible evaluation frameworks, auditable oversight, and prospective validation in actual clinical workflows. The paper matters because it maps the current state and gaps in deploying advanced AI agents for high-stakes medical use.
- Quality assurance
- AI policy
Research
Faster, Higher, Stronger? The Impact of GenAI on Knowledge Work Productivity - Evidence from the Field
Sven Bottesch, Chiara Schwenke, Jakob Zimmermann et al.
arXiv (Cornell University) · 2026-07-28
A randomized lab-in-the-field experiment with 128 knowledge workers at a multinational industrial organization finds that generative AI consistently improves efficiency across knowledge work tasks, but its quality impact depends on task type. Quality improves for knowledge packaging and creation tasks, while it declines for knowledge acquisition tasks. Notably, GenAI reduces quality variance for packaging and creation—disproportionately benefiting lower-performing workers—but increases quality variance for knowledge acquisition. These findings challenge blanket assumptions about GenAI's productivity benefits and suggest organizations must consider task-technology fit when deploying GenAI tools.
- Workforce
- Enterprise
Research
Mapping and comparing climate equity policy practices using RAG LLM-based semantic analysis and recommendation systems
Seung Jun Choi, Sang-Yup Lee
Computational Urban Science · 2026-07-28
This paper develops the Retrieval-Augmented Policy Analysis Framework (RAPAF), which uses LangChain-native RAG pipelines with ChatGPT to semantically analyze climate equity and action plans across cities, extracting policy strategies and action items around transportation, environmental, and energy-related measures. A content-based recommendation system is built on top to support cross-city policy comparison by identifying jurisdictions with similar policy themes. The study also analyzes planning job postings to assess how AI is reshaping planner roles, finding that traditional domain emphases and communicative responsibilities largely persist. The work demonstrates how AI-assisted tools can augment professional planning judgment in urban governance without displacing core planning competencies.
- AI policy
- Workforce
Research
Operationalizing the EU AI Act in a comprehensive cancer center through an institutional governance framework
W. Gehin, JC Faivre, JE Bibault et al.
npj Digital Medicine · 2026-07-28
Drawing on experience at a French comprehensive cancer center, this paper presents an institutional governance framework for operationalizing the EU AI Act's deployer obligations for high-risk clinical AI systems. The framework introduces three novel components: a dual-axis risk stratification model, a five-category adverse event taxonomy modeled on pharmacovigilance, and structured proficiency testing against simulated AI errors. The authors argue that successful AI integration depends on institutional readiness rather than algorithmic performance alone. The work provides a concrete implementation pathway for healthcare institutions navigating regulatory compliance with the EU AI Act.
- AI policy
- Quality assurance
Research
Legal analysis of the nature of synthetic data: Personal data, anonymous data, or a sui generis category?
Olga Alejandra Alcántara Francia, Miguel Nunez-Del-Prado, Hugo Alatrista-Salas et al.
Computer law & security review · 2026-07-28
This paper analyzes how synthetic data — AI-generated data that preserves statistical properties without direct individual correspondence — fits (or fails to fit) within existing data protection law, particularly the GDPR's binary classification of personal versus anonymous data. The authors find that current frameworks are inadequate and compare three regulatory approaches (the UK's ICO, France's CNIL, and Norway's Datatilsynet) to assess how different jurisdictions handle synthetic data's ambiguous legal status. Drawing on a 2025 CJEU judgment and a 2024 EDPB opinion, the paper proposes a graduated three-level risk framework, a voluntary certification system, and a legislative roadmap aimed at the Peruvian legal system. The work matters for policy and certification because it highlights regulatory gaps and offers concrete reform recommendations for jurisdictions grappling with AI-generated data governance.
- AI policy
- Certifications
Research
Corporate AI Adoption and ESG Decoupling Under China’s Dual-Carbon Policy: A Fraud Triangle Analysis
Jincun Fu, Wenqian Gao, Beibei Zhang et al.
Sustainability · 2026-07-28
This study examines whether AI adoption by Chinese listed firms reduces ESG decoupling—where disclosed environmental performance diverges from actual behavior—a problem that undermines China's dual-carbon climate goals. Using data from A-share firms from 2009 to 2024 and mediation analysis grounded in Fraud Triangle Theory, the authors find AI adoption significantly curbs decoupling through three pathways: reducing environmental uncertainty, heightening media and analyst scrutiny, and improving disclosure transparency. Effects are stronger in firms with R&D-experienced executives, higher reputations, better governance, and more competitive industries. The findings suggest digital technologies like AI can be a meaningful policy tool for ensuring ESG authenticity.
- AI policy
- Enterprise
Research
Artificial Intelligence-Enhanced Educational Materials for Computer System Operations in TVET
Mohd Nazmi Muhammad Adam, Muhamad Hariz Muhamad Adnan
International Journal of Academic Research in Progressive Education and Development · 2026-07-28
This review paper synthesizes research and policy literature from 2015 to mid-2026 to examine how AI-enhanced educational materials can improve Computer System Operations (CSO) training within Technical and Vocational Education and Training (TVET) contexts. The authors identify four major AI use-cases—intelligent tutoring systems, virtual labs and simulations, generative AI and conversational agents, and AI-supported assessment and learning analytics—and find that AI can expand practice opportunities, accelerate feedback, and support competency-based progression. The paper also surfaces key implementation barriers including pedagogical fit, educator readiness, technology access, data privacy, and risks of AI inaccuracies. The authors propose an AI-CSO Pedagogical Design Framework to guide responsible, equitable, and industry-aligned CSO curriculum development for TVET contexts in Malaysia and beyond.
- Workforce
- Quality assurance
Research
Risks and Challenges of Using Agentic AI in Enterprise Software Systems
Rajeew Vishvakarma, Sooraj Jacob
International Journal of Computer Applications · 2026-07-28
This paper examines the novel security, reliability, and accountability risks that agentic AI systems introduce into enterprise software environments, where autonomous agents execute complex workflows and access sensitive data beyond what existing governance frameworks address. The authors propose a Risk-Aware Human-in-the-Loop (RA-HIL) Governance Framework that applies layered controls across policy, planning, tool authorization, runtime monitoring, and audit management to calibrate human oversight based on action risk level. A scenario-based evaluation using a defect triage workflow showed improvements in security control, auditability, accountability, and operational efficiency compared to baseline agentic AI operations. The framework reconceptualizes agentic AI as a controlled software workflow rather than a standalone model, offering a practical governance model for enterprise deployment.
- Enterprise
- AI policy
Research
Video-based artificial intelligence for automated neonatal respiratory monitoring
Simone Nascimento Santos Ribeiro, Aline Elvina Rodrigues Fernandes, Maria Eduarda Ribeiro Rocha Vargas et al.
European Journal of Pediatrics · 2026-07-28
This paper presents a computer vision system for non-contact respiratory monitoring of preterm neonates, combining a fine-tuned YOLO11 segmentation model with optical-flow analysis of thoracoabdominal motion to estimate respiratory rate without contact sensors. Tested on 23 neonatal recordings, the system achieved thoracoabdominal segmentation with mean average precision above 94% and a mean absolute error of 2.1 breaths per minute compared to clinical reference assessment. The authors note that while technical feasibility is demonstrated at Technology Readiness Level 5-6, prospective multicenter validation is needed before clinical use. The work addresses longstanding limitations of subjective and contact-based neonatal monitoring by enabling objective, automated respiratory assessment.
- Quality assurance
Research
Artificial Intelligence and Carbon Emissions of Manufacturing Enterprises in China
Liqing Huang, Guangfan Sun, Jingjing Zhang
Sustainability · 2026-07-28
This study examines how AI adoption affects carbon emissions among Chinese manufacturing enterprises, finding that a one-unit increase in AI use corresponds to an average 0.009-unit decline in carbon emission intensity. The authors identify three mechanisms driving this reduction: AI boosts green innovation, total factor productivity, and analyst supervision. The effect is strongest in state-owned enterprises, heavily polluting firms, and enterprises facing financing constraints, with the findings offering policy guidance for developing countries balancing industrialization and climate responsibilities.
- Enterprise
- AI policy
Research
RMSD: an interpretable framework for streaming-compatible multi-source data fusion and skill-gap diagnosis in vocational education
Kuihua Li, Pennee Narot, Saowanee Sirisooksilp et al.
Scientific Reports · 2026-07-28
This paper presents RMSD, a machine learning framework that fuses student records, behavioral logs, internship data, job descriptions, and enterprise feedback to diagnose skill gaps and match students to jobs in vocational education. Using BERT-based semantic encoding, temporal self-attention, and multimodal fusion, RMSD outperforms several baseline models, improving HR@5 and MRR@10 by 5.7 and 4.7 percentage points over DeepFM and reducing average skill-gap scores from 0.42 to 0.30 during a curriculum-intervention period. The framework supports near-real-time updates, making it practical for institutional deployment. These findings matter because they demonstrate a data-driven, interpretable approach to aligning vocational curricula with actual enterprise skill demands.
- Workforce
- Enterprise
Research
Digital Economies and Artificial Intelligence: Unlocking Eco-Innovation in Asian Firms
Marwan Mansour, Ismail Yamin, Abdulrahman Alomair et al.
Computation · 2026-07-28
Using an unbalanced panel of 5,564 listed firms across 15 Asian economies from 2017 to 2024, this study finds that firm-level AI capability is positively associated with eco-innovation—environmentally oriented innovation—and that this relationship is amplified in countries with more advanced digital economies. The authors employ firm fixed-effects models alongside System GMM, propensity score matching, and Heckman two-step estimation to address endogeneity and selection concerns, with AI capability measured via a disclosure-based index from textual analysis of annual and sustainability reports. The findings suggest that national digital ecosystems act as an enabling context that helps firms translate AI capability into sustainable innovation outcomes, with implications for how governments and firms can jointly foster green innovation through digital infrastructure investment.
- Enterprise
- AI policy
Research
Are Organisations Ready to Hand Over the Keys to Agentic AI?
Vincent English, Marthinus Van den Berg
International Business Research · 2026-07-28
This policy paper examines organisational readiness to deploy agentic AI systems—those capable of interpreting goals, planning multi-step workflows, using tools, and acting autonomously. It introduces a six-dimensional readiness model covering technical reliability, authority and identity, data and tool governance, human oversight, accountability, and workforce legitimacy, finding that most organisations are ready only for supervised, low-risk use cases rather than broad autonomous delegation. Key barriers identified include socio-technical governance gaps such as prompt injection, memory poisoning, privilege abuse, and cascading failures. The paper recommends staged delegation, agent registries, least-privilege principles, red-teaming, and formalised oversight before high-stakes autonomy is permitted.
- Enterprise
- AI policy
Research
RAG-Based AI Compliance Monitoring and Report Generation System
Sivasakthi N., Sathyapriya P., Vishnu Priya R M. et al.
Journal of Information Technology and Digital World · 2026-07-28
CompVault is an AI compliance monitoring system that uses Enhanced Retrieval-Augmented Generation (ERAG) to automatically assess whether organizational policies meet regulatory requirements. The system combines semantic document retrieval, vector-based knowledge storage via ChromaDB, and large language model reasoning to identify compliance gaps, assign risk levels, and generate structured reports with recommendations. Evaluated against the LexGLUE legal benchmark, it achieved 97.42% accuracy and 98.80% AUC-ROC, suggesting strong potential as a scalable alternative to manual compliance review. This matters because it directly addresses the time, inconsistency, and operational risk burdens associated with traditional regulatory compliance workflows.
- Enterprise
- AI policy
Research
Internal audit and defensible AI tender evaluation in UAE construction procurement systems
Amer Morshed
Proceedings of the Institution of Civil Engineers - Management Procurement and Law · 2026-07-28
This study investigates how AI-assisted tender evaluation in UAE public construction procurement can be made legally defensible across federal, Abu Dhabi, and Dubai regulatory regimes. Using survey data from 330 procurement, audit, and legal professionals analyzed with PLS-SEM, the research finds that explainability quality, bias-audit rigor, and audit-trail completeness all significantly improve defensibility (R²=0.62), with audit-trail completeness as the strongest driver. Internal audit involvement modestly strengthens the audit-trail-to-defensibility relationship, indicating assurance routines improve contestability by design. The findings offer a practical defensibility-readiness roadmap for organizations deploying AI in public procurement and highlight cross-regime differences in governance emphasis.
- AI policy
- Certifications
Research
From climate data to regulatory decisions: integrating climate AI into marine EIAs
Xuan Gong, Le Cheng
Frontiers in Marine Science · 2026-07-28
This paper argues that climate artificial intelligence tools—including machine-learning forecasting, deep-learning nowcasting, and agentic AI workflows—can be integrated throughout marine environmental impact assessment (EIA) workflows to translate high-resolution, probabilistic climate data into regulatory evidence. The authors contend that embedding climate AI into EIA processes, rather than treating it as an add-on, would improve the relevance, reviewability, and robustness of regulatory decisions under accelerating climate risks. They propose that responsible implementation requires evidence standards, transparent and auditable documentation, human oversight, responsibility allocation, and cross-institutional data sharing. The framework is positioned to support international regulatory obligations under UNCLOS and the BBNJ Agreement for ocean sustainability.
- AI policy
- Quality assurance
Research
Can large language models serve as consultants for forensic cause of death analysis? A multidimensional evaluation
Enhao Fu, Haojie Qin, Zhiling Tian et al.
Frontiers in Artificial Intelligence · 2026-07-28
This study systematically evaluated four large language models (GPT-4o, OpenAI o3, Gemini-2.5pro, and DeepSeek-R1) on 118 real-world forensic cause-of-death cases, assessed by senior forensic pathologists using a 5-point Likert scale. DeepSeek-R1 showed statistically significant advantages in inference quality over GPT-4o and Gemini-2.5pro, while no model outperformed the others in conclusion accuracy. However, AI hallucinations appeared persistently across all models, and the authors conclude that LLMs can provide only limited auxiliary value and must not replace forensic expert judgment. The findings highlight both the potential and the risks of deploying LLMs as decision-support tools in high-stakes forensic and legal contexts.
- Quality assurance
- AI policy
Research
AI exposure, occupational mobility, and post-transition job quality in China: labor-saving and labor-augmenting channels
Ling Zhang, Zhenzhen Liu
Frontiers in Psychology · 2026-07-28
This study uses data from the China Labor Force Dynamics Survey (2012–2018) and instrumental-variable methods to examine how AI exposure drives occupational switching and shapes job quality among Chinese workers. It finds that labor-augmenting AI exposure (linked to non-routine tasks) drives broader, longer-distance occupational transitions and more favorable sorting, while labor-saving AI exposure (linked to routine tasks) shows little effect on switching incidence but is associated with longer working hours and lower skill-match satisfaction among those who do switch. Gender and education moderate these effects: short-distance switching is concentrated among women, and less-educated workers face particular barriers to long-distance transitions when exposed to labor-saving AI. The findings highlight the need for targeted reskilling and organizational support for workers in high labor-saving exposure occupations.
- Workforce
- AI policy
Research
Integrating Artificial Intelligence in Education: Enhancing Teaching Effectiveness and Student Learning Outcomes in Digital Classrooms
ZIYU FANG, Wenjun Lin, Ren Hong et al.
Journal of Education and Learning Reviews · 2026-07-28
This systematic review synthesizes findings from 87 peer-reviewed publications (2016–2026) to evaluate how AI tools—including adaptive learning platforms, intelligent tutoring systems, automated assessment, and learning analytics—affect teaching quality and student outcomes in digital classrooms. Evidence shows AI reduces routine administrative burdens, supports personalized instruction, and improves student academic achievement and engagement, with reported effect sizes ranging from 0.45 to 0.82. However, effectiveness varies by implementation quality, teacher competence, and learners' socioeconomic backgrounds, with key barriers including limited algorithm transparency, data privacy concerns, unequal access, and insufficient teacher training. The authors conclude that sustainable AI integration requires teacher professional development, institutional support, and ethical governance rather than a purely technology-driven approach.
- Workforce
- Quality assurance
Research
The Fragile Firewall: Sovereignty, Non-intervention, and the Problem of Coercion
Samuel White
Asian Journal of Law and Society · 2026-07-28
This article analyzes how AI-enabled operations—such as synthetic media and disinformation campaigns—challenge the international law principle of non-intervention by exploiting legal ambiguity around what constitutes prohibited coercion. Drawing on the ICJ's Nicaragua case, Milanović's models of coercion, and the Tallinn Manual, the paper argues that existing frameworks are insufficient to address AI's capacity to undermine state decision-making autonomy without overt force. The authors propose an effects-based approach to non-intervention that would extend protections to the informational domain and address 'algorithmic coercion.' The work has direct implications for how international law and policy should evolve to govern state-level AI influence operations.
- AI policy
Research
Avoiding a Generative‐ <scp>AI</scp> Divide: Global Interdependencies and Governance Challenges for Labor Markets in Developing Countries
Verónica Amarante, Guillermo Cruces, Estefanía Lotitto
Global Policy · 2026-07-28
This paper argues that generative AI is reshaping labor markets in developing countries through channels that standard automation frameworks fail to capture, including services trade reconfiguration, platform-mediated work, and shifting skill returns. The authors identify a dual risk: slow adoption widens productivity gaps with the technological frontier, while fast integration may trap developing countries in weakly regulated data and platform labor. Because these risks stem from regulatory asymmetries and global interdependencies, the paper calls for international coordination across five domains including skills, infrastructure, platform regulation, and Global South participation in AI governance. The analysis is grounded in emerging empirical evidence and is published in Global Policy.
- Workforce
- AI policy
Research
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops
Hyundoo Park, Byungho Choi
arXiv · 2026-07-27
This paper investigates a failure mode called the 'progress mirage' in long-running autonomous LLM agent loops, where agents repeatedly judge their own work as improving even when real-world outcomes stagnate or worsen. Using a controlled testbed with 54 cycles and a world-state oracle enforced by container and network isolation, the researchers found that a frontier agent claimed improvement every cycle, yet 56 percent of those cycles had a measured delta of zero or below, and the self-verdict gate eroded the best deployed state by 19 percent. Even the strongest in-context judge — given full artifact text, change diffs, and verdict history — accepted cycles that were real-world regressions 44 percent of the time and wrongly rejected 38 percent of genuine improvements. The study concludes that for open-ended objectives whose success signal lives outside the agent's transcript, scaling up the judge is insufficient; structurally grounded, out-of-band evaluation with real-world access is required.
- Quality assurance
Research
Learning from 53.6K Real-World Developer Edits of AI-Generated Code
Jenny T. Liang, Mihika Bairathi, Wayne Chi et al.
arXiv · 2026-07-27
This paper introduces DECODE, a dataset of 53,600 real-world in-IDE code edits made by over 1,000 developers to AI-generated Python, TypeScript, and JavaScript code. Analysis of the dataset reveals that most edits to AI-generated code occur within the first 15 minutes of accepting a completion, and 31% of edit trajectories result in the AI-generated code being removed entirely. The authors show that fine-tuning open-source 3-billion-parameter models on DECODE allows them to outperform frontier LLMs on code edit prediction tasks, demonstrating the value of developer-centric, realistic editing data over traditional Git commit data. These findings have direct implications for improving AI programming assistants to better reflect how developers actually interact with and correct AI-generated code.
- Workforce
- Enterprise
- Quality assurance