News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR
Veronica Chatrath, Bryan Zhu, George Pu et al.
arXiv (Cornell University) · 2026-08-07
CliniCARE-Bench is a new benchmark for evaluating AI clinical agents on retrospective audit tasks using real patient data from MIMIC-IV, covering 750 patient-specific cases across 25 clinician-validated scenarios. Unlike existing medical knowledge benchmarks, it assesses not just verdict accuracy but also evidence grounding, process adherence, calibrated abstention, and efficiency. Across 16 agentic systems tested, four-way verdict accuracy ranged from 65.3–76.1%, but a stricter 'defect-free accuracy' metric—which only credits correct verdicts achieved without prohibited shortcuts—was 4.8–14.8 points lower and reordered system rankings, revealing that raw accuracy substantially overstates real investigation quality. This matters for AI quality assurance and certification because it demonstrates that deployment-ready clinical agents require multidimensional evaluation frameworks beyond simple accuracy metrics.
- Quality assurance
- Certifications
Research
Rethinking data privacy for AI Adoption in African Higher Education: A meta-synthesis of stakeholder perceptions and policy implications
Emmanuel Duncan
Computers and Education Open · 2026-08-07
This qualitative meta-synthesis of 18 studies (2020–2025) examines how students and educators in Sub-Saharan African higher education perceive data privacy in AI-enhanced learning environments. Using the ENTREQ framework and PRISMA procedures alongside Privacy Calculus and Contextual Integrity theories, the study finds that students are primarily worried about data security, consent, surveillance, and misuse, while educators focus on governance, accountability, and ethical oversight. Privacy concerns—compounded by limited digital literacy and weak regulatory infrastructure—erode institutional trust and constrain AI adoption. The authors conclude that strengthening governance frameworks, transparency, and institutional capacity is essential for responsible AI integration across the region.
- AI policy
- Enterprise
Research
IntelliAudit: Using Large Language Models to Evaluate Audit Controls
Allison Wilson, Sina Moradi Sabet, Diar Shakimov et al.
arXiv (Cornell University) · 2026-08-07
IntelliAudit is a retrieval-grounded multi-agent LLM system designed to automate the evaluation of IT audit evidence against security and compliance controls, specifically instantiated on ISO/IEC 27001. The system retrieves relevant organizational artifacts, generates evidence-grounded assessments, challenges adverse findings, and produces auditor-facing recommendations with cited evidence, missing-evidence analysis, and remediation guidance. Expert auditor review and user feedback show it can support control interpretation and audit-preparation workflows, though the study highlights the necessity of human oversight to calibrate sufficiency judgments and correct overly permissive recommendations. The authors conclude such systems should serve as decision-support tools rather than autonomous certification systems.
- Quality assurance
- Certifications
Research
The Problem of Licensing Works for the Training Process of Artificial Intelligence Algorithms
Grzegorz Tylec, Sebastian Kwiecień, Łukasz Sarowski et al.
Studia Iuridica Lublinensia · 2026-08-07
This article analyzes the legal status of using copyrighted materials to train generative AI, arguing that existing EU copyright frameworks—specifically Directive 2019/790 on text and data mining (TDM)—do not adequately cover AI training, which the authors treat as a distinct and currently unregulated form of exploitation. Because identifying which specific works an AI uses during training makes traditional licensing impractical, the authors conclude that most AI training constitutes infringement of creators' economic rights. They propose a statutory compensation model modeled on the private copying levy under Directive 2001/29/EC, which would allow AI developers to legally use copyrighted works while ensuring fair remuneration for rights holders.
- AI policy
Research
From Forensics to Ecosystems: Rethinking Watermarks for Generative AI Oversight
Daniel Susser, John Thickstun, Gili Vidan
arXiv (Cornell University) · 2026-08-07
This paper argues that digital watermarking for AI-generated content is often framed as a forensic tool for identifying individual synthetic artifacts, but that framing exposes watermarks to well-founded criticisms about technical brittleness and epistemic ambiguity. The authors propose an 'ecosystems approach' that repositions watermarks as instruments for measuring the aggregate impact of synthetic content on media ecosystems rather than reliably authenticating individual pieces. They contend this reframing makes governance challenges more tractable and better matches the actual capabilities of watermarking technology, particularly for AI-generated text.
- AI policy
Research
Data Annotation as Measurement
Emma Harvey, Allison Koenecke, Rene F. Kizilcec
arXiv (Cornell University) · 2026-08-07
This paper argues that data annotation for AI systems should be treated as a formal measurement problem rather than simply an agreement-maximization task. Drawing on a review of 132 papers and interviews with 10 annotation team members, the authors develop a framework that maps decision points across annotation processes and identifies five sources of annotation issues: error, ambiguity, impossibility, subjectivity, and annotator identity. The paper translates measurement theory into practical guidance for annotation teams, showing how reliability and validity assessments can go beyond inter-annotator agreement. This matters because annotation quality directly shapes the data underlying AI systems, and better frameworks for diagnosing and correcting annotation issues can improve AI research and practice.
- Quality assurance
Research
Explanation-Guided Metamorphic Testing of Specialized Language Models: An Empirical Study
Xingcheng Chen, Mehmet Besenk, Andrea Stocco
arXiv (Cornell University) · 2026-08-07
This paper presents an empirical study of explanation-guided metamorphic testing for task-specialized language models used in software engineering workflows such as issue triaging and document classification. Across three datasets, four model architectures, and 20 testing configurations, the approach combines attribution-based token prioritization, LLM-driven mutation, and automated semantic verification to generate test variants that expose brittle model behaviors. The method produces 2.30× more verified failure-inducing test cases than heuristic baselines and reveals systematic shortcut behaviors such as over-reliance on named entities and formatting cues. The findings demonstrate that explainability-guided metamorphic testing is a practical and effective technique for robustness evaluation of specialized AI models.
- Quality assurance
Research
Italian Law No. 132/2025 as the First National Implementation Act of the AI Act: A Critical Analysis from a Comparative Legal Perspective
Izolda Bokszczanin-Gołaś
Studia Iuridica Lublinensia · 2026-08-07
This article critically analyzes Italy's Law No. 132/2025, the first national law implementing the EU AI Act, examining its strengths and structural weaknesses. The law introduces sector-specific rules for healthcare, employment, and the judiciary, and creates new criminal provisions around deepfakes and AI-aggravated offenses, but is critiqued for abandoning regulatory autonomy, adopting the AI Act as a maximum standard, and establishing a fragmented, underfunded supervisory framework. The author argues these deficiencies—alongside gaps on environmental impact, civic participation, and data protection—risk generating regulatory asymmetry across the EU Single Market. Comparatively, the Italian law is characterized as an imperfect but instructive test case for the flexibility and limits of the EU AI Act.
- AI policy
Research
Estimating GHG Emissions from AI Use: Framework for Corporate-Level Measurement
John Bistline, Shaena Ulissi, Steven J. Davis et al.
arXiv (Cornell University) · 2026-08-07
This white paper proposes a standardized framework for estimating greenhouse gas (GHG) emissions from corporate AI use, addressing a critical gap where no widely accepted methodology currently exists. The authors note that published per-query emissions estimates vary by orders of magnitude depending on methodology, provider, and grid assumptions, making credible corporate accounting nearly impossible. The framework is designed to be tiered, transparent, and actionable—helping companies set reduction targets as AI-related electricity demand, projected to reach 9–17% of U.S. consumption by 2030, continues to grow. The work is directly relevant to corporate sustainability reporting, regulatory compliance, and the design of decarbonization strategies.
- Enterprise
- AI policy
Research
"Death by a thousand taxonomies?": AI Risk Classification In Practice
Glen Berman, Ned Cooper, Angel Hsing‐Chi Hwang et al.
arXiv (Cornell University) · 2026-08-07
This paper examines how AI risk classification systems—called Sociotechnical Outcome Taxonomies (SOTs)—are developed and used in practice, based on 25 interviews with researchers and practitioners across industry, academia, civil society, and government. The study finds that SOTs are weakly integrated into AI governance processes for two key reasons: users treat taxonomy categories as exhaustive rather than interpretive, and SOTs fail to link harms to specific decision points or responsible actors, making accountability hard to assign. The authors offer design recommendations and argue that realizing the governance potential of SOTs requires new governance infrastructure that does not yet exist. This work matters for AI policy and certification efforts because it reveals structural gaps in the tools currently used to identify, categorize, and act on AI risks.
- AI policy
- Certifications
Research
Determinants of artificial intelligence tools use in research among postgraduate students: Evidence from selected public universities in Tanzania
Jeneth Edward Laizer, Nicholaus Mwalukasa, George Firmin Kavishe
African Journal of Empirical Research · 2026-08-07
This study investigates how and why postgraduate students at two Tanzanian public universities use AI tools in their research, surveying 350 students and applying the extended UTAUT-3 framework. ChatGPT and SciSpace were the most commonly used tools for idea generation and literature searching, while Grammarly and QuillBot were favored for writing and paraphrasing. A regression model explained 33.1% of variance in AI tool use, with performance expectancy, price value, and year of study positively predicting adoption, while effort expectancy had a negative effect. The authors conclude that AI integration in postgraduate research is still nascent and recommend that university research directorates establish targeted training programs to help students use these tools more effectively.
- Workforce
- AI policy
Research
Optimizing productivity and employee retention through advanced human resource strategies in small businesses
Vijay Solanki, Alfaiz Madhiya, S. M. Rezvi et al.
Scientific Reports · 2026-08-07
This paper presents a multi-task AI framework for HR decision-making that combines deep learning (Multi-gate Mixture-of-Experts), explainable AI (SHAP), counterfactual reasoning, and multi-objective optimization (NSGA-II) to predict employee attrition and optimize performance interventions. On benchmark HR datasets, the model achieves an AUC of 0.942, accuracy of 0.935, and F1 score of 0.928, with strong stability confirmed via 5-fold cross-validation. The system goes beyond prediction to generate cost-aware, Pareto-optimal intervention policies that balance retention gains, performance improvement, and intervention cost. The framework is designed to give HR practitioners actionable, interpretable guidance for workforce management in small businesses.
- Workforce
- Enterprise
Research
Machine Learning to Identify Point-of-Care Ultrasound and Evaluate Standardized Documentation: Retrospective Operational Cohort Study
Kevin Nguyen, Zewen Wu, Chu-An Tsai et al.
Journal of Medical Internet Research · 2026-08-07
This retrospective cohort study of over 559,000 clinical encounters at a large academic OBGYN center used machine learning models (LightGBM and BioClinBERT) to automatically identify point-of-care ultrasound (POCUS) procedures in free-text clinical notes and evaluate the impact of a standardized documentation template (ProcDoc). After ProcDoc was introduced, billing recapture rates—charges missed by providers but later identified—dropped from 10.0% to 2.4%, and 75.4% of CPT codes were generated automatically via the template. The findings demonstrate that combining ML-based auditing with standardized documentation workflows can substantially reduce missed procedural charges and manual billing burden without increasing procedure frequency, offering a scalable model for improving documentation and reimbursement accuracy in healthcare settings.
- Enterprise
- Quality assurance
Research
The Capability Ladder: A Curriculum-Modernization Framework for Workforce Readiness in the AI Era
Majid Memari, George Rudolph
arXiv (Cornell University) · 2026-08-07
This paper proposes a curriculum framework called the 'Capability Ladder' to help computing education keep pace with AI-driven changes to the workforce. The authors argue that near-term AI impact is best characterized as task reallocation—routine implementation becomes automated while verification, systems thinking, and AI supervision gain value—rather than wholesale job replacement. The framework organizes AI-augmented work into five autonomy levels and maps these to course updates, assessments, and stackable workforce credentials, illustrated through an exploratory two-semester pilot course. The authors explicitly acknowledge evidence limits, positioning the framework as a structured guide for targeted curriculum modernization rather than a validated prescription.
- Workforce
- Certifications
Research
The Impact of Generative AI Music Systems on Creative Labor, Production, and Industry Structures: A PRISMA-Aligned Systematic Review and Conceptual Framework
Hayder Albayati
Open Science Framework · 2026-08-07
This PRISMA 2020-compliant systematic review synthesizes 102 scholarly publications (2015–2025) to assess how generative AI music systems—including transformer-based and diffusion-based models—are reshaping creative labor, production workflows, and industry structures. The authors introduce a five-stage conceptual framework linking AI capability growth to creative automation, labor displacement, market saturation, and human differentiation in music ecosystems. The review distinguishes empirically supported findings from anticipated developments through evidence mapping and risk-of-bias assessment, and identifies critical gaps in the existing evidence base. Findings carry practical implications for artists, producers, technology developers, and policymakers navigating AI-driven transformation in creative industries.
- Workforce
- AI policy
Research
Adversarial Machine Learning for Secure and Explainable AI Systems: A Comprehensive Review
Hajar Ouazza, Fadoua Khennou, Abderrahim Abdellaoui
Journal of Cybersecurity and Privacy · 2026-08-07
This systematic review of 207 studies examines how adversarial machine learning, reinforcement learning, and explainable AI interact under realistic threat conditions. Key findings include that RL-based attack agents achieve evasion rates of 74–97% against ML-based detectors, while RL-based defenses yield up to 3× robustness gains over static baselines. The review also highlights that explanation methods like LIME, SHAP, and Grad-CAM become unreliable under adversarial perturbation, and no reviewed system certifies that attribution properties hold when inputs are manipulated—raising significant concerns for environments where AI accountability is a legal requirement.
- Quality assurance
- AI policy
Research
Artificial intelligence for food security or digital inequality? Quantifying the global trade-offs between agricultural productivity, social inclusion, sustainable livelihoods, and responsible land stewardship
Moussa El Jarroudi
Frontiers in Sustainable Food Systems · 2026-08-07
This large-scale review synthesizes evidence from 1,276 studies across 112 countries to examine whether AI advances in agriculture translate into broader societal benefits. The authors introduce the AI Food Security Paradox Theory, finding that while AI improved agricultural productivity by 24.8% and resource-use efficiency by 18.9%, gains in food security (8.6%) and social equity (5.4%) were substantially smaller. Governance quality, digital accessibility, and farmer training proved stronger determinants of societal outcomes than algorithmic sophistication. Projections to 2050 suggest inclusive, governance-driven AI pathways could increase societal benefits by over 40%, while widening digital inequality could significantly constrain these gains.
- AI policy
- Workforce
Research
Artificial Intelligence in Scholarly Peer Review: Ethical Considerations, Current Practices, and Future Implications
Aras Bozkurt
The International Review of Research in Open and Distributed Learning · 2026-08-07
This paper critically examines the integration of AI into scholarly peer review, identifying tensions between efficiency benefits and risks such as confidentiality breaches, algorithmic bias, accountability gaps, and erosion of expert judgment. Drawing on policy documents from major publishers and empirical research, the authors find evidence of undisclosed LLM-assisted text appearing in peer review, particularly in conference settings. Major organizations like ICMJE and publishers such as Elsevier and Taylor & Francis have responded with disclosure requirements and prohibitions on uploading unpublished manuscripts to generative AI tools. The authors conclude AI should augment rather than replace human peer review, with robust governance, transparency, and equity evaluation required for responsible use.
- Quality assurance
- AI policy
Research
Beyond AI Adoption: How AI Governance, Technology Trust, and Organizational Culture Shape Employee Empowerment in Local Government
Muhammad Noor, Adam Idris, Annisa Wahyuni Arsyad et al.
F1000Research · 2026-08-07
This study of 687 local government employees in Indonesia uses PLS-SEM to examine how AI Governance Preparedness, Public Service Orientation, Learning Culture, and Technology Trust jointly shape employees' perceived work empowerment in AI-enabled public administration. All four factors showed positive and significant direct effects, with Technology Trust emerging as the strongest predictor and Learning Culture second. Notably, Technology Trust acted primarily as an independent psychological resource rather than a moderator—it did not amplify the effects of governance or public service orientation, and actually weakened the positive relationship between Learning Culture and empowerment. The findings suggest local governments should invest in transparent AI governance, continuous training, and digital competency development to genuinely empower public servants using AI.
- Workforce
- AI policy
Research
Evaluating AI Performance in Systematic Literature Reviews for HEOR: A Case Study
Jing Wang-Silvanto, Mansee Jajoo, He Guo et al.
Journal of health economics and outcomes research · 2026-08-07
This case study evaluates AI performance across key stages of systematic literature reviews (SLRs) in health economics and outcomes research (HEOR), comparing AI-assisted workflows against a traditional human-only SLR benchmark. AI tools showed high accuracy and recall but low precision during title/abstract screening, leading to increased false positives, while full-text screening identified only a minority of studies included in the human review. Mean data extraction accuracy was 72.93%, with no hallucinated outputs detected. The findings highlight both the promise and current limitations of AI for evidence synthesis in HEOR, offering practical recommendations for appropriate human oversight in AI-assisted literature reviews.
- Quality assurance
- Enterprise
Research
From fragmented to integrated surveillance in LMICs: digital pathways for outbreak detection and vaccine intelligence
Delfin Lovelina Francis, Saravanan Sampoornam Pape Reddy
Frontiers in Public Health · 2026-08-07
This structured narrative review examines how digital health tools, AI, and genomic surveillance can transform fragmented, disease-specific surveillance systems in low- and middle-income countries (LMICs) into integrated, real-time platforms for outbreak detection and vaccine deployment. Reviewing 46 studies published between 2019 and 2025, the authors find that electronic immunization registries can cut median reporting delays from 28 to 3 days, reduce stockouts by up to 76%, and raise coverage by 12.3%, though only 14% of AI models have been externally validated in LMIC contexts. Key barriers include non-standardized data formats, intermittent connectivity, algorithmic bias, and data governance gaps. The paper concludes that interoperability standards, equitable data governance, and investment in local capacity are essential for sustainable surveillance innovation.
- AI policy
- Workforce
Research
HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
Ji Wu, Yunshan Peng, Wentao Bai et al.
arXiv · 2026-08-07
HOBA is a hierarchical reinforcement learning framework for online advertising bidding that combines a large language model for hyperparameter inference, a SARSA-based model selector with causal bias correction, and a pool of expert bidding models operating at three different time scales. By confining online learning to discrete expert selection rather than continuous bid optimization, HOBA reduces exploration risk while adapting to non-stationary auction markets without costly manual tuning. Experiments on the AuctionNet benchmark and a large-scale A/B test show consistent gains over baselines, with a real-world deployment achieving a +3.6% increase in target cost. The work demonstrates a practical path to deploying adaptive, hierarchical multi-agent systems in production advertising environments.
- Enterprise
Research
Certifiable Deep Importance Sampling for Rare-Event Simulation of Black Box Systems
Mansur Arief, Yuanlu Bai, Wenhao Ding et al.
Operations Research · 2026-08-07
This paper introduces Deep-PrAE (Deep Probabilistic Accelerated Evaluation), a framework that integrates deep neural network classifiers with importance sampling to improve rare-event simulation for black-box AI systems such as self-driving vehicles. The approach learns rare-event geometries using neural networks to calibrate importance samplers, producing statistically certifiable guarantees of reliability that go beyond conventional rare-event simulation methods. The authors show that existing simulation-based testing methods can dangerously underestimate rare failure probabilities in black-box systems without detectable signals, and their framework addresses this gap. This matters for safety certification of autonomous and AI-embedded systems by enabling more trustworthy pre-deployment risk quantification.
- Certifications
- Quality assurance
Research
Artificial intelligence in academic publishing and the faculty tradeoff between productivity benefits and academic integrity concerns
Mohamed Mekheimer, Walid Abdelhalim
Discover Sustainability · 2026-08-07
This mixed-methods study of 478 faculty members across six universities in Upper Egypt finds that faculty report moderate academic integrity concerns (mean 3.81 out of 5) alongside lukewarm endorsement of AI productivity benefits (mean 3.06) in academic publishing contexts. MANOVA results show small but statistically significant differences by discipline and academic rank, with early-career and STEM faculty leaning more toward AI adoption while senior and humanities faculty express greater integrity concerns. Qualitative follow-up interviews with 118 faculty reveal six behavioral strategies for managing this tradeoff, including principled resistance, pragmatic task segmentation, and active risk management. The authors call for AI-literacy training, disclosure guidance, and verification frameworks to support responsible AI integration in scholarly work.
- AI policy
- Quality assurance
Research
Behavioural Dynamics of AI‐Assisted Project Control: Introducing the Wait Effect
Fredrik Kockum, Martin Kunc, Nicholas Dacre et al.
Systems Research and Behavioral Science · 2026-08-07
Drawing on interviews with 22 project managers, this study identifies a novel behavioural mechanism called the 'wait effect,' in which managers using AI decision-support tools delay corrective actions while awaiting further confirmation from AI outputs. The authors find this tendency amplifies Escalation of Commitment—sticking with a failing course of action—and that the degree of delay is moderated by how much trust managers place in the AI model. The findings are mapped through a Causal Loop Diagram and carry direct implications for how AI-enabled project management systems should be designed to mitigate behavioural biases that can undermine timely intervention and project performance.
- Enterprise
- Workforce