News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
EduPluginBench: Executable Assurance for AI-Generated Educational Plugins
Nizam Kadir
arXiv (Cornell University) · 2026-08-01
EduPluginBench introduces an executable benchmark and staged admission method (P0–P4) for evaluating AI-generated educational plugins in governed software ecosystems. The benchmark tests compliance beyond mere compilation, covering least privilege, telemetry consent, provenance, and lifecycle constraints across 1,440 mutants from 30 specifications. Results show P0–P4 increased release-blocking-defect recall by 74.7 percentage points over a shallower P0–P2 baseline, but frozen generations from two current coding models achieved zero conformance, meaning downstream assurance estimands were undefined. The work highlights a significant gap between AI code generation capability and the safety and compliance requirements needed for deployment in educational plugin ecosystems.
- Quality assurance
- Certifications
Research
Explainable Ensemble Machine Learning for WWTPs: A Systematic Review and Compliance Framework
Yolanda T. Gegana, Pitshou N. Bokoro, Thulane Paepae
Results in Engineering · 2026-08-01
This systematic review examines explainable ensemble machine learning (XEML) applications in wastewater treatment plants (WWTPs) across 43 peer-reviewed studies from 2015–2025, revealing that while tree-based methods like Random Forest (58.1%) and XGBoost (46.5%) dominate and 72.1% of studies use SHAP-based explainability, critical gaps remain: 70% lack temporal validation safeguards, only 14% use local interpretability methods, and no studies address data governance frameworks. To address these gaps, the authors propose the XEML Policy Compliance (XEML-PC) and GEARS frameworks to support auditable, regulation-aligned deployment of AI in WWTPs. The findings matter because WWTPs operate under strict regulatory conditions, and the absence of governance and auditability mechanisms limits the operational trustworthiness of AI systems in these environments.
- Quality assurance
- AI policy
Research
ZK-SR117: A Chunked Zero-Knowledge Attestation Design for Aggregated Fair-Lending Metrics, with a Control Mapping toward Full SR 11-7 Coverage
Mohammad Nasir Uddin, Rahnuma Tabassum Orpita, Eklachur Rahman Bhuiyan et al.
arXiv (Cornell University) · 2026-08-01
This paper presents a zero-knowledge proof system called ZK-SR117 that allows banks to demonstrate to regulators that their ML models satisfy fairness and robustness requirements under U.S. supervisory guidance (SR 11-7, OCC 2011-12) without exposing model weights or customer data. The authors implement and test a chunked zkSNARK circuit design on 32,768 rows of real 2022 HMDA mortgage data, producing 32 verified proofs that attest a demographic-parity fairness gap within 0.0029 of the true value, with per-chunk proving times under 4 seconds. The work also maps nine SR 11-7 control elements to zero-knowledge statements and proposes a nonce-based sampling protocol to prevent cherry-picking, though most of these extensions remain as scoped future work. This matters because it offers a technically feasible path for regulated financial institutions to provide cryptographically verifiable model audits without compromising proprietary or sensitive data.
- Certifications
- AI policy
Research
Auditable Release Control for Pedagogical Leakage in LLM Tutors
Nizam Kadir
arXiv (Cornell University) · 2026-08-01
This paper addresses 'pedagogical leakage'—when AI tutors reveal answers or decisive reasoning before they are authorized to do so—and proposes a formal auditable release control system to prevent it. The system uses a selector, policy gate, renderer, and release function with inspectable checks and replayable traces to enforce disclosure contracts. In controlled experiments on 599 Gemini proposals, strict mediation reduced panel-majority leakage flags from 181 to 0, though at the cost of replacing 581 responses and lowering helpfulness; a replication on 480 attack sequences confirmed the approach reduced flags significantly but did not eliminate all failures. The results demonstrate a measurable safety-utility tradeoff with attributed failure modes, relevant to deploying trustworthy AI tutoring systems in educational settings.
- Quality assurance
- Certifications
Research
A Protocol for Evaluating the Accessibility of AI-Generated Educational Materials: Prompt Configuration, WCAG-Derived Criteria, and Content Overload
Hector R. Amado-Salvatierra
arXiv (Cornell University) · 2026-08-01
This paper introduces a reproducible protocol for evaluating whether AI-generated educational materials—across documents, slides, images, audio, and video—meet Web Content Accessibility Guidelines (WCAG). The protocol tests three prompt conditions (generic, WCAG-configured, and persistent accessibility profile) and combines a per-content-type WCAG rubric with expert heuristic validation to address gaps in automated scanning, including visual and informational overload. In an exploratory application using a single model and single-evaluator scoring, rubric-based compliance rose from a pooled mean of 24.2% under the generic condition to 96.7% under the WCAG-configured condition. The work highlights that AI tools reproduce inaccessible practices by default and provides an open evaluation instrument to help creators and researchers improve accessibility of AI-generated content.
- Quality assurance
- AI policy
Research
Responsible AI and Algorithmic Adoption in Methodology Development for National Statistical Offices
Siu‐Ming Tam
arXiv (Cornell University) · 2026-08-01
This paper examines how national statistical offices can responsibly adopt AI-generated algorithms, focusing on two trust foundations: independent statistical verification and respondent confidentiality protection. The author demonstrates a Mini Max Hierarchical Bayes sampling algorithm constructed with AI assistance, which achieved an 80 percent sample size reduction on a synthetic labour force population while meeting all precision targets, and a 90 percent reduction on 2021 Australian Census microdata with national point estimates accurate to well below 1 percent. The paper concludes with a practical evaluation checklist aligned with the UN Fundamental Principles of Official Statistics and the HLG MOS Quality Framework for Statistical Algorithms, offering concrete governance guidance for official statistics production.
- Quality assurance
- AI policy
Research
Algorithmic wage setting on online labor platforms
Herbert Dawid, Philipp Harting, Michael Neugart
Labour Economics · 2026-08-01
This paper examines how machine-learning algorithms used by firms to set wages on online digital labor platforms affect worker pay. Using a simulation framework where firms apply deep Q-network (DQN) reinforcement learning to update posted wages, the study finds that algorithmic wage-setting can lead to collusive (suppressed) wages when competition is limited and no experience replay is used in the algorithm. The results are robust across varied platform designs, raising concerns about whether workers receive fair compensation in algorithmically governed labor markets.
- Workforce
- AI policy
Research
High-stakes environments and AI regulation: ethical principles, regulatory strategies and legal architecture
Joseph Mante, Blessing Abeji, Harsha Kalutarage et al.
AI and Ethics · 2026-08-01
This paper examines how AI systems deployed in high-stakes environments—such as healthcare, autonomous vehicles, criminal justice, and public sector decision-making—should be regulated to produce effective and enforceable obligations. The authors develop a three-tier Ethics Pyramid framework to guide which ethical principles (safety, fairness, accountability, transparency, human oversight) should be translated into binding legal obligations and at what governance level. They argue that a polycentric, multi-layered legal architecture is better suited to these settings than existing fragmented frameworks, and support this with a comparative analysis of regulatory approaches in the UK, EU, US, and China. Case studies of AI in UK prisons and autonomous vehicles illustrate the contrast between immature and mature regulatory regimes.
- AI policy
- Certifications
Research
A Clinician-in-the-Loop Framework for Validating and Selecting Synthetic Paediatric Dermatology Images
Ali Tariq Nagi, Chiara Bellatreccia, Andrea Borghesi et al.
Information · 2026-08-01
This paper presents a clinician-in-the-loop framework for validating synthetic paediatric dermatology images, addressing data scarcity and skin-tone representation imbalances in medical AI. Four clinicians assessed 93 real and synthetic images across dimensions including visual realism, diagnostic plausibility, and skin-tone relevance, resulting in a curated subset of 18 approved synthetic images. When used to augment a ResNet50 classifier, the clinician-approved images outperformed both unfiltered synthetic and real-only training sets, with the largest gains observed for a dark-skin subgroup. The findings support clinician-guided data curation as a quality-assurance step before synthetic images are used in AI training pipelines, though the authors note results are exploratory given the small sample sizes.
- Quality assurance
- Workforce
Research
Labor Market Implications of Artificial Intelligence
Xin Tang
Selected Issues Papers · 2026-08-01
This IMF Selected Issues Paper examzes the labor market implications of rapid AI adoption among Swedish enterprises, finding that while AI offers substantial potential gains for Sweden's workforce—greater than the EU average—the costs of labor displacement are likely to fall unevenly. ICT professionals, less-educated workers, young people, and immigrants face the greatest risk of displacement. The paper recommends policies that improve labor mobility across regions and sectors and better align the education system with labor market needs to ease the adjustment.
- Workforce
- AI policy
Research
The AI Implementation Gap in Higher Education: Navigating the Disconnect Between Technology Adoption, Policy Awareness, and Institutional Governance
Jonathan Westover
Future of Work The Journal of Labor Transformation Technology Integration and Human Adaptation · 2026-08-01
This study identifies an 'AI implementation gap' in higher education, finding that while 94% of nearly 2,000 surveyed higher education professionals report using AI tools for work, only 54% are aware of relevant institutional policies and more than half have used AI tools not sanctioned by their institutions. The paper situates these findings within frameworks of technology adoption and organizational change, highlighting risks including data privacy concerns, misinformation, skill erosion, algorithmic bias, and unmeasured return on investment. The authors conclude with recommendations for institutional leaders and policymakers to bridge the disconnect between AI adoption and governance structures.
- Workforce
- AI policy
Research
K-12 AI Infrastructure: Findings from Educator and Developer Outreach
Rebecca Griffiths, Vanessa Peters Hinton, Shelton Daal et al.
arXiv · 2026-08-01
This report summarizes findings from outreach to K-12 educators, edtech developers, and other stakeholders about AI tools in education. Educators want AI grounded in local curricula and capable of supporting formative assessment at scale, but current generic tools undermine curriculum coherence and raise concerns about student safety and wellbeing. Existing trust proxies like certifications fail to address privacy and fairness gaps, and evaluation standards for AI-based edtech remain immature. The authors identify five public infrastructure gaps—including shared evaluation benchmarks and child speech recognition datasets—and argue for modular, publicly available infrastructure over proprietary solutions to protect student privacy and serve underserved populations.
- Certifications
- Quality assurance
- AI policy
Research
Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation
William Caban
arXiv (Cornell University) · 2026-08-01
This paper argues that automated benchmarks used to evaluate agentic AI systems are far less trustworthy than commonly assumed due to three compounding layers of measurement failure: language-model-generated tasks with validity flaws, LLM simulators replacing human users with documented inter-simulator variance up to 9 percentage points, and approximately 82% of surveyed papers applying mismatched or absent inter-rater reliability metrics. The authors formalize these failures multiplicatively, showing a pipeline could retain as little as 22–54% validity against its intended construct, and propose eight prescriptions—including mandatory IRR reporting and calibration floors—drawn from psychometric science to address the problem. The findings directly challenge the use of such benchmark scores to justify safety certifications and regulatory compliance claims for AI systems.
- Quality assurance
- Certifications
- AI policy
Research
A Generative AI-Based Framework for Business Process Orchestration in Industrial Enterprises
Galina Ilieva, Yuliy Iliev
Electronics · 2026-08-01
This paper presents an integrated generative AI (GAI) framework designed to orchestrate and govern business processes across four enterprise functions—manufacturing, marketing and sales, accounting and finance, and human resource management—within industrial enterprises. The framework embeds GAI as a structured information systems capability with human-in-the-loop validation, traceability, and auditability rather than as standalone tools. A proof-of-concept validation at an electronics company using maturity-readiness logic showed an overall readiness score increase from 41.60 to 79.08 after the GAI implementation scenario, interpreted as expert-assessed readiness evidence rather than measured operational outcomes. The work contributes a reference architecture aimed at practical, auditable, and human-supervised enterprise-scale AI adoption.
- Enterprise
- Quality assurance
Research
Methodology for Calculating the Human Capital Vulnerability to AI Index
Svitlana Onyshchenko, Олександра Маслій, Alina Yanko et al.
Economies · 2026-08-01
This study develops a composite index to measure and compare how vulnerable human capital is to AI-driven displacement across countries, combining five dimensions—technological readiness, human capital adaptability, institutional protection, labour market vulnerability, and migration risk—using 24 normalised indicators weighted by principal component analysis. Validated on six European countries, the index finds Finland least vulnerable due to strong institutional readiness, while Ukraine is most vulnerable owing to compounded structural, institutional, and migration constraints, with Germany, France, Estonia, and Poland in intermediate positions. The methodology is designed to support cross-country comparisons and inform targeted retraining and social protection policymaking in response to AI-driven labour market change.
- Workforce
- AI policy
Research
Leveraging Technology to Measure the Strength of Demand for Commercial and Trafficked Sex
Michale Shively, Courtney Furlong, Brooke Ruffin et al.
Dignity A Journal of Analysis of Exploitation and Violence · 2026-08-01
This paper examines Transaction Intercept (TI), an AI platform that deploys automated decoy online ads to engage individuals seeking to purchase commercial sex, generating real-time data on demand for sex trafficking. The system has facilitated over 700,000 text exchanges and identified more than 25,000 sex buyers across 75+ U.S. cities, providing law enforcement with jurisdiction-level intelligence at no cost. The article argues that TI's data can inform public policy, improve law-enforcement practice, and potentially be adapted for use in other countries to combat sex trafficking.
- AI policy
- Workforce
Research
Generative AI integration in Australian VET: regulatory constraints and newly qualified trainers
Nanxu Wang, Dr. Michael J. Henderson
International Journal of Training Research · 2026-08-01
This exploratory multiple case study examines how five newly qualified VET trainers in Victoria, Australia used generative AI to prepare, adapt, and check training and assessment materials that must comply with nationally specified units of competency. Findings show that trainers viewed GenAI outputs as plausible but not deployment-ready without verification against vocational knowledge and regulatory requirements, with adoption shaped by unclear regulatory expectations, limited institutional guidance, and unequal tool access. The study frames GenAI integration in VET as a compliance-verification challenge and extends the TPACK framework by clarifying how regulatory, institutional, and tool-access conditions affect trainers' ability to justify GenAI use in high-compliance settings.
- Certifications
- Workforce
Research
The AI Evaluator Gap: Institutional Capacity, Independent Assurance, and the Governance of Frontier AI
Naveen Sundaresan
Zenodo (CERN European Organization for Nuclear Research) · 2026-08-01
This paper identifies what it terms the 'AI evaluator gap'—the mismatch between the rapid advancement of frontier AI capabilities and the slower development of the qualified, independent assurance ecosystem needed to assess them. Using Anthropic's 2026 Advanced AI Framework as a central case, the author argues that governance regimes built around transparency, external evaluation, and deployment-blocking authority will fail not from a lack of rules but from a shortage of evaluators with the requisite technical depth, model access, professional standards, and institutional authority. The paper surveys major governance frameworks—including the NIST AI RMF, EU AI Act, ISO/IEC 42001, and Singapore's MAS and IMDA frameworks—and proposes a global evidence chain linking developer disclosures through independent evaluation to regulatory accountability. It concludes that legislation and standards are necessary but insufficient without concurrent investment in evaluator competence, independence rules, liability structures, and cross-disciplinary assurance methods.
- AI policy
- Certifications
Research
Systemic fragility in European total intravenous anesthesia delivery and opportunities for resilient real-time decision support
Clara M. Ionescu, Erhan Yumuk, Michele Schiavo et al.
Communications Medicine · 2026-08-01
This review examines systemic weaknesses in European total intravenous anesthesia (TIVA) delivery, including reliance on population-based drug models, inconsistent monitoring, and variable training, which together limit safe personalization. The authors outline how interoperable perioperative data, digital patient simulators, patient-adaptive models, multimodal monitoring, and closed-loop decision support tools could improve safety and resilience while keeping clinicians in control. The paper concludes with a European roadmap covering data infrastructure, training, validation, regulation, and evaluation of AI-enabled tools. It matters because it directly addresses how AI and digital health technologies can be responsibly integrated into clinical practice without replacing human oversight.
- AI policy
- Quality assurance
Research
The AI Evaluator Gap: Institutional Capacity, Independent Assurance, and the Governance of Frontier AI
Naveen Sundaresan
Zenodo (CERN European Organization for Nuclear Research) · 2026-08-01
This paper argues that the central bottleneck in governing frontier AI is not the absence of rules but a shortage of qualified, independent evaluators with the technical depth, model access, professional standards, and institutional authority needed to assess advanced systems. Using Anthropic's 2026 Advanced AI Framework as a primary case study, the author introduces the concept of the 'AI evaluator gap'—the mismatch between the rapid advancement of frontier AI capabilities and the slower development of the assurance ecosystem needed to evaluate them. The paper surveys major governance frameworks (NIST AI RMF, EU AI Act, ISO/IEC 42001, Singapore's MAS and IMDA frameworks) and proposes a global evidence chain linking developer disclosures to independent evaluation, regulatory requirements, and clear accountability. It concludes that legislation and standards are necessary but insufficient without investment in evaluator competence, independence rules, liability structures, and cross-disciplinary assurance methods.
- Certifications
- AI policy
Research
The Gomola Framework: A Quantitative Safety Certification Standard for Clinical AI Systems (The DeepSensi Standard)
TOMASZ GOMOLA
Zenodo (CERN European Organization for Nuclear Research) · 2026-08-01
This paper introduces the Gomola Framework (operationalized as the DeepSensi Standard), a proposed quantitative safety certification standard for clinical AI systems built on large language models. The framework defines four certification levels across five pillars—including Evidence Integrity Verification, Deterministic Safety Verification, and Transparent Uncertainty—designed to address LLM hallucination as the dominant clinical failure mode. Certification is two-dimensional, combining an architectural assessment with a probabilistic per-assertion hallucination bound derived from Fault Tree Analysis; the reference implementation achieves a certified worst-case hallucination bound of 3.23 × 10⁻⁶. The standard is vendor-neutral and royalty-free, positioning it as a broadly applicable benchmark for safe clinical AI deployment.
- Certifications
- Quality assurance
Research
WM-Cov: Test Adequacy for Interactive World-Model-Style Autonomous Driving Simulation
Jianxun Cui, Ping Wu, Staniša Perić et al.
arXiv (Cornell University) · 2026-07-31
This paper addresses how to rigorously measure whether autonomous driving tests conducted inside generative 'world model' simulators have produced sufficient, valid evidence to support a testing conclusion. The authors introduce WM-Cov, an evaluation framework that categorizes simulation outputs as requested, realized, or valid evidence and tracks metrics including coverage growth, failure-mode diversity, realism, and artifact suppression. Experiments on TeraSim/SUMO events and a real DriveArena simulator matrix show that raw generated failures can include duplicates, partial realizations, and artifacts that inflate apparent test completeness. The work argues that test adequacy for interactive simulators should be judged by convergence of valid closed-loop evidence under a budget, not by raw failure counts or prompt coverage alone.
- Quality assurance
- Certifications
Research
Exposed by Design: A Dynamic Security Assessment of Internet-Facing MCP Servers at Scale
Nicolás Padilla
arXiv (Cornell University) · 2026-07-31
This paper presents the first large-scale dynamic security assessment of internet-facing Model Context Protocol (MCP) servers, discovering over 21,000 detectable instances and auditing 414 confirmed production servers using a purpose-built 34-module testing framework called Corvus. Researchers uncovered 68 reportable vulnerabilities—including SQL injection, server-side request forgery, prompt template injection, and path traversal—while finding that 91.8% of audited servers lack OAuth authentication and 687 tool instances expose shell execution capabilities without access controls. The finding that 41.6% of confirmed servers disappear within three days suggests rapid, security-review-free deployment cycles. The study highlights systemic security risks in the MCP ecosystem and releases Corvus as an open-source evaluation framework to support responsible disclosure and ongoing assessment.
- Quality assurance
- AI policy
Research
Artificial Intelligence Adoption and Labour Market Outcomes: A Study of Employment Perceptions in Central Europe
Rizwan Arshad
Inverge Journal of Social Sciences · 2026-07-31
This study surveyed 160 working professionals in Hungary and Central Europe to examine how AI adoption in organizations relates to three employment outcomes: overall employment rates, job creation in new sectors, and job displacement in traditional industries. Using correlation and regression analyses, the researchers found statistically significant positive associations between AI adoption and all three outcomes, with the correlation for job creation (r=0.498) being stronger than that for job displacement (r=0.459), suggesting a cautiously optimistic net employment picture. The study is notable for providing individual worker-level perceptual data from Central Europe, a region underrepresented in the largely macroeconomic and Western-focused literature on AI and employment. The authors argue that institutional frameworks — including government policy, education, and business leadership — are key to determining whether AI adoption produces net job gains or losses.
- Workforce
- AI policy
Research
Assurance of Safety-Critical AI-Enabled Capabilities: Integrating Human, Software, and Machine Learning Assurance Paradigms
Randall McCutcheon, Keith F. Joiner, Li Qiao et al.
Preprints.org · 2026-07-31
This paper develops a unified framework for assuring AI-enabled systems in safety-critical environments by systematically comparing three historical assurance paradigms: human-organizational (1980–2000), software (2000–2020), and AI-enabled systems (emerging since 2020). The authors synthesize 20 assurance precepts mapped across a novel Dual Assurance Spiral for AI-Enabled Capabilities (DAS4AIC), incorporating frameworks such as the NIST AI Risk Management Framework. A face-validity workshop with assurance practitioners identified critical gaps in data governance, explainability, and human-autonomy teaming, and the study finds that some human-dominant assurance precepts map more directly to AI systems than through software assurance—helping explain accountability and oversight concerns when AI enters safety-critical roles. The work argues that effective organizational governance underpins auditing, training, and validation of AI systems in operational safety-critical contexts.
- Certifications
- Quality assurance
- AI policy