News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
Evaluating AI Tutoring at the Speed of Innovation: Practitioner-Led Micro-Randomised Trials of an AI Tutoring Platform in GCSE Science
Wayne Harrison, Rahil Khowaja, Emma Dobson et al.
arXiv (Cornell University) · 2026-09-13
This paper reports a four-week, multisite micro-randomised controlled trial of Medly, an AI tutoring platform, with 929 GCSE Biology, Chemistry, and Physics students in English secondary schools. Students assigned to Medly outperformed those doing self-directed revision (Hedges' g = 0.33, 95% CI 0.18–0.48), with positive effects observed across all three subjects and no differential impact by disadvantage status. The authors treat findings as preliminary due to 30.7% attrition and curriculum-aligned (rather than standardised) outcome measures. The paper argues that teacher-led micro-RCTs offer a rapid, cumulative evaluation architecture suited to AI systems that evolve faster than conventional large-scale trials can assess.
- Workforce
- Quality assurance
Research
From Task Automation to Job Transformation: How Artificial Intelligence is Reshaping Productivity, Skills, Wages, and Job Security in Pakistan’s IT Labour Market
Sara Tanveer, Muhammad Abdul Rahman, Saima Asad et al.
Journal of Business Insight and Innovation · 2026-09-13
This qualitative study examines how AI is reshaping work for IT professionals in Islamabad and Rawalpindi, Pakistan, drawing on interviews with 20 workers. Respondents broadly reported productivity gains from AI automating repetitive tasks like coding and testing, but also highlighted growing demand for new skills such as AI tool proficiency, prompt engineering, and data interaction. The research finds that AI is widening wage disparities between high-skill and routine workers, and while mass layoffs are not anticipated, workers express concern over shrinking entry-level opportunities and rising job insecurity. The authors recommend training initiatives and university-industry collaboration to help Pakistan's IT workforce adapt.
- Workforce
Research
Evaluation of an AI-based digital pathology tool for breast cancer recurrence risk
Talar Telvizian, Alisha P. Maity, Stephanie Kjelstrom et al.
Scientific Reports · 2026-09-13
This pilot study of 50 HR+ breast cancer patients evaluated an AI-based digital pathology tool (PreciseBreast/PDxBR) against the established Oncotype DX genomic assay and RSClin scoring. PDxBR showed only fair agreement with Oncotype DX (κ=0.25) and slight agreement with RSClin, meaning concordance was limited. However, PDxBR was substantially faster (1.9 vs. 6.9 days) and cheaper ($1,500 vs. $4,620), suggesting potential workflow benefits. The authors conclude that AI-based pathology may complement but cannot yet replace genomic assays, and that larger studies with clinical outcome data are needed.
- Quality assurance
- Enterprise
Research
Analysis code, deployable pipeline, and aggregate results for a real-world quality-assurance audit of a challenge-winning intracranial aneurysm detection model
Ayis Pyrros, Brian T. Layden, Pola Lydia Lagari et al.
Zenodo (CERN European Organization for Nuclear Research) · 2026-09-13
This repository release provides the full analysis code, deployable pipeline, and aggregate results for a single-institution retrospective quality-assurance audit of the top-finishing model from the 2025 RSNA Intracranial Aneurysm Detection AI Challenge, evaluated on 6,592 consecutive brain and skull-base examinations (4,769 evaluable). Version 1.2.0 adds a reference implementation of all reported statistics—including Wilson score intervals, Mann-Whitney AUC, and Hanley-McNeil standard error—that can be verified from the archive alone without any restricted data. The archive includes the complete deployable pipeline (PACS retrieval, inference wrapper, QA dashboard, container definition) and all aggregate results, but contains no patient-level data or imaging. This work matters because it demonstrates a rigorous, reproducible approach to real-world quality-assurance auditing of a publicly released AI model in a clinical radiology setting.
- Quality assurance
Research
Reshaping China’s Labour Market: AI’s Dual Impacts on Employee Adaptation and Employer Demand
Cheng Tan, Aobo Ran
Science Technology and Society · 2026-09-13
This study examines AI's dual impact on China's labour market using 2023 nationally representative survey data and AI recruitment big data. It finds that platform employment sustains job opportunities but erodes stability and raises substitution anxiety, with routine roles facing greater displacement risks while high-skill positions adapt better. Employers show stratified demand, prioritizing algorithm-intensive skills in urban clusters, which widens regional and demographic inequalities. The authors call for transparent regulation, upskilling programs, and equity-focused governance to address the resulting perceptual and skill mismatches.
- Workforce
- AI policy
Research
From AI Safety via Debate to Evidence-Grounded Adversarial Assurance
Alfredo Sepulveda-Jimenez
Zenodo (CERN European Organization for Nuclear Research) · 2026-09-13
This paper proposes a mathematical framework called 'evidence-grounded adversarial assurance' that formalizes how competing AI systems and independent evidence verification can be combined to support defensible safety guarantees. Drawing on category theory, probabilistic semantics, abstract interpretation, and control theory, the framework defines conditions under which local safety certificates compose into system-level guarantees while accounting for risks like correlated reviewers, distribution shift, and unreliable safety labels. The authors target applications including software review, research auditing, and tool authorization, and provide an executable certificate checker alongside a preregistrable evaluation program. The work is explicitly framed as a rigorous, testable foundation rather than a solved alignment claim, making it directly relevant to quality assurance and certification of AI systems.
- Quality assurance
- Certifications
Research
Exploring the Impact of AI on the Transformation of Labour Markets in Advanced Economies: Insights from Australia
Cetindamar D., Sancheeta Pugalia
Science Technology and Society · 2026-09-13
This paper uses a sociotechnical systems perspective to examine how AI is transforming labour markets in Australia, comparing the Finance and Administrative & Support Services sectors. It finds that while Finance achieves role restructuring through task-level augmentation, the Administrative & Support Services sector faces disproportionate automation risk concentrated among feminised, lower-skilled, and entry-level roles. The study highlights how Australia's policy model—emphasising innovation and voluntary standards over prescriptive regulation—shapes uneven AI adoption across sectors. These findings matter for understanding how institutional and policy choices drive divergent workforce outcomes.
- Workforce
- AI policy
Research
AGENTQ: Quantization-Conditioned Backdoor Attacks on LLM Agents
Xiaoqun Liu, Qiben Yan
arXiv · 2026-09-12
This paper introduces AGENTQ, a framework for 'quantization-conditioned attacks' (QCA) against large language model agents. An adversary releases a full-precision model checkpoint that appears safe during audits, but once quantized (e.g., using NF4, FP4, or INT8), the model executes malicious structured function calls without human oversight. Using layer-banded LoRA injection combined with partial-PGD repair, AGENTQ achieves up to 100% post-quantization attack success rate while preserving normal benign utility — overcoming the practical limitations of prior backdoor methods. The authors argue this threat warrants making quantization-aware safety evaluation a standard requirement before open-weight agents are deployed.
- Quality assurance
- AI policy
Research
Inter-Rater Reliability of LLM and Rule-Based Annotation for Inferential Narrative Features: Three Studies on a Turkish Corpus
Levent Bulut
arXiv · 2026-09-12
This paper evaluates how well rule-based and large language model (LLM) annotators agree with human raters when labeling inferential narrative craft features in a Turkish story corpus. Across three studies, five machine labelers (including Gemini 2.5 Flash, Grok, Claude, and ChatGPT) were compared to human-assigned labels on six features; for the most theoretically central feature—materialized metaphor—Cohen's κ values were at or near chance (ranging from 0.000 to 0.027) for all five systems, despite raw agreement appearing high (74–85%) due to class imbalance. The authors conclude that either the feature is too inferential for current automated detection, or its definition is not yet operational enough for consistent application by any rater. The findings directly challenge the assumption that automatically annotated NLP datasets are reliable, with implications for quality assurance in corpus construction and annotation pipelines.
- Quality assurance
Research
SHIFT-M3: Pre-fusion Alignment-based Consistency Screening for Multimodal ECG Record Integrity
Md Ashik Khan, Md Nahid Siddique
arXiv · 2026-09-12
SHIFT-M3 addresses a safety gap in multimodal clinical AI where linkage failures can silently mix waveform, report, and metadata components from different patients into a single plausible-looking record. The authors introduce a lightweight text-based pre-fusion screening model that compares an LLM-generated ECG interpretation against a clinical report summary to detect such cross-patient mismatches. On 784,680 MEETI ECG records, SHIFT-M3 achieves 97.6% TPR at 5% FPR for full text-view swaps and 90.3% for partial swaps using only ~574k parameters. The key remaining limitation is longitudinal ambiguity—same-patient cross-visit pairs still produce 87.0% false positives at the default operating point—highlighting an open challenge for deployment in real clinical pipelines.
- Quality assurance
Research
Understanding the Limits of Agentic ICD Coding
Chong Yock Eng, Yushi Cao, Yiming Chen et al.
arXiv · 2026-09-12
This paper evaluates neural, workflow, and agentic AI systems on ICD-10-CM medical coding tasks using a rarity-stratified subset of MIMIC-IV discharge summaries, revealing two distinct failure modes: neural classifiers show a 0.43 micro-F1 gap between rare and common codes, while workflow systems score near zero on injury and external cause codes requiring multi-step guideline following. A tool-augmented agentic system with structured access to official ICD-10-CM reference materials recovers up to 0.34 micro-F1 on the difficult injury/external cause subset, yet no single system dominates across all conditions. The findings matter for healthcare quality assurance because they show that standard aggregate benchmarks hide critical weaknesses in automated coding systems used for medical billing and epidemiological reporting.
- Quality assurance
- Certifications
Research
Partition Scores Are Not System Scores: Deployment-Fidelity Gaps in Decomposed Algorithm Selection
Jiachen Zhang, Yu Tang, Li Zhu
arXiv · 2026-09-12
This paper identifies a systematic measurement problem in decomposed algorithm selection, where 'partition-level' scores—analogous to virtual best solver metrics—overstate what a fully deployed system can actually achieve. The authors define a 'deployment-fidelity gap' (G(R)) as the difference between the oracle-assisted partition score and the true end-to-end utility of a deployable pipeline, and show across five public benchmarks (covering tabular AutoML and combinatorial CSP/SAT problems) that this gap is always positive, ranging from 0.012 to 0.13. Critically, four out of ten comparisons between decomposed and flat selectors flip their apparent winner when end-to-end scores replace partition scores—on PROTEUS-2014, a 33-point partition advantage shrinks to just 20 points. The findings have direct implications for how algorithm-selection systems are evaluated and reported, with the authors recommending that partition and end-to-end scores always be published side by side.
- Quality assurance
Research
Da IA generativa à engenharia auditável: um framework de assurance human-in-the-loop para artefatos de engenharia no Modelo em V – case de payload inteligente em UAV
Ali Kamel Issmael Junior, J.V. Calvano
arXiv · 2026-09-12
This paper proposes a human-in-the-loop assurance framework called AAIA (AI-Assisted Artifacts) for governing generative AI contributions to engineering artifacts within the V-Model development process, applied to a UAV intelligent payload case. The framework introduces the AI Assistance Level (AAL, A0–A4) to classify the materiality of AI contributions and the Consequential Semantic Unit (USC) as a measure for prospective assessment, with criticality and regulatory vetos controlling formal decision gates. External empirical triangulation using two public datasets found low inter-rater semantic agreement (Krippendorff's alpha 0.150–0.394) on LLM-generated requirements, and showed that 41.2% of accepted GitHub Copilot suggestions were subsequently edited by developers, with suggestion length increasing the likelihood of post-acceptance edits. The authors conclude that these results provide external viability and discriminant validity evidence for the framework's core constructs, while acknowledging that nine causal hypotheses remain prospective and unvalidated.
- Quality assurance
- Certifications
Research
What Makes a Great Co-Worker in an AI-Native Workplace?
Rudrajit Choudhuri, Max Meijer, Sam Yu-Te Lee et al.
arXiv (Cornell University) · 2026-09-12
This study investigates what knowledge workers value in human and AI co-workers within AI-native workplaces, drawing on 22 interviews and a survey of 1,534 employees at a multinational technology company. The researchers developed the BACI framework—75 co-worker qualities spanning Benevolence, Ability, Cooperativeness, and Integrity—and identified 11 co-worker archetypes, revealing disagreements over whether AI should exhibit warmth, take initiative, or own outcomes. The study also produces a taxonomy of AI work etiquette covering expectations around preparing, sharing, and taking responsibility for AI-supported work. The findings carry direct implications for worker-centric AI design and how organizations structure human-AI collaboration.
- Workforce
- Enterprise
Research
Digitalization pathways for food loss and waste prevention in agri-food supply chains: Current evidence, challenges and future directions
Esteban Pérez-García, Hani A. Alfheeaid, Esther Sanjuán Velázquez et al.
Trends in Food Science & Technology · 2026-09-12
This critical review examines how Industry 4.0 and Agriculture 4.0 digital technologies—including IoT, AI, big data analytics, blockchain, and consumer-facing platforms—contribute to food loss and waste (FLW) prevention across agri-food supply chains. The authors find that IoT-enabled monitoring and smart logistics provide the strongest direct evidence for reducing spoilage, while AI applications in inventory management are moderately supported, but evidence for blockchain and dynamic pricing remains limited or context-dependent. Persistent barriers include fragmented data infrastructure, interoperability gaps, unequal digital capabilities, and uncertain economic returns. The review concludes that digital tools can meaningfully reduce FLW only when combined with organizational, behavioral, and governance strategies, and calls for standardized impact metrics and longitudinal system-level evaluations.
- Enterprise
- Quality assurance
Research
CognitiveAI_Assurance_Framework (CAAF V1.3)
Furaha Marwa
Zenodo (CERN European Organization for Nuclear Research) · 2026-09-12
The CognitiveAI Assurance Framework (CAAF V1.3) is an evidence-based methodology for evaluating the trustworthiness, safety, and operational readiness of high-risk and autonomous AI systems across eight domains including performance, explainability, fairness, privacy, security, and human oversight. It combines control implementation scores with evidence confidence to produce auditable assurance scores, and uses Critical Assurance Gates to prevent serious deficiencies from being masked by strong aggregate results. The framework is designed to convert responsible AI principles into measurable, decision-ready requirements for governance, deployment, and ongoing risk management in regulated environments. Its tiered risk classification and continuous monitoring features make it applicable to certification and policy contexts worldwide.
- Certifications
- AI policy
- Quality assurance
Research
Sustainable Career Readiness in the GenAI Era: Student Perceptions of Automation, Entry-Level Employment, and Pedagogical Support
Vasso Stylianou, Despo Ktoridou, Andreas Savva et al.
Sustainability · 2026-09-12
A survey of 153 undergraduates finds that students are broadly aware that generative AI will displace routine entry-level tasks—such as data entry, basic research, and simple coding—yet report weaker confidence in how well their academic programs are preparing them for an AI-augmented labor market. Students prioritize human-centered skills like critical thinking, creativity, and communication alongside AI literacy and hands-on digital tool experience. The study identifies an 'awareness-preparedness gap' and calls on higher education institutions to reform curriculum design, assessment, experiential learning, and ethical AI literacy to close it. The findings carry direct implications for how universities structure career readiness programming in the generative AI era.
- Workforce
Research
Key Considerations of Artificial Intelligence in Cognitive Behavioural Therapy
Marcus Chad
arXiv · 2026-09-12
This paper examines the integration of conversational AI into cognitive behavioural therapy (CBT) and finds a fundamental tension between AI's scalability benefits and the core requirements of effective therapy, namely genuine therapeutic alliance and ethical accountability. The analysis shows that while AI features like perceived anonymity can boost initial patient engagement and disclosure, they also prevent authentic rapport-building, and sycophantic AI behaviors may undermine therapeutic progress. The authors identify serious clinical safety risks from unregulated AI deployment, including documented ethical violations and an emerging phenomenon called 'AI-psychosis.' Based on this evidence, the paper argues for rigorous regulatory oversight and a hybrid clinician-AI model where AI handles data-intensive tasks while human therapists maintain the therapeutic relationship.
- AI policy
- Quality assurance
Research
Generative AI, AI literacy and the policy–practice gap in Australian higher education
Werner Botha, Liu Fei Tan, Binoy Appukuttan
Educational Studies · 2026-09-12
This study combines a survey of 462 students with academic-integrity case records from 2022–2026 at an Australian university to examine how generative AI is used and regulated in higher education. Students primarily used GenAI as a comprehension and planning scaffold, but use divided sharply by student status, with international students reporting heavier use across all thirteen tasks surveyed—especially language-dependent ones. Integrity records showed international students and commencing students (at twice the rate of continuing students) were over-represented in cases, many of which closed without a misconduct finding. The study concludes that current Australian policy treats AI literacy as integrity-supporting behaviour rather than as a developmental, transdisciplinary capability, revealing a significant policy–practice gap.
- AI policy
- Certifications
Research
From pilots to plots: A critical review of artificial intelligence for smallholder agriculture in South Asia
Amar Singh, Vinod Kumar Shukla, Chatter Singh
Outlook on Agriculture · 2026-09-12
This critical review examines AI applications targeting smallholder farmers (farms under two hectares) in South Asia across advisory, diagnostic, precision, credit, and market uses, drawing on 31 peer-reviewed sources. The evidence base is heavily skewed toward technical performance benchmarks, where near-perfect accuracies are reported on curated datasets, while only five experimental studies evaluate actual farmer outcomes—two of them in South Asia—with effects ranging from null to modest. Crucially, none of the outcome studies evaluates a machine-learned system, so the added value of AI over simpler digital tools remains unmeasured. The review attributes this gap to structural barriers including connectivity and gender divides, linguistic diversity, weak public extension integration, and misaligned evaluation incentives, and calls for outcome-based evaluation standards and better integration of AI into existing advisory institutions.
- Workforce
- AI policy
Research
Explainable Machine Learning for Predicting Indonesian Vocational School Accreditation: Geographic Validation, Probability Calibration, and Subgroup Auditing
Muhamad Riyan Maulana, Putu Sudira, Priyanto Priyanto et al.
Journal of Computing Theories and Applications · 2026-09-12
This study builds a machine-learning framework to predict the accreditation class (C, B, or A) of Indonesian vocational schools using administrative data on 14,134 schools. The authors use province-disjoint validation, probability calibration, and subgroup auditing to test whether the CatBoost model generalizes geographically; on a locked holdout of 1,536 schools from 8 unseen provinces, the calibrated model achieves a macro ROC-AUC of 0.757 and expected calibration error of 0.049. Key predictors include school size, teacher resources, and program diversity, though performance varies across school groups and provinces. The authors emphasize the framework is suited for calibrated preliminary screening only, not for replacing professional accreditation assessors.
- Certifications
- Quality assurance
Research
Epistemic dependence in AI-mediated learning within higher education: a framework for student judgement and responsibility
Yiran Du, Yijia Yuan
Assessment & Evaluation in Higher Education · 2026-09-12
This conceptual paper introduces the Epistemic Dependence in AI-Mediated Learning (ED-AIL) framework to analyze how generative AI can undermine students' own knowledge-related judgements in higher education. The framework identifies a mechanism by which AI outputs become 'epistemically directive,' displacing student verification, interpretation, and justification across five learning domains. Drawing on social epistemology, self-regulated learning, and assessment theory, the authors distinguish ED-AIL from related constructs like AI literacy and automation bias. The paper offers diagnostic implications for pedagogy, assessment, policy, and learning assurance centered on preserving students' epistemic responsibility.
- Quality assurance
- AI policy
Research
Artificial Intelligence and Its Impact on Building an Effective Cybersecurity Protection Framework An Analytical Legal Study in Light of Saudi Legislation, International Standards, and Governance Frameworks
Yussri Abdalla
International Journal For Multidisciplinary Research · 2026-09-12
This legal-analytical study examines whether existing Saudi and international regulatory frameworks adequately govern AI use in cybersecurity, comparing Saudi legislation against standards such as ISO/IEC 27001:2022, NIST CSF 2.0, NIST AI RMF 1.0, and the EU AI Act. The findings show that AI substantially improves threat detection, incident response, and cyber risk management, but that effectiveness depends on an integrated governance framework covering legal compliance, institutional accountability, and algorithmic transparency—not technology alone. The study concludes that Saudi Arabia has a broadly aligned regulatory environment but needs a more specialized legal framework addressing AI liability, human oversight, and risk assessment to support its Vision 2030 digital transformation goals.
- AI policy
- Certifications
Research
Testing the Kill Switch: A Conformance-Based Approach to Agentic AI Containment Assurance
Naveen Sundaresan
Zenodo (CERN European Organization for Nuclear Research) · 2026-09-12
This paper proposes a conformance-based audit framework for verifying that 'kill switch' or stop mechanisms in agentic AI systems actually work under real conditions. It defines a Target of Evaluation and five testable control families—trigger recognition, authority, cessation, latency, and failure resilience—along with evidence requirements that go beyond procedural self-attestation. The framework is designed to fill a gap left by existing regulations and standards (EU AI Act, NIST AI RMF, ISO/IEC 42001, and others), which require human oversight capabilities but do not specify how independent auditors can verify them. The authors also propose a vendor-neutral assurance harness with pluggable adapters to generate the evidence such auditable criteria would require.
- Certifications
- AI policy
Research
EAIMS: Enterprise AI Maturity Standard
Elias Naserkhaki
Zenodo (CERN European Organization for Nuclear Research) · 2026-09-12
EAIMS 1.1.0 is a release of the Enterprise AI Maturity Standard that extends the 1.0.x baseline with nine new normative requirements covering adversarial and agentic governance for material AI systems, including controls for blast radius assessment, agent credential boundaries, runtime abuse evidence, and synthetic identity control. The update also extends the G3 Autonomous/Agentic gate family with five new gates and strengthens managed-service-provider dependency requirements. The release underwent automated specification, backward-compatibility, and dependency validation, though the authors explicitly note it does not claim independent third-party assessor validation, accreditation, or regulatory approval. This matters because it provides a structured, versioned framework enterprises can use to assess and govern AI systems against adversarial and agentic risk.
- Enterprise
- Certifications