News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
5221 items
Research
Rethinking AI in clinical decision support: a framework for reciprocal human-AI interaction
Colin Greengrass
Frontiers in Digital Health · 2026-09-11
This paper introduces BRACE (Bounded Reciprocal Adaptation for Clinician Engagement), a framework for AI-assisted clinical decision support that centers the clinician-AI interaction rather than AI model performance. The framework addresses risks of overreliance and skill degradation—including 'never-skilling,' deskilling, and mis-skilling—by making uncertainty visible, preserving clinicians' reasoning states, and bounding what the AI system may infer or modify about the clinician. The central hypothesis is that the amount of cognitive and metacognitive work preserved during AI-assisted encounters predicts independent clinical capability when AI is unavailable, evaluated through longitudinal within-clinician analyses. The work matters because it directly addresses how repeated AI use may erode clinicians' independent diagnostic reasoning, particularly for ambiguous cases.
- Workforce
- Quality assurance
Research
From DevOps to XOps: an agent-driven reference architecture for autonomous enterprise operations
Mete KÖSE, Ecir Uğur Küçüksille
Scientific Reports · 2026-09-11
This paper proposes XOps, a five-layer reference architecture that unifies fragmented enterprise ML operations (DataOps, MLOps, AIOps) under an agentic orchestration layer with Policy-as-Code governance. Two synthetic case studies evaluate the architecture: a self-healing payment gateway achieving 85.6% action consistency and 99.6% fault classification accuracy with no policy-violating actions executed, and a predictive-maintenance application maintaining R²=0.74 versus 0.29 for a static model. An indicative cost analysis suggests roughly 70% reduction in expected monthly operational costs, though the authors caution these results demonstrate feasibility rather than production-scale performance. The work is relevant to enterprise AI governance and autonomous operations management.
- Enterprise
- Quality assurance
Research
Rethinking Human Capital Development in the Age of AI
Juan M. Lavista Ferres, Frank Nagle
arXiv · 2026-09-11
This paper argues that nontraditional educational providers like LaunchCode and Per Scholas offer a model for adapting workforce training to an AI-driven labor market, where technical skills become obsolete more rapidly. The authors show that shortening the feedback loop between employer skill needs and training programs is central to effective human capital development. They use these two organizations as case studies to derive a roadmap that traditional educational institutions can adopt to scale similar innovations.
- Workforce
Research
Governing with Artificial Intelligence: Use, Ideology, and the Benefits and Risks of AI in State Government
Zachary Baum
arXiv · 2026-09-11
This study surveys U.S. state government professionals to understand how they perceive the benefits and risks of AI in government. It finds a striking asymmetry: frequent AI use strongly predicts perceived benefits, while political ideology—not usage—is the dominant predictor of perceived risk, with more conservative respondents seeing lower AI-related risks. These findings suggest that attitudes toward governmental AI adoption are shaped by distinct and separable factors depending on whether benefits or risks are being assessed. The results have direct implications for how AI policies are likely to be evaluated and adopted across differently ideologically-aligned state governments.
- AI policy
Research
Artificial Intelligence Integrity & Public Access: A Federal Legislative Framework
Terrance J. Chisolm
Zenodo (CERN European Organization for Nuclear Research) · 2026-09-11
This policy paper by Terrance J. Chisolm proposes a federal legislative framework called the Artificial Intelligence Integrity, Accountability, and Public Access Act of 2026, designed to govern AI at the federal level. The framework establishes four core principles: protecting people, preserving public access to AI, requiring evidence-based attribution when AI is alleged to cause serious harm, and maintaining accountability for humans and institutions that deploy or misuse AI. It proposes specific mechanisms including standards for AI incident attribution, privacy-protective forensic accountability for high-risk AI deployments, independent investigation of catastrophic AI incidents, and safeguards against fabricated AI-attribution evidence. The proposal explicitly rejects AI legal personhood and focuses instead on protecting evidence integrity while balancing innovation, competition, civil liberties, and safety.
- AI policy
Research
Current Applications of Artificial Intelligence in Orthopaedic Trauma: A Narrative Review
Ashutosh Yadav, Pushpa ., Sachin Kumar
International Journal of Science and Healthcare Research · 2026-09-11
This narrative review synthesizes evidence on AI applications across the orthopaedic trauma care pathway, including fracture detection, classification, outcome prediction, operative support, and large language model tools. Pooled sensitivity and specificity for fracture detection on plain radiographs reached roughly 0.87–0.91, comparable to specialist clinicians, and clinician sensitivity rose to 0.97 when AI was used as an adjunct; however, fracture classification accuracy remained weaker (60–81%) and outcome-prediction models offered little improvement over conventional regression. Generative language models showed early promise for documentation and patient education but produced clinically relevant errors and did not reach resident-level performance. The authors conclude that while AI has achieved specialist-level accuracy for fracture detection, patient-level benefit has not been demonstrated, and call for representative multicentric datasets, external validation, and prospective clinical-impact trials, especially in low- and middle-income countries such as India.
- Quality assurance
- AI policy
Research
Implications of the new US AI framework in medicine
Antonis A. Armoundas
Communications Medicine · 2026-09-11
This policy analysis examines the March 2026 White House National Policy Framework for Artificial Intelligence and its implications for medical AI governance. The framework is described as pro-deployment and pro-infrastructure, relying on sector-specific oversight rather than a new central regulator, which may accelerate AI adoption in medicine by expanding data access and reducing infrastructure barriers. However, the authors identify a critical governance gap: clinically consequential AI tools that fall outside traditional FDA-regulated pathways remain largely unaddressed, leaving risks around transparency, clinical accountability, and protection of vulnerable populations unresolved. The paper concludes that stronger institutional accountability, consumer-protection mechanisms, and international interoperability are needed to keep pace with rapidly expanding medical AI.
- AI policy
- Quality assurance
Research
Clustering populations for holistic policymaking : promises and limits in British Columbia’s energy transition
Aloysio Kouzak Campos da Paz
cIRcle (University of British Columbia) · 2026-09-11
This thesis develops an AI-based clustering framework—using PCA, k-means, and eta-squared—to group 147 British Columbia municipalities by barriers and enablers to residential and transportation electrification. Results consistently identified one cluster combining remoteness, lower education, smaller populations, and weaker institutional climate capacity, and another less-constrained cluster, suggesting differentiated policy instruments (e.g., regional training hubs and logistical support for remote areas versus income-differentiated rebates elsewhere). The study shows clustering can make policy targeting more holistic than single-variable approaches like latitude or income alone, but also flags that results are sensitive to methodological choices, cluster quality scores were lower than expected, and small communities were excluded due to data gaps. The authors caution that co-occurring barriers do not reveal causation, and that equity trade-offs in data coverage must be managed carefully.
- AI policy
Research
Safety Assurance Methodology for AI-Based Aerospace Systems - Replication package
Alberto Petrucci, Francesco Basciani, Patrizio Pelliccione
Zenodo (CERN European Organization for Nuclear Research) · 2026-09-11
This replication package documents AeroSafe, an ECSS-oriented assurance framework for AI and machine learning software in aerospace and other safety-critical systems. The package includes instruments, response data, and a 52-control catalogue validated across two practitioner review rounds, covering areas such as configuration management, data assurance, model verification, deployment, safety argumentation, and Independent Model Verification and Validation. Practitioner ratings from the second round indicate mean scores of 4.4 for understandability, 4.6 for coverage of expected ECSS-oriented AI/ML assurance obligations, and 4.2 for acceptable effort, suggesting the framework is perceived as clear and appropriately scoped. The materials support researchers and practitioners developing structured assurance approaches for AI-based critical systems, though the authors note the results reflect perceived usability rather than demonstrated defect-detection effectiveness or formal certification compliance.
- Quality assurance
- Certifications
Research
Displacement Without Redundancy: Ricardo's Machinery Chapter, the Acemoglu–Restrepo Task Model, and Four Years of Generative AI
Benjamin Frohman
Zenodo (CERN European Organization for Nuclear Research) · 2026-09-11
This paper reconstructs Ricardo's 1821 machinery argument and Acemoglu-Restrepo's task model to evaluate the first four years of generative AI's labor market effects. While aggregate U.S. employment (~163 million) and unemployment (4.1%) show no broad collapse, the paper identifies concentrated harm: workers aged 22–25 in highly AI-exposed occupations show employment roughly 19 percent below peers in less-exposed work, AI is the leading stated reason for announced U.S. job cuts for five consecutive months in 2026, and the BLS labor-share index has continued to fall. The paper argues that stable headline unemployment figures are the wrong lens—displacement without offsetting reinstatement of labor into new tasks remains the live economic concern, especially for early-career workers—and proposes two monitoring statistics (the canary residual and wage-fund conversion ratio) to track this margin going forward.
- Workforce
- AI policy
Research
A critical assessment of Nepal's digital data protection framework: Legal gaps and reform directions
Tul Bahadur Khadka
Humanities and Social Sciences Journal · 2026-09-11
This qualitative doctrinal study examines Nepal's fragmented legal framework for digital data protection, finding it inadequate for modern challenges including AI, cloud computing, and cross-border data transfers. Drawing on key informant interviews with legal and technology policy experts and comparative analysis of frameworks like the EU's GDPR, the authors identify critical gaps including the absence of an independent Data Protection Authority, weak private-sector obligations, and limited enforcement. The paper argues Nepal needs a comprehensive Digital Data Protection Act alongside institutional reform to protect citizens' informational privacy and sustain digital transformation.
- AI policy
Research
Safety Assurance Methodology for AI-Based Aerospace Systems - Replication package
Alberto Petrucci, Francesco Basciani, Patrizio Pelliccione
Zenodo (CERN European Organization for Nuclear Research) · 2026-09-11
This replication package documents the practitioner validation of AeroSafe, an assurance framework for AI/ML software in aerospace and other safety-critical systems, aligned with the ECSS standard and informed by AMLAS. The package includes instruments, response data, and a final 52-control catalogue validated across two rounds by small groups of practitioners (six and five participants respectively), covering areas such as data management, model verification, deployment, safety traceability, and Independent Model Verification and Validation. Results are descriptive: second-round participants rated the catalogue 4.4/5 for understandability, 4.6/5 for coverage of ECSS-oriented AI/ML assurance obligations, and 4.2/5 for acceptable effort, with no controls receiving a 'Unclear or unjustified' judgment. The work matters because it provides an openly available, structured assurance aid for certifying AI/ML systems in high-stakes aerospace contexts, though the authors note the results do not constitute certification or ECSS compliance.
- Certifications
- Quality assurance
Research
Displacement Without Redundancy: Ricardo's Machinery Chapter, the Acemoglu–Restrepo Task Model, and Four Years of Generative AI
Benjamin Frohman
Zenodo (CERN European Organization for Nuclear Research) · 2026-09-11
This paper argues that stable headline unemployment figures in the U.S. and other Western economies do not settle the deeper debate about AI-driven labor displacement. Drawing on Ricardo's 1821 'On Machinery' and the Acemoglu–Restrepo task model, it distinguishes between economy-wide job loss and narrower, structural displacement: payroll data through June 2026 show employment of workers aged 22–25 in highly AI-exposed occupations running roughly 19 percent below comparable peers in less-exposed roles, AI is cited as the leading reason for announced U.S. job cuts for five consecutive months in 2026, and the BLS labor-share index has continued to fall. The paper introduces two monitoring statistics—the 'canary residual' and the 'wage-fund conversion ratio'—to track whether automation is being offset by new-task creation, concluding that compensation is a contingent mechanism rather than a guaranteed law.
- Workforce
- AI policy
Research
Artificial Intelligence Integrity & Public Access: A Federal Legislative Framework
Terrance J. Chisolm
Zenodo (CERN European Organization for Nuclear Research) · 2026-09-11
This policy paper proposes a federal legislative framework called the Artificial Intelligence Integrity, Accountability, and Public Access Act of 2026, developed by Terrance J. Chisolm. The framework establishes four core principles: protecting people, preserving public access to AI, requiring evidence-based attribution when AI is alleged to cause harm, and maintaining accountability for humans and institutions that deploy or misuse AI. It proposes standards for AI incident attribution, privacy-protective forensic accountability for high-risk deployments, independent investigation of catastrophic AI incidents, and safeguards against fabrication of AI-attribution evidence. The proposal explicitly does not grant AI legal personhood and seeks to balance safety, civil liberties, innovation, and competition in federal AI regulation.
- AI policy
- Certifications
Research
Explainability Assistant: A Conversational XAI Interface for Interpreting Energy Consumption Models
Rodion Krjutškov, Eduard Barbu, Nikos Sakkas et al.
arXiv · 2026-09-10
This paper introduces the Explainability Assistant, an open-source conversational AI system that helps facility managers and building operators interpret complex machine learning models used for energy consumption forecasting. By leveraging the function-calling capabilities of modern Large Language Models, the system achieves 94% intent-parsing accuracy — up from 76.8% in prior approaches like TalkToModel — without requiring task-specific fine-tuning. An evaluation with energy domain specialists found improved usability and consistent task accuracy compared to traditional XAI dashboards, with all experts unanimously preferring the conversational interface for practical use. The work matters because it lowers the technical barrier for non-expert users to interrogate and trust AI-driven energy models in real operational settings.
- Enterprise
- Workforce
Research
Target leakage, not model class, explains reported accuracy in survey-based cardiovascular screening: a leakage-tiered audit of glass-box and tabular foundation models
Raad Bin Tareaf, Murad Al-Rajab, Samia Loucif et al.
arXiv · 2026-09-10
This paper audits ten machine learning classifiers—including linear models, tree ensembles, neural networks, glass-box models, and tabular foundation models—for predicting prevalent myocardial infarction using large national health survey data (over 440,000 respondents). The central finding is that the widely reported ~0.89 AUROC in cardiovascular screening literature is largely driven by target leakage from post-diagnostic features, not genuine model performance: removing just two such features collapses all models into a narrow 0.0045-wide AUROC band. The glass-box explainable boosting machine matched every alternative in discrimination while being roughly 104 times faster than the best foundation model, and its transparency directly enabled fairness repairs and uncertainty calibration. The authors conclude that evaluation methodology and feature set construction—not model capacity—are the binding constraints in this domain.
- Quality assurance
- AI policy
Research
SpecGuard: Inference-Time Backdoor Detection For Free
Rui Wen, Ahmed Salem, Andrew Paverd et al.
arXiv · 2026-09-10
SpecGuard proposes a zero-cost inference-time backdoor detection method for large language models by repurposing speculative decoding, a technique already used to speed up inference. The key insight is that when a backdoor trigger activates a fine-tuned target model, a clean draft model fails to predict the resulting behavioral shift, causing a measurable change in draft-token acceptance rate. The authors formalize when this signal is detectable and show that any attacker who suppresses it must also weaken the backdoor itself. Across diverse backdoor types and model families, SpecGuard reliably detects triggered behavior—including stealthy cases that bypass input-level filters—without requiring extra model computation.
- Quality assurance
Research
The widening evaluation gap in medical large language model research 2023 to 2026
Raad Bin Tareaf, Murad Al-Rajab, Samia Loucif
arXiv (Cornell University) · 2026-09-10
This study examined 11,628 PubMed records on medical large language model (LLM) research published from January 2023 to June 2026, finding that the field has grown 45-fold but that clinical evaluation is struggling to keep up with rapid model development. Only 2.5% of studies used a randomised, controlled, or prospective design, and the 'evaluation lag'—the gap between a model's release and the publication of a study evaluating it—widened from 1.33 to 6.08 quarters over that period. Critically, randomised trials evaluated models a median 4.6 quarters older than other study designs, and 62% of randomised trials evaluated a discontinued model family, revealing a structural tension between research rigour and currency. The authors attribute this gap to model selection choices rather than research timelines, raising serious concerns about the relevance of high-quality clinical evidence by the time it is published.
- Quality assurance
- AI policy
Research
SIRF: A Spec-Internalized Risk Foundation Model for Industrial Content Risk Control
Suwan Wu, Yumeng Lin, Pengcheng Yuan et al.
arXiv · 2026-09-10
SIRF (Spec-Internalized Risk Foundation Model) is a foundation model designed for industrial content risk control that embeds a platform's complex policies directly into model weights through continued pretraining, eliminating the need for additional human annotation. Using techniques called EntiGraph, MAGA rewriting, and account-level chain-of-thought synthesis, the 8B-parameter model achieves 71.3% Black Recall at 95% precision—a 15.1 percentage point improvement over a same-architecture baseline—while operating under second-level latency constraints. When deployed as a tree-model adjudication layer, it recovers 20% more mis-penalized samples and reduces mis-penalization by roughly 70% in a freezing scenario. The work demonstrates that internalizing policy rules into model weights, rather than injecting them at inference time, is a practical path to high-precision, low-latency automated content moderation at industrial scale.
- Enterprise
- Quality assurance
Research
Geospatial AI, Dataverse Metadata, and the Study of Place-Based Government
Danny EBanks, Devika Jain
arXiv (Cornell University) · 2026-09-10
This paper constructs a knowledge graph from Harvard Dataverse's public metadata, organizing 102,650 datasets into a 215,985-node network to make geospatial and policy-relevant research more discoverable. The authors find that 42.9 percent of datasets carry at least one geospatial field, and a keyword search identifies 7,654 datasets (17.4 percent) as directly policy-relevant, with elections and legislatures forming the largest cluster. The paper demonstrates how AI tools—including community language models, stance detection with geographic aggregation, and partisan language bridging—can link political discourse to place. It argues that the graph provides a concrete testbed for AI-driven metadata enrichment and entity resolution, while noting a coverage skew toward American, city-level data.
- AI policy
Research
Who Bears the Risk When Generative AI Enters Transport? A Distributional Sociotechnical Audit of Algorithmic Equity, Synthetic-Data Validity, and Public Trust
Amir Rafe, Subasish Das
arXiv (Cornell University) · 2026-09-10
This paper develops a Distributional Sociotechnical Audit (DSA) framework to measure how generative AI systems create unequal risks across different population groups in transportation contexts, including traveler advisories, synthetic crash-record generation, and policy decision support. Analyzing 5,760 persona-controlled queries across 12 demographic cues, three synthetic crash-record generators, and public attitude data from 4,538 respondents, the study finds that congestion-pricing AI advice shows the highest demographic disparity (mean EDI = 1.96), that CART-based synthetic crash records fail all conditional validity tests, and that existing categorical AI approval frameworks flip their tier assignments 75% of the time under minor weight perturbations. The authors argue that continuous risk indices with sensitivity reporting provide a more defensible basis for transport AI governance than the categorical approval tiers currently in use.
- AI policy
- Quality assurance
Research
Prompt Revision as a Source of Cultural Bias in Text-to-Image Systems
Aleksandra Urman, Elsa Lichtenegger, Salima Jaoua et al.
arXiv · 2026-09-10
This paper investigates how commercial text-to-image systems (DALL-E-3, Imagen-4, GPT-Image-1.5) silently revise user prompts before generating images, and whether that revision step introduces cultural bias. Using WORLDVIEW, a new multilingual benchmark of 8,960 prompts across 15 languages and 31 language-context pairings, the authors find that non-Western and non-Anglophone cultural contexts are marked far more heavily than a US/English baseline, flattened into narrow vocabularies, and reduced to stereotypes. By comparing outputs from original versus revised prompts on models without a revision layer, the study identifies the prompt-revision layer itself as a previously undocumented, causal source of cultural stereotyping. The findings argue that bias audits must examine the full deployed system—not just the generative model—to accurately locate and fix the problem.
- AI policy
- Quality assurance
Research
ActMap: Single-Pass Uncertainty Quantification from Generation-Time Activation Maps
Jacopo Dardini, Roberta Calegari
arXiv · 2026-09-10
ActMap introduces a white-box uncertainty quantification method for large language models that works from a single generation pass, requiring no multiple samples. It compresses the full hidden-state trajectory across every layer and every generated token into a compact fixed-size tensor (96 KiB), which a lightweight classifier then reads to estimate whether a model's answer is correct. Evaluated on short-answer QA, math, and summarization factuality with three 7–8B instruction-tuned models, ActMap consistently outperforms sampling, token-probability, attention, and embedding baselines, and matches a detector trained on tensors 67× larger at essentially the same mean AUROC with lower calibration error on ten of twelve pairs. The resulting correctness score enables abstention, routing, and selective verification, making it a practical tool for scalable oversight of deployed language models.
- Quality assurance
- Enterprise
Research
Characterizing Bluesky Content Moderation Service: From Automation of Service to Landscape of Harms
Pushpdeep Singh, Sayeh Jarollahi, Ayan Majumdar et al.
arXiv (Cornell University) · 2026-09-10
This paper conducts the first large-scale audit of Bluesky's default content moderation system (BMS) by analyzing 10.6 million moderation labels from 2025, made accessible through the platform's transparent, decentralized architecture. The study finds that BMS operates as a human-AI collaborative system—automatically labeling sexual and graphic content within seconds while routing nuanced or high-stakes cases to human reviewers over hours or days. The system achieves high precision (0.837) but low recall (0.222), with human annotators identifying 4.5× more harmful content than the system in a random sample, indicating substantial under-detection. The findings provide a data-driven baseline for improving transparency and effectiveness in automated content moderation.
- AI policy
- Quality assurance
Research
NovGauge: A Fine-Grained Benchmark for Diagnosing LLMs' Capability in Paper Novelty Assessment
Guoqiang Zhang, Kexin Tan, Ming Zhang et al.
arXiv · 2026-09-10
NovGauge is a new benchmark designed to evaluate how well large language models (LLMs) can assess the novelty of research papers across three fine-grained dimensions: task, problem, and method. The benchmark contains 619 paper pairs and 50 multi-paper sets labeled by human experts, and uses a cascading diagnostic pipeline to check not just whether a model's judgment is correct but whether its supporting evidence is logically faithful. Evaluation of 18 LLMs reveals hallucination rates from 0% to 39% and shows that over 70% of non-hallucinated correct-positive judgments cite evidence that fails to logically support the stated reason; even the best model (GPT-5.5) achieves only 43–72% Verified F1 after faithfulness verification. These findings indicate that current LLMs are not yet reliable for scientific novelty assessment, with important implications for AI-assisted peer review at major conferences.
- Quality assurance