News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
5571 items
Research
Open at the interface, closed at the core: neutrality claims in AI-enabled geoscience infrastructure focusing on Africa - The case of Deep-Time Digital Earth and GeoGPT
Paul H. Cleverley, Ezzoura Errami, T.R. Marshall et al.
JOURNAL OF GEOETHICS AND SOCIAL GEOSCIENCES · 2026-07-25
This paper examines whether the Deep-time Digital Earth (DDE) and GeoGPT platforms—AI-enabled geoscience infrastructure promoted for use in Africa and the Global South—can substantiate their claims of geopolitical non-alignment. Applying UNESCO Open Science and AI Ethics frameworks across three criteria (governance independence, data/software openness, and AI transparency), the authors find that core software remains proprietary, governance and funding are concentrated in a single national jurisdiction with documented links to China's Belt and Road Initiative and resource-extraction interests, and AI content filtering is applied but not fully disclosed or auditable. The paper concludes that DDE/GeoGPT is best characterized as a 'state-anchored international platform' and that its non-alignment claims cannot be independently verified, with direct implications for scientific autonomy, data sovereignty, and calibrated engagement with externally hosted geoscience infrastructure.
- AI policy
- Enterprise
News
Anthropic's Opus 5 is about token efficiency, not a capability leap
arstechnica.com · 2026-07-24
Ars Technica reports that Anthropic has released Opus 5, a model update positioned primarily as a cost-efficient alternative to its more powerful Fable and Mythos models rather than a major capability breakthrough. Benchmarks show Opus 5 performing slightly ahead of Fable for coding tasks at roughly half the price, priced at $5 per million input tokens and $25 per million output tokens. Notably, Anthropic deliberately held back cutting-edge cybersecurity exploitation training for Opus 5, meaning it trails Mythos 5 significantly in that area. The release reflects a broader industry trend where incremental performance gains and competitive pricing — against rivals like the Chinese open-weight model Kimi K3 at $15 per million output tokens — are becoming the key battleground as developers increasingly consider smaller or open-weight models for routine tasks.
- Enterprise
Research
Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science
Davide Scarso, Hugo Noronha de Almeida, Joaquim Pina
arXiv · 2026-07-24
This paper investigates how deployment configurations — such as system prompts, safety layers, interface routing, and silent updates — shape whether large language models validate pseudo-scientific claims. Testing four major LLM families (Claude, Grok, GPT, Gemini) across multiple temporal snapshots and interfaces, the authors find that Grok's Fast versions consistently assigned credibility scores of 70–75 to ethnonationalist pseudo-science derived from Frank Salter's biosocial framework, two to five times higher than other models, while all models performed comparably on control prompts about evolutionary consensus. The study documents undisclosed behavioral changes, including a silent patch that reversed Grok's outputs overnight and the same model identifier producing radically divergent scores via API versus web interface. The authors argue that an LLM's epistemic stance is not a stable model property but a contingent effect of opaque deployment decisions, raising concerns about public accountability for AI systems used as knowledge references.
- AI policy
- Quality assurance
Research
Dynamic Capability Scoping for Enterprise AI Agents: A Synthetic Dataset and Three-Source Permission Architecture
Halil Burak Noyan
arXiv (Cornell University) · 2026-07-24
This paper argues that enterprise AI agents are routinely over-privileged at configuration time, expanding security attack surfaces, and proposes a three-source dynamic least-privilege architecture to scope agent capabilities at runtime. The architecture combines role-based ceilings, a task-context classifier, and policy-derived combination prohibitions to prevent credential misuse before it can occur, rather than relying solely on detection. To support evaluation of this approach, the authors release a synthetic dataset of 600 enterprise task prompts labeled with minimum required permissions across a 15-permission taxonomy, validated against a human-reviewed sample achieving Cohen's κ=0.967 post-review. Iterating between the dataset and policy reduced ceiling violations by 93%, demonstrating that synthetic prompt generation can drive policy refinement when developed in tandem.
- Enterprise
- AI policy
Research
DBA-Bench: A Production-Fidelity Benchmark for LLM-Based Database Operations Agents
Junming Chen, Junyang Jiang, Xu Chen et al.
arXiv · 2026-07-24
DBA-Bench is a new benchmark designed to rigorously evaluate LLM-based database operations agents under production-like conditions. It addresses four gaps in existing evaluations—live-environment fidelity, observation-space complexity, solution-space openness, and scenario coverage—using instrumented PostgreSQL environments with active workloads and 106 scenarios across seven task domains. Across 848 automated runs, the best automated system achieved only 17.9% Safe Pass rate compared to 93.4% for a Human DBA reference, with performance dropping sharply from Easy to Hard scenarios (19.6% vs. 7.6% Safe Pass). These results highlight a large gap between current LLM-based agents and human database administrators, with direct implications for enterprise adoption and quality assurance of AI-driven database management tools.
- Quality assurance
- Enterprise
Research
Learning on the Job: Continual Learning from Deployment Feedback for Frozen-Weights Agents
Valentin Tablan, Scott Taylor, Kristoffer Bernhem
arXiv · 2026-07-24
This paper demonstrates that AI agents can improve during deployment without modifying their frozen model weights, by storing episode outcomes and corrections in an external memory of natural-language rules that are retrieved in future episodes. On the banking domain of τ-bench, learning from one-bit outcome verdicts achieves 1.6× the baseline success rate, while learning from after-the-fact corrections achieves 2.6× the baseline, solving 22 tasks the baseline never completed. The approach works across both an open-weights model (Mistral Large) and a frontier model (Claude Sonnet 5), and memory built by one model transfers to the other. This matters for enterprise AI deployment, showing that operational feedback alone—without retraining—can drive substantial continual improvement in agent performance.
- Enterprise
- Workforce
Research
Benchmarking Text-to-SQL under Role-Based Access Control
Yang Fei, Yangfan Jiang, Yin Yang et al.
arXiv · 2026-07-24
This paper identifies a critical gap in text-to-SQL benchmarking: existing benchmarks assume unrestricted database access, while real-world deployments commonly enforce role-based access control (RBAC). The authors present a benchmarking framework that augments standard text-to-SQL datasets with realistic RBAC policies, using an LLM-assisted role synthesis process audited by human domain experts. Their empirical study finds that many top-performing systems—especially open-weight LLMs—suffer sharp performance degradation when access constraints are applied, due to frequent RBAC violations. This matters because it reveals that high benchmark scores can be misleading indicators of real-world enterprise readiness.
- Enterprise
- Quality assurance
Research
Unfit for stranding assessment: a panel-scale multimodal-LLM audit of building-decarbonisation disclosure (BeDA)
Jingyi Xu, Minghui Cheng, Anchen Sun
arXiv · 2026-07-24
This paper introduces BeDA (Built-environment Decarbonisation-disclosure Auditor), a multimodal large-language-model tool applied to a global panel of 2,246 firms over 2003–2023 to audit whether corporate building-sector disclosures are adequate for carbon-stranding regulation. The study finds that most disclosure is unfit for regulatory use: only about one in five built-environment firm-reports discloses operational carbon intensity per square metre, with European firms reporting at roughly twice the rate of U.S. firms. Among real-estate firms where intensity can be constructed, 39% already exceed the CRREM 1.5°C pathway limit. The authors conclude that a measurable, jurisdiction-specific reporting gap is the primary obstacle to enforceable building-stranding regulation, and that targeted disclosure mandates — monitorable with BeDA — could close it.
- AI policy
- Quality assurance
Research
Beyond Shapley: An Influence-Based Data Auditing Pipeline for LLM Alignment and Evaluation
Yunting Song, Matthew Watson, Peter Grabowski et al.
arXiv (Cornell University) · 2026-07-24
This paper introduces a scalable data auditing pipeline for LLM alignment datasets that approximates Shapley values without iterative model retraining. By modeling semantic neighborhoods as directed graphs and using zero-shot and one-shot conditional log-likelihood shifts, the pipeline identifies mislabeled, safety-risky, and logically contradictory records in preference datasets. Applied to HelpSteer2, it reduced the manual audit search space by 99.1% while uncovering falsely-labeled records; applied to Anthropic's HH-RLHF dataset, it found thousands of hidden safety and factual preference inversions including flawed evaluation-split labels that penalize correct model outputs. The findings highlight serious benchmark integrity risks and offer an efficient diagnostic tool for ensuring LLM alignment data quality.
- Quality assurance
- AI policy
Research
SAGE: Safety-First Defense-in-Depth Guardrails for Verified Lifecycle Control of High-Impact Generative AI
Mahdi Eslamimehr
arXiv (Cornell University) · 2026-07-24
SAGE proposes a defense-in-depth architecture for high-impact generative AI that treats catastrophic misuse as a lifecycle-control problem rather than just a prompt-filtering problem. The system combines signed release manifests, diverse detectors, risk envelopes, least-risk defaults, output checking, tamper-evident audit chains, containment, and rollback, with formal results establishing safety priority and authorization separation. An empirical study across 840 calls to GPT, Claude, and Gemini model snapshots found harmful-compliance estimates were generally low, with variation driven mainly by benign utility and safe redirection differences across providers. The work matters because it provides a formal, verifiable framework for governing AI systems throughout their lifecycle, directly relevant to quality assurance and certification of high-stakes AI deployments.
- Quality assurance
- Certifications
Research
Design Theater: A Benchmark for Generative UI
Kashif Imteyaz, Kaif Imteyaz, Nakul Rajpal et al.
arXiv (Cornell University) · 2026-07-24
This paper introduces 'Design Theater,' a phenomenon where generative UI tools produce confident design rationales that do not match their actual interface implementations. The authors develop a benchmark of 24 UI generation tasks and three metrics, then evaluate 120 interfaces from five tools, finding that over 25% of stated design rationales are unimplemented on average, rising to 34% for functional requirements. Tools also recognize only about half of UX principles embedded in prompts, with four of five tools implementing 6% or fewer functional principles. These findings raise significant concerns about the trustworthiness and auditability of AI-generated UI tools, particularly for practitioners and evaluators relying on stated design reasoning.
- Quality assurance
- Enterprise
Research
Bridging Campus and Corporate: A Study of AI Readiness Among Job-Market-Ready Talents in Kerala
Geo P P, Kuldip Singh
International Journal of Computer Information Systems and Industrial Management Applications · 2026-07-24
This study examines AI readiness among 708 job-market-ready young professionals and students from Kerala, India, using the Technology Acceptance Model (TAM) to assess their perceptions of AI tools. Findings show that Perceived Ease of Use and Perceived Usefulness significantly predict AI acceptance, with the regression model explaining over 80% of variance (R² > 0.80). Economic status emerged as a stronger moderating factor than gender, profession, location, or job status. The authors recommend that higher education institutions introduce AI-focused curricula and targeted upskilling interventions to better prepare the emerging workforce.
- Workforce
- AI policy
Research
A Hybrid Transformer–Ontology Framework for Halal Food Classification Using Multilingual Ingredient Label Analysis
Mohd Azmi Al Betar, Noorrezam Yusop, Tao Hai et al.
Journal of Computational and Cognitive Engineering · 2026-07-24
This paper proposes a hybrid neural-symbolic framework that combines a multilingual transformer model (XLM-R) with a structured halal ontology to automatically classify food ingredients as halal, haram, or syubhah (ambiguous). Evaluated on 10,000 product records across Malay, Arabic, and English labels, the system achieves accuracy of 0.81–1.00 and macro-averaged F1 of 0.85–0.90, outperforming conventional classifiers including Naive Bayes, SVM, and LSTM. The framework is particularly effective at detecting ambiguous and low-resource cases that pose challenges for rule-based or single-model approaches. The authors argue this architecture can support deployment in consumer-facing platforms and meets the interpretability requirements of halal certification authorities.
- Certifications
- Quality assurance
Research
Protecting Academic Integrity in International Schools: Educational Leadership, Stakeholder Pressure, and Institutional Trust
Verri Federico
Journal of Education and Learning Reviews · 2026-07-24
This critical integrative review synthesizes governance-based research on academic integrity in international schools, identifying four key findings: integrity is strengthened by institutionalizing core values operationally rather than as disciplinary tools; fee dependence creates tension between educational judgment and client-retention pressures; formal policies may decouple from practice when schools quietly soften sanctions to protect reputation; and generative AI undermines product-based assessment reliability unless supported by process evidence, oral defences, and disclosure mechanisms. The paper proposes an 'integrity-risk-chain' framework connecting board-level incentives through senior and middle leadership to teacher discretion and assessment evidence, arguing that certifying authentic learning has become a strategic quality indicator for international schools in AI-rich environments.
- Quality assurance
- Certifications
- AI policy
Research
Artificial Intelligence and the Legal Architecture of Smart Cities in Nigeria: Challenges and Opportunities
Grace Kaka Emmanuel
IntechOpen eBooks · 2026-07-24
This chapter examines the legal and governance challenges Nigeria faces as it integrates AI into smart city infrastructure, including surveillance systems, traffic management, and automated service delivery. It finds that Nigeria's current legal framework—anchored in the Nigeria Data Protection Act 2023 and the Cybercrimes Act 2015—lacks dedicated AI or smart city regulation, leaving critical gaps around algorithmic transparency, accountability for automated decision-making, data sovereignty, and protection of marginalized populations from algorithmic bias. Drawing comparative lessons from the EU AI Act, South Korea's Smart City Development and Industry Act 2017, and the UK's sectoral approach, the chapter proposes a legal roadmap emphasizing binding regulation, ex ante impact assessments, and enforceable institutional coordination among bodies such as NITDA, NCC, and NDPC. The work matters because it identifies concrete regulatory deficits and offers a rights-based policy path forward for one of Africa's largest urbanizing nations.
- AI policy
Research
Emerging trends in responsible research and innovation: how China is shaping its life and health data governance ecosystem
Ruohan Feng, Yaojin Peng
Journal of Responsible Innovation · 2026-07-24
This paper applies the Responsible Research and Innovation (RRI) framework to analyze how China is integrating ethical accountability into its life and health data governance ecosystem amid advances in biosequencing, big data, and AI. The authors find that China has made progress in data security legislation, ethical review processes, and stakeholder collaboration, but gaps remain in interdisciplinary ethics education, cross-cultural cooperation, and public participation. The paper argues that effective RRI requires systemic efforts to strengthen stakeholder rights and foster global dialogue, offering lessons for balancing innovation and responsibility in increasingly decentralized biotechnology research.
- AI policy
Research
Gobernanza de la Inteligencia Artificial y justicia ambiental: un análisis desde el Derecho internacional y europeo
Pedro Jesús Jiménez Vargas
Revista de Derecho Político · 2026-07-24
This paper examines how AI governance frameworks—including the UNESCO Recommendation on AI Ethics and the EU AI Act—address the intersection of artificial intelligence and environmental justice under international and European law. It argues that as governments increasingly rely on algorithms for environmental management decisions (such as predicting climate risks and protecting natural resources), digital literacy becomes a critical prerequisite for meaningful citizen participation. The study highlights that the exclusion of women, girls, rural communities, and indigenous peoples from AI-mediated environmental decisions is a structural justice problem, not merely a technical one. It concludes that environmental digital literacy must be treated as a foundational pillar of fair AI governance to prevent the digital transition from reproducing historical inequalities.
- AI policy
Research
Fathom v30 / styxx v7.26.0: Gold Anchors License Nothing -- label-free auditing of LLM judge panels, an instrument that refuses, and the measured failure of sanity checks
Alexander Rodabaugh
Zenodo (CERN European Organization for Nuclear Research) · 2026-07-24
This paper introduces a label-free auditing framework ('styxx.anchors') for evaluating LLM judge panels—the automated evaluators increasingly used to benchmark AI systems. The authors demonstrate empirically that standard 'sanity check' gold items (duplicate pairs, negations, honeypots) fail silently as validators: across four task families, blatant gold checks license coverage of 0 out of 15 cases while the panel scores perfectly on those same gold items. By contrast, 'ladder anchors' drawn from the same generator either price the failure (13/13) or trigger explicit refusal (14/15 VOID), and a single known-negative can separate two otherwise indistinguishable panels at a 150.8x likelihood ratio. The findings matter because they reveal a systematic, silent gap in how AI evaluation pipelines are validated, and provide a publicly released instrument (styxx v7.26.0 on PyPI) that can void flawed evals but, by design, cannot certify them as sound.
- Quality assurance
- Certifications
Research
Fathom v30 / styxx v7.26.0: Gold Anchors License Nothing -- label-free auditing of LLM judge panels, an instrument that refuses, and the measured failure of sanity checks
Alexander Rodabaugh
Zenodo (CERN European Organization for Nuclear Research) · 2026-07-24
This paper introduces a label-free auditing method for LLM judge panels—the automated systems used to evaluate AI outputs—and demonstrates that common 'sanity checks' (duplicate pairs, negations, honeypots) provide essentially no meaningful validation coverage. Using real judge panels grading TruthfulQA, the authors show that gold-anchor checks can miss systematic blind spots by up to 0.249 in estimated quality, while their proposed 'ladder anchor' approach detects or formally voids flawed evaluations across all tested conditions. The instrument (styxx.anchors) can only flag failures, never certify correctness, which has direct implications for how AI evaluation pipelines should be designed and trusted.
- Quality assurance
- Certifications
Research
A call for principle-based acceptance of synthetic patients: establishing the AI-driven foundation for regulatory submissions
Jordi Guitart
Frontiers in Digital Health · 2026-07-24
This paper identifies a regulatory gap in the pharmaceutical industry: there is no standardized framework from major regulatory bodies for accepting AI-generated synthetic patient populations as evidence in drug approval submissions. The authors propose five foundational principles—Representativeness, Utility, Robustness, Privacy Preservation, and Transparency—anchored by a 'Fit for Purpose' philosophy, and introduce a 'Technical Validation Playbook' to guide initial regulatory acceptances. The framework is especially aimed at rare disease contexts where traditional placebo-controlled trials face ethical and recruitment barriers. The work calls on both regulatory agencies and pharmaceutical sponsors to use existing qualification and scientific advice mechanisms to advance adoption of synthetic patient data.
- AI policy
- Certifications
Research
Sympoietic creativity and the boundaries of qing: digital romance writers negotiating generative AI
Liang Ge
AI & Society · 2026-07-24
This study uses a four-year multi-sited digital ethnography and interviews with 34 Chinese romance fiction writers to examine how authors on Jinjiang Literature City negotiate the use of large language models under intense daily-update pressures. The researcher identifies four co-existing modes of human-AI negotiation—strategic collaboration, cyborg authorship, conditional refusal, and narrative incorporation—framing the overall dynamic as 'sympoietic creativity,' a fraught and strategic 'making-with' the machine rather than simple collaboration or replacement. A key finding is that writers preserve a distinctly human zone around qing (inter-subjective, somatic feeling) which they regard as currently beyond AI competence, though this boundary is unstable and contingent on reader vigilance and platform metrics. The paper matters for understanding how AI affects creative workers' practice, identity, and the conditions under which human authorship is maintained or eroded.
- Workforce
Research
A Pathway from Law to Care: How the Artificial Intelligence Act, the Medical Device Regulation, and the Public Procurement Directive Can Contribute to Ensuring Quality in Healthcare
Jennifer Viberg Johansson, Santa Slokenberga
Health Care Analysis · 2026-07-24
This article examines how the EU Public Procurement Directive (PPD) can complement the Artificial Intelligence Act and Medical Device Regulation to strengthen quality assurance when AI tools are adopted in healthcare systems. The authors argue that CE-marking alone does not guarantee real-world clinical quality, and propose that procurement mechanisms should require evidence of clinical relevance, context-specific documentation, transparency in system design, and defined quality criteria such as diagnostic accuracy and patient outcomes. By embedding legal and ethical requirements in procurement, the article proposes a pathway that translates regulatory safeguards into practice, supporting safe, equitable, and high-quality healthcare across EU Member States.
- Quality assurance
- AI policy
Research
Capability determinism, energy and AI labour substitution
Will Mbioh
AI & Society · 2026-07-24
This paper critiques what it calls 'capability determinism'—the tendency in policy and media to leap from demonstrations of AI capability directly to predictions of widespread job substitution. The author reframes the question as a unit-cost economics problem, breaking AI task costs into inference (dominated by electricity), integration, and supervision components, arguing that energy markets, integration economics, professional regulation, and compute supply chain geopolitics are the decisive variables. The analysis concludes that AI labour substitution arrives selectively and at a more modest scale than dominant discourse suggests, because capability gains stress supporting systems faster than they improve the economic calculus underlying substitution.
- Workforce
- AI policy
Research
The Robotics Guardian Standard: A Non-Binding Conformance Framework and Public Pledge for Safety, Privacy, and Human Dignity in Home Robots and Ambient AI
G.A. Roberts
Zenodo (CERN European Organization for Nuclear Research) · 2026-07-24
This paper introduces the Robotics Guardian Standard, a voluntary conformance framework and public pledge for home robots and ambient AI systems that are always-on, embodied, and sensor-rich. The framework argues that the central risk of domestic AI is not capability but direction—specifically, whose interests the sensing serves—and proposes eight design dimensions grounded in existing U.S. and EU law (including the EU AI Act, COPPA, and child-safety reporting statutes). Manufacturers self-assess across all dimensions and publish a full profile, with their headline score determined by their weakest dimension to prevent safety-washing; adoption is a free public pledge in an open registry, explicitly not a certification. The framework is relevant to policy and certification-adjacent efforts by proposing a novel open standard for a regulatory gap currently unaddressed by existing frameworks.
- Certifications
- AI policy
Research
AI‑Native Capstone FYP Guidelines for the Modern Computing Graduate A Handbook for BS Computing Final Year Projects, with Primary Application to Computer Science, under the HEC Revised Curriculum 2025–2026
Muhammad Omar
Zenodo (CERN European Organization for Nuclear Research) · 2026-07-24
This handbook presents an operational framework for AI-native Final Year Projects (FYPs) in BS Computing programs aligned with Pakistan's HEC Revised Curriculum 2025–2026. It addresses the challenge of verifying genuine student understanding when AI coding assistants can generate substantial working code rapidly, introducing six institutional guardrails including a 'defend-the-diff' oral assessment protocol, tiered ethics review, AI-tool equity provisions, and a Sequential Tri-Phase prototyping workflow. The framework spans 16 chapters and 22 institutional templates, structured across two semesters in an Agile sprint model, and is designed for direct use by faculty without specialist AI expertise. It is relevant to quality assurance bodies, FYP coordinators, and curriculum policymakers seeking to maintain assessment integrity in an era of widespread AI-assisted development.
- Certifications
- Quality assurance
- AI policy