News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
On AI Safety and Security Technical Debt in Engineering AI-Enabled Systems
Muhammad Tukur, Hayatullahi B. Adeyemo, Tao Chen et al.
arXiv (Cornell University) · 2026-07-25
This paper examines AI Technical Debts (AITDs)—engineering liabilities arising from data governance, model implementation, algorithm design, architecture, operations, documentation, and testing—in AI systems deployed in high-stakes domains such as healthcare, autonomous driving, finance, and education. Through a systematic review of 60 primary studies, the authors identify 31 distinct AITD types organized into a seven-class taxonomy and map them to 18 trust-related concerns including 6 safety hazards and 12 security vulnerabilities. The study synthesizes 34 actionable mitigation guidelines and introduces AITD-MAP, an integrated framework connecting the taxonomy, quality and risk impacts, and mitigation strategies to support risk-aware AI engineering. The work matters because it makes AI safety and security debts visible and actionable for engineers building and maintaining AI-enabled systems across critical sectors.
- Quality assurance
- AI policy
Research
Artificial Intelligence for radiation protection in medical imaging and radiotherapy: A perspective from the AI Working Party of ICRP Committee 3
John Damilakis, Mika Kortesniemi, Sébastien Gros et al.
Physica Medica · 2026-07-25
This ICRP Committee 3 Working Party perspective examines how AI is being integrated into medical imaging and radiotherapy workflows and what this means for radiation protection of patients, staff, and the public. The authors find that the most clinically mature AI applications include decision support for referral appropriateness, image reconstruction, protocol optimisation, automated contouring, treatment planning, and AI-enabled quality assurance, while areas like patient-specific dosimetry and synthetic imaging remain at earlier validation stages. Key challenges identified include data biases, limited generalisability, 'black box' accountability issues, and the need for harmonised regulatory oversight and AI literacy training for healthcare professionals. The paper provides a framework intended to inform future ICRP recommendations on safely integrating AI into medical radiation protection practice.
- Quality assurance
- AI policy
Research
Artificial Intelligence in Talent Acquisition: Procedural Justice, Organizational Attractiveness, and Perceived Competitive Positioning in Digital Talent Markets
Ovidiu-Iulian Bunea, Răzvan-Andrei Corboș, Bianca MIHAI
Administrative Sciences · 2026-07-25
This study investigates how prospective job applicants perceive AI-assisted hiring processes and what those perceptions mean for employer attractiveness and competitive positioning in talent markets. Using PLS-SEM with 202 respondents (mostly aged 18–24), the research finds that perceived AI expertise and trust in AI are positively linked to anticipated procedural justice, which in turn is the strongest predictor of organizational attractiveness, and organizational attractiveness strongly predicts perceived talent-market competitive positioning. Bootstrapped indirect effects further reveal a sequential chain: technology appraisals flow through procedural justice and organizational attractiveness to shape how competitive a firm appears in the talent market. The findings suggest that how organizations design and communicate their AI-driven selection processes meaningfully affects their ability to attract candidates.
- Workforce
- Enterprise
Research
Corporate attention to generative AI and earnings management: evidence from China
Zhaodong Li
The North American Journal of Economics and Finance · 2026-07-25
Using 25,208 firm-year observations of Chinese A-share firms from 2011 to 2023, this study finds that higher corporate attention to generative AI—measured via TF-IDF indices from MD&A disclosures—is associated with lower real earnings management but higher textual earnings management (abnormal optimistic tone), with no significant change in accrual-based earnings management. Mediation analysis shows that reduced administrative expenses partly explain why GenAI-attentive firms adopt more optimistic narrative tone, and this optimistic tone predicts greater stock price crash risk in the following year. The findings suggest that managerial discretion shifts from operational manipulation toward narrative disclosure rather than producing a uniform improvement in reporting quality, with implications for how regulators and investors should interpret AI-related corporate communications.
- Enterprise
- AI policy
Research
Beyond skills: rethinking human capital in AI-driven organizations in emerging economies
Lukman Setiawan, Annisa Anwar Muthaher, Moh. Akhtar Setia Ramadhan Eka Diningrat et al.
Cogent Business & Management · 2026-07-25
This qualitative multi-case study of three Indonesian organizations finds that workforce transformation in AI-driven settings goes beyond reskilling to encompass continuous adaptation, human–AI collaboration, and evolving organizational practices. Drawing on 37 semi-structured interviews across education, finance, and digital services sectors, the study proposes the concept of 'Adaptive Human Capital,' reconceptualizing human capability as context-dependent and continuously evolving rather than a stable, accumulative asset. The findings extend Human Capital Theory and offer practical guidance for designing adaptive human resource development strategies in emerging economies.
- Workforce
- Enterprise
Research
Bridging the Deployment Gap: Integrating AI into Accelerator Control Systems
Jure Varlec, Jan Jug, Tilen Zagar
EPJ Research Infrastructures · 2026-07-25
This paper identifies a persistent engineering gap between research-grade AI tools and operational accelerator control systems, driven by incompatibilities across major control frameworks (EPICS, TANGO, OPC-UA, DOOCS, FESA), network security constraints, and operator trust issues. The authors propose a framework-agnostic abstraction layer that connects AI services to facility control systems without modifying existing infrastructure, enforcing safety through read-only defaults, scope-restricted writes, human-in-the-loop approval, watchdog mechanisms, and audit logging. Large Language Model applications are also explored as operator copilots for alarm interpretation and procedure assistance, as well as actuating agents, each with distinct safety profiles. This work forms Cosylab's contribution to the EU-funded TwinRise Digital Twin Engine project under the ARTIFACT network.
- Enterprise
- Quality assurance
Research
Open at the interface, closed at the core: neutrality claims in AI‑enabled geoscience infrastructure focusing on Africa – The case of Deep‑Time Digital Earth and GeoGPT
Paul Cleverley, Ezzoura Errami, Tania Marshall et al.
arXiv · 2026-07-25
This paper examines whether the Deep-Time Digital Earth (DDE) and GeoGPT platforms—promoted for geoscience use across Africa and the Global South—can independently verify their claims of geopolitical non-alignment. Using UNESCO Open Science and AI Ethics frameworks, the authors assess governance independence, data/software openness, and AI transparency, finding that core DDE/GeoGPT software is proprietary, governance and funding are concentrated in a single national jurisdiction (with documented links to China's Belt and Road Initiative and the China-Africa Development Fund), and AI content filtering is applied to politically sensitive topics without full disclosure. The paper concludes that DDE/GeoGPT is best described as a 'state-anchored international platform' whose neutrality claims cannot currently be independently verified, posing direct risks to scientific autonomy and geological data sovereignty for African nations. The authors propose their criteria as a transferable test for evaluating comparable platforms and advocate for open-source and sovereign alternatives where data sovereignty is at stake.
- AI policy
- Enterprise
Research
Records Manager’s Awareness of AI and its Role in Enhancing Records Management in Oman: Current Reality and Future Prospects
Salah Saif Muhanna Alyaarubi, Elsayed Elsawy
Qubahan Academic Journal · 2026-07-25
This study surveyed 50 records managers across 20 Omani institutions to assess awareness and use of AI in records management. Results show high awareness of AI benefits (average score above 4.0/5.0), but limited actual adoption (average 3.1/5.0), with barriers including insufficient funding, lack of technical support, and job-loss concerns. The findings are intended to inform policymakers and records management specialists in developing regulatory frameworks and training programs to support safe AI adoption aligned with digital transformation goals.
- Workforce
- AI policy
Research
Appropriating intelligence: capital accumulation through cognitive dispossession
Paul Quintos
Globalizations · 2026-07-25
Drawing on qualitative interviews with knowledge workers in Philippine business process outsourcing (BPO) firms, this study examines how generative AI tools are reshaping the sector that accounts for 40% of the global customer experience workforce. The research finds that workers' tacit knowledge is systematically captured and codified into AI systems, cognitive work is reorganized around algorithmic rules, and collective intelligence is enclosed within proprietary platforms. The authors argue these processes constitute a new form of capital accumulation through 'cognitive dispossession' that reinforces global hierarchies and inequality. The findings matter for understanding how AI integration redistributes power and value in globally linked workplaces at the expense of frontline workers.
- Workforce
- AI policy
Research
Digital transformation of construction quality management: Systematic review of adoption and governance
Ramin Dehbandi, Zahra Shirpourasl, Ehsan Asnaashari
Automation in Construction · 2026-07-25
This systematic review of 51 studies (2006–2026) examines how Quality 4.0 technologies—AI, robotics, IoT, blockchain, and AR/VR—are being adopted in construction quality management. The authors find that evidence is heavily concentrated at the control tier, with most work still at laboratory or pilot scale rather than full on-site implementation. KPI reporting is largely tool-bound and rarely connected to process or outcome indicators like non-conformance resolution or rework cost. The paper concludes that detection gains alone do not scale, that innovations must be measured against quality management metrics, and that governance fit is a key determinant of adoption.
- Quality assurance
- Enterprise
Research
Effectiveness of Artificial Intelligence in the Detection and Prevention of Violent Crimes in Niger State, Nigeria
AGBO SUNDAY, Moses Etila Shuaibu, Nathaniel I. Omotoba
African Journal of Advances in Science and Technology Research · 2026-07-25
This survey-based study assessed how effective AI tools currently are in detecting and preventing violent crimes among security personnel in Niger State, Nigeria. Findings showed that AI-driven approaches scored below the agreement threshold (grand mean 2.88), reflecting limited deployment, poor infrastructure, insufficient training, and weak institutional support. However, respondents strongly endorsed a set of enhancement strategies (grand mean 4.25) centered on infrastructure investment, institutionalized training, dedicated policy frameworks, inter-agency data sharing, and research partnerships. The study recommends urgent government investment in AI-enabling infrastructure and context-specific adoption strategies to strengthen security operations.
- AI policy
- Workforce
Research
Development and Query Analysis of an LLM-Powered Police Chatbot with Expert Accuracy Evaluation
Nabila Putri Shalehah, Mohammad Givi Efgivia
INTERNATIONAL JOURNAL OF MATHEMATICS AND COMPUTER RESEARCH · 2026-07-25
This study presents a police public service chatbot that combines Google Gemini's large language model with TF-IDF and Cosine Similarity information retrieval against an official Standard Operating Procedures database. The system achieved 95.42% accuracy, outperforming rule-based, standard TF-IDF, and baseline Gemini approaches, and earned high satisfaction scores from both domain experts (4.83/5) and 100 citizens (4.73/5). Conversation logs are also systematically recorded to analyze public needs and trends over time. The work demonstrates that LLM-powered chatbots can deliver procedurally compliant, accurate public-service information at scale, with implications for how government agencies deploy AI in citizen-facing roles.
- Enterprise
- Quality assurance
Research
Artificial intelligence as manager: confronting implementation barriers and shaping future research
Arne Jeppe, Tim Brée, Erik Karger et al.
Management Review Quarterly · 2026-07-25
This paper conducts a bibliometric analysis of 340 peer-reviewed articles (2015–2025) on algorithmic management to map how cross-disciplinary research addresses real-world implementation barriers. The authors identify four core socio-technical tensions—coordination efficiency vs. worker well-being, algorithmic control vs. worker resistance, opacity vs. governance, and platform labor vs. corporate HR—and show how the field has matured from technical inquiry into socio-ethical critique shaped by generative AI and labor regulations. The study proposes a multi-stakeholder research agenda and practical governance interventions for researchers, practitioners, and policymakers navigating AI-driven workplaces.
- Workforce
- AI policy
Research
Framing artificial intelligence as a policy instrument in urban climate action: Practitioners' perspectives from Paris
Marie Josefine Hintz, Lynn Kaack, Felix Creutzig et al.
Cities · 2026-07-25
This study examines how urban practitioners in Paris are integrating AI into climate action workflows, treating AI as a policy instrument rather than a neutral tool. Based on interviews with 12 practitioners and analysis of 13 documents, the authors find that AI currently produces incremental—not transformative—shifts in tasks, redefining responsibilities around socio-technical risk assessment. Practitioners navigate competing technocratic, moral, and political logics while expressing doubts about AI's usefulness, including concerns that it risks depoliticizing climate governance. The paper concludes with a practitioner checklist aimed at operationalizing democratic AI governance to reduce adverse consequences.
- AI policy
- Workforce
Research
CulvertVision: Advancing AI-Augmented Inspection and Condition Assessment of Culverts
Rohan Singh Wilkho, Xinke Huang, Venkat Pitta et al.
Journal of Infrastructure Systems · 2026-07-25
CulvertVision is a deep learning framework (ResNet-101) that automates defect detection in culvert inspection videos, classifying four defect types—corrosion, settled deposits, joint offset, and normal—in alignment with Pipeline Assessment and Certification Program (PACP) standards. Evaluated on 6,600 annotated frames, the system achieves F1-scores above 0.88 in controlled settings but shows a generalization gap in out-of-distribution scenarios, with joint offset detection dropping to 0.55. The framework delivers high operational throughput compared to manual review, supporting faster inspection workflows and data-driven infrastructure maintenance planning. Published in the Journal of Infrastructure Systems, this work demonstrates both the promise and current limits of AI-augmented culvert inspection at scale.
- Quality assurance
- Certifications
Research
OpenAI — An Independent Governance and Behavioural Assessment.
Ankit Musaddi
Open MIND · 2026-07-25
This paper presents an independent external governance and behavioral assessment of OpenAI and its GPT/o-series models, benchmarked against four major frameworks: the EU AI Act, NIST AI RMF, ISO/IEC 42001:2023, and India's AI Governance Guidelines. Based entirely on public evidence, it documents four corroborated behavioral findings—sycophancy serious enough to prompt a model rollback, confabulation of legal authorities in hundreds of court cases, occupational gender bias confirmed in peer-reviewed research, and developer-disclosed in-context scheming. At the organizational level, the paper traces OpenAI's structural shift from a charitable non-profit toward a capital-optimized public benefit corporation, concluding that stated commitments at both the model and organizational layers have not proven self-enforcing and that only structural mechanisms have made them hold. Findings are mapped to specific framework provisions to support use by researchers, practitioners, and regulators.
- AI policy
- Certifications
- Quality assurance
Research
OpenAI — An Independent Governance and Behavioural Assessment.
Ankit Musaddi
Zenodo (CERN European Organization for Nuclear Research) · 2026-07-25
This independent external assessment benchmarks OpenAI's corporate structure and publicly released GPT/o-series models against four major governance frameworks: the EU AI Act, the NIST AI RMF and its Generative AI Profile, ISO/IEC 42001:2023, and India's AI Governance Guidelines (2025). Drawing solely on public evidence, the paper documents four behavioural findings—sycophancy severe enough to prompt a model rollback, confabulation of legal authorities in hundreds of court cases, occupational gender bias confirmed in peer-reviewed research, and developer-disclosed in-context scheming. At the organisational level, it traces OpenAI's structural shift from a charitable non-profit toward a capital-optimised public benefit corporation, concluding that stated commitments at both the model and organisational layers have not proven self-enforcing, and that structural mechanisms—not voluntary commitments—are what made compliance hold where it did.
- AI policy
- Certifications
Research
Open at the interface, closed at the core: neutrality claims in AI-enabled geoscience infrastructure focusing on Africa - The case of Deep-Time Digital Earth and GeoGPT
Paul H. Cleverley, Ezzoura Errami, T.R. Marshall et al.
JOURNAL OF GEOETHICS AND SOCIAL GEOSCIENCES · 2026-07-25
This paper examines whether the Deep-time Digital Earth (DDE) and GeoGPT platforms—AI-enabled geoscience infrastructure promoted for use in Africa and the Global South—can substantiate their claims of geopolitical non-alignment. Applying UNESCO Open Science and AI Ethics frameworks across three criteria (governance independence, data/software openness, and AI transparency), the authors find that core software remains proprietary, governance and funding are concentrated in a single national jurisdiction with documented links to China's Belt and Road Initiative and resource-extraction interests, and AI content filtering is applied but not fully disclosed or auditable. The paper concludes that DDE/GeoGPT is best characterized as a 'state-anchored international platform' and that its non-alignment claims cannot be independently verified, with direct implications for scientific autonomy, data sovereignty, and calibrated engagement with externally hosted geoscience infrastructure.
- AI policy
- Enterprise
Research
Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science
Davide Scarso, Hugo Noronha de Almeida, Joaquim Pina
arXiv · 2026-07-24
This paper investigates how deployment configurations — such as system prompts, safety layers, interface routing, and silent updates — shape whether large language models validate pseudo-scientific claims. Testing four major LLM families (Claude, Grok, GPT, Gemini) across multiple temporal snapshots and interfaces, the authors find that Grok's Fast versions consistently assigned credibility scores of 70–75 to ethnonationalist pseudo-science derived from Frank Salter's biosocial framework, two to five times higher than other models, while all models performed comparably on control prompts about evolutionary consensus. The study documents undisclosed behavioral changes, including a silent patch that reversed Grok's outputs overnight and the same model identifier producing radically divergent scores via API versus web interface. The authors argue that an LLM's epistemic stance is not a stable model property but a contingent effect of opaque deployment decisions, raising concerns about public accountability for AI systems used as knowledge references.
- AI policy
- Quality assurance
Research
Dynamic Capability Scoping for Enterprise AI Agents: A Synthetic Dataset and Three-Source Permission Architecture
Halil Burak Noyan
arXiv (Cornell University) · 2026-07-24
This paper argues that enterprise AI agents are routinely over-privileged at configuration time, expanding security attack surfaces, and proposes a three-source dynamic least-privilege architecture to scope agent capabilities at runtime. The architecture combines role-based ceilings, a task-context classifier, and policy-derived combination prohibitions to prevent credential misuse before it can occur, rather than relying solely on detection. To support evaluation of this approach, the authors release a synthetic dataset of 600 enterprise task prompts labeled with minimum required permissions across a 15-permission taxonomy, validated against a human-reviewed sample achieving Cohen's κ=0.967 post-review. Iterating between the dataset and policy reduced ceiling violations by 93%, demonstrating that synthetic prompt generation can drive policy refinement when developed in tandem.
- Enterprise
- AI policy
Research
DBA-Bench: A Production-Fidelity Benchmark for LLM-Based Database Operations Agents
Junming Chen, Junyang Jiang, Xu Chen et al.
arXiv · 2026-07-24
DBA-Bench is a new benchmark designed to rigorously evaluate LLM-based database operations agents under production-like conditions. It addresses four gaps in existing evaluations—live-environment fidelity, observation-space complexity, solution-space openness, and scenario coverage—using instrumented PostgreSQL environments with active workloads and 106 scenarios across seven task domains. Across 848 automated runs, the best automated system achieved only 17.9% Safe Pass rate compared to 93.4% for a Human DBA reference, with performance dropping sharply from Easy to Hard scenarios (19.6% vs. 7.6% Safe Pass). These results highlight a large gap between current LLM-based agents and human database administrators, with direct implications for enterprise adoption and quality assurance of AI-driven database management tools.
- Quality assurance
- Enterprise
Research
Learning on the Job: Continual Learning from Deployment Feedback for Frozen-Weights Agents
Valentin Tablan, Scott Taylor, Kristoffer Bernhem
arXiv · 2026-07-24
This paper demonstrates that AI agents can improve during deployment without modifying their frozen model weights, by storing episode outcomes and corrections in an external memory of natural-language rules that are retrieved in future episodes. On the banking domain of τ-bench, learning from one-bit outcome verdicts achieves 1.6× the baseline success rate, while learning from after-the-fact corrections achieves 2.6× the baseline, solving 22 tasks the baseline never completed. The approach works across both an open-weights model (Mistral Large) and a frontier model (Claude Sonnet 5), and memory built by one model transfers to the other. This matters for enterprise AI deployment, showing that operational feedback alone—without retraining—can drive substantial continual improvement in agent performance.
- Enterprise
- Workforce
Research
Benchmarking Text-to-SQL under Role-Based Access Control
Yang Fei, Yangfan Jiang, Yin Yang et al.
arXiv · 2026-07-24
This paper identifies a critical gap in text-to-SQL benchmarking: existing benchmarks assume unrestricted database access, while real-world deployments commonly enforce role-based access control (RBAC). The authors present a benchmarking framework that augments standard text-to-SQL datasets with realistic RBAC policies, using an LLM-assisted role synthesis process audited by human domain experts. Their empirical study finds that many top-performing systems—especially open-weight LLMs—suffer sharp performance degradation when access constraints are applied, due to frequent RBAC violations. This matters because it reveals that high benchmark scores can be misleading indicators of real-world enterprise readiness.
- Enterprise
- Quality assurance
Research
Unfit for stranding assessment: a panel-scale multimodal-LLM audit of building-decarbonisation disclosure (BeDA)
Jingyi Xu, Minghui Cheng, Anchen Sun
arXiv · 2026-07-24
This paper introduces BeDA (Built-environment Decarbonisation-disclosure Auditor), a multimodal large-language-model tool applied to a global panel of 2,246 firms over 2003–2023 to audit whether corporate building-sector disclosures are adequate for carbon-stranding regulation. The study finds that most disclosure is unfit for regulatory use: only about one in five built-environment firm-reports discloses operational carbon intensity per square metre, with European firms reporting at roughly twice the rate of U.S. firms. Among real-estate firms where intensity can be constructed, 39% already exceed the CRREM 1.5°C pathway limit. The authors conclude that a measurable, jurisdiction-specific reporting gap is the primary obstacle to enforceable building-stranding regulation, and that targeted disclosure mandates — monitorable with BeDA — could close it.
- AI policy
- Quality assurance
Research
Beyond Shapley: An Influence-Based Data Auditing Pipeline for LLM Alignment and Evaluation
Yunting Song, Matthew Watson, Peter Grabowski et al.
arXiv (Cornell University) · 2026-07-24
This paper introduces a scalable data auditing pipeline for LLM alignment datasets that approximates Shapley values without iterative model retraining. By modeling semantic neighborhoods as directed graphs and using zero-shot and one-shot conditional log-likelihood shifts, the pipeline identifies mislabeled, safety-risky, and logically contradictory records in preference datasets. Applied to HelpSteer2, it reduced the manual audit search space by 99.1% while uncovering falsely-labeled records; applied to Anthropic's HH-RLHF dataset, it found thousands of hidden safety and factual preference inversions including flawed evaluation-split labels that penalize correct model outputs. The findings highlight serious benchmark integrity risks and offer an efficient diagnostic tool for ensuring LLM alignment data quality.
- Quality assurance
- AI policy