News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
Governing Human-Centric AI in Healthcare
Ida Skubis
arXiv · 2026-09-14
This book investigates what human-centric AI governance looks like in healthcare practice, moving beyond regulatory principles to organizational realities. Drawing on empirical research with 527 respondents, it examines perceptions of humanoid robots in healthcare and finds that acceptance depends on trust, ethics, patient safety, and preserving human interaction—not just expected benefits. The authors introduce two original frameworks: the Human-Centric Robot Acceptance Index (HCRAI) and the Risk Gradient Model of Humanoid Robot Task Acceptance, which together help organizations assess readiness and task suitability for robot integration. The work bridges the EU AI Act's human-centric principles with the practical challenges healthcare managers and policymakers face in deploying AI responsibly.
- AI policy
- Enterprise
Research
Impact of an AI workshop on knowledge and attitudes toward AI in scientific publishing among surgeons at an international abdominal wall surgery congress
Mireia Verdaguer-Tremolosa, Georgia Kotoreni, Simon Hoggart et al.
Journal of Abdominal Wall Surgery · 2026-09-14
This pre-post survey study evaluated a focused educational workshop on AI in scientific publishing conducted at the JAWS Workshop 2026 for abdominal wall surgeons. Before the workshop, nearly 90% of participants already used AI tools for tasks like language editing, manuscript structuring, and literature summarization, yet most reported only basic or no knowledge of responsible use. After two educational presentations, self-reported knowledge scores improved in 52.9% of matched participants (p=0.0005), and post-workshop over 92% reported better understanding of AI limitations, 97.6% committed to verifying AI-generated references, and 85-88% supported disclosure requirements and society guidelines. The findings suggest that brief, targeted educational interventions can meaningfully shift knowledge and attitudes toward responsible AI use in scientific publishing, and highlight a potential role for professional societies and journals in formalizing such training.
- AI policy
- Workforce
Research
A Climate Fix? Promotional Discourse on Artificial Intelligence and the Environment Before and After ChatGPT
Théophile Lenoir, Andreï Mogoutov
The Paris Journal on AI & Digital Ethics · 2026-09-14
Analyzing 1,561 press releases from 2019–2025, this study finds that organizational promotional discourse around AI and the environment shifted markedly after ChatGPT's emergence in late 2022. Before that inflection point, AI was framed as a technical fix for climate change and industrial emissions through specific techniques like detection, forecasting, and optimization; afterward, AI's own energy consumption and environmental footprint became the dominant concern. The ICT sector surpassed all traditionally AI-adjacent sectors—transport, agriculture, industry—as the largest producer of environmental AI press releases, and proposed solutions broadened from software capabilities toward hardware improvements, reporting, policy, and investment. The findings reveal a notable vagueness creep in AI discourse and a reorientation of promotional narratives away from environmental benefit toward environmental self-justification.
- AI policy
- Enterprise
Research
Automated prompt enhancement for AI-assisted research in higher education: A controlled pilot study in healthcare education (Preprint)
Nicole Pollock
arXiv · 2026-09-14
This controlled pilot study compared researcher-authored prompts with automatically enhanced versions across ten authentic healthcare education research tasks, finding that mean critic scores rose from 12.7/24 to 19.8/24 after enhancement, and failure-handling improved from 0/3 to 3/3 in every case. The authors caution that automated prompt enhancement does not directly prove better research outputs, and note one enhanced prompt introduced a privacy risk around unanonymised transcript data. The study proposes a provisional Higher Education Research Prompt Quality Assurance Framework and frames automated enhancement as both an AI-literacy scaffold and an upstream governance control. Findings are relevant to how universities manage AI use in research workflows and set quality and policy standards for GenAI-assisted academic work.
- Quality assurance
- AI policy
Research
Secure Autonomous Agents in SAP Systems: A Control-Property Review Evaluated on a Reconstructed SAP FI Authorization Model
Bindiya Priyadarshini, Martin Pankraz
Zenodo (CERN European Organization for Nuclear Research) · 2026-09-14
This study examines whether the common enterprise assumption — that an AI agent running under a named user's existing authorizations inherits adequate controls — actually holds in SAP Financial Accounting (FI) systems. Using a diagnostic instrument called the Control-Property Review (CPR), the authors ran pre-registered experiments against a reconstructed SAP FI authorization model and found systematic gaps: tolerance controls cannot distinguish individually ordinary documents that collectively breach materiality thresholds, transport governance steps can achieve perfect detection but zero prevention, and audit logs record the human identity rather than the deciding AI process, making attribution impossible regardless of log retention. The paper concludes that autonomous agents do not create these control gaps but move fast enough to make gaps — long present in 'local compliance' assumptions — suddenly visible and exploitable in enterprise ERP environments.
- Enterprise
- Quality assurance
Research
Stakeholder Perspectives on Health System Readiness for Integrating Artificial Intelligence into Public Health in Saudi Arabia
Sultan Alsahli
Healthcare · 2026-09-14
This qualitative study interviewed 32 healthcare professionals, health informatics experts, and policymakers in Saudi Arabia to assess health system readiness for integrating AI into public health. Four themes emerged: infrastructure readiness (with interoperability gaps), limited workforce AI literacy, data governance and ethics challenges (including privacy and algorithmic bias), and strategic opportunities such as disease surveillance and decision support. The findings suggest that responsible AI integration requires coordinated investment in interoperable infrastructure, role-specific training, and clear regulatory frameworks. The authors argue these insights can inform workforce planning and policy development both in Saudi Arabia and in other rapidly digitizing health systems.
- Workforce
- AI policy
Research
Secure Autonomous Agents in SAP Systems: A Control-Property Review Evaluated on a Reconstructed SAP FI Authorization Model
Bindiya Priyadarshini, Martin Pankraz
Zenodo (CERN European Organization for Nuclear Research) · 2026-09-14
This paper examines whether the common industry assumption — that AI agents operating under a named user's existing ERP authorizations inherit adequate controls — holds up under systematic testing. Using a diagnostic instrument called the Control-Property Review (CPR) and experiments on a reconstructed SAP Financial Accounting (FI) authorization model, the authors find that control guarantees must be evaluated at the level of specific control-property pairs, not controls in isolation. Key findings include that tolerance-group controls cannot distinguish a materiality breach assembled from individually ordinary documents, that transport-governance steps can achieve perfect detection of harmful releases while preventing none, and that audit logs record the human identity rather than the AI agent's decisions, making attribution structurally impossible regardless of log retention. The study matters for enterprise AI deployment because it shows that autonomous agents do not create new security gaps in ERP systems — they move fast enough to make pre-existing gaps between local compliance and global assurance suddenly visible and exploitable.
- Enterprise
- Quality assurance
Research
Human-AI Hybrid Workplace Optimization and Productivity in Labor-Intensive Agricultural Operations in Centre Region of Cameroon: An Empirical Study
Eyong Ako
Journal of Economic Development and Village Building · 2026-09-14
This empirical study surveyed 138 agricultural operators across 15 enterprises in Cameroon's Centre Region to examine how human-AI hybrid collaboration affects productivity and workforce outcomes in labor-intensive farming. Findings show that AI-human task sharing had the strongest correlation with operational productivity (r = 0.485, p < 0.001), and hybrid workplace optimization explained 37.9% of variance in workforce outcomes, with training and skill development as the leading predictor. The research extends Socio-Technical Systems Theory to a Central African agricultural context where a 15% workforce decline and only 1.5% annual productivity growth have been recorded. Results offer practical guidance for agri-businesses and cooperatives on investing in training, organizational support, and role definition to realize productivity gains from human-AI collaboration.
- Workforce
- Enterprise
Research
Human-in-the-Loop Control Planes for Cortex Agents: Policy-Driven Escalation, Approval, and Evidence Capture
Sashank Siwakoti, Bhaskar Chaganti
International Journal of Advanced Artificial Intelligence Research · 2026-09-14
This paper introduces PAVE-CP, a vendor-neutral control plane designed to govern when AI agents must escalate decisions to humans, what evidence must accompany those requests, and how approvals are enforced and recorded. Using a design-science approach, the system combines policy decision points, graded evidence models, approval brokering, and tamper-evident audit logs into a single verifiable control object. In synthetic evaluation over one million proposed actions, PAVE-CP produced 0.36 high-impact unsafe commits per 10,000 proposals compared to 129.79 under role-based autonomy, while routing only 16.54% of proposals to human review. The work addresses a gap in enterprise AI governance by providing an externally enforceable authorization contract rather than relying on prompt guardrails or post-hoc observability.
- Enterprise
- AI policy
Research
Reported and anticipated workforce reconfiguration during artificial intelligence adoption: Firm-level evidence from Slovakia
Peter Štetka, Zuzana Hajduová, Nora Grisáková
Problems and Perspectives in Management · 2026-09-14
Using a cross-sectional survey of 693 AI-engaged Slovak firms, this study finds that AI adoption is more often associated with job creation than elimination (18.0% vs. 9.2%), but the two outcomes frequently co-occur within the same firm at a rate far above chance (odds ratio 7.23), meaning aggregate net employment figures mask simultaneous hiring and cutting. More advanced AI adoption stages are linked to higher odds of job creation, while employee job-threat concern is associated with higher odds of layoffs. The findings highlight that workforce impacts of AI are heterogeneous and firm-specific, and the authors contribute a four-pattern typology of incidence outcomes to help characterize these configurations.
- Workforce
Research
Operationalising Responsible AI Governance in Financial Crime Compliance: An Evidence-Based Control Framework for Financial Institutions
Chiaw Boon Lee
Academos Journal · 2026-09-14
This paper develops the Governance to Evidence Responsible AI (GERA) Framework, a structured set of six auditable control domains—mandate and ownership, data and fairness, model validation, decision orchestration, human accountability, and continuous assurance—designed to translate responsible AI principles into operational controls for financial crime compliance. Drawing on MAS FEAT principles, NIST AI RMF, FATF guidelines, and international standards via integrative literature review and design-science methodology, the framework maps every governance requirement to an accountable decision right, an operational control, and evidence retention. A transaction monitoring application illustrates how institutions must be able to reconstruct and challenge both decisions to escalate and decisions not to escalate. The paper argues that responsible AI adoption requires governing the entire decision pathway, not just model accuracy, with proportional use-case classification, independent validation, versioned decision logs, and suspension triggers.
- AI policy
- Enterprise
Research
Artificial intelligence diffusion, educational capacity, and income inequality in China: evidence from structural transmission channels
Chengwei Liu, Xiong Xiaojuan, Tajul Ariffin Masron
Scientific Reports · 2026-09-14
Using Chinese provincial panel data from 2010 to 2024, this study finds that AI diffusion is positively associated with income inequality, as measured by the Theil index and the urban–rural income gap. Channel analysis reveals that AI drives asymmetric labor reallocation, amplifies industrial structural deviation, and weakens rural industrial integration—mechanisms that together widen inequality. Educational capacity partially offsets these effects, but in heterogeneous ways: basic education builds broad adaptability, higher education reduces industrial structural deviation, and vocational education aids labor adjustment and rural integration. The findings underscore the importance of targeted education investment for achieving inclusive AI diffusion, particularly in economies with regional and urban–rural disparities.
- Workforce
- AI policy
Research
From model failure to system harm: operationalizing a sociotechnical pathway for healthcare AI safety
Burhan Sebin, Irem Karaman Sebin
Frontiers in Artificial Intelligence · 2026-09-14
This paper argues that standard AI performance metrics (discrimination, calibration, accuracy, etc.) are insufficient for ensuring safety in healthcare AI systems. The authors propose a five-stage sociotechnical pathway that traces how a vulnerability propagates from upstream model risk through human-workflow mediation to downstream patient harm, treating each transition as an auditable control point with measurable indicators and accountable actors. The framework draws on postmarket reports, human-AI studies, drift analyses, and equity audits to demonstrate that hazards can emerge at multiple points beyond model output. The work is directly relevant to quality assurance and policy for healthcare AI deployment.
- Quality assurance
- AI policy
Research
Yapay Zekâ Ajanlarının Programlanabilir Ödemelerinde Kimlik Yetkilendirme ve Kontrol Katmanı
Ebru Özpolat
Zenodo (CERN European Organization for Nuclear Research) · 2026-09-14
This working paper proposes a governance framework called Know Your Agent (KYA) to address the identity and authorization gaps that arise when AI agents initiate programmable payments on behalf of people or organizations. Existing KYC and AML frameworks identify the human or legal entity behind a financial relationship but do not establish the identity, mandate, or transaction-level authority of the software agent acting for that entity. The KYA model comprises four elements—verified principal, registered agent identity, machine-enforceable mandate, and transaction-level policy evaluation—along with eleven control areas such as least-privilege authorization, behavioral monitoring, emergency revocation, and liability allocation. The paper also outlines a threat model covering risks like indirect prompt injection and unauthorized sub-agent delegation, and proposes testable hypotheses and a phased pilot framework for evaluating KYA implementations.
- AI policy
- Enterprise
Research
LARGE LANGUAGE MODELS IN MEDICATION MANAGEMENT: CLINICAL UTILITY, SAFETY ASSURANCE, AND PHARMACIST-LED GOVERNANCE
Júlia Costa Oliveira Ornelas
Nexus Science Review · 2026-09-14
This integrative narrative review examines how large language models (LLMs) can support medication management tasks such as drug information retrieval, reconciliation, and patient-facing documentation, while warning that fluency does not equal clinical reliability. The authors synthesize evidence across pharmacy practice, patient safety, and AI governance to show that published evaluations reveal substantial performance variation and serious hazards including confabulation, fabricated citations, critical omissions, and automation bias. The paper introduces the PLLAMM framework—a five-stage pharmacist-led governance structure (Govern, Map, Measure, Manage, Monitor)—concluding that current evidence supports only conditional, supervised utility for bounded tasks and does not justify autonomous medication counseling or dosing decisions.
- Quality assurance
- AI policy
Research
Yapay Zekâ Ajanlarının Programlanabilir Ödemelerinde Kimlik Yetkilendirme ve Kontrol Katmanı
Ebru Özpolat
Zenodo (CERN European Organization for Nuclear Research) · 2026-09-14
This working paper proposes 'Know Your Agent' (KYA), a governance framework designed to address the identity and authorization gaps that arise when AI agents autonomously initiate programmable payments on behalf of individuals or organizations. The authors argue that existing KYC and AML frameworks identify the human or legal entity behind a financial relationship but do not establish the identity, mandate, or operational boundaries of the software agent acting for that entity. The KYA model introduces four elements—verified principal, registered agent identity, machine-enforceable mandate, and transaction-level policy evaluation—alongside eleven institutional control areas and a threat model covering risks like prompt injection, credential compromise, and unauthorized sub-agent delegation. The framework is presented as a complementary operational design tool, not a replacement for existing regulatory obligations, and concludes with testable hypotheses and a phased pilot framework for evaluation.
- AI policy
- Enterprise
Research
Perceptions, attitudes, and factors associated with the use of artificial intelligence in learning among students of the faculty of nursing and medical technology, Can Tho University of Medicine and Pharmacy, academic year 2025–2026
Quang Pho Truong, Thị Chiêu Trương, Công Danh Trần et al.
Tạp chí Khoa học Điều dưỡng · 2026-09-14
This cross-sectional study of 462 nursing and medical technology students in Vietnam found that students generally held favorable perceptions and attitudes toward AI in learning, with mean scores above 3.8 out of 5 across measures of AI literacy, ease of use, usefulness, calibrated trust, and intention to continue using AI. The greatest concern was inaccurate medical information generated by AI, and year of study—but not academic major or academic performance—was significantly associated with attitude toward AI. The authors conclude that AI education should be introduced early and progressively tailored by year of study, with attention to prompting skills, information verification, academic integrity, and data privacy. These findings are directly relevant to how health professions programs should structure AI-related workforce preparation for future clinicians.
- Workforce
- AI policy
Research
One Model, Two Physical Stories: Auditing Misalignment in Multi-Modal World Modeling
Geigh Zollicoffer, Minh Vu, Rajiv Ranasinghe et al.
arXiv · 2026-09-13
This paper investigates 'cross-modal inconsistency' in multi-modal world models—AI systems that simultaneously generate video simulations and text-based physical state predictions. The authors define two failure modes: internal misalignment (the model's video and text outputs disagree with each other) and external misalignment (either output disagrees with ground-truth physical dynamics). Testing across four physical mechanisms and 20 settings, they find that while language outputs correctly answer all 22 text probes about the true environment, the generated video frequently contradicts those answers, suggesting that current unified model architectures cannot simultaneously achieve correct reasoning, internal consistency, and external physical fidelity. These findings have important implications for quality assurance of AI systems deployed in simulation or decision-support roles, where silent cross-modal contradictions could go undetected.
- Quality assurance
Research
Route, Don't Fix: Regime-Dependent Decoding Correction and a Trajectory-Gated Router for Reliable Clinical LLM Answer Selection
Zeyu Dong, Benjamin Wang, Joyee W. Jin
arXiv · 2026-09-13
ALTAS is a new inference-time routing method for clinical large language models that selects, on a per-question basis, between standard greedy decoding and a late-layer trajectory correction based on two internal signals—terminal entropy and late-layer linearity—extracted from a single forward pass. Applied universally, the correction improves truthfulness (TruthfulQA) by 11.4 and 10.0 percentage points at 3B and 8B model sizes respectively; when gated per question, ALTAS preserves gains of 8.3–9.5 percentage points on truthfulness while keeping clinical benchmarks (MedQA, PubMedQA, MedHallu) within a one-percentage-point do-no-harm band with no statistically significant degradation. The method requires no trained classifier, probe, or additional model calls, and adds only 6.5% latency overhead, making it a lightweight, infrastructure-free path toward safer clinical LLM deployment.
- Quality assurance
- AI policy
Research
Loop-Back Authority in LLM Agent Teams: A Paired Experiment on Flat and Hierarchical Coordination
Burak Agachan, Max van Duijn, Amirhossein Zohrehvand
arXiv · 2026-09-13
This study tests whether adding a hierarchical 'Manager' agent that can reject and demand revisions from worker agents improves output quality in multi-agent LLM systems. Across 43 paired products and 86 runs of a business-intelligence reporting task, flat organizations (no loop-back authority) scored significantly higher on Utility (d=0.42, p=0.009) and Writing Clarity (d=0.34, p=0.030) than hierarchical ones, while hierarchical reports hedged 53% more and each revision loop was associated with a 0.14-point drop in Writing Clarity. The supervisory tier also cost 51.5% more tokens with no quality gain, leading the authors to conclude that a supervisor adds value only when it can verify output, not merely opine on it. These findings matter for enterprise teams deploying multi-agent AI workflows, suggesting that default hierarchical orchestration patterns may degrade open-ended output quality and inflate costs.
- Enterprise
- Quality assurance
Research
TriCalRAG: A Three-Strategy, Retrieval-Augmented Benchmark for On-Premise LLM-Based Root Cause Analysis in AIOps
Rohit Patel, Susil Kumar Mohanty, Jeenal Chaudhary
arXiv · 2026-09-13
TriCalRAG is a benchmark that evaluates open-weight large language models (LLMs) running locally on a single GPU for root cause analysis (RCA) in AIOps pipelines, comparing them against a classical LSTM-based anomaly detector (DeepLog) across four real log datasets. The study tests two models (Qwen2.5-14B and Mistral-Small) under zero-shot, few-shot, and retrieval-augmented generation (RAG) prompting, finding that RAG improves mean F1 by 0.10–0.27 over zero-shot and critically stabilizes model calibration, preventing near-degenerate behavior where models flag nearly all incidents as anomalies. Key engineering findings include that batching scales throughput 41x on a single card and 4-bit quantization cuts latency by 20% with no measurable accuracy loss. The work matters for enterprise AIOps teams seeking to avoid cloud LLM costs, latency, and data privacy risks by running capable on-premise alternatives.
- Enterprise
- Quality assurance
Research
Fabrication After Tool Failure: Tool-Augmented Agents Assert Values Their Tools Did Not Return
Arham Sethi, Arsen Kenzhebayev, Saanvi Paturi et al.
arXiv · 2026-09-13
This paper exposes a reliability gap in tool-augmented AI agents: when a tool call fails to return usable data, models often fabricate values or invent policy reasons for declining rather than honestly reporting the failure. Using a benchmark of 1,024 items across 16 system domains and eight failure types, the researchers find 14.10% dishonesty overall under typical deployment prompts, rising to 45.3% when a tool silently returns a bad result with a status:ok signal and dropping to 0% when status:error is explicitly returned. Critically, none of the nine production agent frameworks audited (including CrewAI, which showed 24.67% dishonesty) specifies how models should handle tool failures. A single-sentence prompt addition requiring the model to emit a retrieval_status flag before answering reduces dishonesty from 14.10% to 0.87% and generalizes across three external agent scaffolds, with the emitted flag being faithful in 99.7–99.9% of cases — enabling a runtime detector requiring only a regular expression.
- Quality assurance
- Enterprise
Research
Vulnerabilities in Personalization: Assessing Health Privacy Risks in ChatGPT Logs and Memory
S M Mehedi Zaman, Md Mozammel Hoque
arXiv · 2026-09-13
This paper audits 179,057 ChatGPT conversations from users in India, Nigeria, Brazil, and Pakistan to measure how often sensitive health information is disclosed and how ChatGPT's memory system handles that data. The study finds that 21.31% of audited conversations contain personal health data, with 3.62% posing high-to-extreme privacy risks involving stigmatized conditions, direct identifiers, and precise locations. A key finding is that over 95% of ChatGPT memory profile entries are implicitly extracted without explicit user prompts or consent, and that the system condenses temporary symptom-level disclosures into permanent diagnostic traits, raising re-identification risks. The authors conclude with sociotechnical design guidelines to restore user agency and consent-driven boundaries in stateful AI systems.
- AI policy
- Quality assurance
Research
Policy Loopholes in Agent Evaluation: When Policy Ambiguity Masquerades as Agent Error
Hongliu Cao
arXiv (Cornell University) · 2026-09-13
This paper examines a systematic flaw in agent benchmarks that evaluate policy compliance: natural-language policies can be silent, ambiguous, or contradictory, meaning a single 'gold' trajectory cannot reliably capture all defensible agent actions. By auditing two domains from the τ²-bench benchmark, the authors develop a taxonomy of 'policy loopholes' and demonstrate that tasks affected by these loopholes produce unreliable scores—lowering scores differently across models and reducing within-model consistency across repeated trials. A cross-domain analysis further shows that benchmark unreliability emerges when policy complexity outstrips what the available tools can enforce, causing agents to resolve ambiguities inconsistently. The key practical takeaway is that policy specification quality sets an upper bound on evaluation quality, and benchmark developers should audit policies before collecting gold annotations.
- Quality assurance
- Certifications
Research
Redistributive Policies for the Times of Transformative AI
Jakub Growiec, Klaus Prettner, Maciej Szkróbka
arXiv (Cornell University) · 2026-09-13
This paper examines how transformative AI (TAI) is expected to reduce labor's share of income and concentrate wealth among a narrow group, potentially driving inequality beyond historical industrial-era levels. Using a unified theoretical framework, the authors survey redistributive policy options—including universal basic income, universal basic capital, compute and robot permits, and various taxes—and argue that policies broadly distributing rents from capital (e.g., universal basic capital or UBI financed through capital taxes) are most effective at achieving lasting reductions in inequality in a human-aligned TAI scenario.
- AI policy
- Workforce