News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
The invisible hand of generative AI: quasi-agency and knowledge transformation in data-driven decision-making
Matteo Cristofaro, Alexis Bañón-Gomis, Pier Luigi Giardino
Journal of Knowledge Management · 2026-08-31
Drawing on qualitative data from 120 managers at Italian SMEs using generative AI in recurring decision-making activities, this study develops a 'phase-sensitive model of GenAI quasi-agency' showing how GenAI reshapes organizational knowledge through three mechanisms: knowledge coupling, knowledge decoupling, and knowledge hiding. The framework reveals that GenAI systems can influence framing, evidence construction, prioritization, justification, and organizational memory without possessing formal authority or accountability, creating new governance risks rooted in hidden knowledge transformation rather than hidden human action. The findings extend principal-agent theory and argue that organizations need governance mechanisms centered on traceability, reconstructability, contestability, and human justificatory ownership across the decision cycle. This matters for enterprise and policy stakeholders seeking to understand how AI quietly reconfigures accountability structures in data-driven decision-making.
- Enterprise
- AI policy
Research
REGIONAL DIMENSION OF ECONOMIC DEVELOPMENT IN THE CONTEXT OF ARTIFICIAL INTELLIGENCE IMPLEMENTATION IN PUBLIC ADMINISTRATION
Diana SHKUROPADSKA, Larysa Lebedeva, Kateryna NIKOLAIETS et al.
Financial and credit activity problems of theory and practice · 2026-08-31
This study examines how AI readiness in public administration varies across nine world regions and how it correlates with economic development. Using the Government AI Readiness Index and correlation-regression analysis, the authors find that North America leads with a score of 81.51, followed by Western Europe and Eastern Europe, while Sub-Saharan Africa, the Pacific, and Latin America rank lowest. Regression analysis shows that over 90% of variation in regional AI readiness is explained by GDP per capita, though institutional and regulatory factors also play independent roles. The findings support developing differentiated, region-specific national strategies for AI integration in government.
- AI policy
Research
Algorithmic Administrative Authority: Reconstructing the Legal Boundaries of Government Power in the Age of Artificial Intelligence
Laura Dehaibie, Frank C. Maes
LAW & PASS International Journal of Law Public Administration and Social Studies · 2026-08-31
This legal research article examines how AI systems integrated into public administration challenge traditional administrative law, which assumes identifiable human officials who deliberate, give reasons, and remain accountable for decisions. The article argues that the core legal problem is not that algorithms become autonomous authorities, but that they can acquire de facto influence over legally consequential government decisions while remaining outside formal public-authority structures. To address this, the article proposes a five-part framework—covering legal attribution, bounded algorithmic discretion, procedural transparency, institutional responsibility, and effective human and judicial review—scaled so that stronger public-law safeguards apply as algorithmic influence over decisions increases. The framework aims to enable technological innovation while preserving the foundational principle that governmental power must remain subject to law.
- AI policy
Research
Analysis code, deployable pipeline, and aggregate results for a real-world quality-assurance audit of a challenge-winning intracranial aneurysm detection model
Ayis Pyrros, Brian T. Layden, Pola Lydia Lagari et al.
Zenodo (CERN European Organization for Nuclear Research) · 2026-08-31
This archive provides the complete analysis code, deployable pipeline, and aggregate results from a single-institution retrospective quality-assurance audit of the first-place model from the 2025 RSNA Intracranial Aneurysm Detection AI Challenge, applied to 6,592 consecutive brain and skull-base examinations (4,769 evaluable). The release includes a frozen GPT-4.1 report-extraction prompt, study-group classifier, tabulation and figure code, and a full end-to-end pipeline covering PACS retrieval, inference, and a QA dashboard, all pinned to the audited environment. By making every component reproducible and publicly verifiable—with SHA-256 checksums and an exclusions log—the work demonstrates a rigorous framework for real-world clinical QA auditing of challenge-winning AI models before or during deployment.
- Quality assurance
Research
Agentic AI in Security Operations Centers: Transforming Cybersecurity Delivery Through Bounded Autonomy
Erich Barlow
Zenodo (CERN European Organization for Nuclear Research) · 2026-08-31
This paper examines the deployment of agentic AI systems in Security Operations Centers (SOCs), synthesizing controlled and operational studies showing measurable productivity gains: AI-assisted analysts completed investigations 45–61% faster with 22–29% higher accuracy, and generative AI adoption was associated with a 30.13% reduction in incident mean time to resolution. However, the paper cautions that fully autonomous SOC replacement is not yet supported by evidence, citing attack success rates of up to 84.30% against large-language-model agents and immature assurance standards. The authors conclude that the most defensible near-term model is 'bounded agentic operations,' where specialized agents handle high-volume, reversible tasks under human oversight, least privilege, and policy-enforced authority, anchored by frameworks such as ISO/IEC 42001, NIST AI RMF, and MITRE ATLAS.
- Workforce
- AI policy
Research
Transforming Chemical Safety Policy in the AI Era: A Seven-Action Strategy for the Five Chemical Safety Acts Under the Ministry of Climate, Energy, and Environment
Ho-Hyun Kim, Lim Ho-Ju, Hunjoo Lee
Korean Journal of Environmental Health Sciences · 2026-08-31
This paper proposes a seven-action strategy for modernizing chemical safety governance in South Korea by integrating AI and knowledge graph technologies across five acts under the Ministry of Climate, Energy, and Environment. The authors identify fragmented data systems, inconsistent classification schemes, and siloed identifiers as barriers to cross-act risk prediction, drawing parallels to past information-silo failures like 9/11. The strategy progresses from designating high-value datasets and adopting metadata standards, through building domain ontologies and knowledge graphs, to deploying AI-based early warning systems with explainability requirements. The framework is grounded in international benchmarks including EU REACH, OSOA, and US TSCA, and is intended to be implemented in a stepwise, institutionally coordinated manner.
- AI policy
Research
Analysis code, deployable pipeline, and aggregate results for a real-world quality-assurance audit of a challenge-winning intracranial aneurysm detection model
Ayis Pyrros, Brian T. Layden, Pola Lydia Lagari et al.
Zenodo (CERN European Organization for Nuclear Research) · 2026-08-31
This archive provides the full analysis code, deployable pipeline, and aggregate results from a single-institution retrospective quality-assurance audit of the first-place model from the 2025 RSNA Intracranial Aneurysm Detection AI Challenge, evaluated on 6,592 consecutive brain and skull-base examinations (4,769 evaluable). The release includes the frozen GPT-4.1 report-extraction prompt, study-group classifier, tabulation and figure code, and a complete containerized deployment pipeline covering PACS retrieval, inference, and a QA dashboard. By making the full audit infrastructure publicly reproducible—while excluding protected health information—this work demonstrates a rigorous, real-world quality-assurance methodology for validating challenge-winning AI models before or during clinical deployment.
- Quality assurance
Research
Analysis code, deployable pipeline, and aggregate results for a real-world quality-assurance audit of a challenge-winning intracranial aneurysm detection model
Ayis Pyrros, Brian T. Layden, Pola Lydia Lagari et al.
Zenodo (CERN European Organization for Nuclear Research) · 2026-08-31
This archive supports a real-world quality-assurance audit of the top-performing model from the 2025 RSNA Intracranial Aneurysm Detection AI Challenge, applied to 6,592 consecutive brain and skull-base examinations at a single institution (4,769 evaluable). The release includes the full deployable inference pipeline—covering PACS retrieval, a polling agent, inference wrapper, and QA dashboard—alongside the GPT-4.1 report-extraction prompt, analysis code, and aggregate results underlying every reported figure. By making the complete audit infrastructure publicly available with SHA-256 integrity verification and no patient data, the work enables reproducible, institution-level QA evaluation of AI radiology models in clinical deployment contexts.
- Quality assurance
Research
Role of Artificial Intelligence in the Criminal Justice System: A Critical Legal Study of Investigation and Judicial Processes
Bikash Kumar Pandey and Dr. Sanjay Kumar Singh
International Journal of Advanced Research in Science Communication and Technology · 2026-08-31
This legal study critically examines how AI tools—including predictive policing, facial recognition, risk-assessment algorithms like COMPAS, and court-management platforms like SUPACE and SUVAS—are reshaping criminal investigation and adjudication in India and comparatively abroad. Using doctrinal and comparative methodology, the authors find that while AI improves efficiency and case-clearance rates, unregulated deployment risks eroding due process, entrenching discriminatory bias, and undermining accountability. The paper grounds its analysis in Indian constitutional provisions (Articles 14, 20(3), and 21) alongside landmark cases such as State v. Loomis and Puttaswamy v. Union of India. It concludes by recommending a graduated, rights-based regulatory framework modeled on the EU AI Act and emerging Indian data-protection jurisprudence.
- AI policy
Research
Artificial Intelligence for Medical Imaging Diagnosis: From Accuracy to Clinical Reliability through Multimodal Fusion, Validation, and Regulatory Perspectives
Enoch Jacob Dodo, Amos Takai Yayock, Gregory Onwodi et al.
Journal of Science Research and Reviews · 2026-08-31
This narrative survey examines AI-based medical imaging diagnosis, noting that deep learning models frequently report accuracy, sensitivity, and specificity above 90% under controlled conditions, but that translating these results into clinically reliable systems remains a critical challenge. The paper introduces a modality-aware analytical framework organizing literature across imaging modality, data provenance, validation maturity, and model architecture, synthesizing unimodal and multimodal fusion approaches across radiology, pathology, ophthalmology, and multi-source fusion with EHR and genomic data. Key findings highlight that high reported accuracy is strongly contingent on data characteristics and evaluation conditions, with many models lacking real-world generalization evidence, and that significant gaps remain in alignment with regulatory pathways such as FDA 510(k), De Novo, and EU MDR/IVDR. The survey proposes an engineering-oriented deployment framework and a clinical deployment readiness model addressing validation maturity, algorithmic bias, explainability, and regulatory alignment.
- Quality assurance
- Certifications
Research
The Language of the Question Selects the Market: Query Language and Exit IP as Separable Factors in Commercial Recommendations from a Generative Search Interface
Dmitrij Żatuchin
arXiv · 2026-08-30
This study reports a controlled experiment with 234 runs of ChatGPT and the OpenAI API to examine how query language and exit IP address independently shape which market's products a generative search interface recommends. The key finding is that query language — not geographic location — determines whether local suppliers appear at all, while changing only the exit IP shifts which country's brands are named without altering the language of the response. A negative control category with no language effect suggests the mechanism is tied to whether a product category is nationally regulated rather than simply nationally supplied, raising questions about market fairness and algorithmic transparency in AI-driven commercial recommendations.
- Enterprise
- AI policy
Research
Automatic Conversion of NICE Guidelines to an Executable Computational Model Using Large Language Models
Ashvin Gupta, Denys Prociuk, Alessandra Russo et al.
arXiv · 2026-08-30
This paper presents an LLM-based pipeline that automatically converts unstructured NICE clinical guidelines into executable computational models capable of generating patient-specific, explainable recommendations. Applied to pancreatic and lung cancer guidelines, expert review confirmed strong alignment between the source guidelines and the generated models, with most errors being partial omissions rather than incorrect logic, and the pancreatic cancer model achieved an F1 score of 82.5% on 20 patient vignettes. The work demonstrates that scalable, automated generation of computable clinical guidelines is feasible, reducing the need for time-intensive manual encoding. This matters for quality assurance in clinical decision support and for policy efforts to make evidence-based guidelines consistently actionable at scale.
- Quality assurance
- AI policy
Research
Error Detection for PET/CT Radiology Reports: Domain-Specific vs Large Language Models
Hermione Warr, Harry Anthony, Lilli J Freischem et al.
arXiv · 2026-08-30
This study presents the first systematic evaluation of language models for error detection in PET/CT radiology reports, comparing compact domain-specific BERT models against large open-weight LLMs (Qwen3-32B, Gemma-3-27B, and Llama-3.3-70B). Using over 30,000 oncology FDG PET/CT reports, the researchers found that a 15-million-parameter domain-specific model achieved 94.4% balanced accuracy with a 5.8% false-positive rate, outperforming the best prompted LLM at 84.0%. Task-specific fine-tuning of Llama-3.3-70B matched the smaller model's accuracy but required far greater computational resources. The findings suggest that domain-specific training is more important than model scale for radiology report quality assurance, supporting efficient compact models as a practical automated QA tool.
- Quality assurance
Research
Generating Clinical Vignettes that Preserve Cognitive Formulations
Amit Oren, Nimrod Hertz-Palmor, Dean Ariel et al.
arXiv · 2026-08-30
This paper introduces FORMA, a framework that compiles a cognitive model of a psychological disorder into a directed weighted graph to guide large language models in generating clinically accurate case vignettes. Applied to PTSD using the Ehlers and Clark cognitive model, FORMA generated 16,500 vignettes and demonstrated that the underlying cognitive structure is recoverable from full-condition vignettes (MCC = +0.41, AUC = 0.70) but not from zero-shot generation (MCC = +0.01, AUC = 0.50). Clinical experts rated FORMA vignettes substantially higher than zero-shot alternatives, and licensed practitioners judged them as human-written 85% of the time versus 22% for zero-shot. FORMA also reduced demographic disparity in perceived quality by 1.5–7x, showing that theory-grounded generation can serve as an auditable standard for scalable synthetic clinical text.
- Quality assurance
Research
Review Before Trust: Source-Grounded Integrity Gates for AI-Assisted Personal Health Records
Nora Girda, Adrian Groza
arXiv · 2026-08-30
This paper addresses the risk that large language models converting medical documents into structured health records may produce plausible but unsupported outputs, which if persisted could corrupt longitudinal patient data. The authors introduce an evidence-gated trust-promotion model in which a deterministic monitor independently verifies each AI-generated candidate against the source document before it is admitted for downstream use, requiring a unique supporting quotation and correct row-level provenance. Implemented in the Medical DataCloud personal health-record application and evaluated on 102 manually labelled laboratory rows, the system admitted 72 numeric candidates, retained 25 for human review, and passed all 22 conformance and mutation tests. The study demonstrates the technical feasibility of enforcing a verifiable boundary that prevents AI-generated claims from authorizing their own reuse in a longitudinal health record.
- Quality assurance
Research
The Policy Deficit in AI x Social-Emotional Learning Research
Tran Van Cuong, Liu Yihan, Nguyen Van Tuong
arXiv (Cornell University) · 2026-08-30
This paper systematically reviews 65 peer-reviewed studies on the intersection of artificial intelligence and social-emotional learning (SEL), finding a substantial 'policy deficit': nearly three-quarters of those studies mention no policy implications at all. Using a 'WH-question' framework (Who, What, Why, When/Where, and How), the authors show that even when policy implications are discussed, they typically lack the specificity and actor-oriented guidance needed for effective evidence-informed policymaking. The study identifies a 'techno-solutionist' trap in which technical potential is foregrounded while institutional conditions for responsible implementation remain under-specified, and links this pattern to academic publication incentive structures. The authors propose shifting from 'implication-as-afterthought' to 'implication-as-methodology,' offering actionable guidelines for researchers, editors, reviewers, and policymakers to better connect AI-SEL innovation with educational governance.
- AI policy
Research
ContextBias: Controlled Evaluation of Bias Persistence Under Context Shift in Text-to-Image Models
Shaghayegh Kolli, Sina Emami, Moreno D'Incà et al.
arXiv · 2026-08-30
ContextBias introduces a controlled evaluation framework and benchmark (ContextBench) covering 92 professional roles and 1,656 semantically controlled prompts to test whether stereotypical visual associations in text-to-image models persist when roles are placed in unrelated contexts. Evaluating four state-of-the-art models across 66,240 generated images, the study finds that moving a role into a semantically unrelated context does not suppress role-linked visual attributes — demographic cues, characteristic garments, and role-specific tools remain prevalent, and cross-role attribute concentration actually increases (pooled BI +0.047). Only scene composition and camera framing show meaningful sensitivity to context. These findings reveal a form of stereotype persistence that is invisible to context-free bias evaluations, underscoring the need for contextual variation in bias benchmarking of generative AI systems.
- Quality assurance
- AI policy
Research
How You Ask Shapes What You Get: A Theory-Seeded Measurement of Articulation in Advice-Seeking LLM Conversations
Juneha Baek, Suhyeon Lee, Donghyuk Shin
arXiv · 2026-08-30
This paper investigates how the way users phrase advice-seeking requests—their 'articulation style'—affects how large language models respond, independent of the topic being asked about. Analyzing 16,447 prompts from public chat corpora (WildChat, LMSYS, and ShareChat), the researchers identify a small set of stable, replicable latent articulation factors. A key finding is that one common style—long-form but information-poor prompts, appearing in roughly one in six messages—leads models to return shorter, vaguer answers without asking clarifying questions, even though the under-specification would warrant them. The authors argue that AI benchmarks should stratify evaluations by articulation style, not just topic, to avoid systematic blind spots in assessing model behavior.
- Quality assurance
Research
An AI Governance Framework to Address Algorithmic Bias and Educational Equity: A Systematic Literature Review
Muhammad Ruslan Maulani
Telematika · 2026-08-30
This systematic literature review synthesizes evidence from 46 peer-reviewed studies to examine how algorithmic bias manifests in educational AI systems and disproportionately harms students from low-income, racially minoritized, and geographically remote backgrounds. The authors find that existing AI governance frameworks are too generic to address education's unique pedagogical and institutional dynamics. As a primary contribution, the paper proposes the Educational AI Governance Framework (EAGF), a five-layer model grounded in FAIR principles and aligned with UNESCO, OECD, and EU standards, offering a practical roadmap for policymakers, educational institutions, and EdTech developers to operationalize equitable AI governance.
- AI policy
- Certifications
Research
Tax Technology Adoption and Corporate Tax Compliance in the Coretax Era: The Moderating Role of External Consultants
Jonris Hotman Tua, Basyiruddin Nur, Karsam Karsam et al.
Greenation International Journal of Economics and Accounting · 2026-08-30
This study examines how AI and cloud accounting adoption affect corporate tax compliance among Foreign Direct Investment companies in Indonesia under the new Coretax tax administration system. Using PLS-SEM analysis of 150 fiscal leaders in the Bekasi industrial cluster, the researchers find that both AI adoption (β=0.312) and cloud accounting adoption (β=0.285) significantly improve tax compliance. Importantly, external tax consultants amplify these positive effects, suggesting that professional expertise is critical for aligning digital infrastructure with complex regulatory demands. The findings offer practical guidance for corporate fiscal governance as Indonesia modernizes its tax administration.
- Enterprise
- AI policy
Research
AI RELIANCE AND PROFESSIONAL JUDGMENT QUALITY AMONG GOVERNMENT AUDITORS IN ACEH: THE MODERATING ROLE OF PROFESSIONAL SKEPTICISM
Luthfiar Ramiady, Hendri Bin Muhammad Nur
Sumber Informasi Manajemen Bisnis dan Akuntansi (SIMBAN) · 2026-08-30
This mixed-methods study of 110 government auditors in Aceh finds that greater reliance on AI tools has a significant negative direct effect on professional judgment quality, consistent with automation bias theory. However, professional skepticism moderates this relationship by buffering the negative effect: auditors with high skepticism treat AI as a complementary tool and cross-verify its outputs, while those with low skepticism accept AI outputs uncritically—a pattern described by informants as 'thinking laziness.' The findings highlight an urgent need for professional skepticism training integrated with AI literacy to prevent digital transformation from eroding the judgment quality that is central to the audit profession.
- Workforce
- Quality assurance
Research
The Enterprise AI Behavior Control Plane: Governing Longitudinal Drift in Production LLM Systems
Srikanth Devarakonda
International Journal of Computer Information Systems and Industrial Management Applications · 2026-08-30
This paper introduces the Enterprise AI Behavior Control Plane (BCP), a platform-layer architecture designed to detect and constrain behavioral drift in production Large Language Model systems. Drawing on Site Reliability Engineering principles, the BCP establishes behavioral baselines, enforces AI-specific service level agreements, and enables automated rollback strategies. Empirical analysis across production deployments found that behavioral drift affects 23.9% to 62.0% of deployed agents over extended interaction sequences, with task success rates degrading by up to 42.0% in uncontrolled environments. The work frames enterprise AI governance as a continuous control problem rather than a one-time evaluation task, offering a systematic framework for maintaining reliable AI behavior at scale.
- Enterprise
- Quality assurance
Research
An Agentic ERP Governance Framework for Autonomous AI Agent Deployment in Cloud-Based Industrial Management Systems
Venkata Ramachandra Karthik Chundi
International Journal of Computer Information Systems and Industrial Management Applications · 2026-08-30
This paper proposes the Agentic ERP Governance Framework (AEGF), a five-dimension governance instrument designed to guide responsible deployment of autonomous AI agents in cloud-based enterprise resource planning systems. The authors argue that existing governance models built for predictive analytics and rule-based bots are inadequate for agentic AI, which behaves non-deterministically and can produce cascading consequences in live financial environments. The AEGF covers Process Suitability, Autonomy Tiering, Governance and Auditability, Organisational Readiness, and Risk and Continuity Management, grounded in Sociotechnical Systems Theory, the Technology-Organisation-Environment framework, and the NIST AI Risk Management Framework. An illustrative application to accounts payable automation on Oracle ERP Cloud demonstrates the framework's practical use in industrial management contexts.
- Enterprise
- AI policy
Research
Claude-Driven Intelligent Test Automation: A Unified API and UI Framework for Enterprise and Healthcare Compliance
Pratik Dinkar Rane
International Journal of Computer Information Systems and Industrial Management Applications · 2026-08-30
ClinQA is an LLM-orchestrated test automation framework that unifies API and UI testing for enterprise and healthcare environments under a single compliance-aware pipeline powered by Claude. The framework embeds PHI masking, FHIR R4 conformance evidence capture, and FDA 21 CFR Part 11 audit trail generation at the infrastructure level, removing compliance responsibility from individual test authors. Preliminary observations across 12 scenarios show an 87% Cross-Layer Coverage Index, 68% reduction in test authoring effort versus manual methods, and a 94% Compliance Coverage Score, suggesting LLM-driven unified test generation is feasible for regulated environments.
- Enterprise
- Quality assurance
- Certifications
Research
AI-Assisted Test Execution as an Augmentation Layer in Enterprise Quality Engineering
Rejenish Kiran
International Journal of Computer Information Systems and Industrial Management Applications · 2026-08-30
This paper introduces the Probabilistic Augmentation and Governance Model (PAGM), a three-tier framework that formally divides responsibility between AI agents and human testers across autonomous execution, confidence-gated escalation, and human-led verification in regulated enterprise software environments. Applied to a property insurance release pipeline, PAGM enables a full quarterly regression cycle within a five-day SLA while preserving mandatory human accountability for premium calculation and claims workflows. The paper draws on literature showing contextual UI recognition methods achieve a 95% change-handling rate versus 40–80% for conventional tools, and that AI-based classifiers outperform classical approaches by 25–50% across software quality metrics. PAGM positions governance, traceability, and bounded autonomy as core design requirements rather than performance optimizations, offering a governance-ready foundation for AI-assisted testing at enterprise scale.
- Enterprise
- Quality assurance