News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
Collusion with Competitive Marginals: Price-Level Audits Are Blind by Construction
Xin Xu, Chengrui Wu, Jiayu Lu et al.
arXiv · 2026-07-29
This paper demonstrates a fundamental flaw in empirical methods used to detect algorithmic collusion: conspiring agents can coordinate through the joint distribution of their bid components while keeping each individual agent's bid distribution indistinguishable from competitive behavior, rendering single-agent price-level audits blind by construction rather than merely underpowered. The authors validate this theoretically and empirically using language-model bidding agents and Ethereum block-building auction data, showing that residual correlations between co-deployed models exist but are undetectable by standard marginal-price screens. Because lawful multi-identity operation and actual conspiracy are behaviorally indistinguishable to existing detectors, the paper argues the tractable regulatory target is market structure counting—resolving 40 bidding identities into 23 operators raises the Herfindahl index by 247.5%, with behavioral clustering pushing that figure to 324.5%. These findings have direct implications for how regulators and auditors design oversight of AI-driven markets.
- AI policy
- Quality assurance
Research
(Im)Paired Programming: Coding Agents Improve Productivity but Harm Understanding
Nishant Balepur, Connor Baumler, Valerie Chen et al.
arXiv · 2026-07-29
This study examines how AI coding agents (like Cursor) affect developer understanding compared to chatbot-based coding assistance, using a controlled experiment with 54 students building websites. Results show that while coding agents improve initial task completion speed, they significantly harm users' code comprehension and leave them unable to extend their own code without AI help. Low-effort interaction patterns—such as copy-paste prompts and auto-accepted edits—are especially linked to poor understanding, and users prefer agents anyway because of convenience despite acknowledging weaker comprehension. The findings highlight a tension between productivity gains and meaningful human oversight, with implications for how coding agent developers should design for active engagement and learning.
- Workforce
- Quality assurance
Research
Open Security Benchmark: Towards Autonomous Enterprise Cyber Defense
Gal Engelberg, Michael Arenzon, Leon Goldberg
arXiv (Cornell University) · 2026-07-29
The paper introduces Open Security Benchmark (OSB), a framework for evaluating agentic AI systems performing autonomous enterprise cyber defense tasks such as security posture investigation. The authors identify an 'environment data gap'—the lack of shared, queryable, realistic enterprise environments needed to rigorously assess whether AI agents can be trusted for security work—and address it by providing a frozen, holistic synthetic enterprise environment with closed-form ground truth answers. OSB supports evaluation via text-to-SQL queries and native vendor APIs, uses multi-dimensional scoring, and is instantiated with identity-security packs and synthetic organization datasets at multiple scales. This matters because it enables reproducible, trustworthy benchmarking of AI agents before they are deployed to make real security decisions in enterprise environments.
- Enterprise
- Quality assurance
Research
Artificial Intelligence (AI), Audit Quality, and the Future of Professional Judgment: Policy and Governance Challenges in Auditing - A Systematic Literature Review
Geoffrey Odoch
International Journal of Computer Information Systems and Industrial Management Applications · 2026-07-29
This systematic literature review examines how AI integration into auditing is reshaping audit quality and professional judgment, identifying major policy and governance challenges facing the profession. The review finds a core tension between automating audit tasks and preserving professional skepticism, while highlighting risks including algorithmic bias, lack of transparency in AI systems, and regulatory lag. Key policy gaps identified include unresolved liability issues, eroding professional identity, new quality assurance demands, and absent standardization frameworks. The authors call for coordinated action from regulators, standard-setters, firms, and educators to develop governance models and audit methodologies that integrate human and machine intelligence.
- AI policy
- Quality assurance
Research
Lottery Tickets Are Not Deployment Tickets
Bum Jun Kim
arXiv (Cornell University) · 2026-07-29
This paper investigates whether sparse or compressed neural networks (lottery tickets) can directly replace dense incumbent models in deployed systems without reconfiguring downstream decision logic. Across extensive experiments, the authors find that while sparse models often match clean accuracy, they remain behaviorally different from their dense counterparts across calibration, out-of-distribution response, class-level reliability, and representations. Critically, in fixed-threshold policy settings, swapping in a lottery ticket changed 7–10% of accept–review decisions—exactly the kind of churn that drop-in replacement is meant to avoid. The paper concludes that clean-accuracy matching is insufficient for deployment certification, and that behavioral compatibility with a fixed incumbent is a distinct and necessary requirement.
- Quality assurance
- Certifications
Research
Can We Trust AI in 6G? Verifiable and Auditable AI-Driven Trustworthy Wireless Networks
Genze Jiang, Yizhou Huang, Kezhi Wang
arXiv (Cornell University) · 2026-07-29
This paper addresses a critical trust problem in AI-driven wireless networks: there is currently no way to verify that AI functions—such as those handling cell selection and mobility management in 6G—are making decisions for the right reasons rather than exploiting unreliable shortcuts. The authors propose a 'mechanical auditing' approach that inspects AI systems' internal representations and checks them against machine-verifiable 3GPP specifications through a three-step principle: locating protocol-relevant features, verifying their causal role, and diagnosing how adaptation changes their use. They introduce an audit-native network architecture featuring a dedicated verification agent that supports both pre-deployment certification and runtime auditing. The work also identifies open challenges that must be resolved before mechanistic auditing can enter telecommunications standardisation practice.
- Certifications
- Quality assurance
Research
Assurance-Scoped Reliability for Agentic Networks: Capturing the State That Matters
Bilgehan Erman, Andrea Francini, Nikos Papadis
arXiv (Cornell University) · 2026-07-29
This paper identifies failure modes in agentic AI networks—such as acting on stale information, repeating external actions, or quietly relaxing policy enforcement—that conventional reliability measures miss, even when a service appears healthy. The authors propose Reliability Assurance Intelligence (RAI), an architecture that derives a per-service reliability profile specifying what must be checked, recorded, recovered, and audited, then retains durable state in a 'context capsule' for recovery and accountability at runtime. The work matters because it addresses accountability gaps in autonomous, cross-domain AI systems where failures may be invisible to standard monitoring. A methodology for validating RAI's reliability assurances is also proposed, using an agentic lifecycle manager for deterministic network services as a running example.
- Quality assurance
- Certifications
Research
SANDBOX REGULATÓRIO BRASILEIRO COMO LABORATÓRIO CONSTITUCIONAL: LIMITES ESTRUTURAIS DOS MODELOS DO BANCO CENTRAL DO BRASIL E DA COMISSÃO DE VALORES MOBILIÁRIOS DIANTE DE SISTEMAS DE INTELIGÊNCIA ARTIFICIAL E MACHINE LEARNING
Eduarda Hoffmann, Juliano Heinen
Veredas do Direito Direito Ambiental e Desenvolvimento Sustentável · 2026-07-29
This Brazilian legal study examines whether the regulatory sandboxes operated by Brazil's Central Bank (BCB) and Securities Commission (CVM) are constitutionally adequate for overseeing artificial intelligence and machine learning systems. The paper identifies four structural disconnects between the existing regulatory framework and AI/ML realities—temporal, ontological, accountability, and territorial gaps—drawing comparisons with the EU AI Act and the UK Financial Conduct Authority's sandbox model. The authors conclude that the current Brazilian sandbox framework is constitutionally insufficient for AI systems and propose five minimum legitimacy safeguards, including formal legislative reservation, prior fundamental-rights impact assessment, mandatory civil society participation, full disclosure of results, and an independent review body. The findings carry direct implications for how AI regulatory policy should be designed and legitimized under constitutional constraints.
- AI policy
Research
DESAFÍOS JURÍDICOS DE LA INTELIGENCIA ARTIFICIAL EN EL MERCADO LABORAL MEXICANO: ANÁLISIS JURÍDICO-CRÍTICO SOBRE DERECHOS FUNDAMENTALES DEL TRABAJADOR
Liliana C. Becerra-Vargas, Teresa de Jesús Vargas Vega
KAIRÓS REVISTA DE CIENCIAS ECONÓMICAS JURÍDICAS Y ADMINISTRATIVAS · 2026-07-29
This legal analysis argues that Mexico's current legal framework is insufficient to protect workers' fundamental rights in the face of algorithmic management and AI-driven labor practices. The article identifies gaps in the Federal Labor Law and the personal data protection statute (LFPDPPP), particularly regarding dignified work, data privacy, and non-discrimination. Through doctrinal, normative, and comparative analysis, the authors propose reforms to labor legislation and the creation of algorithmic audit mechanisms.
- Workforce
- AI policy
Research
AI application in commerce and services: Digital skills training in SMEs
Trieu Thai Duong, Huynh Thanh Danh, Đỗ Đăng Trình
Journal of Science and Development Economics · 2026-07-29
This study assesses digital transformation and AI adoption in human resource management and workforce training among 89 small and medium-sized enterprises (SMEs) in Vietnam's Mekong Delta commerce and service sectors. Using mixed methods including surveys and interviews with business leaders, managers, and employees, the findings reveal a significant capability gap between leaders and employees in AI application, with employees expressing high demand for AI-integrated digital skills training that current programs have not fully met. The paper highlights collaboration between universities and SMEs as a key mechanism for co-designing training programs to build a digitally skilled workforce.
- Workforce
- Enterprise
Research
Model-based, in-situ, non-destructive qualification and certification of parts made by autonomous additive manufacturing
Dayalan Gunasegaram, T. DebRoy, Paul Greenway et al.
Journal of Physics Materials · 2026-07-29
This paper proposes an integrated framework combining model-based qualification and certification (MBQ&C) with autonomous additive manufacturing (AAM) to address the productivity bottlenecks of traditional post-build inspection and testing for 3D-printed parts. The framework uses high-fidelity machine learning and reduced-order physics models within the Integrated Computational Materials Engineering paradigm to simulate process-structure-property-performance relationships, enabling build-specific fitness assessments using in-situ sensor data rather than generic parameters. Key claimed benefits include faster certification decisions, performance-based defect classification, and reduced reliance on costly post-build computed tomography scanning and destructive testing. The authors argue this approach is especially valuable for high-consequence and mission-critical applications where experimental testing environments are hazardous or impractical.
- Certifications
- Quality assurance
Research
Integrating Artificial Intelligence Technologies into Health Workforce Education: A Scoping Review of Digital Health Tools in Nursing Curricula
Dr S Kanaka Lakshmi
International Journal of Nursing Information · 2026-07-29
This scoping review systematically maps global applications of five AI technologies—intelligent tutoring systems, virtual patient simulations, adaptive platforms, natural language processing, and predictive analytics—in nursing education from 2020 to 2026. Drawing on PRISMA-ScR-guided screening across Scopus, PubMed, and CINAHL, the authors find these tools enhance clinical reasoning, critical thinking, and professional competence, but adoption is geographically uneven and impeded by weak infrastructure, high costs, data privacy concerns, and low faculty digital readiness. The review concludes that AI integration must move from ad-hoc use toward structured, policy-driven curricular frameworks, and offers a strategic benchmark for standardizing digital health competencies across nursing programs worldwide.
- Workforce
- AI policy
Research
Automated software scoring of senior school certificate examination mathematical items in economics using a contextual similarity model
Damilola Daniel Olaoye, H. O. Owolabi, Oluwaseun Tayo Olaoye
Frontiers in Education · 2026-07-29
This study developed an AI software system using a contextual similarity model to automatically score mathematical computation items in secondary school Economics examinations, then validated it against 12 human expert markers using a sample of 1,008 students from South-west Nigeria. The AI system achieved a high intra-class correlation coefficient (ICC r = 0.86, p < 0.01) with human scorers, indicating strong agreement. The findings suggest AI-based automated scoring can deliver consistency comparable to human expert marking while reducing costs and time, and the authors recommend adoption by teachers, evaluators, and examination bodies.
- Quality assurance
- Certifications
Research
‘Light-Touch Rights’ in AI Governance: A Business and Human Rights Analysis of South Korea’s AI Framework Act
Kyoungsic Min
Business and Human Rights Journal · 2026-07-29
This paper analyzes South Korea's AI Framework Act—the first comprehensive national AI legislation in Asia, enacted in January 2025—through the lens of the UN Guiding Principles on Business and Human Rights (UNGPs). The authors argue the Act creates a 'light-touch rights' model in which industrial promotion obligations are legally binding while human rights protections rely on procedural duties and best-effort provisions. Assessed against the three UNGP pillars, the Act lacks a prohibited category for unacceptable-risk AI, disperses due diligence into voluntary obligations, and provides no meaningful remedial pathways for people harmed by AI systems. The authors warn that without reform, South Korea risks exporting a 'rights without remedies' template across the Asia-Pacific region.
- AI policy
Research
How AI Awareness Impacts Employee Adaptive Performance: A Moderated Mediation Model
Hui Li, Yixuan Sun
Behavioral Sciences · 2026-07-29
Using three-wave survey data from 369 manufacturing employees, this study finds that how workers perceive AI at work—as a challenge or a hindrance—has opposite effects on their adaptive performance. Job crafting mediates both relationships, and mindfulness strengthens the positive effect of challenge awareness while buffering the negative effect of hindrance awareness. The findings offer practical guidance for organizations seeking to help employees adjust to AI-enabled workplaces.
- Workforce
Research
Using LLM-generated draft replies to support human experts in responding to stakeholder inquiries in maritime industry: a real-world case study of Industrial AI
Tita Alissa Bach, Aleksandar Babic, Narae Park et al.
Cogent Engineering · 2026-07-29
This real-world case study examines how LLM-generated draft replies can support human experts handling stakeholder inquiries in the maritime industry. Using mixed methods—observations, interviews, surveys, and text similarity analysis—the researchers found that LLM drafts improved efficiency, linguistic consistency, and time management, but frequently required manual adjustments for accuracy and relevance. Text similarity analyses showed only moderate alignment between drafts and final replies, underscoring the need for human oversight, while case handlers displayed a 'skeptical but curious' attitude toward adoption. The authors conclude that LLMs are not yet suitable for independent use in safety-critical maritime settings but can serve as valuable augmentative tools when combined with human expertise.
- Enterprise
- Workforce
Research
Transforming thyroid disease education: AI and virtual technologies in residency training
Shujian Xu, Cui Zhao, Nannan Sun et al.
Frontiers in Endocrinology · 2026-07-29
This narrative review examines how artificial intelligence and virtual reality technologies can improve residency training for thyroid disease specialists. Drawing on 42 studies published between 2015 and 2025, the authors find that AI tools enable real-time, objective assessment of ultrasound skills and support personalized learning, while VR platforms offer immersive, risk-free environments for practicing procedures such as thyroidectomy. Key challenges include data security concerns, limited anatomical fidelity and haptic feedback, high costs, and difficulties integrating these tools into existing curricula. The review concludes that AI and VR are valuable supplements to traditional training but must be implemented within a human-centered framework that preserves clinical and humanistic competence.
- Workforce
- Certifications
Research
Redesigning with Intelligence: Faculty Led Innovation in Nursing Curriculum Using Artificial Intelligence
Taylor Long, Nadine Wodwaski, Phillip Olla
International Journal of Nursing Education · 2026-07-29
This study compared AI-assisted versus manual nursing curriculum redesign, finding that using a customized ChatGPT-4 model reduced redesign time from approximately 41.6 hours to 2.9 hours—a 93% time savings. The AI tool was tailored to align with the American Association of Colleges of Nursing's 2021 Essentials to support accreditation compliance. Qualitative themes from 25 faculty included reduced cognitive workload and interest in AI support, though findings may be affected by self-report and recall bias. The results suggest AI can meaningfully streamline accreditation-driven curriculum work and free faculty for other responsibilities.
- Certifications
- Workforce
Research
Improving Efficiency and Effectiveness in Industrial Support Business Processes through Low-Code Conversational AI: Evidence from a Workflow-Embedded Case Study
Paulo Peças, Diogo Pires, Diogo Jorge
International Journal of Mathematical Engineering and Management Sciences · 2026-07-29
This case study examines how two low-code conversational AI agents—ManuBot and MailBot—were embedded into industrial maintenance-support workflows to reduce administrative burden. MailBot cut supplier-email preparation time from 12–15 minutes down to 2–3 minutes per message, while ManuBot reduced maintenance-data retrieval and querying time by approximately 50%. Users also reported improved access to historical malfunction records and more structured communication routines. The findings suggest that task-specific, workflow-embedded AI agents can deliver measurable efficiency gains, though constraints such as incomplete ERP integration and uneven user readiness limit broader generalization.
- Enterprise
- Workforce
Research
Digital infrastructure, innovation capacity, and AI technology adoption in EU manufacturing
Ruxandra Boghian
Economics of Innovation and New Technology · 2026-07-29
This study analyzes AI adoption across EU-27 manufacturing and service enterprises using Eurostat panel data from 2023 and 2025, finding that digital infrastructure (measured by the Digital Intensity Index) is the strongest predictor of AI uptake, with large firms consistently outpacing SMEs. Central and Eastern European countries lag behind Western Europe overall but are catching up fastest—Romania recorded a 223% increase—and statistical convergence tests confirm narrowing cross-country disparities in AI diffusion. The results carry direct implications for EU industrial policy, quality-management upgrading, and targeted AI-adoption programmes in lagging regions.
- Enterprise
- AI policy
Research
Does Al Readiness Promote Financial Development? Evidence On Institutional Complementarity
Junjie Shu
Journal of Economics Finance and Management Studies · 2026-07-29
This study finds that national AI readiness—measured by the Oxford Insights Government AI Readiness Index across 193 economies from 2020–2024—is associated with a 4.25-percentage-point increase in bank credit to the private sector as a share of GDP per one-standard-deviation rise in AI readiness. Crucially, this effect is conditional on institutional quality: the relationship is statistically insignificant in low-quality institutional environments but grows strongest where institutional quality is high. The authors identify formal entrepreneurship, labor productivity, and structural upgrading as positive transmission channels, and results hold across multiple robustness checks including GMM and instrumental-variable estimates. The findings suggest that AI readiness alone is insufficient to drive financial development without complementary institutional frameworks that ensure credibility, enforceability, and scalability.
- Enterprise
- AI policy
Research
Safety design guidelines for clinician–AI interaction in computer-aided diagnosis systems using system-theoretic framework with explainability validation
Yuki Hagiwara, Katherine Fitch, Mario Trapp
Scientific Reports · 2026-07-29
This paper develops safety design guidelines for AI-assisted computer-aided diagnosis (CAD) systems by applying System-Theoretic Process Analysis (STPA) to identify critical hazards in clinician-AI collaboration, such as automation bias and misinterpretation of AI explanations. The authors translate identified unsafe control actions into actionable design guidelines emphasizing transparency, coherent explanations, and calibrated trust, then operationalize these within a safety-oriented GUI that integrates multiple explainability methods and interactive mechanisms. A novel evaluation approach uses consistency across multiple explanation methods as a quantitative indicator of potentially unreliable or ambiguous AI outputs. The work is relevant to healthcare AI safety and quality assurance, offering a unified framework for improving the reliability of AI-driven diagnostic systems.
- Quality assurance
- AI policy
Research
Artificial intelligence for pediatric fracture detection: impact on diagnostic revisions and patient recall rates in a tertiary emergency setting
Oliver Johannes Deffaa, J Pape, Dominik Schlösser et al.
BMC Emergency Medicine · 2026-07-29
This prospective study evaluated a commercial AI system (TechCare Kids) for pediatric fracture detection during out-of-hours emergency care at a tertiary hospital, enrolling 667 children aged 2–18 years. The AI achieved 95.1% accuracy, and diagnostic revision rates trended lower with AI support (5.7%) versus without (8.6%), but the difference was not statistically significant (risk ratio 0.66; 95% CI 0.37–1.19). No meaningful differences were found in therapeutic changes, senior consultation rates, diagnostic confidence, or emergency department length of stay. The authors conclude that the effect size is too small to justify a larger confirmatory trial for these outcomes in an academic setting.
- Quality assurance
Research
Supervisory XAI Toolkit: A privacy-preserving framework for independent regulatory auditing of artificial intelligence models in financial services
Andrew Moore, Samuel Allen
World Journal of Advanced Research and Reviews · 2026-07-29
This paper proposes a Supervisory Explainable AI (XAI) Toolkit that enables financial regulators to independently audit proprietary AI models—such as credit scoring, AML, and fraud detection systems—without requiring firms to disclose raw data or model weights. The toolkit combines differential privacy, secure multi-party computation, and trusted execution environments with a fairness and robustness metrics engine. An illustrative evaluation shows that firm-reported fairness metrics can materially overstate actual model fairness, and that a modest privacy budget is sufficient to recover most of the audit signal needed for supervisory decisions. The work has direct implications for closing information asymmetries between regulators and regulated firms and for advancing standardization in cross-border regulatory cooperation.
- AI policy
- Certifications
Research
Large language models in intelligent manufacturing and mechanical engineering: a review of robotics, fault diagnosis, design, and engineering knowledge workflows
Sherif Samy Sorour, Anwar Sahbel
Journal of Intelligent Manufacturing · 2026-07-29
This review surveys how large language models (LLMs) are being applied across robotics, fault diagnosis, design, and manufacturing knowledge workflows. The authors find that LLMs are most valuable as semantic and coordination layers—helping engineers navigate documents, data sources, and software environments—rather than as replacements for simulation or numerical tools. The strongest results come from hybrid architectures combining LLMs with retrieval-augmented generation, knowledge graphs, multimodal perception, and digital twins. Key limitations persist, including weak grounding in ambiguous settings, limited numerical reliability, and incomplete integration with trusted engineering software.
- Enterprise
- Workforce