News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
5361 items
Research
A data-driven and theory-guided framework for developing and validating human-robot collaboration training modules for the construction workforce
Ebenezer Omoniyi Olukanni, Abiola Akanmu, Houtan Jebelli
Journal of Information Technology in Construction · 2026-08-26
This study presents a systematic framework for designing and validating training modules that prepare construction workers to collaborate with robots and AI systems. Using natural language processing to augment competency data, an ADDIE instructional design process, and a two-round Delphi expert validation study, the researchers produced 50 validated human-robot collaboration competencies mapped across knowledge, skills, and abilities, which were then organized into seven progressive training modules. A key finding is that assessment requirements differ systematically across competency types, with higher-order and socio-cognitive competencies being harder to translate into measurable outcomes. The work offers construction organizations and educational institutions a scalable, reproducible methodology for workforce planning, training standardization, and supporting broader construction robotics adoption.
- Workforce
- Certifications
Research
How Did Trump's 2025 Trade War Affect the Decoupling of US –China Supply Chains?
Chad P. Bown
Asian Economic Policy Review · 2026-08-26
This paper empirically examines how Trump's 2018–19 and 2025 tariffs reshaped US–China supply chains at the product level. It finds that consumer electronics importers (smartphones, laptops, etc.) largely avoided sharp import declines by pivoting to pre-established alternative supply chains in India and Vietnam, while clothing and footwear importers did not reorient because tariff differentials between China and third countries were too small. The study also shows that growing US sourcing from Taiwan and Mexico partly reflects demand for AI data center inputs rather than pure supply-chain decoupling from China.
- AI policy
- Enterprise
Research
Why Automated Moderation Fails Women and Girls, and What to Do About It
Tomisin Olanrewaju
Zenodo (CERN European Organization for Nuclear Research) · 2026-08-26
This paper argues that AI-powered content moderation systems systematically fail women and girls by under-detecting gendered harms due to biased training data, weak contextual judgment, and a linguistic fluency gap that disadvantages users from the Global South. The authors cite TikTok's disclosure that automated systems actioned 93.8% of violating content without human review, illustrating how scale-driven automation sidelines nuanced, culturally specific abuse. The paper contends that purely technical fixes like retraining models are insufficient without economic incentives, and calls for a multi-layered response combining regulation requiring granular transparency disclosures, civil society-led NLP development for under-resourced languages, and platform-level user controls. The findings have direct implications for how content moderation policy and platform governance should be restructured to address gendered online harms equitably.
- AI policy
- Quality assurance
Research
Physicians’ perspectives on artificial intelligence in electrocardiography in clinical practice: a qualitative study
Gabriel Allgårdh, Gert Helgesson, Ulrik Kihlbom
BMC Medical Ethics · 2026-08-26
This qualitative study interviewed 12 physicians in Sweden about integrating AI-enhanced electrocardiography (AI-ECG) into clinical practice. Physicians saw potential for AI-ECG to improve diagnostic accuracy, support prioritization, and reduce workload, but raised concerns about overreliance, accountability, deskilling, and the management of incidental or prognostic findings. The study concludes that AI-ECG implementation requires real-world validation, protocols for handling unexpected findings, and attention to ethical and professional dimensions beyond technical performance. Because ECGs are extremely common and obtained at low clinical threshold, AI-ECG could affect large numbers of patients and substantially reshape established clinical workflows.
- Workforce
- AI policy
Research
The impact of artificial intelligence (AI) application on marketing performance in small and medium-sized enterprises in Hanoi
Do Hai Hung, Lê Anh Tuấn
Problems and Perspectives in Management · 2026-08-26
This study surveys 237 SME managers and employees in Hanoi, Vietnam to measure how AI-enabled marketing capabilities—customer data analytics, personalization, automation, and interactive communication—affect marketing performance across market, customer, financial, and communication dimensions. Using structural equation modeling grounded in resource-based view and dynamic capabilities theory, the findings show that AI-driven customer data analytics has statistically significant positive effects on all four performance dimensions, while AI-enabled personalization boosts market, financial, and communication performance but is associated with a negative coefficient for customer performance. The results provide empirical evidence from an emerging economy that AI adoption can meaningfully improve SME marketing effectiveness, though effects vary by application type and performance dimension. The study offers practical implications for resource-constrained SMEs seeking to prioritize AI investments in marketing.
- Enterprise
Research
AI washing
Moran Ofir
American Business Law Journal · 2026-08-26
This law review article provides the first comprehensive comparative analysis of how the United States and European Union regulate 'AI washing'—the misrepresentation of AI use in products or services to attract investors or gain competitive advantages. The U.S. relies on market-based enforcement through existing securities laws with penalties up to $225,000, while the EU's AI Act takes a proactive regulatory approach with penalties up to €35 million or 7% of global turnover. The article analyzes SEC enforcement actions, EU implementation patterns, Delaware board oversight duties, and materiality standards for AI disclosures, ultimately recommending selective convergence rather than full harmonization for multinational compliance. The findings matter because they establish regulatory precedents and practical governance frameworks for balancing investor protection with innovation in an era of rapid AI adoption.
- AI policy
- Enterprise
Research
Workforce Readiness in the AI Era: An Empirical Qualitative Study of Higher Education and Upskilling Among Urban Muslim Communities in Indonesia
Reza Fahmi, Prima Aswirna, Adamu Abubakar Muhammad
Al-Madinah Journal of Islamic Civilization · 2026-08-26
This qualitative study examines workforce readiness gaps in Indonesia's higher education system amid AI-driven labor market transformation, focusing on urban Muslim communities. Through interviews, focus groups, and document analysis with university staff, students, HR practitioners, vocational experts, and government officials, the study identifies five key challenges: outdated curricula, limited innovative pedagogy, weak lifelong learning culture, insufficient industry involvement, and slow policy adaptation. The researchers propose an integrated upskilling model combining industry-aligned curricula, project-based learning, micro-credentials, professional certification, and university–industry collaboration. The findings underscore the need for coordinated action among higher education, industry, and government to build a sustainable lifelong learning ecosystem responsive to AI-driven change.
- Workforce
- Certifications
Research
Driverless, not thoughtless? Automated vehicle regulation, the dilemma of control, and the prospect of dual corrigibility
Bård Torvetjønn Haugland
Technological Forecasting and Social Change · 2026-08-26
This article examines Norway's regulatory approach to automated vehicles, analyzing the policy-making process behind the Act Relating to Testing of Self-Driving Vehicles, the Act's content, and interview data from the first post-Act trial. The study finds that Norway employs a two-stage strategy—first allowing trials, then establishing permanent regulation—but argues this strategy rests on a flawed assumption that practical and normative concerns can be addressed separately. Drawing on Collingridge's dilemma of control, the authors propose 'dual corrigibility' (correcting both decisions and normative orientations in parallel) as a better guiding principle for regulating emerging technologies like automated vehicles.
- AI policy
Research
Artificial Intelligence in Investment Services
Filippo Annunziata
arXiv · 2026-08-26
This article analyzes how AI technologies—including robo-advice, algorithmic trading, and automated portfolio management—are being integrated into investment services across the EU, and assesses whether the existing regulatory framework (primarily MiFID II, alongside the AI Act, GDPR, and DORA) is adequate to manage the resulting risks. The study finds that while MiFID II's core duties of care, loyalty, transparency, and governance remain conceptually applicable, they require reinterpretation to address AI-specific challenges such as algorithmic bias, opacity, model oversight, and systemic concentration risks. The author argues that the fragmented evolution of EU digital and financial regulations leaves critical gaps, calling for clearer supervisory standards, harmonized model-governance practices, and a human-centered approach to preserve market integrity and consumer trust.
- AI policy
- Enterprise
Research
Reinventing the echocardiography workflow: from manual quantification to artificial intelligence–driven comprehensive interpretation
Ran Heo, Seung‐Ah Lee, Hyuk‐Jae Chang
Journal of Cardiovascular Imaging · 2026-08-26
This review evaluates the current state of AI integration in echocardiography workflows, showing that AI can reduce examination time, automate measurements, and mitigate sonographer fatigue while improving image quality. AI applications now extend beyond ejection fraction to include myocardial texture, Doppler hemodynamics, and assessments of valvular heart disease, cardiomyopathy, and pericardial disorders. However, the authors identify key gaps including reliance on single-center studies, inconsistent cross-platform performance, and risks of automation bias, and outline requirements for responsible clinical implementation. The findings are directly relevant to healthcare workforce dynamics and quality assurance in diagnostic imaging.
- Workforce
- Quality assurance
Research
ChatGPT-generated rehabilitation programs in sports physiotherapy: an expert evaluation and a mixed-methods study of clinical applicability
Adem Cali, Mehmet Erdem Yörükoğlu, Görkem Açar et al.
Frontiers in Medicine · 2026-08-26
This mixed-methods study had two experienced sports physiotherapists independently evaluate ChatGPT-4.1-generated six-week rehabilitation programs for five sports injury types, using structured rubrics and qualitative analysis. The AI-generated programs scored an overall mean of 3.85 out of 5, with good inter-rater reliability (ICC=0.84), performing best for straightforward, protocol-based injuries like clavicle fracture (5.00/5) and worst for complex postoperative cases like ACL reconstruction (1.88/5). Experts found the programs well-organized with sound progression logic but flagged weaknesses in multifactorial, timing-sensitive clinical scenarios. The authors conclude ChatGPT-4.1 can serve as a clinician-supervised support tool for linear recoveries but should not function as an autonomous decision-maker, particularly in complex rehabilitation contexts.
- Quality assurance
- Workforce
Research
From algorithmic efficiency to democratic legitimacy: rethinking AI governance in the Global South
Aleixandre Brian Duche-Pérez, Marco Tulio Falconi Picardo, Emmanuel Neptalí Augusto Chávez Urquizo et al.
Frontiers in Political Science · 2026-08-26
This article argues that AI in public administration should be evaluated not by technical efficiency or ethical compliance alone, but by 'democratic algorithmic legitimacy'—a framework centered on whether affected persons can publicly contest, influence, and correct AI-driven decisions. Developed specifically for Global South contexts marked by structural inequality and uneven institutional capacity, the framework identifies seven dimensions—transparency, participation, inclusion, accountability, contestability, correctability, and social justice—that distinguish democratic authority from computational performance. The paper warns that algorithmic efficiency can generate new forms of democratic exclusion when deployed in unequal sociotechnical environments lacking robust oversight institutions. It concludes that AI governance must preserve societies' authority to shape, question, and reject AI systems, not merely optimize their outputs.
- AI policy
Research
Meaningful Human Oversight of Artificial Intelligence in Clinical Practice: A Scoping Review Protocol
Kevin Pilger, MATHIAS COMIN, Henrique Moretto Pires et al.
arXiv · 2026-08-26
This scoping review protocol outlines a systematic effort to map how 'meaningful human oversight' of AI in clinical settings is defined, operationalised, and evaluated across the published and grey literature. The authors highlight a critical gap: regulation such as the EU AI Act and WHO guidance mandate human control over high-risk AI systems, yet the term is rarely specified precisely enough to distinguish genuine verification from perfunctory confirmation clicks. The review will address six research questions—covering definitions, mechanisms, responsible parties, outcome measures, empirical vs. normative grounding, and barriers—and aims to produce a typology of oversight arrangements and a research agenda for treating human oversight as a measurable safeguard rather than an assumed one. Findings will be directly relevant to how AI governance frameworks in healthcare are designed, enforced, and evaluated.
- AI policy
- Quality assurance
Research
Rare Diseases, Common Dilemmas: LLMs Prioritize Equal Resource Distribution over Patient Benefit in Decision-Making
Minda Zhao, Xu Han, Rishabh Goel et al.
arXiv (Cornell University) · 2026-08-25
This paper evaluates how 11 state-of-the-art large language models (LLMs) handle ethical decision-making in rare disease care using a benchmark of 208 clinically grounded vignettes that present genuine conflicts between bioethical principles. The researchers found that all tested models consistently prioritized justice—specifically equal resource allocation—over need-based or severity-driven considerations, showing limited responsiveness to clinical context. A strong 'authority-framing effect' was also identified: models shifted toward beneficence or patient autonomy only when decisions were framed as being made by clinicians or patients rather than committees. The findings suggest that institutional pressures around rare disease resource utilization may be silently embedded in LLM-based clinical decision support systems, with nuanced ethical considerations being overlooked.
- Quality assurance
- AI policy
Research
ADeptS-Bench: Measuring the Trustworthiness of Computer Use Agents Across Devices
Joy Chen, Alejandro Castillejo Munoz, Pierluca D'Oro et al.
arXiv · 2026-08-25
ADeptS-Bench introduces a dual-stream benchmark to evaluate the trustworthiness of Computer Use Agents (CUAs)—AI systems that navigate mobile and desktop apps on behalf of users. The Safety stream tests agents against malicious tasks embedded in visual interfaces, while the Disambiguation stream checks whether agents seek clarification on ambiguous instructions. Testing seven models reveals critical failures: no model keeps task success above 80% while holding attack success below 30%, every model blindly confirms a $25K checkout, and none catches a mislabeled 'factory reset' button. An ablation identifies three distinct safety architectures varying in tool dependence, and all models exhibit an over-refusal bias in ambiguous scenarios, raising significant concerns about the reliability of deployed CUAs.
- Quality assurance
- AI policy
Research
Can You Trust Frozen Hematology Foundation Models under Acquisition Shift?
Jai Kumar Sharma, Peeyush Tapadiya
arXiv · 2026-08-25
This paper audits 15 frozen foundation model encoders—covering hematology, pathology, and general vision—on their ability to generalize white blood cell classification across different scanners, sites, stains, and preparation pipelines. While in-domain linear-probe accuracy is near-perfect (macro-F1 of 0.98–0.997), cross-dataset macro-F1 drops by 34–72% and model rankings change substantially, exposing that strong in-domain performance does not predict reliable clinical deployment. Calibration also breaks down off-domain, with expected calibration error rising from 0.004 in-domain to 0.35 out-of-domain, meaning models are confidently wrong in shifted settings. The authors propose Class-Balanced Re-standardization (CBR) as a training-free mitigation and argue that hematology benchmarks must jointly evaluate accuracy, calibration, pretraining exposure, and class-prior robustness before clinical use.
- Quality assurance
- Certifications
Research
SimVerity: When Does Simulated Agent Success Survive Physical Deployment?
Zhonghao Zhan, Yefan Zhang, Krinos Li et al.
arXiv (Cornell University) · 2026-08-25
SimVerity is a framework that quantifies how well AI agent test results from simulations transfer to real physical deployments, specifically in smart home environments. The study found that even when a simulator reported all 240 light-control trials as successful, a physical camera caught 42 sub-second failures that the simulation missed entirely, demonstrating a critical gap between simulated and real-world performance. A risk model trained on measured trials could predict these failures on untested paths better than a baseline, and changing an agent's model configuration improved its scenario-matching rate from 52–88% to 100%. The work matters because it provides a structured method—clear, abstain, or escalate—for deciding when simulated test verdicts can be trusted before deploying AI agents in physical settings.
- Quality assurance
- Certifications
Research
The AI Adaptation Gap in Higher Education: Students, Faculty, and Administrative Staff
Yuriy S. Braun, Salavat M. Khafizov
arXiv · 2026-08-25
This study surveyed 2,121 university members—1,809 students, 250 faculty, and 62 administrative staff—at a large teacher-education institution to examine differences in AI use, attitudes, and institutional readiness. Results revealed a pronounced 'AI adaptation gap': students reported higher AI-use intensity and perceived usefulness than faculty and staff, while faculty and staff expressed stronger academic integrity concerns and endorsement of responsible-use norms. In a pooled OLS trust model, perceived usefulness was the strongest predictor of trust in AI (beta = 0.402), followed by institutional policy clarity (beta = 0.223). The findings highlight that higher education institutions face uneven AI adoption across roles, with implications for how universities develop policies, training, and integrity frameworks.
- AI policy
- Workforce
Research
ARISMA: Guidelines for AI- and LLM-Assisted Systematic Reviews, Scoping Reviews, and Mapping Studies
Mahyar Tourchi Moghaddam, Mina Alipour
arXiv (Cornell University) · 2026-08-25
This paper proposes ARISMA, a guideline framework for integrating AI and large language models into systematic reviews, scoping reviews, and mapping studies. The authors argue that while AI tools are rapidly entering review workflows—covering query formulation, screening, extraction, and reporting—the empirical evidence for their reliability is uneven, making unconstrained automation unjustifiable. ARISMA establishes that every consequential scientific decision must remain human-interpretable, human-auditable, and human-accountable, positioning AI as a validated and reversible assistant rather than an autonomous reviewer. The framework contributes a lifecycle taxonomy, process guidance, a governance and provenance model, a reporting checklist, and a validation matrix, addressing legal, privacy, and sustainability considerations alongside methodological ones.
- Quality assurance
- AI policy
Research
Reading Is Not Using: Retrieval, Judgment, and the Design of AI Financial Research Workflows
Miao Liu, Zhizhe Liu
arXiv · 2026-08-25
This paper identifies a 'retrieval-integration gap' in AI-assisted financial analysis: large language models can accurately retrieve information from financial disclosures like 10-K filings, yet that retrieved information fails to meaningfully influence their investment judgments as context length grows from 2,000 to 128,000 tokens. The authors show this pattern holds across multiple model families and tasks, and that more capable models delay but do not eliminate the problem. Crucially, workflow architecture matters: chunk-and-summarize pipelines lose relevant information, while structured restatements placed adjacent to the decision point restore its influence. The findings warn that retrieval-based evaluations can certify AI analyst systems whose actual investment judgments ignore information those systems demonstrably retrieved.
- Enterprise
- Quality assurance
Research
StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments
Esakkivel Esakkiraja, Denis Akhiyarov, Vikas Yadav et al.
arXiv · 2026-08-25
StarHarness is a framework that improves AI agent performance in enterprise environments by evolving the surrounding 'harness'—including prompt framing, tool interfaces, and agent structure—while leaving model weights unchanged. Tested across three enterprise benchmarks (ITBench SRE, EnterpriseOps-Gym ITSM, and AutomationBench Finance), the approach improves full-benchmark performance by 20–35 percentage points over default harnesses after just 4–12 accepted changes per environment. Gains generalize to tasks excluded from the evolution process and transfer across GPT and Qwen model families without re-evolution. The framework offers a practical path to reducing persistent model-environment mismatch in tool-rich enterprise settings.
- Enterprise
- Quality assurance
Research
Right Diagnoses, Decorative Reasoning:A Perturbation Audit of Medical Chain-of-Thought
Mengzhu Xu, Jifan Gao, Xia Jiang et al.
arXiv · 2026-08-25
This study audits whether chain-of-thought (CoT) rationales in medical large language models actually drive diagnostic answers or are merely decorative text. Using a 30-operator perturbation battery—including severity reversal, negation flips, demographic swaps, and evidence ablation—applied to 14 LLMs across four medical QA benchmarks, the authors find a panel-wide Chain-Decoupling Rate (CDR) of 72.9%: models' answers do not change even when the reasoning chain is meaningfully corrupted, and removing CoT prompting altogether does not reduce accuracy. Two board-certified clinicians validated 197 perturbed questions, confirming 98.5% preserved defensible gold-standard answers. The findings suggest that visible CoT rationales in medical LLMs function as documentation rather than genuine reasoning, raising serious concerns for clinical trust and deployment of AI diagnostic tools.
- Quality assurance
- Certifications
Research
StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing
Zhijie Zheng, Yu Li, Chen Qian et al.
arXiv · 2026-08-25
StepGuard is a step-level guardrail model designed to monitor LLM-based agent actions before they are executed, addressing security risks such as file modification, information leakage, and unauthorized actions that arise when agents interact with external environments through tool invocation. The authors introduce StepGen, an automated data engine that generates paired safe and unsafe trajectories, and Balance-GRPO, a training method that dynamically balances learning between safe and unsafe actions to reduce over-defense and under-defense. In experiments on AgentDojo and AgentDyn benchmarks, StepGuard reduces the mean attack success rate by 77.3% relative to a no-guard setting while dropping mean utility by only 2.8 percentage points, and achieves accuracy comparable to GPT-4 among open-weight guard models. This work matters for enterprise AI deployments where agentic systems must be kept secure without significantly degrading their usefulness.
- Enterprise
- Quality assurance
Research
Confident at the moment of action: belief miscalibration in LLM play under hidden information
Bhushan Kashinath Joshi
arXiv · 2026-08-25
This paper examines whether large language models' stated confidence scores reliably track correctness at the moment they take actions — a critical assumption in agentic AI systems that gate behavior on self-reported confidence. Using a hidden-information chess variant where a royal piece can be secretly relocated, the authors elicit probability estimates from LLMs each turn and score them against ground truth. They find that captures made at high stated confidence (≥0.5) were correct in only 1 of 62 cases, with nearly all of the calibration deficit concentrated in these high-confidence events — a pattern replicated across multiple model configurations and a second provider. Crucially, conventional evaluation metrics like legality, cost, latency, and completion rate can dissociate entirely from belief quality, meaning outcome-only evaluation would not detect this miscalibration, posing risks for deployed agentic systems that rely on model confidence to decide when to act.
- Quality assurance
- Enterprise
News
AI won’t replace radiologists, but it will dramatically change their jobs
arstechnica.com · 2026-08-25
Ars Technica reports that radiology has become the dominant arena for medical AI adoption, with roughly three-quarters of the approximately 1,400 AI-enabled medical devices cleared by the FDA as of early 2026 falling into that specialty. Despite Geoffrey Hinton's 2016 prediction that radiologists would be replaced by AI within five years, the field's workforce is actually projected to grow by 26 percent or more over the next three decades. However, AI tools are genuinely matching or exceeding human performance in some tasks, such as detecting polyps in colonoscopies, where a review of 43 clinical trials found AI-assisted procedures outperform conventional ones. Radiology's experience is being watched as a bellwether for broader AI adoption in expert decision-making across healthcare and other fields.
- Quality assurance
- Workforce