News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
Workforce Readiness in the AI Era: An Empirical Qualitative Study of Higher Education and Upskilling Among Urban Muslim Communities in Indonesia
Reza Fahmi, Prima Aswirna, Adamu Abubakar Muhammad
Al-Madinah Journal of Islamic Civilization · 2026-08-26
This qualitative study examines workforce readiness gaps in Indonesia's higher education system amid AI-driven labor market transformation, focusing on urban Muslim communities. Through interviews, focus groups, and document analysis with university staff, students, HR practitioners, vocational experts, and government officials, the study identifies five key challenges: outdated curricula, limited innovative pedagogy, weak lifelong learning culture, insufficient industry involvement, and slow policy adaptation. The researchers propose an integrated upskilling model combining industry-aligned curricula, project-based learning, micro-credentials, professional certification, and university–industry collaboration. The findings underscore the need for coordinated action among higher education, industry, and government to build a sustainable lifelong learning ecosystem responsive to AI-driven change.
- Workforce
- Certifications
Research
Driverless, not thoughtless? Automated vehicle regulation, the dilemma of control, and the prospect of dual corrigibility
Bård Torvetjønn Haugland
Technological Forecasting and Social Change · 2026-08-26
This article examines Norway's regulatory approach to automated vehicles, analyzing the policy-making process behind the Act Relating to Testing of Self-Driving Vehicles, the Act's content, and interview data from the first post-Act trial. The study finds that Norway employs a two-stage strategy—first allowing trials, then establishing permanent regulation—but argues this strategy rests on a flawed assumption that practical and normative concerns can be addressed separately. Drawing on Collingridge's dilemma of control, the authors propose 'dual corrigibility' (correcting both decisions and normative orientations in parallel) as a better guiding principle for regulating emerging technologies like automated vehicles.
- AI policy
Research
Artificial Intelligence in Investment Services
Filippo Annunziata
arXiv · 2026-08-26
This article analyzes how AI technologies—including robo-advice, algorithmic trading, and automated portfolio management—are being integrated into investment services across the EU, and assesses whether the existing regulatory framework (primarily MiFID II, alongside the AI Act, GDPR, and DORA) is adequate to manage the resulting risks. The study finds that while MiFID II's core duties of care, loyalty, transparency, and governance remain conceptually applicable, they require reinterpretation to address AI-specific challenges such as algorithmic bias, opacity, model oversight, and systemic concentration risks. The author argues that the fragmented evolution of EU digital and financial regulations leaves critical gaps, calling for clearer supervisory standards, harmonized model-governance practices, and a human-centered approach to preserve market integrity and consumer trust.
- AI policy
- Enterprise
Research
Reinventing the echocardiography workflow: from manual quantification to artificial intelligence–driven comprehensive interpretation
Ran Heo, Seung‐Ah Lee, Hyuk‐Jae Chang
Journal of Cardiovascular Imaging · 2026-08-26
This review evaluates the current state of AI integration in echocardiography workflows, showing that AI can reduce examination time, automate measurements, and mitigate sonographer fatigue while improving image quality. AI applications now extend beyond ejection fraction to include myocardial texture, Doppler hemodynamics, and assessments of valvular heart disease, cardiomyopathy, and pericardial disorders. However, the authors identify key gaps including reliance on single-center studies, inconsistent cross-platform performance, and risks of automation bias, and outline requirements for responsible clinical implementation. The findings are directly relevant to healthcare workforce dynamics and quality assurance in diagnostic imaging.
- Workforce
- Quality assurance
Research
ChatGPT-generated rehabilitation programs in sports physiotherapy: an expert evaluation and a mixed-methods study of clinical applicability
Adem Cali, Mehmet Erdem Yörükoğlu, Görkem Açar et al.
Frontiers in Medicine · 2026-08-26
This mixed-methods study had two experienced sports physiotherapists independently evaluate ChatGPT-4.1-generated six-week rehabilitation programs for five sports injury types, using structured rubrics and qualitative analysis. The AI-generated programs scored an overall mean of 3.85 out of 5, with good inter-rater reliability (ICC=0.84), performing best for straightforward, protocol-based injuries like clavicle fracture (5.00/5) and worst for complex postoperative cases like ACL reconstruction (1.88/5). Experts found the programs well-organized with sound progression logic but flagged weaknesses in multifactorial, timing-sensitive clinical scenarios. The authors conclude ChatGPT-4.1 can serve as a clinician-supervised support tool for linear recoveries but should not function as an autonomous decision-maker, particularly in complex rehabilitation contexts.
- Quality assurance
- Workforce
Research
From algorithmic efficiency to democratic legitimacy: rethinking AI governance in the Global South
Aleixandre Brian Duche-Pérez, Marco Tulio Falconi Picardo, Emmanuel Neptalí Augusto Chávez Urquizo et al.
Frontiers in Political Science · 2026-08-26
This article argues that AI in public administration should be evaluated not by technical efficiency or ethical compliance alone, but by 'democratic algorithmic legitimacy'—a framework centered on whether affected persons can publicly contest, influence, and correct AI-driven decisions. Developed specifically for Global South contexts marked by structural inequality and uneven institutional capacity, the framework identifies seven dimensions—transparency, participation, inclusion, accountability, contestability, correctability, and social justice—that distinguish democratic authority from computational performance. The paper warns that algorithmic efficiency can generate new forms of democratic exclusion when deployed in unequal sociotechnical environments lacking robust oversight institutions. It concludes that AI governance must preserve societies' authority to shape, question, and reject AI systems, not merely optimize their outputs.
- AI policy
Research
Meaningful Human Oversight of Artificial Intelligence in Clinical Practice: A Scoping Review Protocol
Kevin Pilger, MATHIAS COMIN, Henrique Moretto Pires et al.
arXiv · 2026-08-26
This scoping review protocol outlines a systematic effort to map how 'meaningful human oversight' of AI in clinical settings is defined, operationalised, and evaluated across the published and grey literature. The authors highlight a critical gap: regulation such as the EU AI Act and WHO guidance mandate human control over high-risk AI systems, yet the term is rarely specified precisely enough to distinguish genuine verification from perfunctory confirmation clicks. The review will address six research questions—covering definitions, mechanisms, responsible parties, outcome measures, empirical vs. normative grounding, and barriers—and aims to produce a typology of oversight arrangements and a research agenda for treating human oversight as a measurable safeguard rather than an assumed one. Findings will be directly relevant to how AI governance frameworks in healthcare are designed, enforced, and evaluated.
- AI policy
- Quality assurance
Research
Rare Diseases, Common Dilemmas: LLMs Prioritize Equal Resource Distribution over Patient Benefit in Decision-Making
Minda Zhao, Xu Han, Rishabh Goel et al.
arXiv (Cornell University) · 2026-08-25
This paper evaluates how 11 state-of-the-art large language models (LLMs) handle ethical decision-making in rare disease care using a benchmark of 208 clinically grounded vignettes that present genuine conflicts between bioethical principles. The researchers found that all tested models consistently prioritized justice—specifically equal resource allocation—over need-based or severity-driven considerations, showing limited responsiveness to clinical context. A strong 'authority-framing effect' was also identified: models shifted toward beneficence or patient autonomy only when decisions were framed as being made by clinicians or patients rather than committees. The findings suggest that institutional pressures around rare disease resource utilization may be silently embedded in LLM-based clinical decision support systems, with nuanced ethical considerations being overlooked.
- Quality assurance
- AI policy
Research
ADeptS-Bench: Measuring the Trustworthiness of Computer Use Agents Across Devices
Joy Chen, Alejandro Castillejo Munoz, Pierluca D'Oro et al.
arXiv · 2026-08-25
ADeptS-Bench introduces a dual-stream benchmark to evaluate the trustworthiness of Computer Use Agents (CUAs)—AI systems that navigate mobile and desktop apps on behalf of users. The Safety stream tests agents against malicious tasks embedded in visual interfaces, while the Disambiguation stream checks whether agents seek clarification on ambiguous instructions. Testing seven models reveals critical failures: no model keeps task success above 80% while holding attack success below 30%, every model blindly confirms a $25K checkout, and none catches a mislabeled 'factory reset' button. An ablation identifies three distinct safety architectures varying in tool dependence, and all models exhibit an over-refusal bias in ambiguous scenarios, raising significant concerns about the reliability of deployed CUAs.
- Quality assurance
- AI policy
Research
Can You Trust Frozen Hematology Foundation Models under Acquisition Shift?
Jai Kumar Sharma, Peeyush Tapadiya
arXiv · 2026-08-25
This paper audits 15 frozen foundation model encoders—covering hematology, pathology, and general vision—on their ability to generalize white blood cell classification across different scanners, sites, stains, and preparation pipelines. While in-domain linear-probe accuracy is near-perfect (macro-F1 of 0.98–0.997), cross-dataset macro-F1 drops by 34–72% and model rankings change substantially, exposing that strong in-domain performance does not predict reliable clinical deployment. Calibration also breaks down off-domain, with expected calibration error rising from 0.004 in-domain to 0.35 out-of-domain, meaning models are confidently wrong in shifted settings. The authors propose Class-Balanced Re-standardization (CBR) as a training-free mitigation and argue that hematology benchmarks must jointly evaluate accuracy, calibration, pretraining exposure, and class-prior robustness before clinical use.
- Quality assurance
- Certifications
Research
SimVerity: When Does Simulated Agent Success Survive Physical Deployment?
Zhonghao Zhan, Yefan Zhang, Krinos Li et al.
arXiv (Cornell University) · 2026-08-25
SimVerity is a framework that quantifies how well AI agent test results from simulations transfer to real physical deployments, specifically in smart home environments. The study found that even when a simulator reported all 240 light-control trials as successful, a physical camera caught 42 sub-second failures that the simulation missed entirely, demonstrating a critical gap between simulated and real-world performance. A risk model trained on measured trials could predict these failures on untested paths better than a baseline, and changing an agent's model configuration improved its scenario-matching rate from 52–88% to 100%. The work matters because it provides a structured method—clear, abstain, or escalate—for deciding when simulated test verdicts can be trusted before deploying AI agents in physical settings.
- Quality assurance
- Certifications
Research
The AI Adaptation Gap in Higher Education: Students, Faculty, and Administrative Staff
Yuriy S. Braun, Salavat M. Khafizov
arXiv · 2026-08-25
This study surveyed 2,121 university members—1,809 students, 250 faculty, and 62 administrative staff—at a large teacher-education institution to examine differences in AI use, attitudes, and institutional readiness. Results revealed a pronounced 'AI adaptation gap': students reported higher AI-use intensity and perceived usefulness than faculty and staff, while faculty and staff expressed stronger academic integrity concerns and endorsement of responsible-use norms. In a pooled OLS trust model, perceived usefulness was the strongest predictor of trust in AI (beta = 0.402), followed by institutional policy clarity (beta = 0.223). The findings highlight that higher education institutions face uneven AI adoption across roles, with implications for how universities develop policies, training, and integrity frameworks.
- AI policy
- Workforce
Research
ARISMA: Guidelines for AI- and LLM-Assisted Systematic Reviews, Scoping Reviews, and Mapping Studies
Mahyar Tourchi Moghaddam, Mina Alipour
arXiv (Cornell University) · 2026-08-25
This paper proposes ARISMA, a guideline framework for integrating AI and large language models into systematic reviews, scoping reviews, and mapping studies. The authors argue that while AI tools are rapidly entering review workflows—covering query formulation, screening, extraction, and reporting—the empirical evidence for their reliability is uneven, making unconstrained automation unjustifiable. ARISMA establishes that every consequential scientific decision must remain human-interpretable, human-auditable, and human-accountable, positioning AI as a validated and reversible assistant rather than an autonomous reviewer. The framework contributes a lifecycle taxonomy, process guidance, a governance and provenance model, a reporting checklist, and a validation matrix, addressing legal, privacy, and sustainability considerations alongside methodological ones.
- Quality assurance
- AI policy
Research
Reading Is Not Using: Retrieval, Judgment, and the Design of AI Financial Research Workflows
Miao Liu, Zhizhe Liu
arXiv · 2026-08-25
This paper identifies a 'retrieval-integration gap' in AI-assisted financial analysis: large language models can accurately retrieve information from financial disclosures like 10-K filings, yet that retrieved information fails to meaningfully influence their investment judgments as context length grows from 2,000 to 128,000 tokens. The authors show this pattern holds across multiple model families and tasks, and that more capable models delay but do not eliminate the problem. Crucially, workflow architecture matters: chunk-and-summarize pipelines lose relevant information, while structured restatements placed adjacent to the decision point restore its influence. The findings warn that retrieval-based evaluations can certify AI analyst systems whose actual investment judgments ignore information those systems demonstrably retrieved.
- Enterprise
- Quality assurance
Research
StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments
Esakkivel Esakkiraja, Denis Akhiyarov, Vikas Yadav et al.
arXiv · 2026-08-25
StarHarness is a framework that improves AI agent performance in enterprise environments by evolving the surrounding 'harness'—including prompt framing, tool interfaces, and agent structure—while leaving model weights unchanged. Tested across three enterprise benchmarks (ITBench SRE, EnterpriseOps-Gym ITSM, and AutomationBench Finance), the approach improves full-benchmark performance by 20–35 percentage points over default harnesses after just 4–12 accepted changes per environment. Gains generalize to tasks excluded from the evolution process and transfer across GPT and Qwen model families without re-evolution. The framework offers a practical path to reducing persistent model-environment mismatch in tool-rich enterprise settings.
- Enterprise
- Quality assurance
Research
Right Diagnoses, Decorative Reasoning:A Perturbation Audit of Medical Chain-of-Thought
Mengzhu Xu, Jifan Gao, Xia Jiang et al.
arXiv · 2026-08-25
This study audits whether chain-of-thought (CoT) rationales in medical large language models actually drive diagnostic answers or are merely decorative text. Using a 30-operator perturbation battery—including severity reversal, negation flips, demographic swaps, and evidence ablation—applied to 14 LLMs across four medical QA benchmarks, the authors find a panel-wide Chain-Decoupling Rate (CDR) of 72.9%: models' answers do not change even when the reasoning chain is meaningfully corrupted, and removing CoT prompting altogether does not reduce accuracy. Two board-certified clinicians validated 197 perturbed questions, confirming 98.5% preserved defensible gold-standard answers. The findings suggest that visible CoT rationales in medical LLMs function as documentation rather than genuine reasoning, raising serious concerns for clinical trust and deployment of AI diagnostic tools.
- Quality assurance
- Certifications
Research
StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing
Zhijie Zheng, Yu Li, Chen Qian et al.
arXiv · 2026-08-25
StepGuard is a step-level guardrail model designed to monitor LLM-based agent actions before they are executed, addressing security risks such as file modification, information leakage, and unauthorized actions that arise when agents interact with external environments through tool invocation. The authors introduce StepGen, an automated data engine that generates paired safe and unsafe trajectories, and Balance-GRPO, a training method that dynamically balances learning between safe and unsafe actions to reduce over-defense and under-defense. In experiments on AgentDojo and AgentDyn benchmarks, StepGuard reduces the mean attack success rate by 77.3% relative to a no-guard setting while dropping mean utility by only 2.8 percentage points, and achieves accuracy comparable to GPT-4 among open-weight guard models. This work matters for enterprise AI deployments where agentic systems must be kept secure without significantly degrading their usefulness.
- Enterprise
- Quality assurance
Research
Confident at the moment of action: belief miscalibration in LLM play under hidden information
Bhushan Kashinath Joshi
arXiv · 2026-08-25
This paper examines whether large language models' stated confidence scores reliably track correctness at the moment they take actions — a critical assumption in agentic AI systems that gate behavior on self-reported confidence. Using a hidden-information chess variant where a royal piece can be secretly relocated, the authors elicit probability estimates from LLMs each turn and score them against ground truth. They find that captures made at high stated confidence (≥0.5) were correct in only 1 of 62 cases, with nearly all of the calibration deficit concentrated in these high-confidence events — a pattern replicated across multiple model configurations and a second provider. Crucially, conventional evaluation metrics like legality, cost, latency, and completion rate can dissociate entirely from belief quality, meaning outcome-only evaluation would not detect this miscalibration, posing risks for deployed agentic systems that rely on model confidence to decide when to act.
- Quality assurance
- Enterprise
Research
The Invisible Editorial Layer: Formalizing Undisclosed Inference-Time Steering, Probability Placement, and the Attribution Problem in Deployed Language Models
Augusto Camargo
arXiv (Cornell University) · 2026-08-25
This paper argues that standard evaluations of large language models overlook a critical gap: modern deployment pipelines can silently modify a model's output probability distributions before token selection, effectively adding an 'invisible editorial layer' between frozen model weights and observed text. The authors formalize three concepts—the Inference Attribution Problem (behavioral bias cannot be attributed to model weights alone), Probability Placement (a hypothetical advertising primitive using probability shifts rather than explicit insertions), and Inference Policy Transparency (a governance principle for auditing deployment-layer interventions). The paper situates these concerns within existing regulatory frameworks including Article 5 of the EU AI Act, the EU Digital Services Act, and FTC doctrines, arguing that undisclosed inference-time steering poses unaddressed governance, security, and economic risks. The work matters because it identifies a structural accountability gap in how deployed AI systems are regulated and attributed, with direct implications for policy and enterprise transparency.
- AI policy
- Enterprise
Research
Beyond Semantic Accuracy: Consequence-Aware Evaluation for Safety-Critical Language Understanding
Yujing Chang, Thinh Pham, Van-Phat Thai et al.
arXiv · 2026-08-25
This paper investigates whether language models can be reliably used in safety-critical settings, specifically air traffic control (ATC), where errors in understanding communication—such as a misread altitude or confused callsign—can have severe consequences. The authors develop a consequence-aware evaluation framework grounded in aviation standards and validated with feedback from 40 air traffic controllers across three countries, then apply it to 8 models. They find a systematic 'semantic-safety gap': conventional NLP metrics like F1 substantially overestimate operational reliability compared to consequence-aware evaluation, even for models that appear strong under standard benchmarks. Risk-aware fine-tuning narrows but does not close this gap, demonstrating that consequence-aware evaluation is a necessary complement to standard metrics before any safety-critical deployment.
- Quality assurance
- Certifications
Research
$\texttt{findr}$: Transparent and Fair Credit Risk Decisions through Semi-Structured Regressions
Victor Medina-Olivares, Stefan Lessmann, Jonathan Crook
arXiv · 2026-08-25
This paper introduces findr (flexible, interpretable deep regression), a semi-structured framework for binary credit risk modelling that splits the logit into an interpretable structured component and an orthogonal neural residual. The orthogonalisation keeps coefficient-based effects separate from nonlinear variation, while an in-processing Wasserstein penalty reduces group score disparities during training. Evaluated in simulations and on eight public credit datasets, findr approaches logistic regression performance when signals are linear and recovers much of the predictive gain of neural models when nonlinearity matters. Built-in diagnostics make accuracy, fairness, and interpretability trade-offs explicit and auditable, supporting practical deployment in credit risk decisions.
- Enterprise
- AI policy
- Quality assurance
Research
When "Must" Becomes "Maybe": Constraint Weakening in LLM Agent Workflows
Yiheng Sun, Huifei Wang, Yancheng Zhu et al.
arXiv · 2026-08-25
This paper investigates a failure mode in multi-agent LLM workflows where critical constraints embedded in 'safety blockers'—conditions that must be resolved before an action can proceed—are systematically weakened as state is passed between workflow stages. Across 1,296 controlled synthetic episodes, the authors show that common handoff transformations such as compression, plan assimilation, and precedent substitution convert binding requirements into mere caveats, with normal handoff compression producing 100% blocker deactivation and 54.2% forbidden action rates. Restoring all four structured state fields (prerequisite, authority, fallback, and execution consequence) brings preservation to 100% and eliminates forbidden actions entirely, demonstrating that the problem is structural rather than informational. The finding—that semantic availability does not guarantee operational preservation—has direct implications for deploying AI agents in enterprise and safety-critical workflows.
- Enterprise
- Quality assurance
Research
Benchmarking LLM Judges for Voice-Agent Evaluation: Reliability, Calibration, and Human Oversight
Anupam Purwar, Shashank Singh, Kritika Srivastava
arXiv · 2026-08-25
This paper benchmarks LLM-based evaluation (GPT-4.1 and GPT-5) against human judgments for assessing conversational voice agents in telecom and retail settings, examining reliability, calibration, and agreement across multiple evaluation configurations and quality/safety dimensions. The study finds that LLM judges can effectively support large-scale voice-agent assessment, but their reliability varies by metric and configuration rather than being uniformly dependable. The authors identify which conversational attributes can be reliably automated and which require human oversight, proposing a hybrid evaluation pipeline where LLMs handle scalable scoring while human evaluators focus on metrics demanding contextual interpretation.
- Quality assurance
- Enterprise
Research
Who is the Agent to Blame? Localizing Faithfulness and Citation Mistakes in Agentic Deep Research
Eran Hirsch, David Wan, Han Wang et al.
arXiv · 2026-08-25
This paper examines 'deep research' (DR) systems—multi-agent AI pipelines that generate long-form reports with citations from web sources—and proposes an evaluation method to pinpoint which specific agent within the pipeline introduced faithfulness or citation errors. The authors develop a four-type error taxonomy (hallucination, uncited input reliance, uncited output, insufficient citations) and apply it to three top-ranked open-source DR systems, finding that nearly every agent makes substantial mistakes except those summarizing a single document. A key finding is that 84.7% of final-report errors in one system (AI-Q) originate at the orchestrator agent, with roughly 31% being hallucinations and the rest citation mistakes. Two simple interventions guided by these diagnostics raise citation recall by 5% without degrading output quality, demonstrating the practical value of localized error attribution.
- Quality assurance
Research
TRACE: An Evidence-Grounded Benchmark for Safety Evaluation of Large Reasoning Models
Zhenyu Wu, Siyuan Chen, Changchun Yang et al.
arXiv · 2026-08-25
TRACE is a new benchmark designed to evaluate how well guardrail models detect unsafe content across the full inference pipeline of Large Reasoning Models (LRMs), including prompts, intermediate reasoning traces, and final responses. The benchmark covers two languages, nine risk categories, and ten attack strategies, with annotations that include not just binary safety labels but also evidence extracted from source text to justify each judgment. Testing 18 guardrail models on TRACE reveals that safety evaluation of reasoning traces is significantly harder than evaluating prompts or final responses, and that current models struggle to accurately localize supporting evidence for their judgments. These findings underscore a critical gap in existing safety tooling and motivate the development of guardrail models capable of reliably detecting unsafe content throughout the entire LRM inference process.
- Quality assurance
- Certifications