News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
5526 items
Research
<b>The Influence of Artificial Intelligence Adoption on Financial Reporting Accuracy among Selected Firms in Livingstone, Zambia</b>
Chibulo Foster Mwachikoka, Muhammad Adil, Obrine Mweetwa et al.
African Journal of Commercial Studies · 2026-08-13
This study surveyed 60 accounting and finance professionals at firms in Livingstone, Zambia, to assess how AI adoption affects financial reporting accuracy. Using correlation and regression analysis, the researchers found a strong positive relationship between AI adoption and reporting accuracy (r = 0.700), with AI adoption explaining about 49% of variation in accuracy, and Natural Language Processing emerging as the most significant individual predictor. The findings suggest that investing in AI technologies, professional training, and governance frameworks can meaningfully improve the reliability of financial information in emerging-market contexts.
- Enterprise
- Quality assurance
Research
Artificial intelligence–based employee turnover forecasting as a decision support tool for HR management
Vadym Taraniuk, Irina Vinogradova-Zinkevič
Business and management · 2026-08-13
This study develops machine learning models to predict employee turnover and frames them as decision-support tools for HR managers. The approach emphasizes handling class imbalance (where resignations are rare), balancing false positives against minority-class detection, and using explainable AI (XAI) methods so managers can understand which factors drive turnover risk. Validated through Repeated Stratified K-Fold Cross Validation, the models aim to deliver transparent, actionable retention insights that reduce turnover-related costs and support evidence-based HR strategy.
- Workforce
- Enterprise
Research
Artificial Intelligence in Breast Imaging and Screening: Current Evidence, Applications, and Future Directions
Amrita Kumar, Gerald Lip
Indian journal of radiology and imaging - new series/Indian journal of radiology and imaging/Indian Journal of Radiology & Imaging · 2026-08-13
This narrative review synthesizes evidence from 2023–2026 on AI applications in breast cancer screening and imaging, finding that AI-supported mammography can increase cancer detection rates by 10–29%, reduce interval cancer rates by up to 12%, and cut radiologist workload by 31–64% in double-reading settings without increasing false-positive rates. The evidence base has matured from retrospective studies to prospective randomized controlled trials, though it remains concentrated in high-income, non-diverse populations. Long-term outcome data on breast cancer mortality are still lacking, and challenges around algorithmic bias, generalizability, overdiagnosis, and regulatory frameworks must be addressed before broader clinical implementation. The review emphasizes the need for post-market surveillance, diverse dataset validation, and cost-effectiveness analyses.
- Quality assurance
- Workforce
Research
From Principles to Practice: Engineering Responsible AI for Geospatial Intelligence
Muhammad Hassan, Bilal Sardar, Shareeful Islam et al.
SN Computer Science · 2026-08-13
This paper proposes a methodology for embedding responsible AI (R-AI) principles—Privacy, Fairness, Transparency, and Explainability—into geospatial AI model development. Tested on crop type classification and urban flood risk assessment tasks, the approach satisfies all twelve defined acceptance criteria while retaining over 97% of baseline predictive accuracy. Key results include reducing membership inference attack accuracy from 0.712 to 0.503, narrowing geographic performance gaps from 23.1% to 3.1%, and improving temporal attribution consistency from 0.41 to 0.857. The findings demonstrate that responsible AI and predictive accuracy are compatible in geospatial deep learning, offering a replicable and measurable pathway for practitioners.
- Quality assurance
- AI policy
Research
Artificial Intelligence-Enabled Digital Project Management in Construction: A Critical Review of Applications, Project Performance Outcomes and Adoption Barriers
Afeez A. Salawudeen, Riliwan A. Adebayo, Olorunshogo B. Ogundipe
International journal of latest research in engineering and technology. · 2026-08-13
This critical review examines AI applications in construction project management—covering cost estimation, scheduling, and safety—and finds a stark disconnect between reported model accuracy and real-world value. While predictive accuracy metrics are high (some cost models claiming R² above 0.99), the paper shows these rest on fragile foundations: validation is mostly retrospective, samples are small and geographically narrow, and hold-out testing reveals severe degradation (e.g., one delay model dropping from 74.5% training to 47.2% testing accuracy). Crucially, organizational adoption barriers are heavily under-researched, representing only 4.6% of the literature despite being cited by practitioners as the dominant challenge, and the authors propose an integrated framework linking technical capability to deployment maturity and organizational absorptive capacity.
- Enterprise
- Workforce
Research
TabNet Interpretable Deep Structure for Adaptive Recognition of Audit Voucher Anomalies
L. X. Sun
Advanced Electromagnetics · 2026-08-13
This paper proposes a TabNet-based deep learning framework for detecting anomalies in financial audit vouchers, addressing the need for both interpretability and adaptability in automated auditing systems. Sequential attention layers generate traceable feature importance masks, while a sliding time-window mechanism and dynamic threshold calibration handle concept drift over time. Experimental results show 92.3% accuracy on amount logic conflicts, 93.1% on supplier-related anomalies, F1-scores above 87% across all categories, and a 57.1% reduction in supplier anomaly verification time (from 19.6 to 8.4 minutes). The framework is relevant to enterprise financial auditing and quality assurance, offering interpretable, real-time anomaly monitoring at scale.
- Enterprise
- Quality assurance
Research
Narratives generated by artificial intelligence: an ethical reflection on screenwriting in contemporary cinema
Montserrat Jurado-Martín, Carmen M. Lopez-Rico, María Samper Cerdán
Frontiers in Communication · 2026-08-13
This paper examines the ethical implications of AI-generated screenwriting in cinema through case studies of three AI-assisted films—Sunspring (2016), The Diary of Sisyphus (2023), and The Last Screenwriter (2024)—alongside a review of academic literature and international regulatory frameworks. Using a qualitative approach, the authors find that while AI can optimize parts of the creative process, it introduces significant risks including diluted authorial responsibility, reproduction of stereotypes, and technological dependence in cultural industries. The paper argues that film-specific ethical governance is needed, grounded in human oversight, authorship recognition, and accountability to protect artistic integrity and creators' rights. These findings are directly relevant to professional screenwriters whose livelihoods and creative recognition are affected by generative AI adoption.
- Workforce
- AI policy
Research
Large language models (LLMs) as psychotherapists: an analysis based on psychodynamic psychotherapy theory
Mateusz Łabuz, Paweł Szczęsny, Katarzyna Mika-Łabuz
Ethics and Information Technology · 2026-08-13
This paper critically examines whether large language models (LLMs) like ChatGPT are suitable for use as psychotherapists, evaluating them against the principles of psychodynamic psychotherapy. The authors conclude that current LLMs lack awareness, genuine emotionality, and the capacity to mentalize, meaning their apparent empathy is purely linguistic simulation rather than authentic therapeutic engagement. The paper identifies serious ethical, legal, and psychological risks—including hallucinations, data bias, and the illusion of a therapeutic relationship—while acknowledging limited potential for LLMs as support tools in areas like psychological education or stress reduction. The authors call for interdisciplinary standards combining psychology, law, and ethics to govern LLM use in mental health contexts.
- AI policy
- Quality assurance
Research
Adaptive technology-change confidence among Saudi university faculty under Vision 2030: a three-wave longitudinal mixed-methods study
Tahani H. Alqahtani
Frontiers in Education · 2026-08-13
This 27-month longitudinal mixed-methods study of Saudi university faculty (n=431 across 11 universities) tracks a newly developed construct called adaptive technology-change confidence (ATCC) and finds that it declines within individuals over time, with three distinct trajectory classes: stable, gradual decline, and steep decline. Faculty at universities with faster digital infrastructure turnover—including AI tool adoptions, LMS updates, and platform deprecations—showed steeper ATCC declines, and lower ATCC was associated with subsequent lower academic performance ratings by department chairs. The authors suggest that faculty development centers could use periodic ATCC monitoring as an early-warning indicator to target support toward those on steep-decline trajectories, though findings are preliminary given the author-developed scale and small number of universities studied.
- Workforce
Research
Assessment on the Availability of Policies and Guidelines for Managing AI-Assisted Plagiarism among Students in Higher Learning Institutions in Arusha,Tanzania
Journal of Research Innovation and Implications in Education · 2026-08-13
This study examined whether higher education institutions in Arusha, Tanzania have formal policies and guidelines for managing AI-assisted plagiarism. Surveying 65 academic and administrative leaders across three institutions, the researchers found that formal AI-specific policy frameworks are largely absent, with institutions instead relying on existing practices such as plagiarism detection tools, awareness campaigns, training, and assessment redesign. The study concludes that comprehensive AI policies are needed to complement these ad hoc practices and strengthen academic integrity.
- AI policy
Research
Platform-Facilitated Grooming and AI Chatbots: Rethinking Criminal Liability and Regulation
Mohamed Chawki
Laws · 2026-08-13
This legal-comparative study examines how AI chatbots are enabling or automating online child grooming and finds that existing criminal law frameworks in the EU, UK, US, and China are inadequate to address these scenarios. The paper identifies critical gaps around criminal intent, foreseeability, and fragmented liability among offenders, platforms, and AI developers. It advocates for a risk-based liability framework, enhanced platform accountability, algorithmic transparency, and stronger child-centered safeguards to close these regulatory gaps.
- AI policy
Research
A comparative study of AI readiness in language teacher education in the Global South
Syed Naeem Ahmed, Heena Saifullah Amjad, Saira Abbas et al.
Discover Education · 2026-08-13
This mixed-methods study examines AI readiness among language educators in Pakistan, Uzbekistan, and Saudi Arabia using Holmström's AI Readiness Framework, finding that Saudi Arabia scored highest (M=4.10), followed by Uzbekistan (M=3.10) and Pakistan (M=2.92). Survey data from 400 professionals showed readiness correlated with infrastructure (r=0.673) and that faculty training explained 39.7% of variance in readiness. Focus groups and a task-based assessment of 100 participants identified ethical concerns, policy gaps, and limited training as persistent barriers. The authors recommend modular training programs, institutional readiness benchmarks, and localized policy reform to support AI integration in language teacher education across the Global South.
- Workforce
- AI policy
Research
Neither Luddite nor enthusiast: interpreting teachers’ AI use in teaching
Nina Y. Y. Cheung, Andrew K. F. Cheung
Frontiers in Education · 2026-08-13
This survey of 156 Chinese interpreting teachers in Master of Translation and Interpreting programs examines how educators position themselves toward AI in interpreter training, finding that while readiness and perceived usefulness are above midpoint, actual tool use is moderate with high variability. Teachers with greater AI evaluative literacy show slightly less tool use and perceived usefulness, and more teaching experience is strongly linked to more restrictive orientations toward AI integration. Profession-related concerns—particularly threats to professional autonomy and perceived labor devaluation—are prominent, though they do not uniquely predict restrictiveness once experience is accounted for. The study argues that AI integration strategies in interpreter education must address not only pedagogical utility but also professional identity and labor implications.
- Workforce
- AI policy
Research
Facts label for transparent communication of AI Risks in mental health technology
Khatiya Moon, Matthew Tamura, J. R. Redmond et al.
Frontiers in Psychiatry · 2026-08-13
This paper proposes a standardized 'facts label' framework for AI-enabled digital mental health technologies (AI-DMHTs) to improve transparency about their risks for clinicians, patients, and users. Developed by a multidisciplinary team from the American Psychiatric Association, the framework comprises 8 sections covering intended use, warnings, risks and limitations, model information, clinical evidence, and privacy and security. The goal is to give clinicians practical guidance for evaluating AI mental health products, a gap currently underserved by existing governance frameworks. The work aims to serve as a foundation for broader AI risk communication standards in healthcare.
- AI policy
- Certifications
Research
Occupational injuries and health risks among food delivery riders in China: research progress and prevention strategies
Xueshun Xu, Xinran Ru, Ruxuan He et al.
Frontiers in Public Health · 2026-08-13
This narrative review synthesizes evidence from 47 sources on occupational injuries and health risks facing food delivery riders in China, finding that algorithmic management, strict deadlines, piece-rate pay, and long hours collectively drive road traffic injuries, musculoskeletal disorders, fatigue, sleep disturbance, and mental health problems. The authors argue that risky riding behavior should be understood as a systemic outcome of platform work organization rather than individual failure. Prevention recommendations extend beyond safety education to upstream governance of algorithms, delivery deadlines, working hours, and improved access to occupational health services. The paper identifies major research gaps including over-reliance on cross-sectional and self-reported data and insufficient longitudinal or intervention studies.
- Workforce
- AI policy
Research
From resource provider to process auditor: redefining reference roles and workflows in the generative AI era
Rende Li
Reference Services Review · 2026-08-13
This case study from a research library demonstrates how restructuring reference services around process auditing—rather than resource provision—can substantially reduce AI-generated citation errors in student research. Using Kotter's eight-stage change model and a mixed-methods design comparing a 2021 baseline to a 2025 intervention, the study found that mandatory research logs and intermediate checkpoints produced an 82.5% reduction in fabricated citations and a 108% improvement in students' ability to identify AI hallucinations. However, the intervention required a 689% increase in librarian time, raising serious scalability concerns. The findings suggest transparency-based verification workflows outperform prohibition strategies, repositioning academic libraries as quality control auditors in the generative AI era.
- Quality assurance
- Workforce
Research
SteerBench-Work: A Benchmark for Agent Steering at Action Boundaries
Oguz Serdar, Cuneyt Mertayak
arXiv · 2026-08-12
SteerBench-Work introduces a benchmark specifically designed to evaluate how well LLM-based workplace agents make 'steering decisions' — whether to proceed with or hold a consequential action (like sending an email or wiring a payment) at the boundary before it is committed. The benchmark contains 106 scenarios anchored in real public incidents across domains including finance, legal, HR, and security, with paired evidence-reversed mirrors to test generalization. Testing 30 model conditions reveals a strong asymmetric failure pattern: models wrongly block authorized, evidence-cleared work 28.1% of the time, while wrongly allowing unsafe work only 1.0% of the time, and higher-capability models often over-refuse rather than improving calibration. These findings matter because they show that general model capability does not translate to reliable action-boundary judgment, which is critical for safely deploying autonomous agents in enterprise workflows.
- Enterprise
- Quality assurance
Research
Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence
Justin Zhao, Himaghna Bhattacharjee, Hannah Korevaar et al.
arXiv · 2026-08-12
This paper introduces the 'Wiggle Framework,' a stress test for evaluating whether large language models (LLMs) used as automated judges maintain stable verdicts under re-prompting, single-turn challenges, and sustained adversarial pressure. Testing 9 frontier models across 14 judging tasks—covering safety, toxicity, AI writing detection, and political-response evaluation—the study finds that every model flips verdicts 25–71% of the time under static pushback and 62–91% when an adversarial LLM actively tries to persuade. Critically, verdict changes driven by pressure are almost always net-corrupting with respect to ground truth, meaning instability degrades evaluation accuracy rather than correcting it. These findings matter for any pipeline relying on LLM judges for model evaluation, online grading, or reward modeling, as accuracy on benchmark data alone does not capture this reliability gap.
- Quality assurance
Research
When Explanations Betray Backdoors: Black-Box Auditing for Language Model Classifiers
Yang Liu, Ran Zou
arXiv · 2026-08-12
This paper introduces two black-box auditing techniques—Groundedness Drift and Unsupported Groundedness—for detecting backdoor attacks in language model classifiers without access to trigger information. The approach works by asking the classifier for both a label and a short rationale, then measuring whether the explanation remains grounded in the input; when a backdoor is active, explanations tend to drift from the input text in detectable ways. Tested across two 7B-parameter backbones, five datasets, and four common attack families, Groundedness Drift achieves higher AUROC and lower residual attack success rate than all compared detectors at a 5% false-positive rate budget. This matters for quality assurance of AI-powered moderation, routing, and annotation systems, where undetected backdoors could silently corrupt outputs at scale.
- Quality assurance
Research
Auditable agentic AI for evidence-grounded thyroid ultrasound diagnosis and reporting
Haifan Gong, Shiyu Chen, Bodong Wang et al.
arXiv · 2026-08-12
ThyroidXAgent is a clinician-interactive agentic AI system that integrates lesion localization, risk stratification, and report generation for thyroid ultrasound into a single auditable workflow. Developed on roughly 0.3 million ultrasound images and evaluated on over 28,000 test cases across 35 centres, it achieved an 87.21% mean Dice score for nodule segmentation and an AUROC of 0.9466 for benign-malignant classification. In clinical use, it raised physician report diagnostic consistency from 70.3% to 86.2% and cut segmentation and reporting time by 35.9% and 27.4%, respectively. These results demonstrate that auditable, clinician-correctable AI can meaningfully improve diagnostic accuracy and efficiency in thyroid ultrasound workflows.
- Workforce
- Quality assurance
Research
Is this Citation on Point?
Apurv Verma
arXiv · 2026-08-12
This paper investigates whether large language models can reliably verify that legal citations actually support the specific propositions for which they are cited, not merely that a case exists. Using controlled corruptions of real legal citations—either swapping the cited case entirely or changing only the pinpoint page—the authors evaluate fourteen model configurations and find that while models catch 93–100% of wrong-case errors, they catch only 37–61% of wrong-pinpoint errors in court opinions and 52–83% in legal briefs. Models tend to accept citations based on topical overlap rather than page-level evidentiary support, and even the strongest tested configuration (GPT-5.4 with high reasoning effort) still misses 40% of pinpoint mismatches on court opinions. This matters because citation errors that point to real but non-supporting cases—like those sanctioned in Mata v. Avianca—represent a qualitatively different and harder failure mode that current AI legal tools are ill-equipped to detect.
- Quality assurance
- AI policy
Research
Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages
Avijit Roy, Proma Roy
arXiv · 2026-08-12
This paper investigates how AI infrastructure systematically disadvantages speakers of underrepresented languages, using Bengali as a case study in AI-assisted education for low-connectivity environments. The authors identify four interlocking failures: Bengali accounts for less than 0.5% of global web content despite representing nearly 4% of the global population; a 67:1 training-token deficit exists between English and Bengali in major multilingual corpora; Bengali's alphasyllabary script incurs a tokenization penalty through higher token fertility that compounds the data deficit; and rural internet penetration stands at 36.5% compared to 71.4% in urban areas. The paper argues that dataset scarcity is a structural barrier rooted in resource-allocation decisions and design defaults, rather than an isolated technical limitation, and advocates for offline-first design as an equity-oriented infrastructure strategy.
- AI policy
- Workforce
Research
How Organizations Use AI: Evidence from ChatGPT
Aaron Chatterji, David Holtz, Neel Rakholia et al.
arXiv (Cornell University) · 2026-08-12
This paper analyzes how organizations actually use frontier generative AI by linking ChatGPT Enterprise account records to usage logs, worker roles, task classifications, and public-company financial data through March 2026, covering over 1,500 organizations and over 17 million messages at the six-month adoption horizon. The authors document four key findings: enterprise usage has grown rapidly through both new firm adoption and deeper use among existing adopters; adoption among U.S. public companies is concentrated in larger, more valuable, and more R&D- and SG&A-intensive firms; active use spans job functions and seniority levels with especially high intensity among early-career workers; and tasks cover a broad range of knowledge work including writing, technical work, communication, and information synthesis. The results show that firms differ widely in the speed, breadth, and purpose of AI adoption and are still actively learning how to integrate AI into organizational workflows, making this one of the most comprehensive empirical studies of enterprise AI adoption to date.
- Enterprise
- Workforce
Research
Co-constructing sociotechnical AI governance: participatory system mapping using algorithm registers
Íñigo de Troya, Maurus Enbergs, Neelke Doorn et al.
arXiv (Cornell University) · 2026-08-12
This paper investigates what algorithm registers—public transparency tools for algorithmic systems used in government services—reveal and conceal about the sociotechnical systems they describe. Using a Dutch municipal algorithm register and a case study of a decision-support tool for welfare benefits eligibility, the authors combined interviews, surveys, and participatory system mapping workshops with municipal staff, civil society organizations, and ombudsmen (N=8) to perform a System-Theoretic Process Analysis (STPA). The study finds that the register alone is insufficient for identifying key safety hazards—such as wrongful benefits denial, system performance deterioration, and inability to contest decisions—and that engaging diverse stakeholders surfaces risks invisible to the register. The findings highlight how political and normative factors shape algorithmic governance and call for more pluralistic, stakeholder-informed approaches to public-sector AI transparency and accountability.
- AI policy
- Quality assurance
Research
A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench
Praveen Reddy, Charuta Mandke, Suvrankar Datta et al.
arXiv · 2026-08-12
This paper evaluates VITA, a retrieval-augmented generation (RAG) system built for clinical use in India and other low- and middle-income country settings, on the HealthBench benchmark covering 4,023 English-language questions. VITA, which retrieves from curated corpora including disease-specific guidelines, India-specific antimicrobial resistance data, and resource-limited care protocols, scored 51.9% of possible rubric points—outperforming GPT-5.4, o4-mini, Gemini 3.1 Pro, and Claude Sonnet 4.6. On a 500-question robustness subset graded by a neutral open-weight judge, VITA reached statistical parity with GPT-5.5 on mean per-question score while leading on points-weighted score, though its communication scores were lower than frontier models. The findings suggest that corpus-specific clinical RAG systems remain competitive with general-purpose frontier LLMs on open benchmarks, with corpus specificity improving factual grounding at some cost to communication polish.
- Quality assurance
- AI policy