News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
5526 items
Research
From Hype to Evidence: Evaluating LLM Reliability in Supply Chain Management
Gökhan Cenk, Tobias Engel, Jonathan Kreßel et al.
Journal of the Association for Information Systems · 2026-08-15
This paper investigates whether Large Language Models (LLMs) can reliably support supply chain management (SCM) tasks such as forecasting and automated decision-making. Using repeated forecasting trials with agentic LLM orchestration, CrewAI, and Retrieval Augmented Generation (RAG), the authors find that LLM-generated forecasts do not outperform traditional algorithmic approaches, highlighting a gap between LLM hype and domain-specific performance. The paper calls for a dedicated SCM benchmarking framework covering accuracy, consistency, contextual fit, and cost efficiency, and specifically addresses how small and medium-sized enterprises (SMEs) can build evaluation capabilities for responsible AI adoption in compliance with data sovereignty requirements.
- Enterprise
- Quality assurance
- AI policy
News
Suspecting court of using AI, man injected prompts in filings to try to win case
arstechnica.com · 2026-08-14
Ars Technica reports that a Connecticut judge has identified what appears to be the first instance in the U.S. of a plaintiff hiding AI-targeted instructions within court filings, formatted to be invisible to human readers but readable by software. The hidden text was designed to manipulate any AI system reviewing the document, directing it to favor the plaintiff's arguments and ignore prior court denials. Judge Walter Spader Jr. confirmed the tactic had no impact on the case's outcome but warned it sets a 'dangerous' precedent as AI tools become more common in court systems.
- AI policy
- Quality assurance
Research
Whose doctor does the AI recommend? An algorithm audit of reputation and demographic signals in large language model-assisted physician choice
Syeda Anshrah Gillani, Mirza Samad Ahmed Baig
arXiv (Cornell University) · 2026-08-14
This paper audits seven large language models acting as physician recommendation assistants, testing whether demographic signals (gender, ethnicity via names) and reputation attributes (ratings, fees) causally influence which doctors the AI recommends. Using 40,068 randomized choice sets, the study finds that reputation signals dominate recommendations—higher ratings and lower fees strongly drive selections—but also detects statistically significant demographic biases favoring female- and minority-signaled names over White-signaled names, effects the models never mention in their own explanations. The findings reveal that LLM self-reported reasoning is an unreliable transparency mechanism, since demographic tilts appeared in fewer than 0.03% of stated reasons, and the authors argue that repeatable behavioral audits are necessary for meaningful AI accountability in high-stakes intermediary roles.
- AI policy
- Quality assurance
Research
Engineering Reliable Coding Agents: Evaluating and Operating the System Around the Model
Stephanie Jarmak
arXiv (Cornell University) · 2026-08-14
This monograph argues that AI coding agents are commonly evaluated as isolated models but deployed as complex systems, and that reliability depends on the broader infrastructure—including execution environments, retrieval, memory, permissions, and observability—not just model capability alone. Drawing on a structured multivocal review of 164 scholarly works, 100 practitioner records, 29 benchmarks, and 17 author-system case records, the authors find that many apparent model failures originate elsewhere in the system, and that improvements at one layer often fail to propagate to end-to-end outcomes. The paper contributes a versioned catalog of 206 reliability records, runnable evaluation protocols, and a dependency-and-repair framework to help practitioners distinguish model capability from infrastructure effects and build systems that recover safely when components fail. This work is directly relevant to enterprise and quality-assurance teams deploying coding agents, providing structured guidance for defensible evaluation and safer operation.
- Enterprise
- Quality assurance
Research
Optimization of organizational processes in healthcare using artificial intelligence and implications for BRICS countries: a systematic review with narrative synthesis
Aleksey V. Cherepov, Laila M. Malikova, M.V. Makarovskaya et al.
The BRICS Health Journal · 2026-08-14
This systematic review of 78 publications examines how AI technologies affect organizational and managerial processes in healthcare, with particular attention to applicability in BRICS countries. The strongest evidence found that predictive analytics for patient flow management reduced waiting times by 18–26%, while machine learning for operating room scheduling cut idle time and improved resource utilization; ambient and generative AI documentation tools were linked to reduced administrative burden and clinician burnout. Evidence for other applications such as revenue cycle management and supply chain optimization was more limited, and major barriers including fragmented infrastructure, interoperability gaps, and governance challenges temper conclusions about scalability and long-term effectiveness.
- Workforce
- Enterprise
Research
AI-mediated relational competence and its limits: Psychedelic-assisted therapy as a stress case and policy signal
Robert McGrath, Everett B. Sackett
Psychedelics · 2026-08-14
This paper examines how AI is entering psychedelic-assisted therapy (PaT) through facilitator training simulators and in-session administrative tools, using PaT as a stress test for a six-part construct of AI-mediated relational competence. The authors find that while empathic-presence and relational-judgment components of that construct largely hold—and may be intensified—in altered-consciousness settings, trustworthiness, transparency, accountability, and equity-awareness components strain significantly given PaT's unique vulnerabilities and documented history of racial and cultural inequity. The paper raises a pointed policy question: without an accreditation infrastructure comparable to medical education, a privately developed AI platform risks becoming the de facto standard for facilitator competency certification by default.
- Certifications
- AI policy
Research
LegacyWorld: Atomicity-Aware Evaluation of GUI Agents for Legacy Workflows
Thilo Reintjes, Sivajeet Chand, Derui Zhu et al.
arXiv (Cornell University) · 2026-08-14
This paper introduces LegacyWorld, a benchmark of 28 Windows GUI workflows designed to evaluate AI agents automating legacy enterprise systems that lack programmable interfaces. The authors assess six multimodal LLM-based computer-use agents using an 'atomicity' criterion—agent runs must either complete the workflow correctly or fail without leaving unintended persistent changes in business or healthcare records. Results show that useful completion, safe failure, and non-atomic side effects represent distinct operational profiles, and that expert-crafted prompts and screen-recording-derived prompts produce meaningfully different outcomes. The authors conclude that workflow capture, state validators, and atomicity-aware acceptance tests should be treated as first-class requirements for deploying AI agents in legacy workflow automation.
- Enterprise
- Quality assurance
Research
Regime-Conditional Verification: Correctness Estimation for Adapting and Monitoring Safety Classifiers
Thiago Sandoval, Ufuk Topcu
arXiv (Cornell University) · 2026-08-14
Regime-Conditional Verification (RCV) is a lightweight wrapper that adapts off-the-shelf safety classifiers for large language models without retraining them, addressing two common failure modes: misalignment with the deployer's desired policy and performance degradation from distribution shift. By estimating correctness probabilities from the classifier's internal representations, RCV selectively corrects likely-wrong predictions and provides a label-free signal for detecting drift. Across three classifiers and two benchmark datasets, RCV improved policy adherence in every combination, catching up to 81% of previously missed unsafe content, and detected every simulated attack campaign in a deployment study. This matters for organizations deploying LLMs at scale who need reliable, maintainable safety filtering without costly retraining cycles.
- Quality assurance
- Enterprise
Research
Causative-Victim Tracing for Integrative Criminal Policy on AI Deepfake Face and Voice Crimes in Indonesia
Muhammad Abdul Azis, Pujiyono
KRTHA BHAYANGKARA · 2026-08-14
This study examines gaps in Indonesia's criminal law framework for addressing AI deepfake crimes involving faces and voices. The authors find that existing regulations—spread across multiple laws on personal data, electronic transactions, the criminal code, and sexual violence—are fragmented and fail to treat deepfakes as a sequence of interconnected crimes involving biometric data, layered actors, and digital evidence. The paper proposes a 'Causative-Victim Tracing' methodology to systematically link the tracing of data sources, technology, and actors with victim mapping, harm assessment, and recovery mechanisms. The work matters for policy by offering an integrative criminal law framework tailored to deepfake-specific harms.
- AI policy
Research
Coupling Coordination between Corporate Digitalization and Green Efficiency under the Data Elements × Action Plan
Shiman Zhou
Journal of innovation and development · 2026-08-14
Using panel data from 2,683 Chinese listed manufacturing firms (2014–2023), this study constructs a composite digitalization index across five technology dimensions and measures green efficiency via a super-efficiency SBM model, then quantifies their dynamic synergy through a coupling coordination degree model. Results show digitalization rose from 0.142 to 0.387 (10.5% annually) and green efficiency from 0.524 to 0.681, with coupling coordination advancing from near-disorder to a primary coordinated stage. Government R&D subsidies, environmental regulation intensity, and firm human capital all significantly drive higher coordination. The findings offer a quantitative framework for policymakers seeking to align digital transformation initiatives with green low-carbon manufacturing goals.
- Enterprise
- AI policy
Research
Mandato: Protocol-Level Enforcement of Digitally Signed Mandates on AI Agent Actions with Cryptographically Chained Audit Trails
Giovanni Racioppi
arXiv (Cornell University) · 2026-08-14
Mandato is a governance proxy system that enforces cryptographically signed authorization mandates on AI agent actions at the protocol level, addressing the gap where current AI tool-calling protocols (like the Model Context Protocol) lack verifiable, auditable authorization infrastructure. The system intercepts every tool call, checks it against a signed mandate specifying permitted tools, parameter constraints, and conditions, blocks non-conforming calls, and records all decisions in a tamper-evident, hash-chained audit log. The paper also maps the mechanism onto key EU regulations including the EU AI Act (Articles 12 and 14), GDPR, NIS2, and eIDAS 2, with a roadmap toward qualified attestation via Qualified Trust Service Providers. This matters because it provides a technical and legally legible accountability layer for AI agents acting on external systems, directly relevant to enterprise governance, regulatory compliance, and policy enforcement.
- AI policy
- Enterprise
Research
A Four-Axis Trustworthiness Benchmark for LLM-as-Judge in Principle-Based Regulation
Dipankar Sarkar
arXiv (Cornell University) · 2026-08-14
This paper introduces Principle-Bench, a benchmark of 168 cryptoasset financial-promotion scenarios mapped to UK FCA regulatory principles, designed to evaluate LLM-based automated judges across four axes: accuracy, paraphrase robustness, adversarial robustness, and calibration. The authors find that no tested method dominates all four axes, and that a 120B-parameter LLM judge drops 47 accuracy points when faced with keyword-stuffed adversarial inputs—a phenomenon they call 'compliance theatre.' The paper argues that any deployment-grade LLM judge used in principle-based financial regulation must report per-principle adversarial deception rates and calibration metrics alongside aggregate accuracy, not just overall performance figures.
- Quality assurance
- AI policy
Research
Determinants of AI Adoption in Banking: Evidence from Tunisia
Rahma Khattab, Sami Bacha
Arab Economic and Business Journal · 2026-08-14
This study examines what drives Tunisian retail banking customers to adopt AI-based banking services, surveying 400 customers and applying regression analysis within a multi-theory framework. Key findings show that perceived usefulness, ease of use, social norms, and trust positively predict AI adoption intent, while perceived risk reduces it; awareness, attitude, and knowledge show no significant effect. Age and education level moderate adoption behavior, underscoring the role of demographic factors. The results provide actionable guidance for banks and policymakers seeking to accelerate digital financial innovation in emerging markets.
- Enterprise
- AI policy
Research
Beyond Authentication: A Three-Pillar Evidentiary Framework for Deepfake Evidence in Court of law
Pournima Inamdar
Journal of Intelligent Decision Making and Information Science · 2026-08-14
This paper argues that existing evidentiary authentication rules in the US, UK, and India are inadequate to handle AI-generated deepfakes in criminal proceedings, as current standards were designed for an era when media manipulation left detectable forensic traces. The authors identify two key risks: admitting fabricated deepfake evidence and wrongly excluding authentic media via false deepfake claims (the 'liar's dividend'). To address these gaps, they propose a three-pillar framework consisting of a burden-shifting authentication protocol, judicial notice of AI generative capability, and a provenance-first doctrine using cryptographic content credentials. The paper concludes with legislative and procedural reform recommendations across all three jurisdictions.
- AI policy
- Certifications
Research
A hybrid intelligence approach to qualitative data analysis combining manual and AI coding with mixed methods analyses of reliability and validity
Julia Boettinger, Christal Bürgel, Anne Bartsch
Quality & Quantity · 2026-08-14
This methodological study compared manual qualitative data analysis (QDA) of 164 interviews (752,388 words) with AI coding under two prompting conditions, assessing reliability and validity through mixed methods. While quantitative inter-coder agreement was low and AI substantially overcoded segments, qualitative analysis showed that AI with collaborative prompt engineering correctly identified most relevant segments and even caught content missed by human coders. The authors conclude AI coding should not replace manual coding but can complement it through a hybrid intelligence approach that combines prompt engineering, quantitative validation, and qualitative validation to augment overall reliability and validity.
- Quality assurance
- Workforce
Research
Acceptance of Generative AI for Supporting Innovative Learning in Vocational Education
Ponprom Chooppawa, Potsirin Limpinan, Thada Jantakoon
World Journal of Education · 2026-08-14
This study surveyed 544 private vocational education instructors in Thailand to identify what drives their acceptance of Generative AI for innovative teaching. Using PLS-SEM and an integrated theoretical model combining TAM, UTAUT, the Information Systems Success Model, and trust perspectives, the authors found that social influence and trust were the strongest predictors of instructors' intention to use Generative AI, while behavioral intention had the strongest effect on actual innovative pedagogy behavior. The model explained 62.2% of variance in behavioral intention and 51.7% in innovative pedagogy behavior. The findings offer practical guidance for policymakers and administrators on building high-quality AI systems, institutional support, and professional development to integrate Generative AI responsibly in vocational education.
- Workforce
- AI policy
Research
Grounding Health AI: Architecture and Evaluation of a Domain-Expert Metabolic Health Agent
Alon Diament, Gal Sapir, Maria Gorodetski et al.
medRxiv · 2026-08-14
This paper presents the HPP Personal Health Agent (PHA), a metabolic health AI system designed to prevent the hallucination of clinical metrics that general-purpose language models routinely produce when generating health reports. The system grounds outputs in four layers: a deep-phenotyped cohort of 13,000+ participants, 21 domain-expert computational tools, declarative behavioral constraints, and 21 automated evaluations covering 8 failure-mode categories. In a 210-report evaluation matrix, the full system raised a form/provenance score from 0.37 to 0.91 on its primary use case, with tools driving numerical accuracy from ~14% to ~90% of reported metrics correct, while declarative skills added gains in citations, completeness, and structure. The authors argue that trustworthy domain-specialized health AI is fundamentally a systems design problem requiring cohort data, expert tools, and eval-driven development working together.
- Quality assurance
- Enterprise
Research
A Review of the TITAN Guideline: Advancing Transparency and Responsible Artificial Intelligence Use in Medical Research and Scholarly Publishing
Arzoo Nazir, Shah Zeb
Electronic Journal of Medical Research · 2026-08-14
This paper reviews the TITAN guidelines, a framework designed to standardize how artificial intelligence use is reported in medical research and scholarly publishing. The guidelines distinguish between minor AI uses (e.g., language assistance) and substantive applications (e.g., research design, data analysis, or interpretation), while affirming that human researchers retain accountability for AI-generated outputs. TITAN is described as technology-neutral and flexible enough to accommodate emerging systems including multimodal and agentic AI. The authors argue that coordinated adoption by journals and research communities is needed to improve transparency, reproducibility, and scientific integrity.
- Quality assurance
- AI policy
Research
A multi-layer social-theoretical framework for AI ethics
Mohammed Fakrudeen, Jim Otieno
AI and Ethics · 2026-08-14
This paper introduces the Multi-Layer Social-Theoretical AI Ethics Framework (MLST-AEF), a structured tool for translating high-level AI ethics principles into context-sensitive operational assessments. The framework combines normative ethical reasoning, stakeholder analysis, institutional context evaluation, and bias and power auditing across four analytical layers, integrated with a configurable scoring mechanism. An illustrative application to facial-recognition technology in UAE policing reveals tensions between public-safety benefits and concerns around rights, fairness, surveillance, and contestability, yielding a mixed ethical profile supporting conditional rather than unconditional deployment. The MLST-AEF is designed to make ethical trade-offs, stakeholder disagreements, and weighting assumptions explicit while preserving human judgement in decision-making.
- AI policy
- Quality assurance
Research
Artificial Intelligence and Predictive Policing in the Indian Criminal Justice System: Methods, Applications and Constitutional Concerns
Sreehari V S
Journal of Intelligent Decision Making and Information Science · 2026-08-14
This paper examines the adoption of AI-driven predictive policing tools in India—including crime-mapping platforms, facial recognition, and forensic AI—comparing India's trajectory to that of the US, UK, and China. The authors evaluate these deployments against India's constitutional guarantees of privacy, equality, and fair trial, finding that the absence of dedicated legislation risks entrenching caste and communal bias while undermining due process. The paper proposes legislative, institutional, and technical safeguards to create an accountable framework for algorithmic law enforcement in India.
- AI policy
Research
Artificial Intelligence-Driven Workforce Optimization: An Operational Research Framework for Strategic Human Resource Management and Organizational Decision-Making
Dr. Alok Kumar Bhargava
Journal of Intelligent Decision Making and Information Science · 2026-08-14
This paper proposes an integrated AI–Operational Research framework that combines machine learning models, explainable AI (SHAP), and Mixed Integer Linear Programming to help organizations predict and reduce employee attrition. Analyzing 16,189 employee records, Random Forest delivered the best predictive performance, and SHAP identified key drivers of attrition. A subsequent optimization model showed that targeting a small proportion of high-risk employees can mitigate a substantial share of overall attrition risk within resource constraints, offering organizations a practical, data-driven approach to strategic workforce planning.
- Workforce
- Enterprise
Research
Valid Compliance Evidence for Fleets of Autonomous Robots: Independent Conformance Monitoring under Regulation (EU) 2023/1230 and Directive (EU) 2024/2853
Marco Galli
arXiv · 2026-08-14
This paper specifies a formal component called the Independent Conformance Monitor (ICM) designed to produce legally valid compliance evidence for fleets of autonomous and humanoid mobile robots operating under three current EU regulations, including the new Machinery Regulation and AI Act. The authors derive eleven normative requirements, each traceable to a specific legal obligation or physical constraint, and provide third-party-executable conformance procedures. Key findings include three closed-form constraints covering response time (placing the monitor outside any reaction loop), calibration bounds (replacing averaging), and a startup-transient vulnerability that can silently disable a check while leaving a conforming-looking record. The work addresses a critical gap: no published type C safety standard yet covers dynamically stable industrial mobile robots, yet the cited regulations impose defined legal consequences for missing behavioral evidence.
- Certifications
- AI policy
- Quality assurance
Research
When algorithms don’t care: physiotherapy and the post-professional turn
David Nicholls
Physiotherapy Theory and Practice · 2026-08-14
This theoretical paper argues that physiotherapy is entering a 'post-professional era' driven by three converging forces: the commodification of health under late capitalism, political challenges to professional authority, and AI-enabled digital disruption. Drawing on Deleuzian theory and updated with a proposed concept of a 'society of indifference,' the paper contends that predictive algorithms—illustrated through the case of Flok Health, described as the UK's first AI physiotherapy clinic—can autonomously assess, diagnose, treat, and discharge patients, threatening not just physiotherapists' tasks but their agency. The paper raises the concern that if the institutional infrastructure supporting professional physiotherapy is dismantled, populations' rehabilitation needs may go unmet when algorithms and individual choice prove insufficient. It concludes by urging physiotherapy to articulate the enduring human 'intensities' of physical therapy practice rather than simply defend professional territory.
- Workforce
- AI policy
Research
The Vigilant Public: Awareness, Skepticism, and Perceived Democratic Impact of AI-Generated Deepfakes Among Indian Voters
Shubham Bhatia
arXiv · 2026-08-14
This survey study of 404 Indian voters examines how awareness of AI-generated political deepfakes and cheapfakes shapes voter attitudes and behavior. The study finds near-universal AI awareness (99.5%) but lower recognition of specific terms like 'deepfake' (85.6%), with respondents strongly agreeing that AI-generated media undermines political trustworthiness (58.4% vs. 31.2%) and reporting heightened personal caution in decision-making (70.8% agreement). A notable third-person effect emerges: respondents feel personally resilient to AI influence while acknowledging broader societal harms to trust and public dialogue (69.8–73.0% agreement). The authors argue these findings have direct implications for platform labeling policies, media literacy programs, and electoral regulation.
- AI policy
Research
Research ethics and publication policies in journals using the Korea Institute of Science and Technology Information (KISTI) academic publishing platform: a content analysis of 181 journals
Jaemin Chung, Eun Jee Lee, Hyejin Lee
Science Editing · 2026-08-14
This content analysis of 181 journals hosted on South Korea's KISTI academic publishing platform found that research ethics and publication policy disclosure is incomplete and uneven, with journals disclosing an average of 8.29 out of 15 identified policy areas. Traditional policies like copyright and duplicate publication were widely disclosed, but emerging governance areas were rare—only 5 journals had an AI policy and 29 had a data sharing policy. The study concludes that standardized policy templates and platform-supported guidance are needed to strengthen publication governance, particularly around AI and data sharing.
- AI policy
- Quality assurance