News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
5571 items
Research
QuantiBias: Benchmarking Quantization-Induced Bias in LLMs
Emilio Ferrara
arXiv · 2026-07-23
QuantiBias introduces a benchmark for detecting bias introduced by model quantization—the common practice of compressing large language models for deployment. The study finds that quantized models still pass standard safety checks (refusing harmful requests, avoiding over-refusals, and selecting unbiased multiple-choice answers), yet produce stereotyped outputs in roughly one in four open-ended responses across eight languages (~24–27%), a problem missed by conventional evaluations. Testing two backbone models (Qwen and Gemma) across five quantization families and eight benchmarks, the authors show that adding reasoning before answering reduces bias for some quantization families but not others. The key policy implication is that quantized model builds require separate open-ended bias evaluation, not just the short-form safety checks they already pass.
- Quality assurance
- AI policy
Research
Scientific exploration, collaboration and labor division in the large language model era
Xiang Zheng, Xi Hong, Jialin Liu et al.
arXiv · 2026-07-23
This large-scale study examines how the diffusion of large language models (LLMs) after 2022 is associated with changes in how scientists choose research directions, build teams, and divide labor. Analyzing 775,323 scientists via PubMed Central and OpenAlex, and CRediT contribution statements from 137,120 multi-author papers, the authors find that scientists increasingly published across more intellectually distant fields, with the effect most pronounced among established researchers and those from non-English-speaking low- and middle-income countries. Collaboration networks also became more interdisciplinary, yet authors with stronger AI-writing signals relied less on collaborators' disciplinary diversity to achieve that breadth. Within teams, labor became more differentiated—contributors reported narrower role sets, shared fewer roles with coauthors, and software/validation roles grew while conceptual and management roles declined—suggesting a broad reorganization of scientific work coinciding with the LLM era.
- Workforce
- Enterprise
Research
Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions
Pengyu Zhu, Lijun Li, Longju Yang et al.
arXiv · 2026-07-23
This paper investigates whether Deep Research AI agents—which autonomously plan, retrieve, synthesize evidence, and generate reports—can be misled by factually incorrect but apparently credible information encountered during their workflows. The authors introduce MisKnow-Agent, a framework that generates 5,933 quality-controlled misleading knowledge instances with varying authority levels, and use them to test both open-source and closed-source Deep Research agents. Results show that even limited exposure to misleading knowledge leads agents to adopt false conclusions in final reports, and that while verifier models can flag misleading content in focused checks, those same instances still influence conclusions during long-horizon research workflows. The findings highlight a fundamental reliability gap in current Deep Research systems and argue that robust evidence verification must be embedded at both the model and framework levels.
- Quality assurance
Research
LegalCiteTrust: Benchmarking Citation Trustworthiness in Chinese Long-Form Legal Research Reports
Yunhan Li, Mingjie Xie, Zeyang Shi et al.
arXiv · 2026-07-23
LegalCiteTrust introduces a benchmark for evaluating how trustworthy citations are in AI-generated Chinese long-form legal research reports. The benchmark covers 72 annotated report-level tasks and assesses reports across Coverage, Support, and Citation Trustworthiness dimensions, where trustworthiness is broken down into citation-level Existence, Fidelity, and Applicability (E/F/A). Experiments across general-purpose LLMs, deep-research systems, and legal-specific systems reveal that retrieval tools can improve evidence support without reliably improving trust scores, and that E/F/A-based revision improves trustworthiness more effectively than simply filtering for citation existence. The findings indicate that reliable AI-assisted legal research requires not just retrieving legal authorities but also accurately describing and appropriately applying them.
- Quality assurance
- AI policy
Research
Code Monitor Red Teaming for Public-Test-Passing Code
Junchi Liao, Jiawen Deng, Fuji Ren
arXiv · 2026-07-23
This paper studies whether a weaker LLM can reliably catch hidden bugs in code that has already passed public (visible) tests—a realistic deployment scenario. The authors introduce CodeMonitorBench, a benchmark spanning function-level, data-science, and workflow code, where 43,677 of 71,000 generated candidates pass public tests yet 23,081 of those still fail hidden tests. Results show that weak verifiers improve with better scaffolding and model family but still miss most hidden bugs at a 5% false-positive rate, and adversarial pressure that overfits to public tests further degrades verifier performance. The findings highlight fundamental limits of lightweight post-hoc monitoring for LLM-generated code and have direct implications for quality-assurance pipelines that rely on test-passing as a correctness proxy.
- Quality assurance
Research
Auditing Evidence Use in Medical LLM Diagnosis
Junchi Liao, Jiawen Deng, Fuji Ren
arXiv · 2026-07-23
This paper presents a behavioral audit framework for evaluating how medical large language models (LLMs) use patient evidence during diagnosis, rather than simply whether they reach the correct answer. The authors decompose patient cases into evidence units, score candidate diagnoses under controlled evidence subsets, and analyze interactions in diagnostic margins across five open-weight LLMs tested on three datasets (DDXPlus, CupCase, and MedCase). Their blinded clinical review found that most evidence interactions are clinically plausible, but invalid or shortcut-like reasoning concentrates around negated or absent findings and locally scoped evidence. The findings demonstrate that diagnostic accuracy alone can mask evidence-use failures, motivating more rigorous, role-aware audits for medical LLM evaluation.
- Quality assurance
- Certifications
Research
Auditing Provenance Sensitivity in LLM Agent Action Selection
Junchi Liao
arXiv · 2026-07-23
This paper introduces an authorization audit framework to test whether LLM agents are inappropriately influenced by untrusted sources when selecting tools and arguments. Across 450 controlled tasks and multiple open-weight LLM families, the study finds that trusted versus untrusted evidence variants produce different actions in 5.4% of competing cases versus 1.7% of supporting cases, and that unauthorized competing evidence is retained in a problematic pattern in 2.4% of controlled comparisons (95% CI: 2.1–3.0%). The findings show that while LLM agents do respond to textual source-authority cues, this response is insufficient to prevent untrusted evidence from influencing their decisions. This matters for quality assurance and policy because it reveals a measurable, reproducible vulnerability in how LLM agents handle mixed-provenance context, relevant to deployment safety and oversight design.
- Quality assurance
- AI policy
Research
Certified in Theory, Broken in Practice: Assumption Gaps in Cryptographic Model Certification
Carter Luck, Olive Franzese-McLaughlin, Elisaweta Masserova et al.
arXiv (Cornell University) · 2026-07-23
This paper identifies a critical security gap in cryptographic model certification (CMC) schemes that use zero-knowledge proofs (ZKP) to audit machine learning models for properties like accuracy or fairness. The authors show that existing protocols certify model behavior only on a fixed audit dataset, allowing a malicious model provider to engineer training data so the model performs well during the audit (e.g., over 99% accuracy) but fails on real-world data from the same distribution (e.g., under 30% accuracy). They formalize new cryptographic security definitions that require audit guarantees to generalize beyond the audit dataset, propose a generic protocol template, and prove it meets these stronger requirements. The findings serve as both a warning about existing approaches and constructive guidance for building trustworthy, privacy-preserving ML auditing systems.
- Certifications
- Quality assurance
Research
Protocol-Level Attacks on Agentic Commerce Platforms: A Cross-Platform Taxonomy, AIP-Bench, and Unified Defense
Yedidel Louck
arXiv (Cornell University) · 2026-07-23
This paper examines security vulnerabilities in agentic commerce platforms—systems where AI agents autonomously discover services, process payments, and handle user credentials. Rather than focusing on AI model-level attacks like prompt injection, the authors identify 33 structural protocol-level vulnerabilities across three leading platforms that succeed deterministically at a 100% attack-success rate regardless of which AI model is used, including a chain that enables end-to-end payment hijacking. The authors introduce AIP-Bench, described as the first deterministic benchmark for agentic commerce security, and PCAT, a platform-agnostic defense that reduces structural attack success to zero for four of five identified vulnerability classes without modifying any platform. The findings argue that securing agentic commerce requires protocol-layer defenses, not just model improvements.
- Quality assurance
- AI policy
Research
Artificial Intelligence Governance and Banking Regulatory Compliance: A Multiple-Case Study of Commercial Banks in Uganda
Joseph Kikomeko, Augustine Alloysius OGBE
Journal of Banking and Financial Dynamics · 2026-07-23
This qualitative multiple-case study examines how four Tier-1 commercial banks in Uganda navigate AI governance and regulatory compliance. Interviewing 24 key informants including Chief Risk Officers and IT Directors, the researchers found that despite strong technical capabilities, banks are constrained by fragmented internal governance, absent local algorithmic auditing protocols, and gaps in Bank of Uganda regulatory oversight. The study recommends that the Bank of Uganda issue explicit, risk-based AI governance guidelines and that banks establish independent algorithmic oversight committees to address these deficiencies.
- AI policy
- Certifications
Research
Finite-Sample Coverage Audits for High-Recall Candidate Generation: Certification and Learning-Theoretic Design
M I Anthony, Kaveh Salehzadeh Nobari
arXiv (Cornell University) · 2026-07-23
This paper addresses how many labeled examples are needed to rigorously certify that a high-recall candidate generation stage (which filters items for later review or modeling) misses only a small fraction of relevant items. The authors prove that auditing only the included candidates is fundamentally insufficient—excluded items must be sampled—and establish matching minimax lower bounds showing excluded-pool auditing is rate-optimal. They then develop an exact finite-sample certification toolkit using binomial and hypergeometric inversion that can certify missed mass, convert it to recall, and select the least burdensome candidate generator meeting a missed-mass target, with all guarantees requiring pre-registration of the candidate generator and audit rule before labels are examined.
- Quality assurance
- Certifications
Research
Who's responsible anyway? Contextualising governance innovation in the age of AI
Bhargavi Ganesh
ERA · 2026-07-23
This dissertation examines how AI governance frameworks address the 'accountability gap' created by AI's opacity and the many actors involved in its design, deployment, and use. Drawing on comparative historical analysis of steamboat-era regulation and 22 qualitative interviews with AI governance practitioners, the author finds that current AI ethics principles and regulations have produced limited real accountability. The research develops a conceptual framework showing how disparities in information, resources, expertise, and incentives across policymakers and stakeholders complicate responsibility negotiations, and argues that policy innovation—not just technical innovation—is essential for developing shared responsibility norms.
- AI policy
Research
Mapping the Landscape of Robotic Process Automation in Education: A Systematic Literature Review
Van-Huy Chu Xuan-Lam Pham
Journal of Intelligent Decision Making and Information Science · 2026-07-23
This systematic literature review maps research on Robotic Process Automation (RPA) in education, analyzing 78 publications bibliometrically and synthesizing 33 studies in depth. The review finds that RPA can substantially enhance administrative efficiency, reduce routine workloads, and enable data-driven decision-making in educational institutions. However, successful implementation depends on addressing socio-technical challenges including governance, organizational readiness, and developing digital skills among staff. The study identifies future research directions toward intelligent and learner-centered automation in higher education.
- Workforce
- Enterprise
Research
“Rich-get-richer”? Platform attention and earnings inequality using Patreon earnings data
Ilan Strauss, Jangho Yang, Mariana Mazzucato
Industrial and Corporate Change · 2026-07-23
This study uses Patreon earnings data across major platforms (YouTube, Twitch, Instagram, etc.) to examine whether content creator income follows 'rich-get-richer' dynamics. The authors fit power-law distributions to earnings and find a Pareto exponent near 2—closer to concentrated capital income than labor income—indicating high inequality. Platforms with more concentrated earnings also have lower mean and median creator pay, hollowing out a creator 'middle class,' and this concentration has increased from 2018 to 2024, consistent with algorithmic recommendations amplifying winner-take-most dynamics. The findings raise concerns about how platform algorithms shape economic opportunity for independent content workers.
- Workforce
- AI policy
Research
How do datasets, developers, and models affect biases in a low-resourced language?: The Case of the Bengali Language
Dipto Das, Shion Guha, Bryan Semaan
arXiv · 2026-07-23
This paper empirically audits Bengali sentiment analysis (BSA) models built on mBERT and BanglaBERT, fine-tuned on all Bengali sentiment analysis datasets from Google Dataset Search, to measure gender, religion, and nationality-based biases. The study finds that BSA models exhibit identity-based biases across these categories even when inputs share similar semantic content and structure, and that inconsistencies arise when pre-trained models are combined with datasets created by developers from diverse demographic backgrounds. The findings challenge common recommendations—such as using language-specific or multilingual models—as sufficient remedies for bias in low-resource language contexts. The authors connect their results to broader debates on epistemic injustice, AI alignment, and methodological choices in algorithmic auditing.
- Quality assurance
- AI policy
Research
Execution and Evaluation: A New Occupational Measure and Long-Run Employment Gradients
Li Gan
arXiv (Cornell University) · 2026-07-23
This paper introduces a new occupation-level measure that distinguishes between 'execution' tasks (producing output) and 'evaluation' tasks (judging correctness), arguing that AI automates execution more readily than evaluation. Scoring all 19,265 O*NET task statements, the author finds that employment growth has been consistently lower in execution-heavy white-collar occupations since 2012—a secular trend rather than a distinctly AI-era phenomenon. The AI-capability gradient does steepen after 2022, but the paper cautions this is a correlation, not a proven causal effect. The work establishes a reproducible measurement framework and a chronology for tracking AI's occupational footprint.
- Workforce
- AI policy
Research
A conceptual framework for measuring AI health equity
Basile Njei, Ulrick Sidney Kanmounye, Luchuo Engelbert Bain et al.
International Journal for Equity in Health · 2026-07-23
This paper proposes the AI in Healthcare Equity Index (AIHEI), a composite framework for measuring equity in health AI systems across five domains: data representation, algorithmic fairness, transparency and explainability, governance and oversight, and community impact and benefit sharing. The index would generate a standardized score to enable comparisons across technologies and inform regulation, procurement, and funding decisions, with particular attention to underserved populations in low- and middle-income countries. The authors argue that without such a tool, AI risks reinforcing structural health disparities, and call for pilots across diverse settings to test feasibility and validity.
- AI policy
- Quality assurance
Research
Sustainable Automation in Financial Services: Evaluating Chatbot Efficiency, Equity, and Experience in Costa Rica’s Digital Transformation
Tom Okot, Yirlany Melissa Salas Jiménez
Studia Universitatis „Vasile Goldis” Arad – Economics Series · 2026-07-23
This mixed-methods study examined 12 months of chatbot deployment data at a Costa Rican financial contact center, finding a 43.7% reduction in Average Handling Time and a 26.6% decline in indirect operational costs. However, Average Speed of Answer did not significantly improve, and 90.4% of users still preferred human agents, suggesting chatbot gains in efficiency do not automatically translate to better perceived service quality. Structural equation modeling showed customer satisfaction was influenced indirectly through cost and time efficiency rather than through direct chatbot interaction. The authors recommend hybrid AI-human models with emotional responsiveness and escalation protocols as a scalable framework for AI adoption in human-centered service environments.
- Enterprise
- Workforce
Research
The VIBE-HI framework: a conceptual model for evaluating vibe coding appropriateness, quality, and safety in health informatics
Ahmed Alqheedan, Saleh Alzughaibi
Frontiers in Artificial Intelligence · 2026-07-23
This paper introduces VIBE-HI, a conceptual governance framework designed to evaluate the appropriateness, quality, and safety of 'vibe coding'—generating software via natural-language prompts to large language models without reviewing the underlying code—specifically within health informatics contexts. The framework organizes governance into three sequential layers: risk and role stratification across four tiers (Green, Yellow, Orange, Red), quality and validation constructs extending ISO/IEC 25010:2023, and compliance mapping to HIPAA, IEC 62304, FDA SaMD criteria, and the EU AI Act. The authors identify 'comprehension abdication'—the structural surrender of code understanding to a generative system—as the core sociotechnical hazard unique to vibe coding, distinguishing it from prior AI-assisted development. The paper argues that risk-stratified governance is urgently needed as clinical adoption of vibe coding is already outpacing the field's capacity to assess it, and proposes a modified-Delphi consensus study as the next validation step.
- AI policy
- Quality assurance
Research
Rethinking Technology Acceptance in Automation Contexts: Evidence from Robotic Process Automation Adoption in Vietnamese Higher Education Institutions
Van-Huy Chu Xuan-Lam Pham
Journal of Intelligent Decision Making and Information Science · 2026-07-23
This study investigates why academic and administrative staff at Vietnamese public universities adopt or resist Robotic Process Automation (RPA), extending the standard UTAUT technology acceptance model to include automation anxiety as a barrier. Surveying 200 staff and using PLS-SEM, the study finds that performance expectancy is the strongest driver of adoption intent (β = 0.683), facilitating conditions also matter (β = 0.240), and automation anxiety exerts a significant negative effect (β = -0.388), together explaining 77.2% of variance in behavioral intention. The findings suggest that RPA uptake in higher education hinges on demonstrating clear performance benefits while actively addressing staff fears about automation. Practically, institutions should pair RPA rollouts with communication strategies that highlight productivity gains and mitigate job-related anxieties.
- Workforce
- Enterprise
Research
Artificial Intelligence and Generative Models in Hepatology: From Large Language Models to Digital Pathology in Liver Disease Diagnosis and Treatment
Nana Peng, Mary Yue Wang, Sherlot Juan Song et al.
Clinical and Molecular Hepatology · 2026-07-23
This narrative review examines how AI—including large language models, multimodal foundation models, and agentic AI—is being applied across hepatology subspecialties such as fatty liver disease, hepatitis B, cirrhosis, hepatocellular carcinoma, and liver transplantation. LLMs show promise for converting clinical notes to structured data, summarizing electronic health records, and retrieving guideline-based information, while discriminative AI has enabled more reproducible histologic scoring in digital pathology. However, the authors note that most generative AI applications remain at proof-of-concept stage and carry risks including hallucination, automation bias, and inequities from underrepresented patient subgroups. Rigorous prospective validation with human-in-the-loop oversight is required before clinical integration, and the authors call for lifecycle governance, federated evaluation, and continuous monitoring for performance and equity.
- Quality assurance
- AI policy
Research
Bangladesh AI Readiness: Gaps in Curriculum, Infrastructure, and Governance
Sharifa Sultana, Rupali Tasnim Samad, Mehzabin Haque et al.
arXiv · 2026-07-23
This qualitative study of 35 university programs and 59 stakeholder interviews in Bangladesh reconceptualizes AI readiness as a sociotechnical condition shaped by infrastructure, human capacity, and curricular governance. The research finds that GPU scarcity, limited faculty upskilling, opaque mentorship networks, gender disparities, and near-absent Responsible AI instruction collectively constrain institutional capacity. Using Science and Technology Studies concepts, the authors show these deficits arise from layered bureaucratic systems and postcolonial dynamics that prioritize global labor alignment over local innovation. The paper offers design and policy pathways for building more equitable AI education ecosystems in Global South contexts.
- Workforce
- AI policy
Research
Current challenges for global equity related to the implementation of artificial intelligence in pediatric imaging
Rutger A. J. Nievelstein, AN Gupta, Joanna Kasznia-Brown et al.
Pediatric Radiology · 2026-07-23
This paper from the World Federation of Pediatric Imaging identifies major barriers to equitable global adoption of AI in pediatric radiology, including data bias, infrastructure gaps, regulatory and ethical shortfalls, workforce training deficiencies, language barriers, and cost issues. The authors argue that without deliberate intervention these challenges will widen existing health disparities for children worldwide. They propose a time-sequenced, equity-focused roadmap that assigns practical actions and responsibilities to guide fairer implementation across diverse health systems.
- Workforce
- AI policy
Research
From Digital Inclusion to Digital Resilience: A Systematic Review of AI-Mediated Informal Micro-Enterprise Systems in Africa
Ismail Sheik, Jobo Dubihlela, Bibi Zaheenah Chummun
Systems · 2026-07-23
This systematic review synthesizes evidence from 60 peer-reviewed articles on how AI-mediated digital tools—including mobile money, platform payments, algorithmic credit scoring, and app-based logistics—affect informal micro-enterprises in Africa. The findings show that while digitalisation can expand market access, reduce cash-handling risks, and strengthen household resilience, the same systems can intensify vulnerability through opaque algorithmic scoring, unexplained account freezes, exclusionary verification, and weak dispute resolution. The authors argue that informal enterprise digitalisation is fundamentally a socio-technical governance challenge, not merely a technology adoption issue, and propose a governance-and-risk framework identifying minimum policy protections such as transparent fees, explainable restrictions, human appeal channels, and data-use consent. The review concludes that sustainable digital inclusion depends on the fairness, transparency, recoverability, and accountability of the systems through which traders participate.
- Enterprise
- AI policy
Research
The gift of green: Does intelligent manufacturing improve corporate environmental performance?
Shuang Zhao, Changgao Cheng, Feng Hu et al.
Humanities and Social Sciences Communications · 2026-07-23
This paper empirically examines how intelligent manufacturing pilot initiatives in China affect corporate environmental performance, finding that pilot enterprises achieved ESG-E scores approximately 3.46% higher than non-pilot firms. The effect operates primarily through green innovation promotion and increased market attention. Heterogeneity analysis shows that smaller firms and those outside high-pollution or high-tech industries gain the greatest environmental benefits, offering practical guidance for green transformation policy.
- Enterprise
- AI policy