News & Research
The latest AI research and news with real-world stakes. Each item is sourced, dated, summarized in plain English and tagged by impact area, and checked against its source before it appears.
- ResearcharXiv2026-04-24QCP
Make Mechanistic Interpretability Auditable: A Call to Develop Guidelines via Continuous Collaborative Reviewing · Michael Lan, Narmeen Fatimah Oozeer, Chaithanya Bandi et al.
This position paper argues that mechanistic interpretability (MI) research lacks the standardized auditing infrastructure needed for adoption in high-stakes applications such as medical AI and autonomous systems. The authors illustrate the problem concretely: conflicting findings across studies on the same neural network behaviors have gone unresolved due to methodological inconsistencies, leaving stakeholders unable to certify the validity of MI results. To address this, the paper proposes a three-part framework — a continuous collaborative reviewing platform, expert-verified guidelines derived from that platform, and source-based auditing that traces how claims depend on prior arguments — designed to complement traditional peer review. The authors contend that auditing MI itself is a prerequisite for its responsible use in AI safety, industry, and governance.
- ResearcharXiv2026-04-24EP
How Supply Chain Dependencies Complicate Bias Measurement and Accountability Attribution in AI Hiring Applications · Gauri Sharma, Maryam Molamohammadi
This paper examines how the complex supply chains underlying AI hiring systems—spanning data vendors, model developers, platform providers, and deploying organizations—make it difficult to measure algorithmic bias and assign accountability. The authors show that bias can emerge from interactions among individually compliant components (e.g., a resume parser that is unbiased alone but discriminatory when combined with specific ranking algorithms), and that information asymmetries leave deploying organizations legally responsible for systems they cannot fully inspect. Drawing on literature review and regulatory analysis of frameworks such as the EU AI Act, NYC Local Law 144, and Colorado's AI Act, the paper argues that current technical and regulatory approaches fail to address these fragmented responsibility structures. The authors propose multi-layered interventions including system-level audits, vendor disclosure guidelines, continuous monitoring, and chain-wide documentation to enable meaningful governance of distributed AI development.
- ResearcharXiv2026-04-24WQ
Rethinking XAI Evaluation: A Human-Centered Audit of Shapley Benchmarks in High-Stakes Settings · Inês Oliveira e Silva, Sérgio Jesus, Iker Perez et al.
This paper evaluates eight Shapley value variants used in explainable AI (XAI) through a large-scale human-centered study involving professional analysts and 3,735 fraud-detection case reviews. The researchers find a fundamental misalignment: standard quantitative metrics like sparsity and faithfulness do not predict how useful or clear explanations are to human decision-makers. Critically, while no Shapley formulation improved analysts' objective performance, explanations consistently increased their decision confidence—raising a serious risk of automation bias in high-stakes settings. The findings challenge current XAI evaluation practices and provide guidance for selecting explanation methods in operational risk workflows.
- ResearcharXiv2026-04-24WP
Measuring and Mitigating Persona Distortions from AI Writing Assistance · Paul Röttger, Kobi Hackenburg, Hannah Rose Kirk et al.
This paper investigates how AI writing assistance distorts the perceived identity, beliefs, and personality of writers—what the authors call 'persona distortions.' Across three large-scale experiments with nearly 3,000 writers and over 11,000 readers, AI-assisted text consistently made writers appear more opinionated, competent, and positive, while shifting perceived demographics toward more privileged groups. Attempts to mitigate these distortions through reward model training were partially successful but reduced user acceptance, suggesting a tension between faithful representation and what users find desirable. The findings carry significant implications for public discourse, political persuasion, and democratic deliberation as AI writing tools scale to hundreds of millions of users.
- ResearcharXiv2026-04-24WP
Smaller, Younger, and More Impactful: How AI-Assisted Writing Transforms Research Teams · Haoyang Wang, Mingze Zhang, Yi Bu et al.
Analyzing 147,074 full-text publications from the PLoS family and Nature portfolio since 2020, this study finds that research teams using AI-assisted writing tend to be smaller and composed of younger researchers compared to those that do not. Crucially, these more compact, junior-leaning teams show a higher probability of producing highly impactful publications, suggesting AI writing tools can offset traditional advantages of large, senior teams. The authors use multiple statistical methods—including propensity score matching—to support these findings, and conclude that the trend warrants policy changes in research evaluation, funding, and training.
- ResearcharXiv2026-04-24WQ
Learning-augmented robotic automation for real-world manufacturing · Yunho Kim, Quan Nguyen, Taewhan Kim et al.
This paper introduces Learning-Augmented Robotic Automation (LARA), a hybrid system combining learned task controllers and a neural 3D safety monitor with conventional industrial robot workflows. Deployed on a live electric-motor production line, the system automated deformable cable insertion and soldering — tasks previously done manually — using less than 20 minutes of real-world training data per task. Over a continuous 5-hour 10-minute run producing 108 motors, it achieved a 99.4% pass rate on product-level quality-control tests, maintained near-human takt time, and reduced variability in solder-joint quality and cycle time without physical safety fencing. The results demonstrate that learning-based robot control can meet industrial standards for reliability, quality, and human safety outside of laboratory settings.
- ResearcharXiv2026-04-24QP
Calibration Drift Under Reasoning: How Chain-of-Thought Budgets Induce Overconfidence in Large Language Models · Prakul Sunil Hiremath, Harshit R. Hiremath
This paper investigates how chain-of-thought (CoT) reasoning affects the calibration of large language models (LLMs), finding that extending reasoning beyond a task-specific threshold can cause systematic overconfidence — a phenomenon the authors term Calibration Drift Under Reasoning (CDUR). Using Llama-3.1-8B and Llama-3.3-70B evaluated on 47 reasoning-trap questions across four reasoning budgets and three seeds (1,368 API calls; 574 valid responses), they show that Expected Calibration Error follows a non-monotonic pattern: it first decreases as reasoning corrects errors, then increases as longer reasoning produces internally consistent but incorrect explanations. To address this, the authors propose CABStop, a calibration-aware stopping rule that halts reasoning when confidence diverges from an auxiliary accuracy estimate. These findings indicate that deeper reasoning does not reliably improve model reliability, which has direct implications for the safe deployment of LLMs in high-stakes settings.
- ResearcharXiv2026-04-24QC
An LLM-Native Psychometric Instrument Reveals a Self-Report--Behavior Gap Across 25 Models · Juan Manuel Contreras
This paper investigates whether the well-known gap between LLMs' self-reported personality traits and their actual behavior is simply an artifact of applying human-derived psychological categories to models. The researchers built the first psychometric instrument grounded in LLM behavior itself, administering 300 items to 25 LLMs across 17 model families and identifying five reliable behavioral factors (Responsiveness, Deference, Boldness, Guardedness, and Verbosity). Even with these LLM-native constructs, self-report still failed to predict how models actually behaved as rated by 151 human observers, showing the self-report–behavior gap is not merely a category-mismatch problem. Critically, the study finds that LLM self-report items and LLM judges share a source of variance invisible to human raters—a confound that standard within-ensemble reliability checks cannot detect, posing concrete risks for LLM-as-judge evaluation pipelines widely used in model assessment.
- ResearcharXiv2026-04-24QP
Behavioral Canaries: Auditing Private Retrieved Context Usage in RL Fine-Tuning · Chaoran Chen, Dayu Yuan, Peter Kairouz
This paper introduces 'Behavioral Canaries,' an auditing framework designed to detect unauthorized use of legally protected retrieved documents in Reinforcement Learning Fine-Tuning (RLFT) of large language models. Standard auditing methods like verbatim memorization checks and membership inference fail for RL-trained models because RL shapes behavioral style rather than retaining specific facts. The proposed framework plants specially crafted preference data—pairing document triggers with distinctive stylistic feedback—so that if those documents are used in training, a detectable behavioral signal emerges. Empirical results show a 67% detection rate at a 10% false-positive rate (AUROC = 0.756) at a 1% canary injection rate, establishing behavioral canaries as a viable tool for auditing training-time influence that manifests as distributional behavioral change.
- ResearcharXiv2026-04-24QP
Estimating Tail Risks in Language Model Output Distributions · Rico Angell, Raghav Singhal, Zachary Horvitz et al.
This paper addresses a critical gap in AI safety evaluation: current methods focus on what inputs cause harmful outputs but ignore how probable those harmful outputs actually are. The authors propose an importance-sampling approach that creates 'unsafe' versions of a target language model to efficiently estimate the probability of harmful outputs, achieving 10–20x sample efficiency gains over brute-force Monte Carlo methods and enabling probability estimates as low as 10^-4 with just 500 samples. The method also reveals model sensitivity to input perturbations and can predict deployment risks, making it practically useful for real-world safety assessments. Because language models are queried billions of times daily, even very rare harmful behaviors will occur frequently in aggregate, making accurate tail-risk estimation essential for responsible deployment.
- ResearcharXiv2026-04-24QP
PrivSTRUCT: Untangling Data Purpose Compliance of Privacy Policies in Google Play Store · Bhanuka Silva, Anirban Mahanti, Aruna Seneviratne et al.
PrivSTRUCT is a new encoder-decoder framework that analyzes privacy policies in Android apps by preserving the document's logical hierarchy—such as section headings—rather than treating the text as flat, uniform content. Applied to 3,756 Google Play Store apps, it extracts more than twice as many data item and purpose excerpts compared to the leading tool PoliGrapher. The study reveals a critical transparency gap: developers who rely on globally defined purposes rather than locally scoped disclosures are 20.4% more likely to overstate data purposes for first-party collection and 9.7% more likely for third-party sharing. Sensitive data flows, such as sharing financial data for analytics, are frequently obscured under generic categories, pointing to systemic failures in how app developers disclose data practices.
- ResearcharXiv2026-04-24Q
Reliable Self-Harm Risk Screening via Adaptive Multi-Agent LLM Systems · Meghana Karnam, Ananya Joshi
This paper presents a statistical framework for multi-agent LLM pipelines used in behavioral health screening tasks such as self-harm risk assessment. The framework models each agent as a stochastic categorical decision and introduces tighter confidence bounds, a bandit-based adaptive sampling strategy, and logarithmic regret guarantees for multi-agent systems. Evaluated on two labeled behavioral health datasets (AEGIS 2.0, N=161; SWMH Reddit posts, N=250), the adaptive sampling approach reduces false positive rates by roughly 40% compared to single-agent models (0.095 vs. 0.159 on AEGIS 2.0) without sacrificing recall. These results suggest that principled adaptive decision-making can meaningfully improve reliability in safety-critical AI screening applications.
- ResearcharXiv2026-04-24P
When AI Speaks, Whose Values Does It Express? A Cross-Cultural Audit of Individualism-Collectivism Bias in Large Language Models · Pruthvinath Jeripity Venkata
This study audited three major AI systems (Claude Sonnet 4.5, GPT-5.4, and Gemini 2.5 Flash) for cultural value bias by presenting them with personal dilemmas framed for users across 10 countries, 5 continents, and 7 languages (840 scored responses), then comparing the advice against World Values Survey Wave 7 data on what people in each country actually believe. All three systems consistently gave Western-style, individualist advice even to users from collectivist societies, with a statistically significant mean gap of +0.76 on a 1–5 scale (t=15.65, p<0.001) and gaps as large as +1.85 for Nigeria. The models also differed in mechanism—Claude shifts toward collectivism in native-language prompts, Gemini shifts more individualist, and GPT-5.4 responds only to stated country identity—while Japan was over-stereotyped as more group-oriented than surveys indicate, revealing encoding of outdated cultural assumptions. The findings suggest frontier AI systems are systematically homogenizing human values toward a Western-individualist baseline, with implications for policy governing AI deployment across diverse populations.
- ResearchAmerican Journal of Health-System Pharmacy2026-04-24ECP
Balancing opportunities and risks of artificial intelligence in drug policy and regulation · Pineal Bareamichael, Tinglong Dai, Mariana P. Socal et al.
This paper examines how artificial intelligence is reshaping five domains of prescription drug policy—drug discovery, regulatory efficiency, coverage decisions, pricing, and supply chains—while also raising concerns about access, equity, and ethical risks. The authors highlight that AI-designed drug candidates have shown early promise, such as one therapy entering phase 2 trials in roughly one-third the conventional development time, though no fully AI-designed drug has yet reached market approval. Without deliberate policy guidance and transparency requirements, the authors warn that AI could reinforce existing inequities and shift priorities away from public health needs. Pharmacists are identified as key advocates for ensuring AI tools remain transparent, equitable, and patient-centered.
- ResearcharXiv (Cornell University)2026-04-24EQP
The Security Cost of Intelligence: AI Capability, Cyber Risk, and Deployment Paradox · Sukwoong Choi
This paper develops an analytical model examining how firms decide on AI deployment and cybersecurity investment when organizational governance controls have not kept pace with AI capability. The central finding is a 'deployment paradox': in high-loss environments, more capable AI can actually lead firms to deploy less, because greater capability requires broader authority exposure (data access, delegated workflows) that increases breach risk under weak governance. The model shows optimal deployment falls below a no-risk benchmark, with the gap worsening as breach-loss magnitude and authority exposure grow, and that governance maturity is a prerequisite—not just a constraint—for translating AI capability into productive use. This has direct implications for enterprise AI adoption strategies and the design of organizational and policy frameworks around AI governance.
- ResearchArtificial Intelligence Review2026-04-24EQ
Explainable artificial intelligence techniques for interpretation of food models: a review · Leonardo Arrighi, Ingrid Alves de Moraes, Marco Zullich et al.
This review examines how Explainable AI (XAI) techniques—such as SHAP and Grad-CAM—can improve transparency and reliability of AI models applied to food quality and safety tasks in Food Engineering. The authors argue that XAI remains underutilized in this domain, limiting trust and adoption of AI-driven assessments like contaminant detection and freshness evaluation via spectral imaging. The paper presents a taxonomy for classifying food quality research by data types and explanation methods, and identifies trends, challenges, and opportunities to encourage broader XAI adoption. This matters for quality assurance because it enables food quality control inspectors to understand and verify AI-generated predictions rather than treat them as black boxes.
- ResearchElectronics2026-04-24EQP
Bias in Large Language Models: Origin, Evaluation, and Mitigation · Yufei Guo, Muzhe Guo, Juntao Su et al.
This review paper systematically examines bias in large language models (LLMs), categorizing biases as intrinsic or extrinsic and surveying evaluation methods at the data, model, and output levels alongside pre-model, intra-model, and post-model mitigation strategies. The authors highlight that biased LLMs pose ethical and legal risks in high-stakes real-world domains such as healthcare and criminal justice. The work serves as a resource for researchers and practitioners seeking to understand, detect, and reduce bias in order to build fairer and more responsible AI systems.
- ResearchCESifo2026-04-24WE
The Organizational Transmission of AI: The Role of Managers on AI Adoption and Impact · Christos Makridis
Using longitudinal survey data from roughly 10,000 U.S. workers tracked annually from 2023 to 2025, this study finds that managerial trust and clear communication are the strongest predictors of whether employees adopt generative AI at work, outweighing factors like income, occupation, and sector. Employees who adopt AI in high-trust, well-communicated workplace environments show markedly higher engagement than peers adopting AI under weaker managerial conditions. The findings suggest that productivity gains from AI depend not just on the technology itself but critically on the organizational culture in which it is deployed, placing managers at the center of technology diffusion within firms.
- ResearchAssessment & Evaluation in Higher Education2026-04-24QCP
On AI glasses and wearable AI in assessment · Thomas Corbin, Sue Sharpe, Phillip Dawson
This paper examines how AI-enabled smart glasses and wearable AI devices—capable of displaying AI-generated text, processing speech, and reading exam materials without detection—undermine the physical exclusion strategies that higher education institutions use to ensure academic integrity. The author introduces the concept of 'dual transparency' to describe how wearable AI erodes the separability and observability conditions that invigilated exams and oral assessments rely upon. The paper warns that attempting to enforce physical exclusion under these conditions risks creating a 'bodily adjudication' regime that disproportionately burdens students with disabilities, health conditions, and religious dress practices. The findings are significant for assessment policy and certification integrity in higher education.
- ResearchIbnosina Journal of Medicine and Biomedical Sciences2026-04-24QCP
Advances in Regulatory Review Pathways in the Middle East: A Narrative Review of Reliance, Expedited Approvals, and Digitalization in Medicine Registration · Mohammad Nammas
This narrative review examines how Arab countries are modernizing medicine registration through expedited review pathways, reliance on stringent regulatory authority assessments, conditional approvals, and digital submission systems like the electronic Common Technical Document. The paper finds that Saudi Arabia has the most comprehensive suite of accelerated tools, while Egypt, Jordan, Kuwait, the UAE, Bahrain, Qatar, and Oman have also introduced reliance-based and accelerated pathways, with the Gulf Cooperation Council providing a regional work-sharing mechanism. Ongoing challenges include regulatory heterogeneity, limited institutional capacity, and insufficient real-world evidence guidance, while early AI initiatives—such as national AI authorities and automated dossier-screening—offer further opportunities to improve regulatory efficiency. The authors call for harmonized standards, stronger capacity, and expanded digitalization to ensure timely, equitable access to medicines across the region.
- ResearchInformation Systems Frontiers2026-04-24EQP
An Explainable AI Multi-Agent Recommender System for Financial Document Access Control · Kanellos Toudas, Konstantinos I. Roumeliotis, Dimitrios Κ. Nasiopoulos et al.
This paper presents a multi-agent AI system for classifying financial documents into four sensitivity levels—Public, Internal, Confidential, and Restricted—using fine-tuned models (FinBERT, BERT-base-uncased, GPT-4.1-mini) orchestrated by GPT-5.1, which generates natural language explanations for each access-control decision. The overall system achieves 83.71% accuracy, while cases where all three agents agree (78.8% of cases) reach 92.28% accuracy, outperforming any single model. By providing interpretable justifications rather than opaque classifications, the system addresses transparency and accountability concerns in AI-driven financial document security. The work matters for enterprise compliance and quality-assurance processes where explainability of automated access-control decisions is both an ethical and regulatory concern.
- ResearcharXiv2026-04-23QP
Mathematical Modelling of Ethical AI Use in Higher Education: A Coordination Game Framework for Future-Facing Learning · Ndidi Bianca Ogbo, Zhao Song, Shatha Ghareeb et al.
This paper models student use of generative AI in higher education as a coordination game, arguing that collective norms around responsible versus opportunistic AI use are shaped more by peer expectations and assessment design than by individual compliance. Using an evolutionary game-theoretic framework with finite-population simulations, the authors show that small, well-calibrated changes in reflective assessment incentives can trigger rapid, threshold-driven shifts toward responsible AI-use norms, while weak or misaligned incentives allow opportunistic behavior to persist. These non-linear dynamics help explain why policy statements alone frequently fail to change student behavior, whereas modest assessment redesigns can have disproportionate effects. The work offers institutions an analytically grounded approach to AI governance that avoids surveillance or punitive enforcement in favor of pedagogy-led design.
- ResearcharXiv2026-04-23QP
Reliability Auditing for Downstream LLM tasks in Psychiatry: LLM-Generated Hospitalization Risk Scores · Shevya Panda, Shinjini Bose, Ananya Joshi
This study proposes a reliability auditing framework for LLMs used in psychiatric hospitalization risk assessment, testing four models (Gemini 2.5 Flash, LLaMa 3.3 70b, Claude Sonnet 4.6, GPT-4o mini) against synthetic patient profiles with varying prompt designs and clinically insignificant inputs. The results show that including medically irrelevant variables statistically significantly increased both mean predicted hospitalization risk scores and output variability across all models and prompts, indicating reduced predictive stability as contextual noise increased. Prompt framing alone also independently shifted model instability in model-dependent ways. The findings underscore the need for systematic attribution and uncertainty evaluations before deploying LLM-based psychiatric risk tools in clinical settings.
- ResearcharXiv2026-04-23CP
A Systematic AI Adoption Framework for Higher Education: From Student GenAI Usage to Institutional Integration · Michael Neumann, Lasse Bischof, Maria Rauschenberger et al.
This study investigates how students in computer science-oriented programs use generative AI tools and proposes a structured framework to help higher education institutions adapt their regulations and curricula accordingly. A case study at the University of Applied Sciences and Arts Hannover (Germany) combined document analysis with an online survey of 151 students in Business Information Systems and E-Government programs. Findings show that GenAI adoption—particularly ChatGPT—is widespread, but many students were unaware of or uncertain about institutional regulations, and document analysis revealed regulatory gaps, ambiguous terminology, and inconsistencies between formal rules and teaching practices. In response, the authors propose the AI Adoption Framework for Higher Education, an iterative model integrating empirical observation, document analysis, and targeted updates to governance, assessment validity, and academic integrity policies.
- ResearcharXiv2026-04-23QP
When Cow Urine Cures Constipation on YouTube: Limits of LLMs in Detecting Culture-specific Health Misinformation · Anamta Khan, Ratna Kandala, Deepti et al.
This paper examines the limits of large language models in detecting culture-specific health misinformation by using YouTube discourse around gomutra (cow urine) in India as a case study. Analyzing 30 multilingual transcripts across three LLMs (GPT-4o, Gemini 2.5 Pro, DeepSeek-V3.1) with varied prompt tones, the researchers find that culturally embedded health misinformation blends sacred traditional language with pseudo-scientific claims in ways that LLMs trained predominantly on Western corpora are systematically ill-equipped to analyze. The study also identifies that cultural obfuscation extends to gendered rhetoric and prompt design, compounding analytical unreliability, and concludes that cultural competency cannot be retrofitted through prompt engineering alone. These findings matter for quality-assurance and policy efforts relying on AI tools to moderate health misinformation on social media in the Global South.