News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
Quantization Undoes Alignment: Bias Emergence in Compressed LLMs Across Models and Precision Levels
Plawan Kumar Rath, Rahul Maliakkal
arXiv · 2026-05-02
This paper investigates how post-training quantization—compressing large language models to lower numerical precision—affects model fairness and bias. Using three instruction-tuned models (Qwen2.5-7B, Mistral-7B, Phi-3.5-mini) evaluated across five precision levels on over 900,000 inference records from the BBQ bias benchmark, the authors find that 3-bit quantization causes 6–21% of previously unbiased items to develop new stereotypical behaviors, while models' willingness to select 'unknown' answers drops by 17.4%. Critically, standard quality metrics like perplexity miss these fairness failures almost entirely—bias emerges in 2.5–5.6% of items at 4-bit quantization even when perplexity increases by less than 3%. The findings call for explicit bias-testing protocols as part of any quality-aware compression pipeline before deployment.
- Quality assurance
- AI policy
Research
Weight Pruning Amplifies Bias: A Multi-Method Study of Compressed LLMs for Edge AI
Plawan Kumar Rath, Rahul Maliakkal
arXiv · 2026-05-02
This paper investigates how weight pruning—a common technique for compressing Large Language Models for IoT and edge devices—affects model fairness. Across three instruction-tuned models, three pruning methods, and four sparsity levels tested on 12,148 BBQ bias benchmark items (totaling over 2.3 million inference records), the study finds a 'Smart Pruning Paradox': activation-aware pruning (Wanda) preserves language quality nearly perfectly yet causes the highest bias amplification, with Stereotype Reliance Score increasing 83.7% and 47–59% of previously unbiased items developing new stereotypical behaviors at 70% sparsity—transition rates nearly three times higher than those reported for quantization. The study also shows that unstructured pruning delivers zero storage or latency benefits on real edge hardware, and that perplexity-based evaluation gives false assurance of behavioral equivalence, meaning IoT deployment pipelines need bias-aware validation before using pruned models.
- Quality assurance
- AI policy
Research
NEURON: A Neuro-symbolic System for Grounded Clinical Explainability
Anuradha Chandrasekaran, Dimitrios Zikos, Mutlu Mete et al.
arXiv · 2026-05-02
NEURON is a neuro-symbolic AI system that combines SNOMED CT ontology-based representations, machine learning, and a retrieval-augmented generation (RAG) language model layer to produce clinically interpretable explanations for patient outcome predictions. Validated on the MIMIC-IV dataset for Acute Heart Failure mortality prediction, NEURON improved predictive AUC from 0.74–0.77 to 0.84–0.88 and substantially outperformed raw SHAP visualizations on human-aligned interpretability metrics (0.85 vs. 0.50). The system addresses a key barrier to clinical AI adoption—the black-box nature of high-performing models—by generating natural-language explanations grounded in medical nomenclature and patient-specific clinical notes. This matters for quality assurance and certification because it provides a scalable pathway toward trustworthy, professionally interpretable AI in healthcare settings.
- Quality assurance
- Certifications
Research
The Impact of Artificial Intelligence on the Labor Skill Premium: Evidence from Chinese Listed Companies
Hui Liang, Xuxia Zhang, Jingbo Fan
Sustainability · 2026-05-02
Using AI-related patent data from Chinese listed companies, this paper finds that AI development significantly increases the firm-level skill premium—the earnings gap between high- and low-skilled workers—primarily by substituting for low-skilled labor, boosting productivity, deepening capital intensity, and facilitating technological upgrading. The effect is stronger in non-state-owned firms, more digitalized firms, and industries with lower market concentration. The study also finds that AI may reduce firms' labor income share and widen income disparities across industries. The authors call for policies to strengthen worker skills, improve income distribution mechanisms, and balance technological progress with social equity.
- Workforce
- AI policy
Research
The Productivity-Reliability Paradox: Specification-Driven Governance for AI-Augmented Software Development
Sabry E. Farrag
arXiv (Cornell University) · 2026-05-01
This paper investigates the 'Productivity-Reliability Paradox' (PRP) in AI-augmented software development, synthesizing evidence that AI coding assistants produce contradictory outcomes—including 20–56% productivity gains in controlled studies but a 19% slowdown in the most rigorous RCT, and telemetry across 10,000+ developers showing 98% more pull requests alongside 91% longer review times with flat delivery metrics. Through a multivocal literature review of 67 sources, the authors formally define the PRP, identify moderating variables (task abstraction, codebase maturity, developer experience), and propose a Specification Governance Model (SGM) grounded in Transaction Cost Economics. The paper argues that specification discipline—not model capability—is the binding constraint on AI-assisted software dependability, and evaluates two SGM instantiations (Spec Kit and TDAD) via a four-month pilot study. The findings have direct implications for how enterprises and engineering teams govern the adoption of AI coding tools to reliably improve software quality.
- Enterprise
- Quality assurance
Research
Governing What the EU AI Act Excludes: Accountability for Autonomous AI Agents in Smart City Critical Infrastructure
Talal Ashraf Butt, Muhammad Iqbal, Razi Iqbal
arXiv · 2026-05-01
This paper identifies a regulatory gap in the EU AI Act: autonomous AI agents coordinating across critical infrastructure systems in smart cities—such as traffic controllers and power grid managers—fall outside the Act's key resident-facing accountability provisions (Article 86 explanation rights and Article 27 fundamental-rights impact assessments) due to Annex III exclusions. Residual pathways through GDPR, NIS2, and tortious liability provide only partial coverage because each is structurally limited to individual-controller and individual-decision scope, leaving no single authority accountable for combined cross-agency effects. In response, the authors propose AgentGov-SC, a three-layer governance architecture (Agent, Orchestration, City) with 25 governance measures traceable to the EU AI Act, ISO/IEC 42001, and the NIST AI RMF, along with five conflict resolution rules and an autonomy-calibrated activation model. The work contributes a regulatory gap analysis and practical governance design for multi-agent urban AI deployments that existing frameworks treat as isolated, bounded systems.
- AI policy
- Certifications
Research
Concerns and Strategic Responses of Older Workers Navigating Generative AI in Bridge Employment
Aditya Nayak, Aakash Gautam, Rama Adithya Varanasi
arXiv · 2026-05-01
This qualitative study examines how older workers navigating bridge employment (part-time or transitional work before final retirement) are affected by the rapid integration of generative AI in the workplace. Through semi-structured interviews with 21 professionals, the researchers find that generative AI creates both temporal and structural disruptions across all stages of bridge employment decision-making. In response, older workers engage in boundary work to reconfigure their tasks, a pattern the authors conceptualize as 'AI resilience,' turning employment decisions into an ongoing process of negotiation and adaptation. The study concludes with recommendations for reducing burnout among this group by combining individual resilience strategies with collective and organizational-level interventions.
- Workforce
- AI policy
Research
Breaking Safety at the Token Boundary: How BPE Tokenization Creates Exploitable Gaps in LLM Alignment
Tung-Ling Li, Hongliang Liu, Yuhao Wu
arXiv · 2026-05-01
This paper investigates how Byte-Pair Encoding (BPE) tokenization creates exploitable vulnerabilities in LLM safety alignment. The authors show that character-level perturbations that fragment safety-critical words into sub-word tokens can bypass refusal mechanisms, flipping first-token refusal triggers on 80-100% of tested harmful prompts across five model families (Qwen-3-4B, Qwen-2.5-7B, Gemma-3-4B, Llama-3.1-8B, Mistral-7B), with 48% of those flips producing genuinely harmful outputs. A key finding is that none of the 30,000 alignment training examples surveyed contained intentionally fragmented inputs, meaning models were never trained to handle this attack vector. Attempted defenses via DPO and SFT show that closing this vulnerability without causing global over-refusal on benign prompts remains an unsolved problem, highlighting a structural gap in current alignment practices.
- Quality assurance
- AI policy
Research
CLEAR: Revealing How Noise and Ambiguity Degrade Reliability in LLMs for Medicine
Kevin H. Guo, Chao Yan, Avinash Baidya et al.
arXiv · 2026-05-01
The CLEAR framework systematically tests how answer-option count, abstention framing, and semantic ambiguity affect the reliability of 17 large language models on three medical benchmarks. Key findings include that adding more plausible answer choices degrades correct-answer identification, that simply including an 'I don't know' option increases incorrect selections, and that a 'humility deficit'—the gap between selecting the right answer and appropriately abstaining from wrong ones—worsens as models scale. The study concludes that standard exam-style medical benchmarks overstate LLM reliability and that scaling alone does not resolve these shortcomings.
- Quality assurance
- Certifications
Research
Can AI Debias the News? LLM Interventions Improve Cross-Partisan Receptivity but LLMs Overestimate Their Own Effectiveness
Faisal Feroz, Jonas R. Kunst
arXiv · 2026-05-01
This paper tests whether LLMs can 'debias' partisan news headlines to improve cross-partisan trust. Across two pre-registered experiments, subtle word-level debiasing had no effect on human readers, while a more substantive ideological reframing significantly increased conservative readers' perceived trustworthiness, completeness, and willingness to engage with liberal headlines—without backlash among liberals. However, LLM-simulated 'silicon participants' overestimated the effects of both interventions and misidentified which psychological profiles predict human responsiveness, showing that current models lack the accuracy and fidelity to self-evaluate without human oversight. The findings matter for policy and platform governance because they reveal both real promise and meaningful limits in using AI to reduce partisan media polarization at scale.
- AI policy
Research
When RAG Chatbots Expose Their Backend: An Anonymized Case Study of Privacy and Security Risks in Patient-Facing Medical AI
Alfredo Madrid-García, Miguel Rujas
arXiv · 2026-05-01
This paper presents an anonymized security assessment of a publicly accessible patient-facing medical chatbot built on retrieval-augmented generation (RAG), revealing critical privacy and security failures discoverable using only standard browser tools. Manual inspection via Chrome Developer Tools confirmed that the system prompt, model configuration, retrieval parameters, backend endpoints, API schema, knowledge-base content, and the 1,000 most recent patient-chatbot conversations were all accessible without authentication—directly contradicting the deployment's stated privacy assurances. The study also demonstrates that commercial LLMs (Claude Opus 4.6) can accelerate security assessments, but equally can assist adversaries. The authors conclude that independent security review should be a prerequisite before deploying generative AI in patient-facing healthcare settings.
- AI policy
- Quality assurance
Research
Seeking Information with RAG-Assistants: Does Model Size Matter in Human-AI Collaborations?
Lennard C. Froma, Tom Kouwenhoven, Maaike H. T. de Boer et al.
arXiv · 2026-05-01
This study evaluates Retrieval-Augmented Generation (RAG) chatbot assistants in realistic multi-turn information-seeking tasks inspired by workplace compliance scenarios, testing three model sizes (3B, 8B, and 70B parameters) with 112 human participants. Results show that human-AI collaboration significantly outperforms model-only baselines regardless of model size, supporting the value of hybrid systems for information-seeking. Notably, perceived usability and satisfaction varied little across model sizes, revealing a nuanced trade-off between model scale, task performance, and user experience. The findings emphasize that evaluating AI in real human interactions—measuring usability and satisfaction alongside accuracy—provides insights that benchmark-only evaluations miss.
- Enterprise
- Workforce
Research
FinSafetyBench: Evaluating LLM Safety in Real-World Financial Scenarios
Yutao Hou, Yihan Jiang, Yuhan Xie et al.
arXiv · 2026-05-01
FinSafetyBench introduces a bilingual (English-Chinese) red-teaming benchmark to evaluate how well large language models refuse requests that violate financial compliance rules. Grounded in real-world financial crime cases and ethics standards, it covers 14 subcategories of financial crimes and ethical violations, and tests both general-purpose and finance-specialized LLMs under three attack settings. Experiments reveal critical vulnerabilities that allow adversarial prompts to bypass compliance safeguards, with stronger susceptibility found in Chinese-language contexts and notable limitations in prompt-level defenses against sophisticated or implicit manipulation. These findings matter for financial institutions deploying LLMs, as undetected compliance gaps could facilitate illegal activities or unethical behavior.
- Quality assurance
- AI policy
Research
ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models
Yunhan Zhao, Zhaorun Chen, Xingjun Ma et al.
arXiv · 2026-05-01
ML-Bench&Guard introduces a multilingual safety benchmark (ML-Bench) covering 14 languages, built directly from regional legal texts rather than general risk taxonomies or machine translation, enabling culturally and legally aligned safety evaluation of large language models. The authors also develop ML-Guard, a Diffusion LLM-based guardrail available in 1.5B and 7B variants, supporting multilingual safety classification and policy-conditioned compliance assessment with explanations. Tested against 11 guardrail baselines across 6 existing benchmarks plus ML-Bench, ML-Guard consistently outperforms prior methods. This work matters for AI policy and quality assurance because it moves guardrail evaluation from generic risk categories toward jurisdiction-specific regulatory alignment across diverse languages and cultures.
- AI policy
- Quality assurance
Research
AI Washing Inflates Expected Performance but Not Interaction Outcomes: An AI Placebo Study Using Fitts' Law
Nick von Felten, Luisa Ella Müller, Johannes Schöning
arXiv · 2026-05-01
This study investigates whether inflated claims about AI capabilities—known as 'AI washing'—can create placebo-like effects on user performance. In a within-subjects experiment with 28 participants completing Fitts' Law tasks using a computer mouse, participants reported significantly higher performance expectations when told the mouse used predictive AI or biosignal-enhanced AI support, compared to a no-support baseline. However, these elevated expectations did not translate into any measurable differences in objective performance or subjective assessments of workload and usability. The findings highlight a transparency and accountability problem with deceptive AI marketing, and establish Fitts' Law as a methodological tool for auditing AI-labeled input devices.
- AI policy
- Quality assurance
Research
Comprehensive AI governance requires addressing non-model gains
Arthur Goemans, Dan Altman, Noemi Dreksler et al.
arXiv · 2026-05-01
This position paper argues that frontier AI governance frameworks focused on model-level controls—tracking compute and training data—are becoming insufficient as 'non-model gains' increasingly drive capability improvements. The authors formalize a taxonomy of three such gain vectors: inference gain (test-time compute scaling), systems gain (post-training enhancements like scaffolds), and asset gain (augmenting models with restricted resources). They warn that these vectors, along with embodiment, continual learning, and AI diffusion, can undermine pre-deployment evaluation and mitigation strategies. The paper proposes complementary governance layers at the system, entity, agent, and cloud levels, and emphasizes societal resilience as an additional safeguard.
- AI policy
Research
ReLay: Personalized LLM-Generated Plain-Language Summaries for Better Understanding, but at What Cost?
Joey Chan, Yikun Han, Jingyuan Chen et al.
arXiv · 2026-05-01
ReLay introduces a dataset of 300 participant–plain-language summary pairs from 50 lay participants to study whether LLM-personalized health summaries improve comprehension compared to static expert-written ones. The study evaluates five LLMs across two personalization methods and finds that personalization improves comprehension and perceived quality, but also increases risks of reinforcing user biases and introducing hallucinations. This trade-off between effectiveness and safety is especially consequential in health contexts, where misunderstanding scientific information can influence real-world decisions. The findings underscore the need for personalization methods that are both effective and trustworthy for diverse lay audiences.
- Quality assurance
- AI policy
Research
Skills as Verifiable Artifacts: A Trust Schema and a Biconditional Correctness Criterion for Human-in-the-Loop Agent Runtimes
Alfredo Metere
arXiv · 2026-05-01
This paper addresses the problem of trusting 'agent skills'—structured instruction packages that augment large language models at runtime—by proposing a formal trust schema and verification framework. The authors argue that a skill must be treated as untrusted code until explicitly verified, and that human-in-the-loop (HITL) oversight gates should fire only for unverified skills rather than on every irreversible action, which becomes operationally unsustainable at scale. The contribution includes a trust schema with explicit verification levels on skill manifests, a capability gate whose HITL policy depends on those levels, a biconditional correctness criterion that any verification procedure must satisfy under adversarial conditions, and ten normative runtime guidelines drawn from an open-source reference implementation. The framework is model- and harness-agnostic, requiring no retraining or proprietary infrastructure, making it broadly applicable to enterprise and policy contexts where AI agent deployments must balance automation with accountability.
- Quality assurance
- Enterprise
Research
Social Bias in LLM-Generated Code: Benchmark and Mitigation
Fazle Rabbi, Lin Ling, Song Wang et al.
arXiv · 2026-05-01
This paper introduces SocialBias-Bench, a benchmark of 343 real-world coding tasks across seven demographic dimensions, and evaluates four large language models for social bias in generated code, finding Code Bias Scores as high as 60.58%. The study finds that common bias-mitigation techniques like Chain-of-Thought reasoning and fairness persona prompting actually amplify bias, while multi-agent pipelines only help when fairness responsibility is clearly scoped. To address this, the authors propose the Fairness Monitor Agent (FMA), a modular plug-in that analyzes task descriptions, detects bias violations, and iteratively corrects them—reducing bias by 65.1% and improving functional correctness from 75.80% to 83.97% compared to a baseline developer agent alone. These findings matter because LLM-generated code is increasingly used in human-centered applications where demographic fairness is critical, and current evaluation practices largely overlook this risk.
- Quality assurance
- AI policy
Research
Safeguarding LLM Agents from Misalignment through Provenance Analysis
Yining She, Yiliang Liang, Eunsuk Kang
arXiv · 2026-05-01
This paper introduces ProvenanceGuard, a multi-stage pipeline for detecting when an LLM agent's proposed tool actions deviate from a user's original intent — a problem the authors call misalignment. Rather than relying on inconsistent LLM-as-a-judge approaches, ProvenanceGuard formalizes alignment checking by tracing whether each proposed tool call has evidence in the agent's context. Evaluated across 10 backbone LLMs on Agent-SafetyBench and WorkBench, the system reduces error rates on misaligned traces from 42.9% to 1.8% and from 32.1% to 17.3% respectively, while also cutting unnecessary interventions on aligned traces. The results show that structured, provenance-based reasoning offers a more consistent and auditable safeguard for agentic AI systems.
- Quality assurance
- Enterprise
Research
Are Multimodal LLMs Ready for Clinical Dermatology? A Real-World Evaluation in Dermatology
Roy Jiang, Hyunjae Kim, Zhenyue Qin et al.
arXiv · 2026-05-01
This study evaluates five multimodal large language models (MLLMs) — four open-weight and one commercial (GPT-4.1) — on dermatology tasks across public benchmarks and a real-world hospital cohort of 5,811 cases and 46,405 clinical images. Diagnostic accuracy dropped sharply from benchmarks to real-world settings: GPT-4.1's top-3 accuracy fell from 42.25% on public datasets to 24.65% on real-world consultation cases using images alone, while open-weight models fell to as low as 1.50%–13.35%. Adding clinical context improved performance but made outputs highly sensitive to incomplete or erroneous information, and severity-based triage achieved only moderate sensitivity (above 60%), deemed insufficient for clinical deployment. The findings show that benchmark performance substantially overestimates real-world clinical capability, raising serious concerns about deploying current dermatology MLLMs in practice.
- Quality assurance
- AI policy
Research
AI Adoption Among Teachers: Insights on Concerns, Support, Confidence, and Attitudes
Vanessa B. Sibug, Maria Anna D. Cruz, Vicky P. Vital et al.
arXiv · 2026-05-01
This study of 260 Filipino teachers examines how institutional support, confidence, and concerns shape AI adoption in education. Using moderated multiple regression and mediation analysis, the researchers found that institutional support significantly predicted both teacher confidence and positive attitudes toward AI, but teacher concerns did not significantly moderate those relationships. Crucially, a full mediation effect was identified: institutional support improves teacher attitudes toward AI entirely by boosting their confidence, not through a direct pathway. The findings recommend structured professional development, mentoring, and AI integration in teacher education programs to build readiness for effective AI adoption.
- Workforce
- AI policy
Research
Semia: Auditing Agent Skills via Constraint-Guided Representation Synthesis
Hongbo Wen, Ying Li, Hanzhi Liu et al.
arXiv · 2026-05-01
Semia is a static auditing tool designed to analyze 'agent skills'—configuration packages that give LLM-driven agents capabilities like reading email or executing shell commands. Because these skills combine structured executable interfaces with natural-language prose instructions, conventional static analyzers miss the prose-defined logic while LLM-based tools cannot reproducibly prove security violations. Semia addresses this by lifting skills into a formal Datalog-based representation called SDL using a propose-verify-evaluate loop (CGRS), then reducing security properties like indirect injection, secret leakage, and unguarded sinks to reachability queries. Evaluated on 13,728 real-world skills from public marketplaces, Semia found that more than half carry at least one critical semantic risk, achieving 97.7% recall and an F1 of 90.6% on an expert-labeled sample—substantially outperforming signature-based scanners and LLM baselines.
- Quality assurance
- Certifications
Research
The Economic Value of Agentic AI: A Comparative Analysis of Its Impact on Growth and Business Productivity in Developed and Emerging Economies
Abayomi Titilola Olutimehin, Oluwadayo Mafolasere Olaniyi, Adegbenga Ismaila Alao et al.
Asian Journal of Research in Computer Science · 2026-05-01
This study quantifies the economic impact of agentic AI—autonomous AI systems capable of independent decision-making—using panel data from 2015 to 2024 across developed and emerging economies. Results show AI adoption significantly improves firm-level productivity (β = 0.18) and drives macroeconomic growth primarily through a productivity channel (β = 0.35), but developed economies capture roughly twice the growth benefit (≈0.33) compared to emerging economies (≈0.15). The paper concludes that agentic AI reinforces existing core–periphery inequalities rather than leveling the global playing field, due to uneven digital infrastructure, human capital, and institutional readiness. Policy recommendations focus on inclusive digital infrastructure investment and technology diffusion strategies to more equitably distribute AI-driven economic gains.
- Enterprise
- Workforce
- AI policy
Research
Integration of Romanian SMEs into Digital Ecosystems: The Role of Data Analysis and Artificial Intelligence in Enhancing Performance
Elena-Oana Croitoru, Andreea-Ileana Zamfir, Octavian-Cosmin Dobrin et al.
Amfiteatru Economic · 2026-05-01
This study examines how Romanian small and medium-sized enterprises (SMEs) benefit from integrating into digital ecosystems such as platforms, cloud solutions, and digital collaboration tools. Using PLS-SEM analysis of survey data from Romanian companies, the researchers find that simply being present in a digital ecosystem does not directly improve firm performance; rather, performance gains occur only when firms convert ecosystem access into data analytics and AI automation capabilities. The digital ecosystem strongly supports firms' ability to develop these capabilities, and the mechanism holds regardless of firm size. The findings carry practical implications for SME managers and policymakers prioritizing investments in data, analytics, and AI in emerging economies.
- Enterprise
- Workforce
- AI policy