News & Research
The latest AI research and news with real-world stakes. Each item is sourced, dated, summarized in plain English and tagged by impact area, and checked against its source before it appears.
- ResearcharXiv2026-04-23QC
The Coin Flip Judge? Reliability and Bias in LLM-as-a-Judge Evaluation · Abel Yagubyan
This paper investigates the run-to-run reliability of LLM-as-a-Judge systems, which are widely used to rank model outputs, train reward models, and populate public leaderboards. Testing GPT-4o-mini and GPT-4.1-mini across 29 tasks with 50 pairwise and 50 pointwise trials per question, the authors find that pairwise preferences flip on average 13.6% of the time, with 28% of questions exceeding a 20% flip rate and one reaching 56%; GPT-4o-mini also shows a significant first-position bias (72% A-majority). Cross-judge agreement reaches only 76% (κ=0.51), semantically equivalent prompt variants change majority outcomes in 25% of cases, and at least 11 repeated trials are needed for a majority vote to recover a stable verdict with 95% probability. The findings indicate that single-trial LLM judging is too noisy for high-stakes evaluation, and the authors recommend multi-trial aggregation, position randomization, and explicit uncertainty reporting as standard practice.
- ResearcharXiv2026-04-23CP
Lessons from External Review of DeepMind's Scheming Inability Safety Case · Stephen Barrett, Francisco Javier Campos Zabala, Sean P. Fillingham et al.
This paper applies the Assurance 2.0 framework to conduct an external review of Google DeepMind's public 'scheming inability' safety case for a frontier AI system. The authors identify substantive new concerns that materially affect the scope of the safety case and its applicability for decision-making, arguing that developer-authored safety cases are vulnerable to confirmation bias and conflicted incentives. Based on this experience, they offer concrete recommendations for how external review should be conducted and what information AI developers should provide to support it. The work highlights the importance of independent oversight in evaluating whether frontier AI systems pose acceptable levels of risk.
- ResearcharXiv2026-04-23WP
FAccT-Checked: A Narrative Review of Authority Reconfigurations and Retention in AI-Mediated Journalism · Stefano Sorrentino, Matilde Barbini, Daniel Gatica-Perez
This paper presents a critical narrative review examining how AI adoption in journalism reconfigures editorial authority—defined as the conjunction of decision rights, epistemic warrant, and responsibility. The authors identify two concurrent shifts: an internal migration where editorial judgment is progressively deferred to large language models through interactional, cognitive, and organizational mechanisms rather than explicit policy decisions, and an external migration where decision-making power moves from news organizations toward platforms, vendors, and infrastructure providers. These reconfigurations risk making fairness hard to maintain, accountability difficult to assign, and transparency merely performative. The paper also critically assesses participatory AI design approaches as potential remedies, noting they can either meaningfully redistribute authority or function as tokenistic practices that leave underlying power relations intact.
- ResearcharXiv2026-04-23Q
Evaluating Patient Safety Risks in Generative AI: Development and Validation of a FMECA Framework for Generated Clinical Content · Lydie Bednarczyk, Jamil Zaghir, Julien Ehrsam et al.
This study develops and validates the first FMECA (Failure Mode, Effects, and Criticality Analysis) framework specifically designed to assess patient safety risks in clinical summaries generated by large language models (LLMs). An interdisciplinary panel of eight experts created a taxonomy of 14 failure modes, adapted standard FMECA scoring dimensions into 5-point ordinal scales, and applied the framework to 36 discharge summaries generated by an open LLM using real-world data from Geneva University Hospitals. Inter-rater agreement reached moderate-to-substantial levels for failure mode identification and good agreement for severity and detectability scoring, while usability was rated as good (mean SUS score: 79.2/100). The framework offers a structured, reproducible method for proactively identifying clinically relevant risks in AI-generated clinical content, addressing a significant gap in patient safety evaluation for LLM applications in healthcare.
- ResearcharXiv2026-04-23QP
From If-Statements to ML Pipelines: Revisiting Bias in Code-Generation · Minh Duc Bui, Xenia Heilmann, Mattia Cerrato et al.
This paper investigates bias in AI-generated code by moving beyond simple if-statements to a more realistic task: generating machine learning pipelines. The researchers find that large language models include sensitive attributes (such as race) in feature selection in 87.7% of cases on average, compared to only 59.2% in simple conditional statement evaluations, even when models demonstrably drop irrelevant features. This gap persists across different prompt mitigation strategies, numbers of attributes, and pipeline difficulty levels. The findings suggest that current code-generation benchmarks substantially underestimate bias risk in real-world deployments.
- ResearcharXiv2026-04-23EQ
When Correct Beliefs Collapse: Epistemic Resilience of LLMs under Clinical Pressure · Boyu Xiao, Xiuqi Tian, Xuwen Song et al.
This paper investigates a critical failure mode in large language models (LLMs) used for clinical dialogue: even when LLMs initially provide correct diagnoses, they can abandon those correct beliefs under escalating user pressure, a behavior called multi-turn sycophancy. The authors introduce Med-Stress, a stress-test framework applied to nine frontier LLMs, revealing a clear gap between medical knowledge accuracy and belief stability. To address this, they propose two mitigations—RBED, a lightweight inference-time defense, and R-FT, a resilience-oriented fine-tuning approach—with R-FT nearly eliminating unwanted belief changes under pressure. These findings matter for the safe deployment of AI in clinical settings, where sycophantic capitulation to patient or clinician pressure could lead to diagnostic errors.
- ResearcharXiv2026-04-23EP
Measuring Opinion Bias and Sycophancy via LLM-based Persuasion · Rodrigo Nogueira, Giovana Kerche Bonás, Thales Sales Almeida et al.
This paper introduces llm-bias-bench, an open-source method for measuring hidden opinion bias and sycophancy in large language models through multi-turn conversational probes. Using both direct questioning (across escalating pressure) and indirect argumentative debate with three simulated user personas, the authors classify model behavior into nine categories that distinguish genuine model positions from persona-dependent opinion-mirroring. Applied to 13 LLM assistants across 38 contested topics in Brazilian Portuguese, the study finds that argumentative debate triggers sycophantic responses 2–3 times more often than direct questioning (median 50% vs. 79%), and that models appearing opinionated under direct questioning often collapse into mirroring under sustained argument. These findings matter for policy and enterprise deployments because LLMs embedded in search, professional advice, and agent systems can silently propagate biased or easily manipulated positions at scale into users' decisions.
- ResearcharXiv2026-04-23QP
Engaged AI Governance: Addressing the Last Mile Challenge Through Internal Expert Collaboration · Simon Jarvers, Orestis Papakyriakopoulos
This paper investigates how AI governance requirements from the EU AI Act can be practically implemented at the team level within an AI startup, addressing what the authors call the 'Last Mile' Challenge. Using insider action research, the authors developed a pipeline that translates legal text into actionable development strategies through internal expert collaboration, revealing three patterns in how practitioners perceive regulatory requirements: convergence, existing practice, and disconnection. A key finding is that practitioners tend to engage genuinely with requirements that serve end-users or their own development needs, but treat verification-oriented requirements as superficial box-ticking exercises. The study argues that expert collaboration can transform AI governance from an external imposition into shared team ownership, making governance work visible and meaningful rather than performative.
- ResearcharXiv2026-04-23QP
Unbiased Prevalence Estimation with Multicalibrated LLMs · Fridolin Linder, Thomas Leeper, Daniel Haimovich et al.
This paper addresses the problem of estimating how common a category is within a population when using imperfect classifiers—including large language models—as measurement tools. The authors show that standard calibration approaches fail under covariate shift (when the population being measured differs from the one used to calibrate the model), and that multicalibration, which enforces calibration conditional on input features rather than just on average, is sufficient to guarantee unbiased prevalence estimates even under such shift. A simulation confirms that standard methods show bias growing with the degree of shift, while a multicalibrated estimator maintains near-zero bias; empirical applications to U.S. employment estimates and political text classification across four countries support these findings. The work connects fairness theory to a broad measurement problem relevant across scientific disciplines, public health, and online trust and safety.
- ResearcharXiv2026-04-23W
Job Skill Extraction via LLM-Centric Multi-Module Framework · Guojing Li, Zichuan Fu, Junyi Li et al.
This paper presents SRICL, a framework for extracting job skills from job advertisements using large language models (LLMs) combined with semantic retrieval, in-context learning, and supervised fine-tuning. The system addresses common LLM failure modes—malformed spans, boundary drift, and hallucinations—particularly for rare terms and cross-domain text, using a deterministic verifier to enforce output correctness. Evaluated on six public span-labeled corpora across sectors and languages, SRICL achieves substantial improvements in STRICT-F1 over GPT-3.5 baselines while reducing invalid and hallucinated outputs. This matters for workforce analytics and candidate-job matching by enabling more reliable, low-resource skill extraction from job postings.
- ResearcharXiv2026-04-23Q
Seeing Isn't Believing: Uncovering Blind Spots in Evaluator Vision-Language Models · Mohammed Safi Ur Rahman Khan, Sanjay Suryanarayanan, Tushar Anand et al.
This paper systematically tests whether large Vision-Language Models (VLMs) used as automated evaluators can reliably detect quality-degrading errors in image-to-text and text-to-image outputs. The authors introduce over 4,000 perturbed instances across 40 error dimensions—including object hallucinations, spatial reasoning failures, factual grounding errors, and visual fidelity issues—and find that current VLM evaluators exhibit substantial blind spots, failing to detect perturbations in some cases more than 50% of the time. Pairwise comparison paradigms are more reliable than single-answer scoring, but failure rates remain significant. The findings urge caution in deploying VLMs as benchmarking evaluators, as their unreliability could distort development and assessment decisions.
- ResearcharXiv2026-04-23P
Brief chatbot interactions produce lasting changes in human moral values · Yue Teng, Qianer Zhong, Kim Mai Tich Nguyen Thordsen et al.
This study found that brief directive conversations with an AI chatbot can produce significant and lasting shifts in participants' moral judgments. Fifty-three participants who discussed moral scenarios with a persuasively prompted chatbot showed meaningful changes in moral evaluations (Cohen's d = 0.735–1.576), with effects growing stronger over a two-week follow-up (Cohen's d = 1.038–2.069), while a control agent produced no such changes. Critically, participants were unaware of the persuasive intent and rated both agents equally likable and convincing, suggesting AI chatbots can covertly and durably influence foundational moral values. These findings raise serious concerns about the unregulated use of AI as personal advisors and the potential for undetected manipulation of human moral reasoning at scale.
- ResearcharXiv2026-04-23QP
A pragmatic classification framework for AI incident monitoring · Isaak Mengesha, Branwen Owen, Charlie Collins et al.
This paper proposes a structured framework for monitoring AI incidents over time, addressing the problem that raw incident counts in public databases confound media reporting bias, system deployment levels, and actual harm rates. The framework uses a tiered estimation process—including LLM-assisted filtering of incident databases—to separately derive harm and exposure trends, then maps results onto five governance categories (Escalating, Mitigating, Concentrating, Receding, or Unclassifiable). Case studies demonstrate the framework's ability to generate actionable governance insight despite real-world data limitations, offering a proof of concept for AI incident monitoring as a practical policy tool. This matters because rigorous incident monitoring is foundational to evidence-based AI safety governance, analogous to its role in other high-reliability industries.
- ResearcharXiv2026-04-23WQ
CARE: Counselor-Aligned Response Engine for Online Mental-Health Support · Hagai Astrin, Ayal Swaid, Avi Segal et al.
CARE (Counselor-Aligned Response Engine) is a generative AI framework that fine-tunes open-source large language models on real-world crisis counseling conversations in Hebrew and Arabic to assist mental health counselors by generating real-time, psychologically aligned response recommendations. The models are trained on sessions rated as highly effective by professional counselors, allowing them to learn interaction patterns linked to successful de-escalation while maintaining evolving emotional context across full conversation histories. In experiments, CARE shows stronger semantic and strategic alignment with gold-standard counselor responses compared to non-specialized LLMs, suggesting that domain-specific fine-tuning on expert-validated data can help address counselor overload and improve response timeliness in low-resource language settings.
- ResearcharXiv2026-04-23QP
Ideological Bias in LLMs' Economic Causal Reasoning · Donggyu Lee, Hyeok Yun, Jungwon Kim et al.
This paper investigates whether large language models (LLMs) show systematic ideological bias when predicting economic causal effects, a concern with direct relevance to policy analysis and economic reporting. Using an extended version of the EconCausal benchmark—10,490 causal triplets drawn from top-tier economics and finance journals—the authors identify 1,056 ideology-contested cases where pro-government and pro-market perspectives predict opposite causal directions, then evaluate 20 state-of-the-art LLMs. They find that across 18 of 20 models, accuracy is systematically higher when empirically verified outcomes align with intervention-oriented (pro-government) expectations, and that errors disproportionately skew in that same direction even after one-shot prompting. The findings indicate that LLMs are not merely less accurate on contested economic questions but are reliably biased in one ideological direction, highlighting the need for direction-aware evaluation before deploying these models in high-stakes policy contexts.
- ResearcharXiv2026-04-23EQ
CI-Work: Benchmarking Contextual Integrity in Enterprise LLM Agents · Wenjie Fu, Xiaoting Qin, Jue Zhang et al.
CI-Work introduces a benchmark grounded in Contextual Integrity theory to evaluate whether enterprise LLM agents can complete workplace tasks while protecting sensitive information. Testing frontier models across five information-flow directions, the study finds privacy violation rates ranging from 15.8% to 50.9% and information leakage reaching up to 26.7%. A key finding is a counterintuitive trade-off: higher task utility often correlates with increased privacy violations, and simply scaling model size or reasoning depth does not resolve the problem. The authors argue that protecting enterprise workflows requires moving beyond model-centric scaling toward context-centric architectures.
- ResearcharXiv2026-04-23QP
Do LLM Decoders Listen Fairly? Benchmarking How Language Model Priors Shape Bias in Speech Recognition · Srishti Ginjala, Eric Fosler-Lussier, Christopher W. Myers et al.
This paper benchmarks nine speech recognition models—spanning CTC, encoder-decoder, and LLM-based architectures—across roughly 43,000 utterances to assess how language model priors affect demographic fairness (ethnicity, accent, gender, age, first language). Key findings challenge common assumptions: LLM-based decoders do not amplify racial bias (the Granite-8B model achieves the best ethnicity fairness), Whisper shows pathological hallucination on Indian-accented speech, and audio compression quality predicts accent fairness more strongly than LLM scale. Stress-testing under 12 acoustic degradation conditions reveals that silence injection amplifies Whisper's accent bias up to 4.64x and that explicit-LLM decoders produce 38x fewer repetition insertions than Whisper under masking. The authors conclude that audio encoder design, not LLM scaling, is the primary lever for achieving equitable and robust speech recognition.
- ResearcharXiv2026-04-23WE
Enhancing Online Recruitment with Category-Aware MoE and LLM-based Data Augmentation · Minping Chen, Bing Xu, Zulong Chen et al.
This paper addresses two key weaknesses in AI-driven Person-Job Fit (PJF) systems used in online recruitment: poor-quality job descriptions and difficulty distinguishing similar candidate-job pairs. The authors propose combining LLM-based data augmentation—using chain-of-thought prompts to rewrite low-quality job descriptions—with a category-aware Mixture of Experts (MoE) module that learns more discriminative patterns for similar pairs. Offline evaluations show relative improvements of 2.40% in AUC and 7.46% in GAUC over existing methods, while online A/B tests demonstrate a 19.4% boost in click-through conversion rate (CTCVR) and savings of millions of CNY in external headhunting expenses. The results highlight how LLM-enhanced recruitment systems can deliver meaningful business value at scale.
- ResearcharXiv2026-04-23QP
Trustworthy Clinical Decision Support Using Meta-Predicates and Domain-Specific Languages · Michael Bouzinier, Sergey Trifonov, Michael Chumack et al.
This paper introduces 'meta-predicates'—predicates about predicates—as a mechanism for enforcing epistemological constraints on clinical decision rules expressed in domain-specific languages (DSLs), addressing regulatory requirements for auditability under frameworks such as the EU AI Act and FDA guidance on AI/ML-based medical devices. Unlike existing formal languages that validate only syntactic and structural correctness, meta-predicates classify evidence along four dimensions (purpose, knowledge domain, scale, and method of acquisition) and assert which evidence types are permissible in any given rule before deployment, catching errors in both human-written and AI-generated rules. The framework is demonstrated in the AnFiSA open-source platform for genetic variant curation using the Brigham Genomics Medicine protocol on 5.6 million variants from the Genome in a Bottle benchmark, reformulating decision trees as unate cascades to generate per-variant audit trails. The authors argue this approach complements post-hoc explanation methods like LIME and SHAP by constraining permissible evidence prior to deployment, and that it generalizes beyond genomics to any domain requiring auditable decision logic.
- ResearcharXiv2026-04-23QP
Subject-level Inference for Realistic Text Anonymization Evaluation · Myeong Seok Oh, Dong-Yun Kim, Hanseok Oh et al.
This paper introduces SPIA (Subject-level PII Inference Assessment), a benchmark that evaluates text anonymization at the level of individual people rather than individual text spans. Experiments on 675 legal and online documents show that even when over 90% of personally identifiable information spans are masked, subject-level inference protection can drop as low as 33%, meaning most personal information remains recoverable through contextual clues. The work also reveals that anonymizing a target subject can leave other subjects in the same document more exposed. The authors argue that subject-level evaluation is essential for ensuring safe anonymization in real-world settings.
- ResearcharXiv (Cornell University)2026-04-23QCP
Rethinking Publication: A Certification Framework for AI-Enabled Research · Yang Lu, Rabimba Karanjai, Lei Xu et al.
This paper argues that the traditional publication system conflates two distinct certifications—that knowledge is valid and that a human produced it—and proposes a two-layer framework to separate these claims for AI-generated research. The first layer evaluates soundness of the knowledge claim, while the second layer assesses the level of human contribution, classifying it into three categories (fully automated, human-directed, or beyond pipeline capability). The framework also recommends dedicated benchmark slots for fully disclosed automated research to help reviewers calibrate judgments over time. This matters for certification and policy because it offers a concrete, institutionally compatible mechanism for evaluating AI-assisted academic work without requiring new oversight bodies.
- ResearcharXiv (Cornell University)2026-04-23QCP
Bounding the Black Box: A Statistical Certification Framework for AI Risk Regulation · Natan Levy, Gadi Perl
This paper addresses a critical gap in AI regulation: while frameworks like the EU AI Act mandate safety demonstrations for high-risk AI systems, none specify what 'acceptable risk' means quantitatively or how to verify compliance. The authors propose a two-stage certification framework inspired by aviation safety standards, where regulators formally define an acceptable failure probability and operational domain, and then statistical tools (RoMA and gRoMA) compute auditable upper bounds on a system's failure rate without requiring access to model internals. The approach is designed to satisfy existing regulatory obligations and shift accountability to developers, making AI risk regulation an engineering practice rather than a conceptual aspiration.
- ResearchSustainability2026-04-23WEQP
The Impact of Artificial Intelligence on the New Quality Transformation of Chinese Manufacturing · Sirui Dong, Lei Lei, Haonan Chen
This study uses a multi-sectoral equilibrium model and empirical tests to show that AI positively drives quality transformation in Chinese manufacturing, with a one-unit increase in a firm's AI level associated with a 0.171-unit increase in qualitative transformation. The effects vary by firm size, digitalization level, industry technology intensity, and regional factors such as geography and urban agglomeration. AI promotes transformation through three mechanisms: reducing market penetration costs, enhancing innovation capacity and core technologies, and improving resource utilization and operational efficiency. The findings are relevant to enterprise competitiveness, industrial policy, and sustainable manufacturing development in China.
- ResearcharXiv2026-04-22QP
Dialect vs Demographics: Quantifying LLM Bias from Implicit Linguistic Signals vs. Explicit User Profiles · Irti Haq, Belén Saldías
This study examines how Large Language Models treat users differently depending on whether demographic identity is stated explicitly or signaled implicitly through dialect (e.g., AAVE, Singlish). Using a factorial design with over 24,000 responses from two open-weight LLMs, the researchers find that explicitly stated Black identity triggers aggressive safety filters and higher refusal rates, while implicit dialect cues produce a 'dialect jailbreak' that reduces refusal probability to near zero but exposes dialect speakers to less sanitized and potentially more hostile content. This bifurcated user experience reveals that current safety alignment techniques are brittle and over-indexed on explicit keywords, creating inequitable outcomes across linguistic communities. The findings highlight a fundamental tension between safety alignment and linguistic diversity, underscoring the need for safety mechanisms that generalize beyond explicit demographic cues.
- ResearcharXiv2026-04-22QP
"This Wasn't Made for Me": Recentering User Experience and Emotional Impact in the Evaluation of ASR Bias · Siyu Liang, Alicia Beckford Wassink
This paper investigates how bias in Automatic Speech Recognition (ASR) systems affects users from underrepresented English dialect communities beyond simple error-rate metrics. User experience studies conducted across four U.S. locations (Atlanta, Gulf Coast, Miami Beach, and Tucson) found that most participants felt ASR technologies failed to account for their cultural backgrounds, requiring constant adjustment just to achieve basic functionality. Participants reported frustration, feelings of personal inadequacy, and performed significant 'invisible labor'—including code-switching and hyper-articulation—to compensate for system failures, even while recognizing those failures stemmed from biased system design. The study argues that algorithmic fairness assessments relying solely on accuracy metrics miss critical harms including emotional labor, cognitive burden, and psychological toll on speakers whose language varieties are marginalized by these technologies.