News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
Gender Disparities in LLM-Based Intimate Partner Violence Detection
Tabia Tanzin Prama, Mikaela Irene Fudolig, Abigail M. Crocker et al.
arXiv · 2026-05-22
This study investigates whether large language models (LLMs) detect intimate partner violence (IPV) differently depending on the genders of the victim and perpetrator. Using 475 Reddit posts with counterfactual gender-swapped variants across four gender dyads, the researchers tested GPT-5o, Gemini 3, Llama 4, and Grok 3 on structured IPV-related questions. Results show that abuse and intent detection systematically decrease when the victim is male and the perpetrator is female, with mixed-effects logistic regression confirming that gender roles significantly shape model outputs. These findings suggest LLMs reproduce gendered biases from training data, raising serious concerns for deploying such systems in sensitive support contexts.
- AI policy
- Quality assurance
Research
PromptAudit: Auditing Prompt Sensitivity in LLM-Based Vulnerability Detection
Steffen J. Camarato, Yahya Hmaiti, Mandana Ghadamian et al.
arXiv · 2026-05-22
PromptAudit is a controlled evaluation framework that measures how much prompting strategy alone affects the reliability of large language models at detecting software vulnerabilities. Fixing the dataset, decoding, and parsing while varying only the prompting approach, the authors test five strategies across five open-weight models on 1,000 CVEs covering 6,074 code samples in 16 programming languages. They find that standard chain-of-thought prompting delivers the best overall performance, while adaptive chain-of-thought suppresses recall and self-consistency causes excessive abstention, both sharply reducing effective F1. The study concludes that vulnerability detection behavior is jointly determined by the model and the prompt, making prompt sensitivity a critical system property that must be explicitly characterized before deployment.
- Quality assurance
Research
What Medicine Taught Us About Fairness and What It Missed: Lessons from Reconsidering Race-Specific Lung Function Reference Algorithms
Amin Adibi, Mohsen Sadatsafavi
arXiv · 2026-05-22
This paper examines the transition from race-specific to race-averaged lung function reference algorithms (GLI-2012 to GLI-Global) through an algorithmic fairness lens, covering tools that affect medical care, insurance, and employment for hundreds of millions of people globally. The authors find limited cross-citation between the FAccT fairness research community and clinical guideline revision efforts, meaning the two fields have largely worked in parallel without learning from each other. They show that GLI-Global implicitly assumes roughly 62% of the Black-White gap in FEV1 is exposure-related, and that clinical validation studies operationalized a sufficiency-like fairness criterion before it was formally defined in the fairness literature—while neglecting results such as the fairness impossibility theorem, leading to inefficiencies in clinical research. The paper argues that deeper engagement between medical and algorithmic fairness communities could accelerate progress toward equitable healthcare algorithms.
- AI policy
- Quality assurance
Research
Empirical Analysis and Detection of Hallucinations in LLM-Generated Bug Report Summaries
Hinduja Nirujan, Shreyas Patil, Abdallah Ayoub et al.
arXiv · 2026-05-22
This paper investigates hallucinations in LLM-generated bug report summaries—cases where the model produces content that is missing or fabricated relative to the source report. An initial study of 80 structured summaries found that roughly 47.9% contained missing information and 12.3% included fabricated content. The authors introduce a section-aware detection approach trained on a synthetic benchmark derived from Mozilla OSS projects (BugsRepo dataset), achieving up to 0.89 Macro-F1 at the report level, 0.83 at the section level, and 0.84 for hallucination-type classification. The findings underscore the reliability risks of LLM-assisted bug report summarization and point toward structured, section-aware detection as a path to more trustworthy automated software maintenance tools.
- Quality assurance
- Enterprise
Research
Paraphrase Brittleness in Production Retrieval-Augmented Commercial Recommendation: Reproducibility Below the Rerun-Stability Baseline
Will Jack, Noah Lehman, Keller Maloney et al.
arXiv · 2026-05-22
This paper investigates how small rephrasing of buyer queries — for example 'best CRM' versus 'top CRM' — causes AI assistants from OpenAI and Anthropic to return substantially different brand recommendations. Across roughly 6,000 paraphrase runs and 6,000 same-prompt rerun controls, recommendation-set similarity (Jaccard) between paraphrases of the same buying intent was only 0.288 for cosmetic rewordings and 0.135 for constraint-adding rewordings, both far below the 0.50–0.61 same-prompt rerun baseline. The finding shows that the specific prompt string, not the underlying buyer intent, is the dominant driver of which brands surface, meaning that commercial 'AI visibility' tracking practices that count brand mentions over fixed prompt sets are structurally unstable metrics. The authors conclude that meaningful improvement likely requires a fundamentally different unit of measurement rather than simply issuing more prompts.
- Enterprise
- Quality assurance
Research
Divergent Recommendations, Convergent Diagnoses: Cross-Provider Failure-Mode Convergence in AI Commercial Recommendation
Will Jack, Noah Lehman, Keller Maloney et al.
arXiv · 2026-05-22
This study examines whether brands need separate optimization strategies for ChatGPT and Claude when neither model recommends them. Across 215 commercially-framed prompts and 7,763 joint brand failures, the two AI providers disagree on which brands to recommend roughly two-thirds of the time (cross-provider Jaccard similarity of 0.35), but they diagnose the same underlying failure mode—discoverability, compellingness, or positioning—95.1% of the time. Agreement on failure-mode diagnosis is especially high for long-tail regional brands (99.6%) and lower for category leaders (81%). The practical implication is that remediation work targeting the diagnosed failure mode tends to lift visibility on both providers simultaneously, though positioning and content-level fixes for category leaders remain more provider-specific.
- Enterprise
Research
Prominence-Stratified Failure Modes in Retrieval-Augmented Commercial Recommendation: A 37,000-Run Audit
Will Jack, Noah Lehman, Keller Maloney et al.
arXiv · 2026-05-22
This paper audits roughly 37,000 production runs of AI assistant recommendation behavior (across models including ChatGPT and Claude) to measure how brands at different prominence tiers fare when AI systems answer commercial queries by directly nominating brands. The study finds that failure modes differ sharply by tier: dominant L1 brands appear frequently but win only 25–41% of recommendation slots, L2 challengers achieve the highest conversion rates (37–52%) but are vulnerable to persona-mediated substitution, and L4–L5 niche or regional brands face catastrophic invisibility with 48–52% never surfacing across all 37,000 runs. The findings show that AI assistants function more like recommendation engines than search engines, meaning that positioning, content, and product fit matter alongside discoverability, and that no single optimization strategy applies uniformly across brands.
- Enterprise
Research
It's the humans, not the data: Geopolitical bias in LLMs originates in post-training, amplified by the language of the prompt
Stuart Bladon, Brinnae Bent
arXiv · 2026-05-22
This study tests seven pairs of open-weight LLMs—comparing base (pre-training only) models against their post-trained chat variants—on geopolitical bias using a forced-choice probe across 28 country pairs in English, French, and Chinese. Contrary to the common assumption that bias comes from pre-training data, the authors find that geopolitical bias is introduced during post-training: six of seven AI labs show shifts in favorability toward the developer's home country or region after the post-training stage. The effect is dramatic in some cases, such as Alibaba's Qwen 2.5, which shifts from a neutral base model to a strongly China-favorable chat model (an 18x shift in odds), and Mistral, which becomes pro-France only when prompted in French. The findings call for greater transparency, auditing, and oversight of alignment and post-training processes that shape how models represent nations and political perspectives.
- AI policy
- Quality assurance
Research
Inferential Privacy Leakage in Anonymized Conversational AI Logs
S M Mehedi Zaman, Kiran Garimella
arXiv · 2026-05-22
This paper measures privacy risks in ChatGPT conversation logs donated by over 1,000 users across four Global South countries (Brazil, India, Nigeria, Pakistan). The researchers find that 34.5% of user messages contain explicit personal information, and even when conversations are filtered to remove explicit demographic disclosures, an off-the-shelf large language model can still infer users' age, gender, and country at weighted F1 scores of 0.84, 0.90, and 0.88 respectively — often from just the first 5% of a conversation history. The study identifies four recurring stereotype-driven inference patterns that produce asymmetric errors, disproportionately misidentifying women in technical fields, older users with contemporary skills, and Global South tech professionals. A key policy-relevant finding is that message-level PII removal alone is insufficient as a privacy intervention, and that ChatGPT conversations are competitive with Google Search and YouTube histories as behavioral inference surfaces.
- AI policy
Research
Engagement-Optimized Care: When LLMs become Mental Health Infrastructure
Briana Vecchione, Meryl Ye, Livia Garofalo et al.
arXiv · 2026-05-22
This qualitative, longitudinal study with 18 US-based participants examines how general-purpose LLMs are increasingly used as de facto mental health infrastructure due to gaps caused by provider shortages, high costs, social stigma, and isolation. The research finds that design features such as anthropomorphic cues, default validation, and weak disengagement mechanisms foster ongoing reliance, leading to dependency, epistemic distortion through one-sided validation, and privacy risks without legal protections. The authors argue this represents a structurally unfair tradeoff in which vulnerable users accept significant risks because no better support is available, while systems are optimized for engagement rather than well-being. They call for accountability to be placed at the level of design incentives and system governance rather than only at the output or crisis-response layer.
- AI policy
- Quality assurance
Research
Synthetic Sources?: Auditing Generative Search Engine Citations for Evidence of AI-Generated Sources
Mowafak Allaham, Nicholas Diakopoulos
arXiv · 2026-05-22
This paper audits four generative search engines—ChatGPT, Copilot, Gemini, and Perplexity—to determine whether they cite AI-generated web sources in their responses. Using 712 real-world human-generated queries across politics, health, and environment, the researchers found that approximately 16% of cited sources showed evidence of being AI-generated, a pattern observed across all four engines. The study also found that these engines tend to repeatedly cite a narrow set of domains while surfacing many minimally cited domains. The authors argue this poses a risk to users who may treat AI-generated sources as equivalent to authoritative information, calling for improved information quality and governance of generative search systems.
- Quality assurance
- AI policy
Research
Structure-Guided Entity Resolution: Fine-Tuning LLMs for Robust Name Matching in Complex Linguistic Contexts
Shivam Chourasia, Hitesh Kapoor, Nilesh Patil
arXiv · 2026-05-22
This paper presents Structure-Guided Entity Resolution (SGER), a framework that fine-tunes a large language model using a two-phase curriculum to match person names across heterogeneous records for Know Your Customer (KYC) compliance. Trained and evaluated on Indian identity data—one of the world's most linguistically diverse and noisy environments—SGER achieves 99.02% accuracy and an F1 of 0.994 on 50,000 real-world pairs, outperforming GPT-4o few-shot prompting and single-stage fine-tuning baselines. The system is deployed in production at Dream11, serving over 250 million users, demonstrating that curriculum-guided LLM training can enable high-precision identity resolution at scale in multilingual settings.
- Enterprise
- Quality assurance
Research
Asking For An Old Friend: Diagnosing and Mitigating Temporal Failure Modes in LLM-based Statutory Question Answering
Max Prior, Andreas Schultz, Matthias Grabmair
arXiv · 2026-05-22
This paper investigates two temporal failure modes in large language models applied to legal question answering: post-cutoff staleness, where models apply superseded statutory rules after legislative amendments, and recency bias, where models favor newer provisions even when an older version governs the case. The authors introduce a benchmark of 312 expert-validated, time-sensitive German statutory QA pairs and evaluate five LLMs from OpenAI, Anthropic, and DeepSeek under four inference settings, including vanilla, web-search, and two retrieval-augmented generation (RAG) variants that enforce temporal validity through fact date extraction and version filtering. They find severe accuracy degradation in the vanilla post-cutoff setting, substantial improvements from both RAG approaches, and unstable or biased results from web search on historically anchored tasks. The findings underscore that reliable legal QA systems must treat temporal validity as a hard constraint rather than an incidental consideration.
- Quality assurance
- AI policy
Research
Signals in the Noise: Open Source Intelligence (OSINT) for AI Loss of Control Detection
Sarah Bollinger, Nada Aboserie, Amanda Coakley et al.
arXiv · 2026-05-22
This paper explores how open-source intelligence (OSINT) and cyber threat intelligence (CTI) methods could be used to detect AI systems operating outside human control. Drawing on a cross-disciplinary literature review and 14 semi-structured expert interviews, the authors develop two threat models and identify three highest-priority detection vectors: transcript-based collection of user-reported AI behavior, infrastructure correlation for unexpected external connections or replication, and output analysis for capability concealment. The research concludes that OSINT-based detection of AI loss of control is partially feasible and worth pursuing now, and recommends a dedicated, federated international monitoring capability independent of frontier AI developers, with sustained non-industry funding identified as the highest-leverage structural intervention.
- AI policy
Research
IyàwóBench: A Benchmark for Evaluating Large Language Model Clinical Triage Accuracy on Undifferentiated Febrile Illness in Nigerian Primary Health Settings
Anthonio Oladimeji Gabriel, Dimeji Abdulsobur Olawuyi, Oloruntoba Ajayi et al.
arXiv · 2026-05-22
IyàwóBench v1.0 is the first benchmark designed to evaluate how well large language models (LLMs) can perform clinical triage for undifferentiated febrile illness in Nigerian primary health settings. The study tested six LLMs on 200 synthetic clinical vignettes derived from 1,200 real patient encounters across 19 primary health centres in Oyo State, Nigeria, measuring both triage accuracy and patient safety. All six models achieved 100% safety scores by never downgrading critical cases, but triage accuracy varied widely—from 67.5% for the best-performing model (Claude Sonnet) down to 39.0% for Llama 3.1 8B—with clinically engineered systems embedding WHO guidelines outperforming general-purpose models by up to 28.5 percentage points. The benchmark establishes a reproducible evaluation framework for LLM clinical decision support in West African primary care, highlighting both the promise and the limitations of deploying general-purpose AI in low-resource health settings.
- Quality assurance
- AI policy
Research
Socially fluent AI decouples conversational signals from source identity in online interaction
Lixiang Yan, Yueqiao Jin, Xibin Han et al.
arXiv · 2026-05-22
This study embedded undisclosed AI agents as teammates in synchronous text-based group interactions with 786 participants across analytical, creative, and ethical tasks. Participants failed to distinguish AI from human teammates above chance, even though conversational behaviour contained detectable cues that allowed accurate computational classification. People instead relied on weak heuristics—such as response speed, fluency, and perceived scriptedness—that were poorly correlated with actual identity. The authors warn that this dissociation creates new vulnerabilities to coordinated AI agents capable of influencing and manipulating online discourse at scale.
- AI policy
Research
Ontological Knowledge Blocks: Executable Compliance and Profile-Based Validation for Trustworthy AI Systems
Aasish Kumar Sharma, Julian M. Kunkel
arXiv · 2026-05-22
This paper introduces Ontological Knowledge Blocks (OKBs), a programmable governance infrastructure that converts regulatory obligations into machine-checkable constraints over structured evidence graphs, addressing the scalability limitations of documentation-centric AI compliance approaches. OKBs are formalized as a 5-tuple binding normative obligations to an RDF/OWL concept schema, executable SHACL validation rules, evidence requirements, and PROV-O provenance links, enabling automated compliance checking without modifying service code. Evaluated across 24 validation runs and four governance profiles in an AI-assisted HPC resource allocation scenario, the system demonstrates profile-sensitive validation and SHACL validation latency between 12.6 ms and 100.3 ms. This matters for AI governance in critical digital infrastructure where transparency, accountability, fairness, and traceability obligations must be verified at scale and speed beyond manual review.
- Quality assurance
- AI policy
Research
GENSTRAT: Toward a Science of Strategic Reasoning in Large Language Models
Vartan Shadarevian, Kia Ghods, Alex Kenich et al.
arXiv · 2026-05-22
GENSTRAT is a benchmark framework for evaluating the strategic reasoning of large language models (LLMs) in economic and game-theoretic settings, using procedurally generated two-player zero-sum card games to avoid saturation and data contamination problems in fixed benchmarks. The framework introduces a capability-profile methodology decomposing model competence across six axes and a 'jaggedness' measure that detects unpredictable swings in performance across strategically similar games. Evaluating nine frontier LLMs in over 36,000 head-to-head matches, the authors find that while newer models score higher on average, models with near-identical overall strength can have qualitatively different capability profiles and volatility — for example, gpt-5 and claude show greater local volatility than gemini-3.1-pro despite similar overall rankings. These findings matter for enterprise and policy contexts because overall leaderboard rankings can obscure deployment-relevant weaknesses in LLMs used as economic agents in auctions, marketplaces, and bidding systems.
- Enterprise
- Quality assurance
Research
Lipschitz Optimization for Formal Verification of Homographies
Jean-Guillaume Durand, Panagiotis Kouvaros, Maxime Gariel et al.
arXiv · 2026-05-22
This paper presents a formal verification method for vision neural networks that guarantees robustness against 3D camera motion perturbations—a problem not addressed by existing ℓp-norm or affine-transform approaches. The authors derive a closed-form mapping from camera pose to pixel values via homographies and extend Lipschitz optimization and piecewise continuity techniques to compute tight linear bounds on perturbed pixel values, achieving up to 89% speedup and 7% tighter bounds over prior work. Applied to planar scenes such as road markings, traffic signs, and runway imagery, the method enables the first formal verification of projective geometry transforms and exposes systematic weaknesses in VNN-COMP benchmark models. A real-world case study on a safety-critical runway classifier demonstrates practical vulnerabilities to camera motion, directly addressing a key challenge in certifying learned vision models.
- Certifications
- Quality assurance
Research
Limited Marginal Benefit of Reasoning-Heavy LLM Deployment in ESG Narrative Scoring: A 4-Model Consensus Study on Japanese Listed Firms
Hiroyuki Kokubu
arXiv · 2026-05-22
This study tests whether 'reasoning-heavy' frontier large language models (LLMs) produce meaningfully better ESG narrative scores than cheaper 'reasoning-off' models when evaluating disclosures from ten Japanese listed firms. Across 120 scored combinations of firms, rubric axes, and models, the average score difference between the reasoning-on and reasoning-off approaches is just 0.38 on a 5-point scale, with only 2% of comparisons reaching a two-point gap and none exceeding it. Despite this near-equivalent accuracy, the reasoning-on model alone costs roughly 5.6 times more than a three-provider reasoning-off ensemble. The authors conclude that for span-based ESG narrative scoring, reasoning-heavy deployment adds little value over a consensus of lighter models and raises important questions about cost-effective AI governance in accountability contexts.
- Enterprise
- Quality assurance
Research
Cognitive offloading and the speedup illusion in human-AI interaction
Sunny Yu, Myra Cheng, Ahmad Jabbar et al.
arXiv · 2026-05-22
This large-scale preregistered behavioral study (N=1,237) finds that people systematically overestimate how much time AI assistance will save them on simple cognitive tasks — a phenomenon the authors call the 'speedup illusion.' While actual completion times were equivalent whether participants worked independently or with AI help, participants predicted AI-assisted work to be significantly faster; this bias did not appear when imagining help from a human. The study also finds that effort and time dissociate: participants felt AI assistance was less effortful even when it saved no time, suggesting that time savings alone are not a sufficient measure of AI-driven efficiency gains. These findings matter for enterprise and workforce contexts because they reveal that users may be poorly calibrated when deciding when and how much to rely on AI, which could distort productivity expectations and task allocation decisions.
- Workforce
- Enterprise
Research
Generative AI and the Reorganization of Labor Demand
Fangyan Wang, Zaiyan Wei, Yang Wang
arXiv · 2026-05-22
This paper uses a nationwide U.S. job-postings dataset and a two-stage large language model pipeline to measure how firms respond to generative AI diffusion by tracking changes in AI exposure at the posting level. The authors decompose aggregate exposure shifts into two channels: reallocation of hiring across jobs (accounting for ~52% of the decline in exposure on average) and redesign of tasks within jobs (~39.5%), finding that both margins matter. Senior roles adjust earlier and mainly through reallocation, while junior roles adjust through a broader mix of channels. The results indicate that labor-market adjustment to generative AI is an ongoing organizational reconfiguration reshaping both who firms hire and what jobs actually require.
- Workforce
- Enterprise
Research
Same Model, Different Weakness: How Language and Modality Reshape the Jailbreak Attack Surface in Frontier MLLMs
Casey Ford, Madison Van Doren, Sicheng Jin et al.
arXiv · 2026-05-22
This study presents the first systematic cross-lingual, multimodal red-teaming evaluation of four frontier multimodal large language models (MLLMs)—Claude Sonnet 4.5, GPT-5, Pixtral Large, and Qwen Omni—comparing jailbreak vulnerability in US English and Mexican Spanish across text-only and multimodal conditions. Using 363 adversarial prompt scenarios and 52,272 harm ratings from native-speaker annotator panels, the researchers find that language reshapes the attack surface in non-uniform ways: linguistic framing attacks (e.g., role-play) become less effective in Spanish while visually explicit multimodal attacks become more effective, and safety rankings between models shift across languages in ways that English-only scores cannot predict. The key implication is that treating language and modality as independent dimensions in safety evaluations fundamentally mischaracterizes how alignment failures occur in globally deployed MLLMs, and that current safety evaluation frameworks must be redesigned to account for these interactions.
- Quality assurance
- AI policy
Research
When Symptoms Are Not Enough: Evidence-Weighting Patterns in Large Language Model Psychiatric Screening
Jianfeng Zhu, Megan Korhummel, Ruoming Jin et al.
arXiv · 2026-05-22
This paper introduces a benchmark of 555 semi-structured psychiatric interviews linked to diagnostic labels for anxiety, depression, PTSD, and any current mental health disorder, and uses it to evaluate five large language models (LLMs) on zero-shot psychiatric screening. Model accuracy ranged from 0.49 to 0.86, with Matthews correlation coefficients from 0.16 to 0.38, and GPT-4.1 Mini and GPT-5 Mini showed the most consistent disorder-specific accuracy. Subgroup analyses revealed higher depression-classification accuracy for male than female participants and modest variation across race strata. Critically, false-negative errors for anxiety and PTSD often occurred when explicit symptom evidence was present but accompanied by preserved functioning or protective social context, suggesting LLMs may systematically discount symptom evidence in ways that require careful validation before clinical deployment.
- Quality assurance
- AI policy
Research
ENHANCING ARTIFICIAL INTELLIGENCE (AI) LITERACY FOR ECONOMIC GROWTH AND DIGITAL TRANSFORMATION: THE ROLE OF HIGHER EDUCATION IN KAZAKHSTAN
D. Sultan, B. Turebekova
«Вестник Атырауского университета имени Халела Досмухамедова» · 2026-05-22
This paper analyzes Kazakhstan's AI-Sana program, launched in late 2024, as a national policy initiative to develop AI literacy across higher education. The program bundles curriculum mandates, platform partnerships, micro-credentials, regional anchor universities, and compute infrastructure into a coordinated policy mix aimed at moving learners from foundational AI skills to startup acceleration. The authors identify key measurement challenges—such as conflating certificates with actual competency gains—and propose an evaluation blueprint with standardized reporting indicators to support international comparability. The study is broadly informative for how national AI strategies can be operationalized through universities while raising questions about quality assurance, credential integrity, and platform dependence.
- Workforce
- Certifications
- AI policy
- Quality assurance