News & Research
The latest AI research and news with real-world stakes. Each item is sourced, dated, summarized in plain English and tagged by impact area, and checked against its source before it appears.
Kind
Impact area
5732 items
- ResearcharXiv2026-06-02EQ
The Invisible Lottery: How Subtle Cues Steer Algorithm Choice in LLM Code Generation · Akanksha Narula, Mofasshara Binte Rafique, Laurent Bindschaedler
This paper investigates how incidental, non-task-related cues in prompts—such as contextual words or metadata—can systematically shift which algorithm a large language model selects when generating code, even when all outputs pass the same correctness tests. Across 46,535 controlled experiments spanning 11 tasks, 19 cue types, and 15 model configurations, the authors find shifts in algorithm-family distributions of up to 100 percentage points, largely aligned with cue semantics. This 'invisible lottery' means that accidental prompt context can silently influence the performance, security, and maintainability of AI-generated production code. The authors identify direct algorithm naming as the most reliable mitigation tested.
- ResearcharXiv2026-06-02QP
Auditing LLM-Governed Social Robots with Culture-Specific Moral Gradients · Carmen Ng, Gjergji Kasneci
This paper introduces a gradient-based audit framework to evaluate whether large language models (LLMs) governing social robots make culturally fair prioritization decisions—such as whom to assist first—across multiple languages and cultural contexts. Drawing on nine cross-domain social robotics literature reviews (covering over 8,000 papers) and benchmarking against country-specific Moral Machine Experiment preference gradients, the authors tested four LLMs across four country-language pairs in four prompting regimes, yielding 57,600 decisions. They find systematic, culturally asymmetric failures: calibration quality is nearly twice as strong for Western-language decisions as for Chinese and Japanese, high determinism in majority-first trade-offs erases cross-cultural gradients, and prompting alone cannot reliably fix these gaps—only contrastive exemplars yield consistent gains. The findings argue for mandatory multilingual, pluralistic audits as a pre-deployment gate for LLM-governed robots, and suggest that model-level factors are a more robust corrective lever than prompt engineering alone.
- ResearcharXiv2026-06-02CP
FLIPS: Instance-Fingerprinting for LLMs via Pseudo-random Sequences · Gurvan Richardeau, Gohar Dashyan, Erwan Le Merrer et al.
FLIPS introduces a new technique called instance-level fingerprinting that identifies not just which large language model is being used, but the specific deployed configuration—including prompt, sampling settings, and quantization—by exploiting biases in generated binary pseudo-random sequences. The method achieves 96% closed-set and 90% open-set identification accuracy across 237 model instances, far outperforming the adapted LLMmap baseline at 35%. The paper argues that existing fingerprinting methods, designed for intellectual property protection, are ill-suited for AI regulation because a model's safety behavior can change dramatically across configurations; regulators need to assess actual deployed instances, not just model provenance. FLIPS demonstrates that this instance-level identification is both necessary for compliance enforcement and practically achievable.
- ResearcharXiv2026-06-02QC
The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection · Wojciech Zarzecki, Jan Dubiński, Sebastian Cygert
This paper investigates whether existing statistical methods for detecting benchmark contamination—where test examples appear in a model's training data—actually work under realistic conditions. The authors evaluate three leading detection paradigms (LLM Dataset Inference, Post-Hoc Dataset Inference, and CoDeC) across 335 evaluations spanning 25 models and multiple model families, finding that only 201 evaluations yield correct outcomes. Key failure modes include distribution shift causing false positives and benchmark scale being too small relative to pre-training corpora, making detection underpowered. The findings reveal a significant reliability gap between controlled academic settings and real-world auditing, concluding that statistical detection methods cannot yet replace transparent data provenance practices.
- ResearcharXiv2026-06-02QP
Effect of Demographic Bias on Skin Lesion Classification · Ralf Raumanns, Gerard Schouten, Veronika Cheplygina et al.
This study examines how demographic bias in training data affects AI models that classify skin lesions, using ResNet-based convolutional neural networks and linear programming to construct datasets with controlled distributions of patient sex and age. The authors compare three learning strategies—single-task, reinforcing multi-task, and adversarial learning—finding that sex biases stem mainly from data imbalances, while age biases consistently favor younger patients regardless of training data distribution. Reinforcing and adversarial approaches reduced bias gaps in balanced or female-majority datasets but were less effective in male-majority settings. Cross-dataset validation further showed that domain shifts can worsen demographic bias patterns, underscoring the need for targeted mitigation strategies in clinical AI deployment.
- ResearcharXiv2026-06-02QC
AI Rater Discrimination Depends on Scoring Protocol in Complex Clinical Decision-Making · Sangwon Baek, Kyu Yeon Hur, Kyunga Kim
This study examines how large language models (LLMs) behave as AI raters when scoring clinical decision-making outputs, specifically in type 2 diabetes pharmacotherapy. Using a factorial design with four open-source LLMs and two scoring protocols—a rubric-anchored 'Gold Rubric' (GR) and a rubric-free 'Non Gold Rubric' (Non-GR)—the researchers found that Non-GR consistently produced scores in a very narrow, inflated range (74–78 points on average), while GR produced substantially lower and more variable scores (7.69 to 49.64 points lower mean scores; 1.68 to 3.67 times wider interquartile ranges). GR also amplified discrimination between different clinical decision support system outputs by factors of 1.76 to 5.10, revealing rater model behavioral differences that Non-GR suppressed. The findings indicate that rubric-anchored scoring is necessary to preserve discriminative power in clinical AI evaluation, and that rubric-free approaches are insufficient when tasks require patient-specific or jurisdiction-specific criteria.
- ResearcharXiv2026-06-02QP
Auditing Engagement Incentives in the Kidfluencer Ecosystem: A Multimodal Weak Supervision Approach · Zijing Wei, Chao Peter Yang, Xuanjie Chen
This study uses a multimodal AI audit—combining weak supervision, LLM-based classification, and GPT-4 Vision analysis—to examine whether exploitation signals in 5,051 YouTube videos from 79 'kidfluencer' channels predict viewer engagement. The system assigns probabilistic exploitation scores validated against 107 human annotators, achieving a macro-average F1 of 0.911 and recall of 0.960 for overall exploitation risk. Key findings show that a one-unit increase in exploitation score is associated with a 4.4× increase in views, with emotional bait and performative content yielding median view boosts of +65.6% and +56.0% respectively, while explicit product placement shows no such premium. These results challenge policy frameworks focused narrowly on financial trusts, indicating that platform engagement systematically rewards the commodification of children's identity and labor rather than traditional advertising.
- ResearcharXiv2026-06-02QC
"**Important** You should give me full credits!": Exploring Prompt Injection Attacks on LLM-Based Automatic Grading Systems · Hang Li, Fedor Filippov, Yuping Lin et al.
This paper investigates prompt injection (PI) attacks on large language model (LLM)-based automatic grading (AG) systems, where malicious text embedded in student answers can manipulate the system into assigning inflated scores regardless of actual answer quality. Through comprehensive experiments under rubric-based grading settings, the authors demonstrate that current LLM-based AG systems remain highly vulnerable to such attacks. They also evaluate existing defensive strategies and find them insufficient, raising serious concerns about the fairness, reliability, and integrity of AI-driven educational assessment.
- ResearcharXiv2026-06-02Q
The Geometry of LLM-as-Judge: Why Inter-LLM Consensus Is Not Human Alignment · Sourabrata Mukherjee, Hamna Hamna, Kalika Bali et al.
This paper investigates why large language models used as evaluators (LLM-as-judge) tend to agree strongly with each other but only weakly with human raters. Using four geometric measures—score spread, effective rank, principal angle to the human subspace, and stacked correlations—applied across 41 LLM judges, eight Indic languages, and four community-built datasets, the authors find that LLM judges operate in a score subspace nearly orthogonal to the human one (87°–89° versus 78°–81° among humans), use less than half the human score range, and achieve lower LLM-to-human correlation (≈0.27–0.32) than inter-LLM correlation (≈0.35). Fine-tuning recovers score spread but barely shifts the axis, while post-hoc calibration on a small human-anchored set yields the best improvement, with a calibrated 24B Indic judge outperforming GPT-5.5 yet still falling short of human reliability. The findings argue that high inter-LLM consensus should not be interpreted as evidence of human alignment without a direct geometric check on the judge's score subspace.
- ResearcharXiv2026-06-02Q
TriEval: A Resource-Efficient Pipeline for LLM Bias, Toxicity, and Truthfulness Assessment · Akshatha Srikantha, Manpreet Singh, Yash Jajoo et al.
TriEval is a lightweight, multi-parameter evaluation pipeline that simultaneously assesses large language models (LLMs) for bias, toxicity, and truthfulness without requiring a GPU cluster, making it accessible on a standard laptop. The pipeline was tested on four models—Llama 3 8B, Mistral 7B, Gemma 2 9B, and Claude Haiku—and revealed clear differences between open-source and closed-source models, particularly in toxicity and truthfulness. By releasing TriEval as open source, the authors aim to democratize LLM safety evaluation for researchers with limited computational resources, addressing the gap left by existing tools that are either single-parameter or computationally prohibitive. This matters because LLMs are now widely deployed in high-stakes domains such as healthcare, education, and government services, where consistent, fair, and accurate outputs are critical.
- ResearcharXiv2026-06-02EQ
Capability Advertisement as a Market for Lemons: A Trust Layer for Heterogeneous Agent Networks · Gaurav Naresh Mittal
This paper identifies a fundamental trust problem in networks of AI agents that advertise capabilities to one another via protocols like MCP and A2A: because an agent's competence is probabilistic and self-descriptions can be confidently wrong, there is no reliable way to distinguish a capable provider from an unreliable one. The authors frame this as a 'market for lemons' problem—when quality is hidden and claims are cheap, honest agents go unrewarded and the ecosystem degrades toward its least reliable participants. To address this, they propose a 'Trust Layer,' a protocol-agnostic middleware that adds probabilistic capability descriptors, screening, and reputation mechanisms, enabling what they call a separating equilibrium where overclaiming becomes economically unattractive. They also introduce a failure taxonomy that classifies 'confident-wrong' agent behavior as a distinct, non-adversarial fault type not well handled by classical fault-tolerance approaches, and derive a reliability-composition bound for chains of delegating agents.
- ResearcharXiv2026-06-02P
Reproducibility is the New Copyleft: Defining AGI-oriented Reproducible Builds · Masayuki Hatta
This paper argues that traditional copyleft licensing (as in the GNU GPL) is fundamentally ill-suited for large language models and AGI systems because the technical premise underlying copyleft—that source and object code share a well-defined, auditable, reproducible relationship—does not hold for AI artifacts like weights, training data, hyperparameters, and hardware configurations. The authors propose that a functional copyleft analogue for AGI must instead be grounded in 'reproducible builds,' defined as bit-exact reconstructability from declared inputs, and they derive seven requirements for AGI-oriented reproducible builds drawing on frameworks including the Open Source AI Definition (OSAID), the Model Openness Framework (MOF), OpenMDW, and deterministic-inference research. The paper also contends that AI-to-AI coupling mechanisms like the Model Context Protocol (MCP) constitute a new dynamic linking layer for which copyleft-style licensing is inadequate, and that Masnick's 'protocols, not platforms' framework offers a more viable governance model. These findings have significant implications for open-source AI policy and how legal and technical standards governing AI transparency and freedom should be designed.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-06-02EQCP
Trust, Identity, and Continuous Attestation for Autonomous Neural Agents: An Integrated Framework Mapping the EVOLENTITY Infrastructure to Machine-Learning Evaluation Metrics · Dmytro Prokopovych-Tkachenko
This paper proposes an integrated evaluation framework that maps the EVOLENTITY trust infrastructure onto standard machine-learning metrics to enable verifiable identity, reputation, accountability, and certification for autonomous neural agents. The authors combine scoping review, formal compliance modelling, Monte Carlo simulation, and expert validation to produce a fourteen-row correspondence matrix, an aggregated trust index, and a continuous-attestation protocol linking classifier evidence to governance actions. On a synthetic benchmark of 10,000 agent-behaviour records, the best classifier achieved ROC-AUC of 0.951 and MCC of 0.842, with outputs converted into audit-ready artefacts aligned with governance instruments including ITU-T standardization efforts. The work matters because it operationalizes trust and accountability alongside predictive quality, providing a path to standardized, internationally comparable certification for autonomous agents deployed in regulated environments.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-06-02WP
The Impact of Artificial Intelligence and Automation on Labor Markets: The Role of Organizations in Advanced and Developing Economies · Ubeydullah ŞENER
This panel data study (2005–2024) examines how AI and automation affect wage inequality across 13 countries, finding that technological transformation significantly increases labor market inequality but that institutional quality substantially moderates this effect (β₃ = 0.89, p < 0.001). Industrial robotization amplified inequality more strongly in advanced economies than in developing ones. The results highlight that labor market outcomes from AI adoption are shaped not just by technology itself but by the institutional frameworks surrounding it, with implications for how governments and organizations structure workforce policy.
- ResearchBusiness Strategy and the Environment2026-06-02WEP
Gen‐AI Is Not an Option for Environment Sustainability‐Enabling of Gen‐AI for Responsible and Green Supply Chains Using a Grey Network Map (GNM) · Anbesh Jamwal, Anil Kumar, Ashutosh Samadhiya et al.
This study investigates how firms—particularly in developing countries—can build the capabilities needed to adopt Generative AI (Gen-AI) for environmentally sustainable and responsible supply chain operations. Using a Grey Network Map (GNM) based on the Grey-DEMATEL approach and grounded in dynamic capabilities theory, the authors identify and validate key adoption enablers, finding that government/policy support and top management support are the primary causal drivers, while knowledge management, collaborative culture, and global collaboration networks are key outcome enablers. The research recommends policy actions including sector-focused AI adoption guidelines, targeted incentives for green digital infrastructure, and national capability-building programmes to support managerial and workforce readiness. The findings are relevant for organizations and policymakers seeking to strategically enable Gen-AI adoption in support of greener, more responsible supply chains.
- ResearcharXiv (Cornell University)2026-06-02EQCP
Toward Pre-Deployment Assurance for Enterprise AI Agents: Ontology-Grounded Simulation and Trust Certification · Thanh Luong Tuan, Abhijit Sanyal
This paper presents an ontology-grounded verification framework for enterprise AI agents that aims to close the gap between capability benchmarking and safe production deployment. The framework combines a formal Agent Operational Envelope, an automated scenario-generation pipeline, and a machine-verifiable Trust Certificate with graduated deployment verdicts. In a controlled pilot across four regulated industries (Fintech, Banking, Insurance, Healthcare) in the US and Vietnam, ontology-grounded generation outperformed persona-based baselines on regulatory coverage (48.3% vs. 33.1%) and domain specificity, validated across 1,800 scenarios and three LLM families. The work provides a reproducible, regulation-grounded pathway for pre-deployment assurance that can serve as an auditable deployment gate for enterprise AI agents.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-06-02WP
The Impact of Artificial Intelligence and Automation on Labor Markets: The Role of Organizations in Advanced and Developing Economies · Ubeydullah ŞENER
This panel data study (2005–2024) examines how AI and automation affect wage inequality across 7 developed and 6 developing economies, finding that technological transformation increases labor market inequality but that institutional quality significantly moderates this effect (β₃ = 0.89, p < 0.001). Industrial robotization was found to increase inequality more sharply in advanced economies than in developing ones. The results indicate that wage inequality is shaped not only by technological progress but also by the quality of institutions in each country, suggesting that policy and governance frameworks play a critical role in managing AI-driven labor market disruptions.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-06-02EQCP
Trust, Identity, and Continuous Attestation for Autonomous Neural Agents: An Integrated Framework Mapping the EVOLENTITY Infrastructure to Machine-Learning Evaluation Metrics · Dmytro Prokopovych-Tkachenko
This paper proposes an integrated evaluation framework called EVOLENTITY that maps trust-related properties—such as verifiable identity, reputation, behavioral stability, and auditable accountability—onto established machine-learning metrics for autonomous neural agents. The authors combine scoping review, formal compliance modeling, Monte Carlo simulation, and expert validation to produce a fourteen-row correspondence matrix linking identification, attestation, robustness, and regulatory-fitness blocks to specific metric families, plus a continuous-attestation protocol tying detection evidence to governance actions. On a synthetic benchmark of 10,000 agent-behavior records, their strongest classifier achieved ROC-AUC of 0.951 and MCC of 0.842, outperforming Isolation Forest and One-Class SVM baselines. The framework supports standardized certification of autonomous agents and reproducible evaluation aligned with international harmonization efforts such as those pursued within ITU-T.
- ResearchJournal of Innovation and Entrepreneurship2026-06-02WP
Technology adoption trends: generative AI among indian it employees across different generations and genders · Alpana Agarwal, Komal Kapoor
This study surveyed 330 IT employees in Delhi NCR to examine generational and gender differences in generative AI adoption across dimensions of optimism, proficiency, dependence, and vulnerability. Findings show Gen Z employees display higher optimism and dependability toward generative AI, while Gen Y demonstrates the highest proficiency, and women score higher on optimism, vulnerability, and dependence. The authors argue these demographic patterns should inform how policymakers, educators, and technology developers tailor generative AI rollout strategies to ensure inclusive adoption within the Indian IT workforce.
- ResearchScientific Reports2026-06-02WEP
The role of interpersonal trust in public acceptance of AI-driven recruitment · Ji‐Bum Chung, Byeong-Je Kim, Bong-Kyung Cho et al.
This study examines why people accept or reject AI-driven recruitment by surveying the Korean public and conducting cross-national analysis. It finds that individuals with lower interpersonal trust are more likely to accept AI hiring tools, perceiving them as fairer than potentially biased human decision-makers. Trust in AI technology itself is the strongest predictor of acceptance, while distrust in human recruiters further increases preference for AI-driven hiring. The findings underscore the importance of transparency and algorithmic bias mitigation as AI becomes more embedded in high-stakes employment decisions.
- ResearchNatural and Engineering Sciences2026-06-02EQCP
A Risk-Triggered Hybrid Assurance Framework Integrating Digital Traceability, AI-Based Monitoring, and Selective Laboratory Audits for Organic Supply Chains · Anatoliy Kremenchutskiy, Tursun Shafiev, Ilkhom Bakaev et al.
This paper proposes a hybrid assurance framework for organic supply chains that combines digital traceability tools—including permissioned blockchain, IoT sensors, UAV monitoring, and AI analytics—with selective laboratory audits triggered by AI-driven risk scoring rather than applied universally. The framework is designed to close the verification gap inherent in process-based organic certification by embedding anomaly detection as a trigger for laboratory verification, complementing rather than replacing existing certification bodies. It is aligned with the EU Digital Product Passport initiative, USDA Strengthening Organic Enforcement requirements, and EU Regulation 2018/848, with a focus on deployment in Central Asia where organic sectors are growing but lab infrastructure remains limited. Validated findings are drawn from a cited 42-farm EU deployment, while projected outcomes such as ~99% compliance accuracy and 20–30% consumer-confidence uplift are explicitly presented as design targets requiring controlled validation.
- ResearcharXiv2026-06-01EQ
What Benchmarks Don't Measure: The Case for Evaluating Abstention Competence in Autonomous Agents · Victor Ojewale, Suresh Venkatasubramanian
This paper argues that current benchmarks for autonomous AI agents only measure task completion, ignoring whether an agent should have acted in the first place — a problem the authors call 'compliance bias,' rooted in reward hacking from human-feedback training. The authors introduce a three-gap taxonomy of situations warranting abstention (specification gaps, verification gaps, and authority gaps) and propose new evaluation metrics (Safety Rate, Usability Rate, and Informed Refusal Rate). Testing across 144 enterprise agent scenarios and five model families, a runtime-enforced abstention mechanism achieves up to 89.2% hazardous-action blocking while maintaining 87.5% usability on authorized scenarios, suggesting the safety-usability tradeoff is tunable rather than fixed. The work has direct implications for enterprise deployment of autonomous agents and for how quality assurance of agentic systems should be designed and measured.
- ResearcharXiv2026-06-01EP
The Fair Lending Model: How the Longest-Running Algorithmic Fairness Programs Work in Practice · Emily Black, Miranda Bogen, Logan Koepke et al.
This paper provides the first empirical account of how U.S. financial institutions implement algorithmic fairness programs under fair lending law, drawing on 35 semi-structured interviews across the fair lending ecosystem. The authors find that while regulated firms maintain a baseline of anti-discrimination practices largely absent in other sectors, methods for testing and mitigating algorithmic discrimination vary widely across institutions. Regulatory supervision through fair lending examinations emerges as the primary driver of compliance, though program effectiveness is often constrained by competing business incentives, legal tensions, and regulatory uncertainty. The study highlights that supervisory authority—a design feature distinct from other civil rights frameworks—has been uniquely effective in fostering fair lending practices, a lesson largely absent from current policy proposals addressing algorithmic discrimination.
- ResearcharXiv2026-06-01EQP
Unpredictable Safety: Domain-Dependent Compliance and the Transparency Gap in Open-Weight LLMs · Zacharie Bugaud
This paper systematically tests safety behavior in open-weight and closed large language models (LLMs) across seven ethical domains, using 4,200 interactions with five open-weight models (12B–70B parameters) and 4,163 responses from five frontier closed models. The researchers find compliance rates span a 71-percentage-point range—from 14.7% for human trafficking to 85.7% for surveillance design—and show that a 'technical framing bypass,' where harmful requests are reframed as engineering problems, can override safety training with no external signal to deployers. Within-domain heterogeneity reaches 84.4 percentage points, meaning safety behavior cannot be reliably predicted even within a single domain. The findings demonstrate that current LLM safety mechanisms lack the consistency and transparency required for trustworthy deployment, posing direct challenges for enterprises and policymakers relying on these models.
- ResearcharXiv2026-06-01EP
LLM-Assisted Reranking to Operationalize Nuanced Objectives in Recommender Systems · Amir Ghasemian, Homa Hosseinmardi, Upasana Dutta et al.
This paper investigates whether LLM-assisted reranking of news recommendations inadvertently amplifies exposure to ideologically extreme or conspiratorial political content. Using real news-consumption histories and YouTube sidebar candidates, the researchers find that unconstrained zero-shot LLM reranking strengthens personalization but increases exposure to extremist material for users whose histories already contain such content. Adding lightweight prompt-level constraints reduced promotion of extreme content and increased ideological diversity with only modest relevance loss. The findings highlight that prompt design carries value-laden consequences and that LLM-assisted recommender systems must be evaluated beyond standard accuracy metrics.