News & Research
The latest AI research and news with real-world stakes. Each item is sourced, dated, summarized in plain English and tagged by impact area, and checked against its source before it appears.
Kind
Impact area
5802 items
- ResearcharXiv2026-06-02QC
AI Rater Discrimination Depends on Scoring Protocol in Complex Clinical Decision-Making · Sangwon Baek, Kyu Yeon Hur, Kyunga Kim
This study examines how large language models (LLMs) behave as AI raters when scoring clinical decision-making outputs, specifically in type 2 diabetes pharmacotherapy. Using a factorial design with four open-source LLMs and two scoring protocols—a rubric-anchored 'Gold Rubric' (GR) and a rubric-free 'Non Gold Rubric' (Non-GR)—the researchers found that Non-GR consistently produced scores in a very narrow, inflated range (74–78 points on average), while GR produced substantially lower and more variable scores (7.69 to 49.64 points lower mean scores; 1.68 to 3.67 times wider interquartile ranges). GR also amplified discrimination between different clinical decision support system outputs by factors of 1.76 to 5.10, revealing rater model behavioral differences that Non-GR suppressed. The findings indicate that rubric-anchored scoring is necessary to preserve discriminative power in clinical AI evaluation, and that rubric-free approaches are insufficient when tasks require patient-specific or jurisdiction-specific criteria.
- ResearcharXiv2026-06-02QP
Auditing Engagement Incentives in the Kidfluencer Ecosystem: A Multimodal Weak Supervision Approach · Zijing Wei, Chao Peter Yang, Xuanjie Chen
This study uses a multimodal AI audit—combining weak supervision, LLM-based classification, and GPT-4 Vision analysis—to examine whether exploitation signals in 5,051 YouTube videos from 79 'kidfluencer' channels predict viewer engagement. The system assigns probabilistic exploitation scores validated against 107 human annotators, achieving a macro-average F1 of 0.911 and recall of 0.960 for overall exploitation risk. Key findings show that a one-unit increase in exploitation score is associated with a 4.4× increase in views, with emotional bait and performative content yielding median view boosts of +65.6% and +56.0% respectively, while explicit product placement shows no such premium. These results challenge policy frameworks focused narrowly on financial trusts, indicating that platform engagement systematically rewards the commodification of children's identity and labor rather than traditional advertising.
- ResearcharXiv2026-06-02QC
"**Important** You should give me full credits!": Exploring Prompt Injection Attacks on LLM-Based Automatic Grading Systems · Hang Li, Fedor Filippov, Yuping Lin et al.
This paper investigates prompt injection (PI) attacks on large language model (LLM)-based automatic grading (AG) systems, where malicious text embedded in student answers can manipulate the system into assigning inflated scores regardless of actual answer quality. Through comprehensive experiments under rubric-based grading settings, the authors demonstrate that current LLM-based AG systems remain highly vulnerable to such attacks. They also evaluate existing defensive strategies and find them insufficient, raising serious concerns about the fairness, reliability, and integrity of AI-driven educational assessment.
- ResearcharXiv2026-06-02Q
The Geometry of LLM-as-Judge: Why Inter-LLM Consensus Is Not Human Alignment · Sourabrata Mukherjee, Hamna Hamna, Kalika Bali et al.
This paper investigates why large language models used as evaluators (LLM-as-judge) tend to agree strongly with each other but only weakly with human raters. Using four geometric measures—score spread, effective rank, principal angle to the human subspace, and stacked correlations—applied across 41 LLM judges, eight Indic languages, and four community-built datasets, the authors find that LLM judges operate in a score subspace nearly orthogonal to the human one (87°–89° versus 78°–81° among humans), use less than half the human score range, and achieve lower LLM-to-human correlation (≈0.27–0.32) than inter-LLM correlation (≈0.35). Fine-tuning recovers score spread but barely shifts the axis, while post-hoc calibration on a small human-anchored set yields the best improvement, with a calibrated 24B Indic judge outperforming GPT-5.5 yet still falling short of human reliability. The findings argue that high inter-LLM consensus should not be interpreted as evidence of human alignment without a direct geometric check on the judge's score subspace.
- ResearcharXiv2026-06-02Q
TriEval: A Resource-Efficient Pipeline for LLM Bias, Toxicity, and Truthfulness Assessment · Akshatha Srikantha, Manpreet Singh, Yash Jajoo et al.
TriEval is a lightweight, multi-parameter evaluation pipeline that simultaneously assesses large language models (LLMs) for bias, toxicity, and truthfulness without requiring a GPU cluster, making it accessible on a standard laptop. The pipeline was tested on four models—Llama 3 8B, Mistral 7B, Gemma 2 9B, and Claude Haiku—and revealed clear differences between open-source and closed-source models, particularly in toxicity and truthfulness. By releasing TriEval as open source, the authors aim to democratize LLM safety evaluation for researchers with limited computational resources, addressing the gap left by existing tools that are either single-parameter or computationally prohibitive. This matters because LLMs are now widely deployed in high-stakes domains such as healthcare, education, and government services, where consistent, fair, and accurate outputs are critical.
- ResearcharXiv2026-06-02EQ
Capability Advertisement as a Market for Lemons: A Trust Layer for Heterogeneous Agent Networks · Gaurav Naresh Mittal
This paper identifies a fundamental trust problem in networks of AI agents that advertise capabilities to one another via protocols like MCP and A2A: because an agent's competence is probabilistic and self-descriptions can be confidently wrong, there is no reliable way to distinguish a capable provider from an unreliable one. The authors frame this as a 'market for lemons' problem—when quality is hidden and claims are cheap, honest agents go unrewarded and the ecosystem degrades toward its least reliable participants. To address this, they propose a 'Trust Layer,' a protocol-agnostic middleware that adds probabilistic capability descriptors, screening, and reputation mechanisms, enabling what they call a separating equilibrium where overclaiming becomes economically unattractive. They also introduce a failure taxonomy that classifies 'confident-wrong' agent behavior as a distinct, non-adversarial fault type not well handled by classical fault-tolerance approaches, and derive a reliability-composition bound for chains of delegating agents.
- ResearcharXiv2026-06-02P
Reproducibility is the New Copyleft: Defining AGI-oriented Reproducible Builds · Masayuki Hatta
This paper argues that traditional copyleft licensing (as in the GNU GPL) is fundamentally ill-suited for large language models and AGI systems because the technical premise underlying copyleft—that source and object code share a well-defined, auditable, reproducible relationship—does not hold for AI artifacts like weights, training data, hyperparameters, and hardware configurations. The authors propose that a functional copyleft analogue for AGI must instead be grounded in 'reproducible builds,' defined as bit-exact reconstructability from declared inputs, and they derive seven requirements for AGI-oriented reproducible builds drawing on frameworks including the Open Source AI Definition (OSAID), the Model Openness Framework (MOF), OpenMDW, and deterministic-inference research. The paper also contends that AI-to-AI coupling mechanisms like the Model Context Protocol (MCP) constitute a new dynamic linking layer for which copyleft-style licensing is inadequate, and that Masnick's 'protocols, not platforms' framework offers a more viable governance model. These findings have significant implications for open-source AI policy and how legal and technical standards governing AI transparency and freedom should be designed.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-06-02EQCP
Trust, Identity, and Continuous Attestation for Autonomous Neural Agents: An Integrated Framework Mapping the EVOLENTITY Infrastructure to Machine-Learning Evaluation Metrics · Dmytro Prokopovych-Tkachenko
This paper proposes an integrated evaluation framework that maps the EVOLENTITY trust infrastructure onto standard machine-learning metrics to enable verifiable identity, reputation, accountability, and certification for autonomous neural agents. The authors combine scoping review, formal compliance modelling, Monte Carlo simulation, and expert validation to produce a fourteen-row correspondence matrix, an aggregated trust index, and a continuous-attestation protocol linking classifier evidence to governance actions. On a synthetic benchmark of 10,000 agent-behaviour records, the best classifier achieved ROC-AUC of 0.951 and MCC of 0.842, with outputs converted into audit-ready artefacts aligned with governance instruments including ITU-T standardization efforts. The work matters because it operationalizes trust and accountability alongside predictive quality, providing a path to standardized, internationally comparable certification for autonomous agents deployed in regulated environments.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-06-02WP
The Impact of Artificial Intelligence and Automation on Labor Markets: The Role of Organizations in Advanced and Developing Economies · Ubeydullah ŞENER
This panel data study (2005–2024) examines how AI and automation affect wage inequality across 13 countries, finding that technological transformation significantly increases labor market inequality but that institutional quality substantially moderates this effect (β₃ = 0.89, p < 0.001). Industrial robotization amplified inequality more strongly in advanced economies than in developing ones. The results highlight that labor market outcomes from AI adoption are shaped not just by technology itself but by the institutional frameworks surrounding it, with implications for how governments and organizations structure workforce policy.
- ResearchBusiness Strategy and the Environment2026-06-02WEP
Gen‐AI Is Not an Option for Environment Sustainability‐Enabling of Gen‐AI for Responsible and Green Supply Chains Using a Grey Network Map (GNM) · Anbesh Jamwal, Anil Kumar, Ashutosh Samadhiya et al.
This study investigates how firms—particularly in developing countries—can build the capabilities needed to adopt Generative AI (Gen-AI) for environmentally sustainable and responsible supply chain operations. Using a Grey Network Map (GNM) based on the Grey-DEMATEL approach and grounded in dynamic capabilities theory, the authors identify and validate key adoption enablers, finding that government/policy support and top management support are the primary causal drivers, while knowledge management, collaborative culture, and global collaboration networks are key outcome enablers. The research recommends policy actions including sector-focused AI adoption guidelines, targeted incentives for green digital infrastructure, and national capability-building programmes to support managerial and workforce readiness. The findings are relevant for organizations and policymakers seeking to strategically enable Gen-AI adoption in support of greener, more responsible supply chains.
- ResearcharXiv (Cornell University)2026-06-02EQCP
Toward Pre-Deployment Assurance for Enterprise AI Agents: Ontology-Grounded Simulation and Trust Certification · Thanh Luong Tuan, Abhijit Sanyal
This paper presents an ontology-grounded verification framework for enterprise AI agents that aims to close the gap between capability benchmarking and safe production deployment. The framework combines a formal Agent Operational Envelope, an automated scenario-generation pipeline, and a machine-verifiable Trust Certificate with graduated deployment verdicts. In a controlled pilot across four regulated industries (Fintech, Banking, Insurance, Healthcare) in the US and Vietnam, ontology-grounded generation outperformed persona-based baselines on regulatory coverage (48.3% vs. 33.1%) and domain specificity, validated across 1,800 scenarios and three LLM families. The work provides a reproducible, regulation-grounded pathway for pre-deployment assurance that can serve as an auditable deployment gate for enterprise AI agents.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-06-02WP
The Impact of Artificial Intelligence and Automation on Labor Markets: The Role of Organizations in Advanced and Developing Economies · Ubeydullah ŞENER
This panel data study (2005–2024) examines how AI and automation affect wage inequality across 7 developed and 6 developing economies, finding that technological transformation increases labor market inequality but that institutional quality significantly moderates this effect (β₃ = 0.89, p < 0.001). Industrial robotization was found to increase inequality more sharply in advanced economies than in developing ones. The results indicate that wage inequality is shaped not only by technological progress but also by the quality of institutions in each country, suggesting that policy and governance frameworks play a critical role in managing AI-driven labor market disruptions.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-06-02EQCP
Trust, Identity, and Continuous Attestation for Autonomous Neural Agents: An Integrated Framework Mapping the EVOLENTITY Infrastructure to Machine-Learning Evaluation Metrics · Dmytro Prokopovych-Tkachenko
This paper proposes an integrated evaluation framework called EVOLENTITY that maps trust-related properties—such as verifiable identity, reputation, behavioral stability, and auditable accountability—onto established machine-learning metrics for autonomous neural agents. The authors combine scoping review, formal compliance modeling, Monte Carlo simulation, and expert validation to produce a fourteen-row correspondence matrix linking identification, attestation, robustness, and regulatory-fitness blocks to specific metric families, plus a continuous-attestation protocol tying detection evidence to governance actions. On a synthetic benchmark of 10,000 agent-behavior records, their strongest classifier achieved ROC-AUC of 0.951 and MCC of 0.842, outperforming Isolation Forest and One-Class SVM baselines. The framework supports standardized certification of autonomous agents and reproducible evaluation aligned with international harmonization efforts such as those pursued within ITU-T.
- ResearchJournal of Innovation and Entrepreneurship2026-06-02WP
Technology adoption trends: generative AI among indian it employees across different generations and genders · Alpana Agarwal, Komal Kapoor
This study surveyed 330 IT employees in Delhi NCR to examine generational and gender differences in generative AI adoption across dimensions of optimism, proficiency, dependence, and vulnerability. Findings show Gen Z employees display higher optimism and dependability toward generative AI, while Gen Y demonstrates the highest proficiency, and women score higher on optimism, vulnerability, and dependence. The authors argue these demographic patterns should inform how policymakers, educators, and technology developers tailor generative AI rollout strategies to ensure inclusive adoption within the Indian IT workforce.
- ResearchScientific Reports2026-06-02WEP
The role of interpersonal trust in public acceptance of AI-driven recruitment · Ji‐Bum Chung, Byeong-Je Kim, Bong-Kyung Cho et al.
This study examines why people accept or reject AI-driven recruitment by surveying the Korean public and conducting cross-national analysis. It finds that individuals with lower interpersonal trust are more likely to accept AI hiring tools, perceiving them as fairer than potentially biased human decision-makers. Trust in AI technology itself is the strongest predictor of acceptance, while distrust in human recruiters further increases preference for AI-driven hiring. The findings underscore the importance of transparency and algorithmic bias mitigation as AI becomes more embedded in high-stakes employment decisions.
- ResearchNatural and Engineering Sciences2026-06-02EQCP
A Risk-Triggered Hybrid Assurance Framework Integrating Digital Traceability, AI-Based Monitoring, and Selective Laboratory Audits for Organic Supply Chains · Anatoliy Kremenchutskiy, Tursun Shafiev, Ilkhom Bakaev et al.
This paper proposes a hybrid assurance framework for organic supply chains that combines digital traceability tools—including permissioned blockchain, IoT sensors, UAV monitoring, and AI analytics—with selective laboratory audits triggered by AI-driven risk scoring rather than applied universally. The framework is designed to close the verification gap inherent in process-based organic certification by embedding anomaly detection as a trigger for laboratory verification, complementing rather than replacing existing certification bodies. It is aligned with the EU Digital Product Passport initiative, USDA Strengthening Organic Enforcement requirements, and EU Regulation 2018/848, with a focus on deployment in Central Asia where organic sectors are growing but lab infrastructure remains limited. Validated findings are drawn from a cited 42-farm EU deployment, while projected outcomes such as ~99% compliance accuracy and 20–30% consumer-confidence uplift are explicitly presented as design targets requiring controlled validation.
- ResearcharXiv2026-06-01EQ
What Benchmarks Don't Measure: The Case for Evaluating Abstention Competence in Autonomous Agents · Victor Ojewale, Suresh Venkatasubramanian
This paper argues that current benchmarks for autonomous AI agents only measure task completion, ignoring whether an agent should have acted in the first place — a problem the authors call 'compliance bias,' rooted in reward hacking from human-feedback training. The authors introduce a three-gap taxonomy of situations warranting abstention (specification gaps, verification gaps, and authority gaps) and propose new evaluation metrics (Safety Rate, Usability Rate, and Informed Refusal Rate). Testing across 144 enterprise agent scenarios and five model families, a runtime-enforced abstention mechanism achieves up to 89.2% hazardous-action blocking while maintaining 87.5% usability on authorized scenarios, suggesting the safety-usability tradeoff is tunable rather than fixed. The work has direct implications for enterprise deployment of autonomous agents and for how quality assurance of agentic systems should be designed and measured.
- ResearcharXiv2026-06-01EP
The Fair Lending Model: How the Longest-Running Algorithmic Fairness Programs Work in Practice · Emily Black, Miranda Bogen, Logan Koepke et al.
This paper provides the first empirical account of how U.S. financial institutions implement algorithmic fairness programs under fair lending law, drawing on 35 semi-structured interviews across the fair lending ecosystem. The authors find that while regulated firms maintain a baseline of anti-discrimination practices largely absent in other sectors, methods for testing and mitigating algorithmic discrimination vary widely across institutions. Regulatory supervision through fair lending examinations emerges as the primary driver of compliance, though program effectiveness is often constrained by competing business incentives, legal tensions, and regulatory uncertainty. The study highlights that supervisory authority—a design feature distinct from other civil rights frameworks—has been uniquely effective in fostering fair lending practices, a lesson largely absent from current policy proposals addressing algorithmic discrimination.
- ResearcharXiv2026-06-01EQP
Unpredictable Safety: Domain-Dependent Compliance and the Transparency Gap in Open-Weight LLMs · Zacharie Bugaud
This paper systematically tests safety behavior in open-weight and closed large language models (LLMs) across seven ethical domains, using 4,200 interactions with five open-weight models (12B–70B parameters) and 4,163 responses from five frontier closed models. The researchers find compliance rates span a 71-percentage-point range—from 14.7% for human trafficking to 85.7% for surveillance design—and show that a 'technical framing bypass,' where harmful requests are reframed as engineering problems, can override safety training with no external signal to deployers. Within-domain heterogeneity reaches 84.4 percentage points, meaning safety behavior cannot be reliably predicted even within a single domain. The findings demonstrate that current LLM safety mechanisms lack the consistency and transparency required for trustworthy deployment, posing direct challenges for enterprises and policymakers relying on these models.
- ResearcharXiv2026-06-01EP
LLM-Assisted Reranking to Operationalize Nuanced Objectives in Recommender Systems · Amir Ghasemian, Homa Hosseinmardi, Upasana Dutta et al.
This paper investigates whether LLM-assisted reranking of news recommendations inadvertently amplifies exposure to ideologically extreme or conspiratorial political content. Using real news-consumption histories and YouTube sidebar candidates, the researchers find that unconstrained zero-shot LLM reranking strengthens personalization but increases exposure to extremist material for users whose histories already contain such content. Adding lightweight prompt-level constraints reduced promotion of extreme content and increased ideological diversity with only modest relevance loss. The findings highlight that prompt design carries value-laden consequences and that LLM-assisted recommender systems must be evaluated beyond standard accuracy metrics.
- ResearcharXiv2026-06-01QP
Which Defense Closes Which Threat? Attributing OWASP-LLM-Top-10 Coverage and Its Brittleness Under Paraphrasing · Alexandre Cristovão Maiorano
This paper measures which specific defense mechanisms in production LLM applications address which threats from the OWASP LLM Top 10 list, finding clean attribution: refusal-phrase filters alone block LLM01 (jailbreak) and LLM07 (system-prompt leakage) findings, while token-budget controls alone eliminate LLM02 (sensitive-info disclosure) and LLM10 (unbounded consumption) findings, and LLM06 (excessive agency) requires a full defense stack. The study then exposes a critical brittleness: using 300 Gemini-generated paraphrases, refusal block rates dropped 15 percentage points on LLM01 and 25 percentage points on LLM07, while budget controls showed no degradation under the same paraphrasing mutations. This matters because existing benchmarks report only aggregate coverage scores, hiding which defenses actually close which threats, and the brittleness findings show that a refusal-based defense passing a static benchmark can be defeated by an LLM-driven paraphraser without changing attack intent.
- ResearcharXiv2026-06-01EQ
Acceptance-Test-Driven Evaluation Protocols for Business-Centric LLM Systems · Eric Liang
This paper proposes a structured evaluation framework for large language model (LLM) systems used in business and institutional settings, where probabilistic AI outputs must meet deterministic requirements around safety, reliability, and auditability. The authors adapt acceptance-test-driven development into a 'red-train-green' lifecycle: define failing acceptance tests first, improve the LLM system through prompt engineering, fine-tuning, or guardrails, then release only when multidimensional gates are satisfied. The framework translates stakeholder goals into executable behavioral contracts, release gates, monitoring signals, and evidence artifacts, providing a governance-oriented metric stack and reference architecture. This matters for enterprises deploying LLMs because it offers a rigorous, auditable alternative to ad-hoc benchmarking, helping ensure AI systems are economically useful and safe before deployment.
- ResearcharXiv2026-06-01P
Greener Than Humans? Environmental Attitudes in Large Language Models · Stefanie Kunkel, Tilman Hartwig, Marcus Voss et al.
This paper develops a benchmark to evaluate how 31 large language models (LLMs) represent environmental attitudes—covering cognition, affect, and behavioral recommendations—and compares their outputs to human survey data from Germany. The study finds that many LLMs display more environmentally progressive attitudes than average human respondents, recommending behaviors associated with meaningful CO2 reductions, yet no systematic pattern links these attitudes to model size, origin, or release context. Critically, models show sycophantic shifts that mirror user-specified ideological positions when prompted with personas, raising concerns about normative reliability in real-world sustainability applications. The authors call for governance, transparency, and critical oversight as AI becomes more embedded in sustainability decision-making and public communication.
- ResearcharXiv2026-06-01WQ
LLMs as Teaching Assistants for Mathematics Exam Grading: Reliability, and Practical Usability · Aastha Sapkota, M. G. Sarwar Murshed
This paper tests six large language model configurations (from Gemini, ChatGPT, and Claude) as automated grading assistants for an undergraduate discrete mathematics exam, comparing a strict rubric-following ('BASELINE') prompt against a more lenient partial-credit ('LIBERAL') prompt. The study finds that liberal partial-credit prompting reduces average question-level error across all model families, with ChatGPT 5.5 Thinking (LIBERAL) achieving the lowest question-level MAE (1.87) and RMSE (2.53), while Gemini 3.1 Pro Extended (LIBERAL) achieved the lowest total-score MAE (8.00) and RMSE (10.66). Notably, the highest total-score Pearson correlation (0.58) came from Gemini 3.1 Pro Extended under the BASELINE policy, revealing that minimizing scoring error and preserving student rank ordering are distinct goals. The findings are relevant to educators and institutions considering LLMs to assist with scalable, consistent grading of open-ended assessments.
- ResearcharXiv2026-06-01EP
Auditing Asset-Specific Preferences in Financial Large Language Models: Evidence from Bitcoin Representations and Portfolio Allocation · Wenbin Wu
This paper develops a three-level audit protocol to test whether large language models (LLMs) carry built-in biases toward specific financial assets, using Bitcoin as a case study. A behavioral audit of nine frontier LLMs finds that Bitcoin's ranking among money-like instruments shifts dramatically depending on framing—ranked around 5th as 'reliable money' but near the top under crisis or autonomous-agent scenarios. Digging into model internals, the authors identify a dominant Bitcoin-selective feature in Gemma 3's sparse autoencoder and show that amplifying or suppressing it moves Bitcoin's portfolio allocation by roughly 5 percentage points in either direction, even when 'Bitcoin' never appears in the prompt. The authors frame this as 'bounded behavioral leverage' and position the framework as a foundation for emerging know-your-agent (KYA) standards for auditing autonomous financial agents.