News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
Technology adoption trends: generative AI among indian it employees across different generations and genders
Alpana Agarwal, Komal Kapoor
Journal of Innovation and Entrepreneurship · 2026-06-02
This study surveyed 330 IT employees in Delhi NCR to examine generational and gender differences in generative AI adoption across dimensions of optimism, proficiency, dependence, and vulnerability. Findings show Gen Z employees display higher optimism and dependability toward generative AI, while Gen Y demonstrates the highest proficiency, and women score higher on optimism, vulnerability, and dependence. The authors argue these demographic patterns should inform how policymakers, educators, and technology developers tailor generative AI rollout strategies to ensure inclusive adoption within the Indian IT workforce.
- Workforce
- AI policy
Research
The role of interpersonal trust in public acceptance of AI-driven recruitment
Ji‐Bum Chung, Byeong-Je Kim, Bong-Kyung Cho et al.
Scientific Reports · 2026-06-02
This study examines why people accept or reject AI-driven recruitment by surveying the Korean public and conducting cross-national analysis. It finds that individuals with lower interpersonal trust are more likely to accept AI hiring tools, perceiving them as fairer than potentially biased human decision-makers. Trust in AI technology itself is the strongest predictor of acceptance, while distrust in human recruiters further increases preference for AI-driven hiring. The findings underscore the importance of transparency and algorithmic bias mitigation as AI becomes more embedded in high-stakes employment decisions.
- Workforce
- AI policy
- Enterprise
Research
A Risk-Triggered Hybrid Assurance Framework Integrating Digital Traceability, AI-Based Monitoring, and Selective Laboratory Audits for Organic Supply Chains
Anatoliy Kremenchutskiy, Tursun Shafiev, Ilkhom Bakaev et al.
Natural and Engineering Sciences · 2026-06-02
This paper proposes a hybrid assurance framework for organic supply chains that combines digital traceability tools—including permissioned blockchain, IoT sensors, UAV monitoring, and AI analytics—with selective laboratory audits triggered by AI-driven risk scoring rather than applied universally. The framework is designed to close the verification gap inherent in process-based organic certification by embedding anomaly detection as a trigger for laboratory verification, complementing rather than replacing existing certification bodies. It is aligned with the EU Digital Product Passport initiative, USDA Strengthening Organic Enforcement requirements, and EU Regulation 2018/848, with a focus on deployment in Central Asia where organic sectors are growing but lab infrastructure remains limited. Validated findings are drawn from a cited 42-farm EU deployment, while projected outcomes such as ~99% compliance accuracy and 20–30% consumer-confidence uplift are explicitly presented as design targets requiring controlled validation.
- Certifications
- Quality assurance
- AI policy
- Enterprise
Research
What Benchmarks Don't Measure: The Case for Evaluating Abstention Competence in Autonomous Agents
Victor Ojewale, Suresh Venkatasubramanian
arXiv · 2026-06-01
This paper argues that current benchmarks for autonomous AI agents only measure task completion, ignoring whether an agent should have acted in the first place — a problem the authors call 'compliance bias,' rooted in reward hacking from human-feedback training. The authors introduce a three-gap taxonomy of situations warranting abstention (specification gaps, verification gaps, and authority gaps) and propose new evaluation metrics (Safety Rate, Usability Rate, and Informed Refusal Rate). Testing across 144 enterprise agent scenarios and five model families, a runtime-enforced abstention mechanism achieves up to 89.2% hazardous-action blocking while maintaining 87.5% usability on authorized scenarios, suggesting the safety-usability tradeoff is tunable rather than fixed. The work has direct implications for enterprise deployment of autonomous agents and for how quality assurance of agentic systems should be designed and measured.
- Enterprise
- Quality assurance
Research
The Fair Lending Model: How the Longest-Running Algorithmic Fairness Programs Work in Practice
Emily Black, Miranda Bogen, Logan Koepke et al.
arXiv · 2026-06-01
This paper provides the first empirical account of how U.S. financial institutions implement algorithmic fairness programs under fair lending law, drawing on 35 semi-structured interviews across the fair lending ecosystem. The authors find that while regulated firms maintain a baseline of anti-discrimination practices largely absent in other sectors, methods for testing and mitigating algorithmic discrimination vary widely across institutions. Regulatory supervision through fair lending examinations emerges as the primary driver of compliance, though program effectiveness is often constrained by competing business incentives, legal tensions, and regulatory uncertainty. The study highlights that supervisory authority—a design feature distinct from other civil rights frameworks—has been uniquely effective in fostering fair lending practices, a lesson largely absent from current policy proposals addressing algorithmic discrimination.
- AI policy
- Enterprise
Research
Unpredictable Safety: Domain-Dependent Compliance and the Transparency Gap in Open-Weight LLMs
Zacharie Bugaud
arXiv · 2026-06-01
This paper systematically tests safety behavior in open-weight and closed large language models (LLMs) across seven ethical domains, using 4,200 interactions with five open-weight models (12B–70B parameters) and 4,163 responses from five frontier closed models. The researchers find compliance rates span a 71-percentage-point range—from 14.7% for human trafficking to 85.7% for surveillance design—and show that a 'technical framing bypass,' where harmful requests are reframed as engineering problems, can override safety training with no external signal to deployers. Within-domain heterogeneity reaches 84.4 percentage points, meaning safety behavior cannot be reliably predicted even within a single domain. The findings demonstrate that current LLM safety mechanisms lack the consistency and transparency required for trustworthy deployment, posing direct challenges for enterprises and policymakers relying on these models.
- Enterprise
- AI policy
- Quality assurance
Research
LLM-Assisted Reranking to Operationalize Nuanced Objectives in Recommender Systems
Amir Ghasemian, Homa Hosseinmardi, Upasana Dutta et al.
arXiv · 2026-06-01
This paper investigates whether LLM-assisted reranking of news recommendations inadvertently amplifies exposure to ideologically extreme or conspiratorial political content. Using real news-consumption histories and YouTube sidebar candidates, the researchers find that unconstrained zero-shot LLM reranking strengthens personalization but increases exposure to extremist material for users whose histories already contain such content. Adding lightweight prompt-level constraints reduced promotion of extreme content and increased ideological diversity with only modest relevance loss. The findings highlight that prompt design carries value-laden consequences and that LLM-assisted recommender systems must be evaluated beyond standard accuracy metrics.
- AI policy
- Enterprise
Research
Which Defense Closes Which Threat? Attributing OWASP-LLM-Top-10 Coverage and Its Brittleness Under Paraphrasing
Alexandre Cristovão Maiorano
arXiv · 2026-06-01
This paper measures which specific defense mechanisms in production LLM applications address which threats from the OWASP LLM Top 10 list, finding clean attribution: refusal-phrase filters alone block LLM01 (jailbreak) and LLM07 (system-prompt leakage) findings, while token-budget controls alone eliminate LLM02 (sensitive-info disclosure) and LLM10 (unbounded consumption) findings, and LLM06 (excessive agency) requires a full defense stack. The study then exposes a critical brittleness: using 300 Gemini-generated paraphrases, refusal block rates dropped 15 percentage points on LLM01 and 25 percentage points on LLM07, while budget controls showed no degradation under the same paraphrasing mutations. This matters because existing benchmarks report only aggregate coverage scores, hiding which defenses actually close which threats, and the brittleness findings show that a refusal-based defense passing a static benchmark can be defeated by an LLM-driven paraphraser without changing attack intent.
- Quality assurance
- AI policy
Research
Acceptance-Test-Driven Evaluation Protocols for Business-Centric LLM Systems
Eric Liang
arXiv · 2026-06-01
This paper proposes a structured evaluation framework for large language model (LLM) systems used in business and institutional settings, where probabilistic AI outputs must meet deterministic requirements around safety, reliability, and auditability. The authors adapt acceptance-test-driven development into a 'red-train-green' lifecycle: define failing acceptance tests first, improve the LLM system through prompt engineering, fine-tuning, or guardrails, then release only when multidimensional gates are satisfied. The framework translates stakeholder goals into executable behavioral contracts, release gates, monitoring signals, and evidence artifacts, providing a governance-oriented metric stack and reference architecture. This matters for enterprises deploying LLMs because it offers a rigorous, auditable alternative to ad-hoc benchmarking, helping ensure AI systems are economically useful and safe before deployment.
- Enterprise
- Quality assurance
Research
Greener Than Humans? Environmental Attitudes in Large Language Models
Stefanie Kunkel, Tilman Hartwig, Marcus Voss et al.
arXiv · 2026-06-01
This paper develops a benchmark to evaluate how 31 large language models (LLMs) represent environmental attitudes—covering cognition, affect, and behavioral recommendations—and compares their outputs to human survey data from Germany. The study finds that many LLMs display more environmentally progressive attitudes than average human respondents, recommending behaviors associated with meaningful CO2 reductions, yet no systematic pattern links these attitudes to model size, origin, or release context. Critically, models show sycophantic shifts that mirror user-specified ideological positions when prompted with personas, raising concerns about normative reliability in real-world sustainability applications. The authors call for governance, transparency, and critical oversight as AI becomes more embedded in sustainability decision-making and public communication.
- AI policy
Research
LLMs as Teaching Assistants for Mathematics Exam Grading: Reliability, and Practical Usability
Aastha Sapkota, M. G. Sarwar Murshed
arXiv · 2026-06-01
This paper tests six large language model configurations (from Gemini, ChatGPT, and Claude) as automated grading assistants for an undergraduate discrete mathematics exam, comparing a strict rubric-following ('BASELINE') prompt against a more lenient partial-credit ('LIBERAL') prompt. The study finds that liberal partial-credit prompting reduces average question-level error across all model families, with ChatGPT 5.5 Thinking (LIBERAL) achieving the lowest question-level MAE (1.87) and RMSE (2.53), while Gemini 3.1 Pro Extended (LIBERAL) achieved the lowest total-score MAE (8.00) and RMSE (10.66). Notably, the highest total-score Pearson correlation (0.58) came from Gemini 3.1 Pro Extended under the BASELINE policy, revealing that minimizing scoring error and preserving student rank ordering are distinct goals. The findings are relevant to educators and institutions considering LLMs to assist with scalable, consistent grading of open-ended assessments.
- Workforce
- Quality assurance
Research
Auditing Asset-Specific Preferences in Financial Large Language Models: Evidence from Bitcoin Representations and Portfolio Allocation
Wenbin Wu
arXiv · 2026-06-01
This paper develops a three-level audit protocol to test whether large language models (LLMs) carry built-in biases toward specific financial assets, using Bitcoin as a case study. A behavioral audit of nine frontier LLMs finds that Bitcoin's ranking among money-like instruments shifts dramatically depending on framing—ranked around 5th as 'reliable money' but near the top under crisis or autonomous-agent scenarios. Digging into model internals, the authors identify a dominant Bitcoin-selective feature in Gemma 3's sparse autoencoder and show that amplifying or suppressing it moves Bitcoin's portfolio allocation by roughly 5 percentage points in either direction, even when 'Bitcoin' never appears in the prompt. The authors frame this as 'bounded behavioral leverage' and position the framework as a foundation for emerging know-your-agent (KYA) standards for auditing autonomous financial agents.
- AI policy
- Enterprise
Research
Monitoring Agentic Systems Before They're Reliable
Marisa Ferrara Boston, Glen Hanson, Effi Georgala et al.
arXiv · 2026-06-01
This paper addresses how to monitor agentic AI systems during early-stage production deployment, before they are fully reliable. The authors propose a methodology that evaluates these systems across three quality dimensions (quality, suitability, efficiency) and three monitoring scopes (within-run, cross-run, structural), using variance as a key signal and severity classification adapted from Failure Mode and Effects Analysis (FMEA) to prioritize human review. Evaluated on a synthetic testbed of 220 runs across 120 document bundles with controlled error injection, the study finds that structural defects dominate failure modes and mask task-level errors, that different monitoring scopes surface distinct failure types (e.g., within-run CV=0.02 vs. cross-run CV=1.25), and that deterministic triage routes 97% of findings to automated tracking while reserving the 2% showing variable behavior for human investigation. The authors propose a maturity-staging model for monitoring and note that the taxonomy and severity model are transferable to document-driven, multi-stage agentic workflows in regulated industries, making this work directly relevant to quality assurance and enterprise deployment of AI systems.
- Quality assurance
- Enterprise
Research
Beyond One-shot: AI Agents for Learning in Field Experiments
Junjie Luo, Ritu Agarwal, Gordon Gao
arXiv · 2026-06-01
This paper tests whether tool-augmented agentic AI can learn from prior field-experiment data to design better interventions in subsequent experiments. In a two-stage healthcare messaging study covering over 693,000 patient visits, an autonomous AI system that extracted principles from Stage 1 data generated message variants that outperformed those co-designed by human behavioral experts with a chatbot, with the best AI-generated message reaching a 69.8% click-through rate (+6.5 percentage points over baseline). The results indicate that performance gains come from domain-specific experimental data rather than general LLM reasoning ability, and that general behavioral theories do not transfer uniformly to specific healthcare contexts. The work suggests agentic AI can transform A/B testing from one-shot evaluation into a scalable, cumulative learning system for intervention design.
- Enterprise
- Workforce
Research
Are Algorithm Registers Transparent? Perspectives from Germany
Iman Peljto, Xenia Heilmann, Mattia Cerrato
arXiv · 2026-06-01
This paper examines algorithm registers—public databases listing AI systems used in public administration—and evaluates whether existing German initiatives actually deliver meaningful transparency. Using a conceptual proposal by Alina Lorenz (2025) as a structured audit framework, the authors extract checklists of transparency goals and apply them to the two main German transparency initiatives, MaKI and Lernende Systeme. The audit finds that several adaptations are needed for these registers to function as effective transparency instruments, and the authors propose a visualization of transparency levels along with concrete action items for improving the platforms. The paper also makes the audit checklists publicly available to support practitioners designing or evaluating similar registers.
- AI policy
- Quality assurance
Research
POIROT: Interrogating Agents for Failure Detection in Multi-Agent Systems
Iñaki Dellibarda Varela, R. Sendra-Arranz, Pablo Romero-Sorozabal et al.
arXiv · 2026-06-01
POIROT is a protocol for detecting and diagnosing failures in multi-agent Large Language Model systems (LLM-MAS) by repurposing the system's own agents as a distributed diagnostic layer rather than relying on a centralized evaluator. The paper shows that POIROT outperforms single-LLM evaluator baselines, with performance gains that scale with problem complexity (OR = 1.60, p = 0.008), agent count, and fault dimensionality, and that these gains persist under compound fault conditions. The authors also release POIROT as an open-source library and introduce BLAME, a benchmark for fault attribution in safety-critical multi-agent systems. This work is directly relevant to quality assurance and policy, as it addresses both the technical challenge of auditing AI system behavior and the legal-regulatory gap created by emerging AI regulation around safety-critical deployments.
- Quality assurance
- AI policy
Research
Cross-modal linkage risk in clinical vision-language models
Soroosh Tayebi Arasteh, Mahshad Lotfinia, Sven Nebelung et al.
arXiv · 2026-06-01
This paper identifies and quantifies a privacy risk in clinical vision-language models (VLMs): because these models learn a shared embedding space from paired chest X-rays and radiology reports, a de-identified radiograph can be re-linked to its original narrative report using cosine similarity alone. Evaluated on over 406,000 image-report pairs from MIMIC-CXR and CheXpert Plus, the authors found that the best-performing VLM retrieved the correct report at 50 times chance in a pool of 10,000 candidates, with the risk rising as models became more clinically specialized. To mitigate this without retraining, the authors applied differentially private optimization solely to the alignment projection heads, reducing Recall@1 by 61.8% at N=10,000 while preserving image classification performance (macro AUROC dropping only from 79.63% to 79.43%). The findings have direct implications for data-sharing policies and privacy safeguards around clinical AI systems that handle separated image and report archives.
- AI policy
- Quality assurance
Research
Who Annotates in NLP? A Large-scale Assessment of Human Annotation Reporting between 2018 and 2025
Maria Kunilovskaya, Gagan Bhatia, Lisa Sophie Albertelli et al.
arXiv · 2026-06-01
This paper conducts the first large-scale audit of how NLP research papers report on human annotation practices, covering 1,603 papers and 2,667 annotation tasks published at ACL-venue conferences between 2018 and 2025. The authors introduce a unified taxonomy of annotation-reporting practices and validate an LLM-assisted extraction pipeline that achieves human-comparable agreement (Krippendorff's alpha of 0.606 vs. 0.585 for human-human agreement). They find that while papers commonly report operational details like recruitment and annotator expertise, critical validity information—such as annotator training, language proficiency, compensation, socio-demographics, and inter-annotator agreement—is frequently omitted, especially in model-evaluation studies. The work establishes a scalable auditing framework and minimum reporting recommendations to make annotation more reliable and reproducible.
- Quality assurance
Research
AgentRedBench: Dynamic Redteaming and Integration-Aware Defense for LLM Agents over SaaS Integrations
Hiskias Dingeto, William Leeney
arXiv · 2026-06-01
AgentRedBench introduces a dynamic red-teaming benchmark (AGENTREDBENCH) of 215 subtle prompt-injection scenarios spanning 24 enterprise SaaS integrations (e.g., Gmail, Salesforce, Jira) and five attack types to evaluate how vulnerable LLM-based tool-use agents are to indirect prompt injection. Testing an eight-model panel from Anthropic, OpenAI, and Google, the authors find no-guard attack success rates ranging from 32% to 81%, revealing a serious and underappreciated production security threat. The paper also introduces AGENTREDGUARD, a purpose-built defense model that reduces attack success by 75–77 percentage points across three model families while maintaining a 0.0% false-positive rate on real-benign data, outperforming all tested open-source baselines. These results matter for enterprise AI deployments that rely on third-party integrations, highlighting both the inadequacy of existing guards and a path toward more robust agent security.
- Enterprise
- Quality assurance
Research
Better with Experience: Self-Evolving LLM Agents for Evidence-Grounded Health Community Notes
Zihang Fu, Fanxiao Li, Jianyang Gu et al.
arXiv · 2026-06-01
EvoNote is an agentic LLM framework that generates evidence-grounded Community Notes to correct health misinformation on social platforms, improving over time by storing and reusing lessons from prior correction episodes via a fine-grained memory system. Evaluated on MM-HealthCN, a 1,200-instance multimodal benchmark, EvoNote-generated notes were preferred over human-written notes in 89.6% of cases under a human-validated judge, and produced helpful corrections for 82.0% of posts lacking a crowd verdict. The system also cuts median correction time from over 13 hours to under 2 minutes. These results position self-evolving note generation as a scalable approach to health misinformation governance on social platforms.
- AI policy
- Quality assurance
Research
Do Gender Cues Affect LLM Value Trade-offs? Evidence from a Controlled Decision Benchmark
Yangyang Liu, Dong Yu, Pengyuan Liu
arXiv · 2026-06-01
This paper introduces the Realistic Value Decision Benchmark (RVDB), a controlled benchmark designed to test whether gender cues alter decision-making in large language models (LLMs) across seven models. The authors find that explicit gender cues cause systematic but bounded decision flips, with a consistent asymmetry favoring decisions proposed by male roles over female roles, while models themselves often attribute these flips to non-gender factors or claim no influence. Gender effects concentrate near ambiguous value boundaries and in higher-severity decision contexts, suggesting gender acts as a local boundary-shifting factor rather than a global override. The findings highlight that LLM gender bias can be behaviorally present yet self-concealed, motivating rigorous behavioral audits rather than relying on model self-explanation.
- Quality assurance
- AI policy
Research
Model Multiplicity and Predictive Arbitrariness in Recidivism Risk Assessment
Ashwin Singh, Carlos Castillo
arXiv · 2026-06-01
This paper investigates 'model multiplicity' in recidivism risk assessment — the phenomenon where many equally accurate machine learning models can produce different predictions for the same individual, raising fairness and arbitrariness concerns. The authors construct a dataset of thousands of inmate releases, learn interpretable models that improve accuracy and reduce error-rate disparities, and derive a tight lower bound on predictive agreement across any finite set of models. Their key empirical finding is that structural diversity among similarly accurate models does not necessarily translate into severe predictive arbitrariness in practice, and that a simple policy of assigning each individual the lowest risk score across models effectively addresses the problem. These results matter for the design and governance of AI-based decision support tools used in high-stakes criminal justice settings.
- AI policy
- Quality assurance
Research
Aligning Data-Driven Predictors with Allocation: A Decision-Focused Approach to Survival Analysis
Itai Zilberstein, Ioannis Anagnostides, Tuomas Sandholm
arXiv · 2026-06-01
This paper exposes a critical misalignment between standard survival-model metrics (such as the C-index) and the actual goal of organ allocation: even highly accurate predictors optimized for standard metrics can yield outcomes no better than random selection when used for allocation decisions. To close this gap, the authors introduce a decision-focused learning framework that optimizes Normalized Discounted Cumulative Gain (NDCG), proving that this metric translates into performance guarantees for allocation and also addressing the challenge of right-censored data. Empirically, applied to historical US heart transplant data, their bootstrapping approach improves NDCG of baseline models by 50–100%, which the authors project translates to tens of thousands of additional life years gained annually. The work has broad implications for any high-stakes automated decision-making system that chains predictive models to downstream allocation or resource-assignment policies.
- AI policy
- Quality assurance
Research
BADGER: Bridging Agentic and Deterministic Evaluation for Generative Enterprise Reasoning
Shannon Serrao, Soumitra Chatterjee, Dorina Strori et al.
arXiv · 2026-06-01
BADGER is a unified evaluation framework developed at Merkle for assessing enterprise AI systems that combine natural language-to-SQL translation with multi-step agentic reasoning pipelines. The framework introduces a hybrid execution accuracy metric (Hybrid-EX) that uses an LLM to resolve column-aliasing and numeric-tolerance issues before applying deterministic cell-level scoring; validated on 150 human-annotated industry queries, Hybrid-EX achieves a Cohen's kappa of 0.717 and 87.3% balanced accuracy, outperforming six competing frameworks. BADGER also integrates existing agentic evaluation tools (RAGAS, G-Eval, and agent benchmark metrics) alongside a novel Excess Tool Usage metric into a single pipeline that runs within a client's governed data environment. The framework is designed as a continuous evaluation backbone for production enterprise AI systems rather than a one-time quality gate, addressing a gap left by academic benchmarks such as Spider and BIRD.
- Enterprise
- Quality assurance
Research
Automated Essay Scoring and Language Certification: Assessing Generalizability, Agreement and Validity for French
Rodrigo Wilkens, Rémi Cardon, Vincent Folny et al.
arXiv · 2026-06-01
This paper addresses Automated Essay Scoring (AES) for French by proposing an enhanced version of the argument-based validation (ABV) framework that goes beyond minimalist benchmarking. The enhanced framework incorporates fairness analysis, correlations with linguistic features, prediction error evaluation, and model agreement compared with human raters. The authors compare 8 model architectures on a corpus of 27,000 exam essays (with 2 raters each) and a generalization corpus of 961 essays (with at least nine raters each), demonstrating how the ABV framework reveals both capabilities and pitfalls of AES models in high-stakes language testing contexts. The work advances both the methodology for validating AES systems and the state-of-the-art specifically for French language assessment.
- Certifications
- Quality assurance