News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
5608 items
Research
The Paternalistic Filter: Epistemic Injustice and Differential Refusal in LLM-Mediated History Education for Marginalized Romanian Students
Alexis Popovici, Andrei Ionascu, Adrian-Marius Dumitran
arXiv · 2026-07-13
This study conducts a systematic API audit of four LLMs acting as history tutors, evaluating 1,800 responses about the 1989 Romanian Revolution across five student personas varying by ethnicity and socio-economic tier. The researchers uncover four patterns of epistemic paternalism: safety-aligned models blocked 76.7% of educational requests from low-tier students, marginalized learners received three times less access to geopolitical complexity, models like LLaMA produced a five-times higher victimization-to-politics vocabulary ratio for Roma students compared to elite peers, and AI tutors disproportionately withheld epistemic confidence from low-resource demographic profiles. The authors argue that current safety alignment acts as a paternalistic filter that transforms conversational AI into an agent of narrative segregation — a form of hermeneutical injustice — with urgent implications for how AI tutoring systems are designed, audited, and deployed in educational settings.
- Workforce
- AI policy
- Quality assurance
Research
Automated Textbook Auditing with Multi-Agent LLM Systems
Ciprian Cristescu, Adrian-Marius Dumitran, Angela-Liliana Dumitran et al.
arXiv · 2026-07-13
This paper presents AI Textbook Auditor, a multi-agent LLM pipeline that automatically audits educational textbooks for factual accuracy, technical correctness, and linguistic quality. The system uses specialized LLM agents for a Factual and Technical Track and a Grammar Track, with a Judge Agent filtering false positives before surfacing findings to human reviewers. Demonstrated on two Romanian upper-secondary textbooks, the system found 56 technical findings (62.5% expert-validated precision) in a CS textbook and 72 findings in a history and social sciences textbook. It is designed as a triage tool to reduce manual review effort, with human expert validation required before any editorial action.
- Quality assurance
- Enterprise
Research
Beyond AI-Generated Labels: Watermarking, Co-Creation, and Conflation of AI-Generation with Disinformation
Federico Germani, Giovanni Spitale
arXiv · 2026-07-13
This paper critically examines watermarking as a policy tool for identifying AI-generated content, arguing that invisible watermarks encode only model origin and that converting them into visible 'AI-generated' labels creates a misleading binary that says nothing about truthfulness or deceptive intent. The authors contend that such labels may unfairly stigmatize legitimate creative uses of generative AI while fostering misplaced trust in unmarked content. As an alternative, they advocate for process transparency and information literacy as more effective responses to the epistemic and ethical challenges posed by AI-generated disinformation. The work has direct implications for platform regulation, content labeling policy, and public understanding of AI-assisted authorship.
- AI policy
- Quality assurance
Research
Evaluating Nonuniform Dependability Across Response Conditions: A Conditional Generalizability Framework Illustrated in Automated Essay Scoring
Yi Gui
arXiv · 2026-07-13
This paper introduces a conditional generalizability framework for evaluating whether automated essay scoring (AES) systems are equally reliable across different types of student responses. Rather than reporting a single aggregate reliability estimate, the authors condition dependability on entropy-defined strata of responses—grouping essays by linguistic complexity—and show that reliability declines modestly but consistently across strata (Phi = 0.88, 0.87, 0.84) even when aggregate dependability appears adequate (Phi ≈ 0.76). The framework treats AI scoring configurations, such as encoder architectures and scoring-head families, as a universe of admissible measurement conditions, enabling more precise diagnosis of where a scoring design may fall short. This matters for high-stakes language assessment because it reveals that the most complex responses require more scoring conditions to achieve the same level of dependability, with direct implications for how automated scoring systems should be validated and deployed.
- Quality assurance
- Certifications
Research
When the Target Domain Changes: AI-Mediated Construct Drift in High-Stakes English Language AssessmenW
Yi Gui
arXiv · 2026-07-13
This conceptual paper argues that high-stakes English proficiency tests face a validity crisis as generative AI becomes embedded in academic communication. The author introduces 'AI-mediated construct drift' to describe the growing misalignment between what unaided test performance measures and the AI-assisted communicative competencies actually required in real academic settings. To address this, the paper proposes 'bounded AI mediation' as a design principle—giving all test-takers access to the same institutionally controlled AI assistant under standardized, logged conditions—so that tests can better capture relevant abilities. The paper concludes that score interpretations should be narrowed and supplemented when used to make claims about AI-mediated academic communication readiness.
- Certifications
- Quality assurance
- AI policy
Research
Inference Economics of Enterprise Coding Agents: A Case Study of Cloud vs. On-Premise LLMs
Sheng-Wei Peng, Yi-Hsun Lin, Yi-Pei Lee
arXiv · 2026-07-13
This longitudinal case study compares API-based frontier LLMs (Claude Opus 4.7/4.8) against on-premise quantized open-weights models (GLM-5.1/5.2) for autonomous coding agents over two 28-day periods on a production monorepo. The study finds that prompt caching at a 99.3% hit rate cuts effective API costs by 88.6% to $0.57 per million tokens, which can fall below the amortized on-premise unit cost depending on GPU utilization. However, the on-premise configuration was associated with a substantially higher defect-repair burden—a Fix Commit Ratio of 74.9% versus 45.9%, with 2.6–4.9 times higher odds of a repair commit—suggesting meaningful quality-assurance and developer-experience trade-offs that monetary comparisons alone do not capture. The findings indicate that the choice between cloud and on-premise LLM deployment for enterprise coding agents involves a cost-quality frontier rather than a clear winner, with hybrid routing as a potential middle ground.
- Enterprise
- Quality assurance
- Workforce
Research
AMT-X: Phase-Structured Multi-Turn Red-Teaming with Checklist-Gated Evaluation
Yi Ting Shen, Kentaroh Toyoda, Alex Leung
arXiv · 2026-07-13
AMT-X introduces a structured multi-turn red-teaming framework for evaluating the safety of large language models, addressing gaps left by single-turn attack datasets and single-judge scoring. The system models attacks as an explicit multi-phase state machine guided by semantic signals from the target model, and uses a multi-role jury with phase-conditioned checklists to distinguish partially actionable outputs from fully operational harmful content. Tested against six frontier LLMs across seven harm subcategories, AMT-X achieves overall attack success rates of 97.6–100% under lenient scoring, but only 66.7–78.6% under stricter criteria requiring complete and operational detail—revealing a gap of up to 33 percentage points. This finding matters for AI safety policy and quality-assurance processes, showing that commonly reported success rates can significantly overstate actual risk containment.
- Quality assurance
- AI policy
- Certifications
Research
Do LLMs Fabricate Legal Citations? A Bilingual Benchmark on Saudi Data Protection Law and the GDPR
Noura Suliman Alrajeh
arXiv · 2026-07-13
This paper introduces a bilingual (Arabic/English) benchmark of 120 questions to test whether freely accessible large language models fabricate legal article citations under two data-protection regimes: the EU GDPR and the Saudi Personal Data Protection Law (PDPL). Testing three models, the authors find near-perfect citation accuracy on the GDPR (94–100%) but majority fabrication on the Saudi PDPL (60–77%), with the highest fabrication rates (67%) stemming from confusion between the statute and its implementing regulations. Critically, 91% of fabricated citations are asserted with high confidence (≥0.8), meaning model self-confidence offers no protection against errors. The authors conclude that verbatim-verification safeguards—not model confidence—must gate any institutional reliance on LLMs for compliance screening.
- AI policy
- Enterprise
- Quality assurance
- Certifications
Research
AgentCheck: A Reproduce-Intervene-Mitigate Workbench for LLM Agents over MCP
Aritra Mazumder, Nusrat jahan Lia
arXiv · 2026-07-13
AgentCheck is an open-source web workbench that helps developers test and debug tool-using LLM agents by simulating real-world tool failures such as timeouts, stale data, and poisoned descriptions. It records live tool responses, then replays them with injected faults across 12 failure types, letting developers toggle mitigations and verify fixes against identical fault conditions before deployment. Across five agents tested on 120 scenarios, results ranged from 77 to 105 passes, with the most common failure mode being silent, confident use of incorrect tool outputs rather than crashes. The workbench demonstrates that retry mitigations can raise timeout-fault success rates from 30% to 100%, while stale-data faults remain stubbornly low regardless of mitigation strategy.
- Quality assurance
- Enterprise
Research
NVAITC AI Scientist: A Governed End-to-End Research System -- A Hypertension GWAS Case Study
Eddie Huang, Ken Liao, Iven Fu et al.
arXiv · 2026-07-13
NVAITC AI Scientist (NAIS) is a governed, end-to-end agentic research system designed to support scientific workflows in institutional biomedical settings while keeping protected data within privacy boundaries. Validated on a real-world hypertension genome-wide association study (GWAS) using hospital-linked genotype and electronic health record data from 286,422 individuals, NAIS orchestrated cohort extraction, GWAS execution, quality-control summaries, and publication-oriented outputs under an aggregate-only data policy. Human-AI review identified phenotype discrepancies, enabling iterative refinement that allowed the system to reproduce established hypertension loci—including FGF5, ATP2B1, CNNM2, FTO, and GRB14—with the strongest signal at FGF5 reaching −log10(p) ~70. The results demonstrate that governed agentic systems can support scalable AI-assisted biomedical discovery while producing outputs comparable to expert-led workflows.
- Enterprise
- Quality assurance
- AI policy
Research
Same Stories, Different Journeys: From Social Comparison to Sensemaking in AI-Mediated Peer Career Exploration
Pengping Tan, Baoquan Zhao, Zhenhui Peng
arXiv · 2026-07-13
This paper presents JobMate, an interactive system that converts real social media career posts into conversational AI agents (personas) to help young job seekers explore career possibilities. In a between-subjects study (N=24 across three disciplines), JobMate was compared against native social media browsing on RedNote; results showed that the AI-mediated dialogue shifted users away from potentially harmful upward social comparison toward constructive self-reframing and active sensemaking. Users still valued the authenticity of real peer content for emotional grounding, highlighting a tension between AI mediation and genuine human experience. The findings offer design implications for AI systems that augment user-generated content consumption in social comparison contexts.
- Workforce
- Enterprise
Research
Dimensionality in Satisfaction Ratings
Andrew Hong, Jason Potteiger
arXiv · 2026-07-13
This paper uses GPT-4.1 to annotate roughly 9,000 customer support conversations at a consumer-goods firm, decomposing satisfaction into five dimensions (overall, agent, outcome, product, and customer effort) and validating these LLM-generated scores against customers' self-reported ratings. Four of the five axes correlate strongly with self-reported satisfaction, with overall, agent, and outcome near 0.65 and effort at -0.54, and the overall correlation rises to 0.914 when the most divergent sessions are excluded. Crucially, applying LLM annotation to every contact rather than only surveyed contacts reveals markedly lower satisfaction (2.91 vs. 3.62 on a five-point scale), showing that traditional surveys systematically oversample satisfied customers. The methodology's value lies not in incremental prediction but in attribution, coverage, and identifying nuanced drivers of customer experience across the full population of interactions.
- Enterprise
- Quality assurance
Research
When the Reward Suite Is Leaky: A Preregistered Causal Contrast of Natural Verifier False Positives in RLVR
Chuyifei Zhang
arXiv · 2026-07-13
This paper investigates a specific failure mode in Reinforcement Learning from Verifiable Rewards (RLVR) for code generation: test suites used as reward signals contain 'natural false positives'—persistent, per-task errors that incorrectly accept wrong programs every time they appear. Through a preregistered causal experiment comparing GRPO trained on original MBPP tests versus hardened MBPP+ tests, the authors find that while the average held-out performance gap is small (0.20 pt, bounded below 0.75 pt), roughly 47.57% of rewarded false positives correspond to genuinely wrong code rather than suite artifacts, meaning the reward signal is paying for real bugs. The mechanism appears to involve selection of pre-existing error modes rather than learned exploitation, as false-positive incidence does not grow during training and untrained base models already produce the same wrong outputs under the leaky filter. A cheap static leakiness audit computed before training correlates strongly with rewarded false-positive mass (Spearman 0.80), offering a practical tool for identifying and hardening vulnerable reward suites before training begins.
- Quality assurance
- Enterprise
Research
QwenPaw-Data: Bridging Facts, Methodology, and Execution for Autonomous Enterprise Data Analytics
Tianjing Zeng, Yuntao Hong, Zhongjun Ding et al.
arXiv · 2026-07-13
QwenPaw-Data is an agentic data system designed to automate enterprise data analytics by integrating heterogeneous data sources—including warehouses, dashboards, documents, and interaction logs—into a unified, evolvable framework. Its architecture comprises three collaborative subsystems: DataBridge for semantic grounding, Skill-Hub for codifying expert analytical methods into reusable skills, and a Host runtime that executes end-to-end analytical workflows from natural-language requests through to report generation and decision support. The system features a self-evolving 'asset flywheel' where semantics, methods, traces, and feedback are continuously fed back into the system to improve over time. Experiments on public benchmarks and real-world industrial business intelligence workloads demonstrate improvements in both verifiable data access and higher-level analytical quality, offering a practical foundation for reliable and traceable enterprise data agents.
- Enterprise
- Workforce
- Quality assurance
Research
The Role of Artificial Intelligence in Enhancing Women's Economic Participation: Opportunities and Challenges for Sustainable Development
ideal research review
Zenodo (CERN European Organization for Nuclear Research) · 2026-07-13
This qualitative, descriptive study synthesizes peer-reviewed literature and international organization reports to examine how AI affects women's economic participation in the context of SDG 5 (gender equality). The analysis finds that AI expands women's access to employment, entrepreneurship, and digital platforms through flexible work, improved market access, and AI-enabled business tools, while also reducing unpaid care burdens via health and education applications. However, the study identifies significant challenges including algorithmic bias, occupational displacement, wage disparities, precarious gig work, and structural barriers such as limited digital skills and restricted financial access. The authors conclude that inclusive AI policies, digital capacity-building, and gender-responsive governance are essential to ensure AI contributes to equitable development outcomes for women.
- Workforce
- AI policy
Research
Artificial Intelligence, Social Capital, and Sustainable Employment in Peripheral SMEs: A Biocultural Reading from Eastern Macedonia and Thrace, Greece
Eugenia P. Bitsani, Αντώνιος Κώστας, Vasileios Kapilidis et al.
Sustainability · 2026-07-13
This qualitative study examines how AI adoption affects employment and sustainable development in small and medium-sized enterprises (SMEs) in Eastern Macedonia and Thrace, one of the EU's least developed regions. Through thematic analysis of twelve semi-structured interviews with SME owners and managers, the authors find that knowledge deficits and financial constraints are the primary barriers to AI adoption, while technology partnerships, targeted education, and economic incentives act as enablers. The study argues that without parallel investment in digital literacy, organizational culture, and inter-firm networks, AI risks deepening rather than reducing employment inequalities in peripheral economies. The findings carry direct implications for EU Cohesion policy and Sustainable Development Goals related to education, decent work, industry, reduced inequalities, and partnerships.
- Workforce
- Enterprise
- AI policy
Research
The Role of Artificial Intelligence in Enhancing Women's Economic Participation: Opportunities and Challenges for Sustainable Development
ideal research review
Zenodo (CERN European Organization for Nuclear Research) · 2026-07-13
This study examines how AI affects women's economic participation in the context of SDG 5 (gender equality), using a qualitative review of peer-reviewed literature and international organization reports. The analysis finds that AI creates opportunities through flexible work, entrepreneurship tools, improved market access, and indirect benefits via health and education applications that reduce unpaid care burdens. However, significant challenges persist, including algorithmic bias, occupational displacement, wage disparities, and structural barriers such as limited digital skills and inadequate infrastructure. The study concludes that inclusive AI policies, digital capacity-building, and gender-responsive governance are necessary for AI to contribute to equitable development outcomes for women.
- Workforce
- AI policy
Research
AI, blockchain, and cloud accounting in smart governance: How real-time audit reporting enhances transparency and decision quality
Hamood Mohammed Al‐Hattami
Array · 2026-07-13
This study examines how digital accounting innovations—AI, blockchain, and cloud accounting—shape real-time audit reporting (RTAR) and its downstream effects on organizational transparency and decision-making quality in Yemen. Using PLS-SEM with 188 accounting and auditing professionals, the findings show that these digital technologies significantly enable RTAR capabilities, which in turn mediate improvements in governance outcomes by increasing the timeliness, reliability, and traceability of financial information. The research positions RTAR as a critical mechanism through which digital transformation generates organizational value, offering actionable guidance for auditors, managers, and policymakers in resource-constrained economies seeking to advance smart governance.
- Enterprise
- Quality assurance
- AI policy
Research
Ensuring Data Integrity in Official Financial Statistics: A Review of Hybrid AI and XAI Methods in the Context of Market Efficiency and Value Investing
Krzysztof Podgórski
Journal of Official Statistics · 2026-07-13
This review paper examines how hybrid AI methods—such as ARIMA-LSTM and autoencoder-based GANs—combined with Explainable AI (XAI) techniques like SHAP and LIME can improve data quality assurance, anomaly detection, and imputation in official financial statistics. The authors find that hybrid frameworks can outperform single-model approaches in detecting nonlinear data manipulations and producing high-fidelity datasets, which is critical for market efficiency and value investing decisions. The paper also addresses the 'black box' opacity problem, arguing that XAI tools support but do not independently guarantee the interpretability needed by regulatory and statistical agencies. The study concludes that combining predictive AI power with XAI transparency is a valuable component for modern market supervision and institutional accountability.
- Quality assurance
- AI policy
- Enterprise
Research
A framework for developing university policies on generative AI governance: a cross-national comparative study
Ming Li, Qin Xie, Ariunaa Enkhtur et al.
Studies in Higher Education · 2026-07-13
This cross-national study analyzes generative AI (GAI) governance policies issued by leading universities in the United States, Japan, and China, finding notable differences in policy orientation: U.S. institutions emphasize faculty autonomy and adaptability, Japanese universities prioritize ethics and risk management in alignment with government guidance, and Chinese universities reflect a centralized model focused on technology application. Using an extended Technology Acceptance Model, the authors identify 20 themes across five domains and synthesize them into a proposed University Policy Development Framework for Generative AI (UPDF-GAI). The framework is designed to help universities balance innovation and risk, assess policy priorities, and build institutional capacity for sustainable AI governance in higher education. The findings matter because they offer a structured, comparative basis for institutions worldwide to develop or refine their own generative AI policies.
- AI policy
- Enterprise
Research
Operational Evidence Gaps for LLMs in Fraud Detection and Trust-and-Safety Workflows
Keyur Gabani
arXiv · 2026-07-12
This survey examines the gap between how large language models (LLMs) are evaluated in research versus how they actually perform as operational components in fraud detection and trust-and-safety workflows. Analyzing 49 operationally relevant sources, the authors find a significant evidence imbalance: while fraud detection supplies the largest share of task-specific literature, none of the 18 fraud and investigation sources report per-decision latency, per-decision cost, or calibration evidence — relying instead on offline task performance or case-study accuracy. The paper introduces FORTE, a framework for categorizing LLM roles in pipelines (classifiers, retrieval interfaces, agents, etc.), and a minimum deployment-evidence checklist covering latency, cost, decision thresholds, explanation integrity, and adversarial pressure. The findings matter for enterprise and quality-assurance contexts because they highlight what evidence is still missing before LLMs can be responsibly deployed in high-stakes operational fraud and content moderation systems.
- Enterprise
- Quality assurance
- AI policy
Research
The Hitchhiker's Guide to Monoculture
Gordon Burtch
arXiv · 2026-07-12
This paper examines whether AI coding assistants are causing homogenization in software artifacts by analyzing Kaggle contest submissions from 2019 to mid-2026. The study finds substantial syntactic homogenization—individual submissions have grown more alike in literal syntax and code structure, and the latent dimensionality of syntactic variation has narrowed—including widespread convergence toward the random seed value 42. However, the paper finds little evidence of semantic homogenization: average semantic distance remains essentially flat and the conceptual span of problem-solving approaches has remained stable or even modestly expanded. These findings suggest AI coding assistants are standardizing implementation details without yet producing convergence in the underlying approaches and strategies developers employ, with important implications for software diversity and developer autonomy.
- Workforce
- Enterprise
- Quality assurance
Research
LOGOS: A Living Logic for AI Agent Teams That Evolve With Humans
Yuma Ichikawa, Yamato Arai, Kosaku Kimura et al.
arXiv · 2026-07-12
LOGOS introduces a governance and self-evolution layer for multi-agent AI systems that enforces human oversight over how agents change over time. The system compiles diverse inputs into versioned 'agent packs' containing agents, tools, knowledge, tests, permissions, and policies, and requires that any learned prompt, memory, skill, or workflow remain an untrusted candidate until verified by execution evidence and explicit human authorization. This 'verifiable human-agent loop engineering' approach ensures agents can improve at machine speed while humans retain control over approvals and irreversible actions. The architecture matters for enterprise and policy contexts because it directly addresses accountability, auditability, and governance in continuously operating AI agent teams.
- Enterprise
- AI policy
- Quality assurance
- Certifications
Research
How Do Practitioners Build SE Agents? Insights from a Mixed-Methods Study
Yunbo Lyu, David Williams, Jieke Shi et al.
arXiv · 2026-07-12
This mixed-methods study—combining semi-structured interviews with 20 practitioners from 12 organizations and an online survey of 80 practitioners—is the first to systematically examine how software engineering (SE) agents built on large language models are actually developed in practice. The researchers find that as implementation costs fall, bottlenecks shift rather than disappear: non-coding work such as requirements, coordination, review, and deployment becomes more prominent, while evaluating agent output emerges as a new central activity. The paper characterizes a seven-stage development workflow and a shift toward evaluation-driven development, in which evaluation guides iteration and specifications become versioned artifacts read by both humans and agents. Six key challenges are identified, including unreliable evaluation signals, comprehension debt as code outpaces understanding, and behavioral changes introduced by provider-side model updates—findings with direct relevance to enterprise adoption, quality assurance, and workforce practices around AI-assisted software engineering.
- Workforce
- Enterprise
- Quality assurance
Research
Return of the solo author: The changing division of labor in science in the age of generative AI
Akira Matsui
arXiv · 2026-07-12
Analyzing over 300 million works across 26 fields, this study finds that the decades-long decline in solo authorship in science halted and partially reversed following ChatGPT's public release in late 2022. The reversal is strongest in fields where coauthors' tasks are more readily replaceable by AI, and is driven by authors who previously only wrote collaboratively, including established researchers and newcomers alike. Solo papers produced in this period stay close to authors' existing coauthored work but narrow in scope and shift toward computational topics, suggesting that generative AI is substituting for specific collaborative labor rather than simply enabling larger teams. This provides empirical evidence of a reconfiguration of cognitive labor within research, with direct implications for how scientific workforce dynamics and collaboration norms are being reshaped by AI tools.
- Workforce
- AI policy