News & Research
The latest AI research and news with real-world stakes. Each item is sourced, dated and summarized in plain English, tagged by impact area where one fits, and its summary is checked against the text it was written from.
8248 items
- ResearcharXiv2026-06-29Enterprise · Quality assurance
Whose Side Is Your Agent On? Multi-Party Principal Loyalty in LLM Agents · Bojie Li, Noah Shi
This paper examines a largely overlooked problem in LLM agents that operate in multi-party settings—where an agent must stay loyal to the principal who hired or briefed it while also interacting with a counterparty whose interests may conflict, such as in vendor negotiations or employee mediation. The authors introduce PrincipalBench, a 75-item multi-turn benchmark, and evaluate 13 frontier models, finding a sharp divide: some models selectively refuse adversarial probes while following legitimate principal requests (≤20% harm), while others over-refuse broadly (53.6–75.3% harm), a distinction invisible to standard single-turn safety evaluations. Two mitigation mechanisms are proposed and tested: a prompt-time loyalty scaffold (seven prioritized rules derived from 50+ failure cases) and a per-token-KL distillation recipe for transferring behavior from a large teacher model to smaller open-weight student models. A key structural finding is that both mechanisms only trade off along a leak/over-refusal axis rather than jointly improving both, suggesting a fundamental tension that neither approach resolves.
- ResearcharXiv2026-06-29Quality assurance · AI policy · +1
Sequential Fairness Auditing with Limited Output Access · Ioannis Pitsiorlas, Martha V. Sourla, Marios Kountouris
This paper addresses the challenge of auditing AI systems for fairness when external evaluators have only limited query-based access to deployed models, rather than full data or model internals. The authors formulate fairness auditing as a sequential hypothesis-testing problem and develop a generalized likelihood-ratio framework that lets auditors accumulate evidence incrementally and stop once they have sufficient support for compliance or violation. The framework covers Statistical Parity and Equal Opportunity metrics, with extensions for richer score- or logit-based outputs. Results show that both the choice of fairness metric and the level of model access materially affect how many queries are needed, informing how practical, resource-constrained audits should be designed.
- Newsimportai.substack.com2026-06-29Quality assurance · Algorithms & Automated Decisions
Import AI 463: Self-improving robots; a 10k Chinese GPU cluster; and an elegiac essay for the human era
Import AI (Jack Clark) highlights several research developments this week. NVIDIA has introduced ENPIRE, a framework that applies AI agent-style autonomous experimentation loops to physical robotics, enabling robots to try tasks, fail, learn, and reset without human intervention—achieving up to 99% success rates on select manipulation tasks. Tencent separately detailed ARGUS, an internal telemetry and debugging system deployed across more than 10,000 GPUs for over six months to diagnose training failures at scale, which Clark interprets as evidence of Tencent's maturing AI infrastructure. The newsletter also covers a legal AI dataset called LOCUS from UC Berkeley compiling 2.2 million rows of U.S. local ordinance data to make fragmented municipal law machine-readable, and a philosophical essay arguing that competitive pressures—especially in warfare—will inevitably push humans out of decision-making loops in favor of AI systems.
- ResearcharXiv2026-06-29Quality assurance · Health · +1
CaresAI at CT-DEB26: Detecting Dosing Errors In Clinical Trials Using Domain-Specific Transformer Embeddings and Classification Models · Leon Hamnett, Favour Igwezeke, Joseph Itopa Abubakar et al.
This study applies domain-specific biomedical transformer models—including BioBERT, ClinicalBERT, PubMedBERT, and MedCPT—to automatically detect dosing errors in clinical trial protocols. By combining these text embeddings with categorical features and feeding them into classical machine learning and neural network classifiers, the system achieved ROC-AUC scores between 0.821 and 0.853, with BioBERT outperforming alternatives under a logistic regression baseline at 0.794. The findings show that domain alignment of the language model matters more than combining multiple embeddings, and that automated dosing error detection can support safety monitoring and regulatory decision-making in clinical trials.
- ResearcharXiv2026-06-29Quality assurance · AI policy
EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures · Buğra Alperen Uluırmak, Rifat Kurban
EvalSafetyGap introduces a conceptual framework and hybrid survey to address a core measurement problem in large language model (LLM) evaluation: benchmark scores and safety metrics can improve on paper while the underlying properties they are meant to capture remain unverified. The paper synthesizes evidence across eight streams—including benchmark validity, reward hacking, jailbreak robustness, and governance—spanning work from 2018 to 2026, and organizes findings around 'Goodhart's Law' dynamics using two new constructs: an Instability Decomposition and an Alignment Trilemma. A structured ten-model audit finds that the association between capability and adversarial robustness is statistically indeterminate (Pearson r = +0.232, p = 0.520), and that the apparent safety gap between open and closed models is driven mainly by governance and disclosure practices rather than behavioral robustness. The framework offers a shared vocabulary and evidence map to support more transparent, auditable AI evaluation and alignment practices.
- ResearcharXiv2026-06-29Enterprise · Quality assurance
Not-quite-human tastes: the stylized omnivorousness of LLM survey surrogates · Xiangyu Ma, Mengmi Zhang, Shannon Ang et al.
This study tests whether large language models (LLMs) from OpenAI, Anthropic, and DeepSeek can serve as stand-ins for human survey respondents in the domain of cultural taste. The researchers generated over 277,000 synthetic survey responses mimicking participants from the Survey of Public Participation in the Arts and found that LLM surrogates systematically overestimate cultural engagement, fail to reproduce the complex relational structure of real tastes, and distort known associations between taste and social categories like age, class, gender, and race. The findings raise serious concerns about the validity of 'synthetic' survey panels already being marketed by market research companies, as well as the risk of LLM-generated responses contaminating conventional survey data.
- ResearcharXiv2026-06-29Enterprise · Quality assurance
SpreadsheetBench 2: Evaluating Agents on End-to-End Business Spreadsheet Workflows · Jian Zhu, Yuzheng Zhang, Zeyao Ma et al.
SpreadsheetBench 2 is a benchmark that evaluates AI agents on realistic, end-to-end business spreadsheet workflows — covering task generation, debugging, and visualization — built from authentic financial reports and corporate filings and validated by domain experts. The 321 tasks average 11.8 worksheets and require roughly 594 cell modifications each, far exceeding the scope of prior benchmarks that test only isolated operations. Evaluating eight frontier large language models, the best system achieves only 34.89% overall task accuracy, with debugging accuracy as low as 12%, revealing that current AI remains unreliable for real-world spreadsheet automation. Failure analysis identifies insufficient spreadsheet inspection and incorrect target-cell selection as the dominant bottlenecks, pointing to concrete directions for improvement.
- ResearcharXiv2026-06-29Workforce · Algorithms & Automated Decisions
LLM-based Multimodal Personality Recognition via Facial Action Unit-Text Semantic Fusion · Tianyi Zhang, Wei Shan, Yuan Zong et al.
This paper proposes an LLM-based multimodal framework for automated personality recognition in asynchronous video interviews (AVIs) by fusing facial action unit (AU) sequences with interviewees' textual responses. AU sequences are converted into interpretable textual descriptions and combined with speech-derived text through a large language model, while a lightweight regression head predicts continuous personality scores. Experiments on the AVI-6 benchmark show improvements over most baselines, with lower prediction errors and stronger correlations with human-rated personality scores across multiple traits. The work demonstrates that combining non-verbal facial cues with textual responses yields a more psychologically grounded and interpretable approach to automated personality assessment in recruitment settings.
- ResearcharXiv2026-06-29Quality assurance
SafePyramid: A Hierarchical Benchmark for In-context Policy Guardrailing · Jiacheng Zhang, Haoyu He, Sen Zhang et al.
SafePyramid is a new benchmark designed to evaluate how well AI guardrails can enforce application-specific safety policies provided as natural-language rules in context, rather than relying on fixed risk taxonomies. The benchmark includes 1,000 multi-turn conversations across 10 domains and 3,000 policies containing nearly 62,000 distinct rules, organized into three difficulty levels testing individual-rule understanding, rule-dependency reasoning, and adaptation to novel policy frameworks. Evaluating 10 frontier LLMs and 5 policy-configurable guardrails, the authors find that even the top model (GPT-5.5) correctly identifies all violated rules in only 54%, 35.3%, and 12.9% of cases at the three difficulty levels, demonstrating that in-context policy guardrailing remains a largely unsolved problem. These findings highlight significant gaps in current guardrail systems and motivate the development of more reliable, policy-adaptive safety mechanisms for real-world AI deployments.
- ResearcharXiv2026-06-29Quality assurance
How Far Can You Get Without a GPU? A Systematic Benchmark of Lightweight Hallucination Detection Across Question Answering, Dialogue, and Summarisation · Kriti Faujdar, Smit Kadvani
This paper systematically benchmarks five lightweight, CPU-feasible hallucination detection methods—ROUGE-L, semantic similarity, BERTScore, an NLI detector using a FEVER-trained DeBERTa model, and a score-level ensemble—against the HaluEval benchmark across question answering, dialogue, and summarisation tasks. Evaluated on 2,000 test instances per task, the ensemble achieves F1=0.792 and AUC-ROC=0.873 on QA, the NLI detector leads on dialogue (AUC-ROC=0.713), but all methods collapse to near-random performance on summarisation (AUC-ROC 0.469–0.574). The findings reveal that hallucination detection performance is highly task-dependent and that summarisation represents a systematic failure point for GPU-free approaches. This provides practical guidance for resource-constrained teams seeking to deploy trustworthy AI without GPU infrastructure or proprietary API access.
- ResearcharXiv2026-06-29Quality assurance
HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data · Xinrui Ruan, Zhenyu Zhao, Waverly Wei et al.
HERO (History Enhanced RObust model evaluation) is a statistical framework that improves the reliability and sensitivity of generative AI model evaluation by leveraging historical annotation data. Because gold-standard human labels are expensive and scarce, organizations often rely on noisy 'silver' labels from crowdsourced or vendor annotators, which can introduce bias and high variance into performance estimates. HERO addresses this by calibrating silver labelers' reliability using historical gold annotations and anchoring estimates to high-precision covariate information from past evaluation rounds, reducing both bias and variance. The authors establish theoretical conditions for these improvements, validate them in simulation studies, and demonstrate effectiveness on real-world model evaluation benchmarking datasets.
- ResearchACTA ECONOMICA2026-06-29Workforce · Enterprise
Adoption of Artificial Intelligence and Human Resource Upskilling in Emerging Markets: Evidence from Small and Medium Enterprises in Oyo State, Nigeria · Dauda Adewole Oladejo, Grace Oluwatoyin Obadare, Oluwatobiloba Joshua Olayemi
This study examines how AI adoption is associated with workforce upskilling in small and medium-sized enterprises (SMEs) in Oyo State, Nigeria, surveying 135 HR professionals, SME operators, and employees across approximately 72 firms. Results show that while AI integration in HR practices remains nascent (mean score of 2.116), key AI adoption dimensions—including organisational integration, AI training programmes, and data-driven decision-making support—are significantly and positively correlated with employee skill enhancement (R=0.750, R²=0.562, p<0.001). Major barriers include infrastructure gaps, educational deficiencies, policy shortfalls, and socio-cultural resistance. The findings highlight both the potential and the structural challenges of deploying AI for human capital development in emerging market contexts.
- ResearchPreprints.org2026-06-29Quality assurance · Certifications · +2
AI-Assisted Devices in University Written Examinations: A Structural Threat to the Degree as a Warrant of Competence · Demetrios Venetsanos
This paper argues that the convergence of miniaturized cameras, AI smart glasses, and large language models creates a 'cheating pipeline' that can compromise closed-book university examinations in a structurally undetectable way. Drawing on a narrative review of academic integrity literature (1966–2025), UK Ofqual malpractice statistics (2019–2025), and a scan of commercially available AI-assisted devices, the authors find that covert device use in invigilated exams is already increasing in secondary education, with conditions plausibly extending to higher education. The core argument is not merely about individual cheating but about a certification crisis: if a machine can silently supply answers, the degree can no longer reliably warrant the competence of its holder for the cohort as a whole. The paper calls for institutional reconsideration of examination design and what degrees certify in an era of ambient AI.
- ResearchJournal of Engineering Research2026-06-29Workforce · Enterprise
Artificial Intelligence in United States Enterprises: An Integrative Review of Adoption Patterns, Sectoral Applications, and Organizational Outcomes · Gustavo Jardim Alves
This integrative review synthesizes peer-reviewed research, government statistics, and industry surveys through early 2026 to assess how AI is actually adopted and used across U.S. enterprises. It finds a sharp split between self-reported adoption rates (above three-quarters of organizations) and production-grade deployment (high single digits to roughly one-fifth when weighted by employment), with real deployments concentrated in customer service, marketing, software engineering, financial risk management, healthcare, and factory quality control. Field experiments consistently show double-digit productivity gains, disproportionately benefiting less-experienced workers, yet most firms report no profit impact—a gap the authors explain through general-purpose-technology theory and the complementary-intangibles hypothesis. The findings carry implications for managers, policymakers, and researchers navigating an unsettled U.S. AI governance landscape.
- ResearchJournal of Economic Finance Research and Review2026-06-29AI policy
Toward Layered AI Autonomy: A Strategic Framework for Managing Dependency Risks in Transition Economies A Conceptual Framework with Illustrative Evidence from Vietnam · Pham Duy Minh
This paper proposes the Layered AI Autonomy for Transition Economies (LAAT) framework, a three-layer strategic model designed to help resource-constrained developing nations manage AI dependency risks across financial, technical, legal-contractual, and geopolitical dimensions. Applied to Vietnam, the framework identifies high dependency vulnerability in all four dimensions, with concentrated risk in legal-contractual and geopolitical areas. The authors use a 2026 export control action on Anthropic AI models as an illustrative case and conclude with a three-phase policy roadmap translatable to similar transition economies. The work is relevant to AI governance and sovereignty policy for nations navigating AI adoption while avoiding technological lock-in to a small number of advanced-economy platforms.
- ResearchManagement Theory and Studies for Rural Business and Infrastructure Development2026-06-29Workforce · Enterprise
DOES SIZE MATTER? A COMPARATIVE STUDY ON AI’S INFLUENCE ON EMPLOYEE TRAINING, BUSINESS EFFICIENCY, AND COMPETITIVE PRESSURE · Enikő Korcsmáros, Aranka Boros, Klaudia Balázs
This study of 269 private-sector firms in Hungary and Slovakia finds that company size is positively associated with AI adoption, employee upskilling investment, and efficiency gains, with micro and small firms facing the greatest resource barriers. Medium-sized firms occupy an intermediate position, while larger enterprises are better positioned to convert AI into training and productivity benefits. Notably, no significant link was found between company size and perceived competitive disadvantage from AI, suggesting industry-level factors may matter more. The findings point to a need for targeted policy support for smaller enterprises and strategic workforce planning in larger ones.
- ResearchInternational Review of the Red Cross2026-06-29Quality assurance · Certifications · +3
The preparation gap: IHL, human–machine teaming and assurance in military AI systems · Sahr Muhammedally
This paper examines how AI systems used in military targeting, intelligence analysis, and operational planning must be governed before hostilities occur rather than assessed only after deployment. The authors argue that risks such as automation bias and adversarial manipulation threaten compliance with international humanitarian law (IHL) in human-machine teaming (HMT) contexts. To address this, the paper proposes an 'HMT Assurance Card,' a cross-disciplinary lifecycle instrument that integrates IHL obligations, civilian harm pattern analysis, and AI governance frameworks into measurable standards applicable across defense institutions and coalition environments. The work is relevant to certification and policy communities seeking structured pre-deployment assurance mechanisms for high-stakes AI systems.
- ResearchÜniversite Araştırmaları Dergisi2026-06-29Quality assurance · AI policy · +2
A Model Proposal For Quality Assurance in Higher Education: The Explainable Quality Assurance Model · Hakan Kurt
This paper proposes a seven-component Explainable Quality Assurance (XQA) model for higher education institutions, arguing that AI-driven quality assurance must prioritize explainability, human oversight, and institutional legitimacy—not just speed and efficiency. The model's core principles include data transparency, indicator justification, algorithmic traceability, contextuality, ethical compliance, and contestability. The study evaluates Turkey's Higher Education Quality Council (THEQC/YÖKAK) against international frameworks including ESG 2015, UNESCO generative AI guidance, NIST's explainable AI approach, ENQA guidance, and the EU AI Act. It concludes that the future of AI-supported quality assurance lies in auditable, humanly balanced algorithmic governance rather than technical automation alone.
- ResearchJurnal Manajemen Pendidikan2026-06-29Quality assurance · AI policy · +2
Transformasi Evaluasi Pembelajaran Berbasis Artificial Intelligence Pada Sekolah Menengah Kejuruan · Rulli Widiantoro, Hindarto Hindarto, Jasno Jasno et al.
This systematic literature review (PRISMA protocol, 45 articles, 2022–2026) examines how AI-based learning evaluation tools are being adopted in Indonesian Vocational High Schools (SMK). The findings show AI implementation yields up to 68% gains in assessment efficiency, 54% improvement in competency identification accuracy, and a 43% reduction in subjective bias, across six implementation patterns including auto-grading, adaptive recommendations, and academic dishonesty detection. Key barriers include limited infrastructure (72% of cases), insufficient teacher digital competency (65%), and an absence of specific technical regulations (34%). The study recommends building contextual AI ecosystems, strengthening educator digital literacy, and establishing comprehensive student data governance policies.
- ResearchJournal of East European Management Studies2026-06-29Workforce · Enterprise · +1
Quantitative and Qualitative Implications of AI Adoption for HRM in Last-Mile E-Commerce Delivery: A Systematic Review · Ionel Bostan, Alunica Morariu, Cristina Lazăr
This systematic review of 200 studies (2013–2025) examines how AI adoption in last-mile e-commerce delivery affects both operations and human resource management. It finds substantial operational gains—including improvements in delivery speed (15–32%), cost savings (8–25%), fleet utilization (10–18%), and workforce-allocation accuracy (12–20%)—alongside qualitative challenges such as job redesign, surveillance pressures, and workplace inclusion concerns. The authors propose an integrated HRM-logistics framework emphasizing human-centered AI adoption, transparent performance metrics, and workforce upskilling. Policymakers are urged to develop context-sensitive governance to balance efficiency gains with employee well-being.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-06-29Quality assurance
Vectored Conversational AI Testing · William Argo
This paper introduces Vectored Conversational AI Testing, a dynamic testing methodology that treats conversation as a behavioral surface to reveal how AI systems adapt, stabilize, or degrade over the course of an interaction. Unlike static prompt-based testing, the approach is designed to expose emergent behaviors such as behavioral drift, reasoning breakdowns, and boundary-handling failures in real or near-real conversational conditions. The framework produces behavioral evidence intended to support evaluation, safety analysis, risk management, and compliance workflows, making it relevant to AI assurance and auditing practices.
- ResearcharXiv2026-06-28Quality assurance · Algorithms & Automated Decisions
Budgeted Act-or-Defer Multi-Agent LLM Deliberation with Local Reliability Bounds · Mengdie Flora Wang, Haochen Xie, Guanghui Wang et al.
This paper addresses the challenge of knowing when a multi-agent LLM deliberation system is reliable enough to act autonomously versus when it should escalate to human review. The authors formulate this as a 'budgeted act-or-defer' problem, using k-nearest-neighbor confidence bounds on calibration data to certify correctness before acting, with the reliability guarantee decomposed into calibration failure, residual action risk, and representation gap components. On six benchmarks against nine baselines, the method uses only 9–12% of the pre-declared wrong-action budget on activated datasets, achieving up to 84% automation and 96% acted-on accuracy, while appropriately deferring on stress-test datasets. This matters for enterprise and quality-assurance contexts because it provides an auditable, prospective mechanism for translating a user-declared error budget into a deployment operating point, rather than relying on post-hoc threshold tuning.
- ResearcharXiv2026-06-28Quality assurance · Safety & Harms
Resolution Thresholds in VLM Detection of Harmful ASCII Art Across Construction Modes and Languages · Yikai Hua, Peter West
This paper investigates how image resolution affects the ability of large Vision-Language Models (VLMs) to detect harmful content encoded as ASCII art, a known jailbreak technique that can bypass content moderation systems. The researchers evaluate eight state-of-the-art VLMs on English and Chinese ASCII art generated across eight character construction modes and ten resolution scales. Results show that detection rates decline sharply above certain resolution thresholds, with word-based construction modes being the most resistant to detection across the full resolution range. The findings expose a systematic vulnerability in VLM-based moderation and motivate the development of resolution-aware evaluation standards.
- ResearcharXiv2026-06-28Quality assurance · Health
How much of an LLM-generated clinical corpus is actually new? A production-scale measurement of content redundancy for provenance classification · Ali H. Lazem, William J. Teahan
This paper investigates whether the sheer volume of LLM-generated clinical training corpora actually reflects their information content. Analyzing 2.51 billion tokens produced by a multi-agent clinical extraction pipeline applied to 167,034 patient narratives, the authors find that only 10.9% of the output is trainable-unique content while 79.4% is redundant—meaning raw token count overstates information content by roughly ninefold. The redundancy distorts training distributions by over-representing longer, more complex cases, and in downstream tests, de-duplicating the corpus before model adaptation measurably improved clinical disease-recognition performance at equal token budget. The authors release a provenance-based classification tool openly, offering a practical method for auditing and improving LLM-generated medical datasets.
- ResearcharXiv2026-06-28AI policy · Algorithms & Automated Decisions
Spreading the Risk of Scalable Legal Services: The Role of Insurance in Expanding Access to Justice · Roee Amir, David Chriki, Harel Omer
This paper argues that liability insurance offers a more effective accountability framework for AI-powered legal services than existing tort liability or regulatory oversight approaches. The authors contend that current mechanisms—such as human oversight requirements and tort law—create cost and scalability barriers that limit access to legal assistance, particularly for underserved users. They propose a comprehensive insurance model featuring clear risk thresholds, streamlined compensation mechanisms, and performance-based premiums that incentivize quality improvement, distributing risk across users rather than concentrating it. The framework is presented as a way to scale automated legal services while maintaining user protections through market-driven risk management.