News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
Scaling Learning-based AEB with Massive Unlabeled Data
Xiangyu Wang, Yang Zhan, Mengxiang Hao et al.
arXiv · 2026-06-17
This paper presents a semi-supervised learning framework called MF-SSL (meta-feedback semi-supervised learning) for scaling automatic emergency braking (AEB) systems using massive amounts of unlabeled fleet driving data. The method uses a teacher model to generate pseudo-labels for unlabeled data, updated via a small labeled anchor set, and includes two stabilization mechanisms—Noise-Aware Decoupling and kinematics-gated pseudo-labeling—to suppress errors that arise from anchor ambiguity and labeled-unlabeled data mismatch. Trained on up to 1 billion data windows, the resulting student model was deployed to hundreds of thousands of vehicles and validated over 10^9 km of driving, achieving a positive-to-false activation ratio exceeding 100:1 and a 35% improvement in accident-free driving mileage compared to a production rule-only baseline. The work demonstrates that semi-supervised learning can substantially improve real-world automotive safety systems at scale while maintaining driver comfort.
- Enterprise
- Quality assurance
Research
Improving Human-Robot Teamwork in Urban Search and Rescue Through Episodic Memory of Prior Collaboration
Taewoon Kim, Emma van Zoelen, Mark Neerincx
arXiv · 2026-06-17
This paper investigates whether robots in Urban Search and Rescue (USAR) scenarios can leverage recorded past collaboration patterns—stored as knowledge-graph episodic memories—to become better teammates from the very start of a new interaction. Using the MATRX simulation environment, the researchers had human participants externalize their teamwork strategies via a chat and reflection interface, then applied graph representation learning to automatically select the most effective prior collaboration pattern to initialize the robot before a new episode. Across 20 participants and 160 round-level observations, initializing the robot with a single automatically selected prior collaboration pattern increased rescue success from 25.7% to 41.3% and reduced average task time by 283 seconds, with the strongest gains appearing early in the interaction. These findings suggest that episodic memory of prior collaboration can meaningfully accelerate human-robot team performance, offering a practical path toward robots that adapt to human partners more quickly and effectively.
- Workforce
- Enterprise
Research
GateMem: Benchmarking Memory Governance in Multi-Principal Shared-Memory Agents
Zhe Ren, Yibo Yang, Yimeng Chen et al.
arXiv · 2026-06-17
GateMem is a new benchmark designed to evaluate how well LLM-based memory agents handle shared, multi-user environments such as hospitals, workplaces, campuses, and households. Unlike prior benchmarks that assume a single user, GateMem tests three dimensions simultaneously: utility for legitimate long-horizon requests, access control across different roles and authorization boundaries, and reliable deletion ('active forgetting') of information after explicit requests. Tested across diverse baselines and backbone models, no existing method achieves strong performance on all three dimensions—long-context prompting offers the best governance scores but at high token cost, while retrieval-based methods reduce cost yet still leak unauthorized or deleted information. The findings indicate that current memory agents are not yet reliable enough for shared institutional deployment.
- Enterprise
- Quality assurance
Research
ProfiLLM: Utility-Aligned Agentic User Profiling for Industrial Ride-Hailing Dispatch
Tengfei Lyu, Zirui Yuan, Xu Liu et al.
arXiv · 2026-06-17
ProfiLLM is an agentic LLM pipeline that generates behavioral user profiles for ride-hailing dispatch at industrial scale, specifically targeting DiDi's production platform. It addresses three core challenges: log data exceeding LLM context windows, sparse data for long-tail users, and profiles that are fluent but not predictively useful. The system uses tool-augmented knowledge mining and utility-aligned profile exploration with DPO fine-tuning to produce profiles that measurably improve downstream prediction and business outcomes. Deployed in a 14-day online A/B test, ProfiLLM achieved +0.47% GMV, +0.33% Completion Rate, and -0.82% Cancel-Before-Accept rate, demonstrating that LLMs can serve as practical semantic feature extractors in latency-sensitive production matching systems.
- Enterprise
Research
EARS: Explanatory Abstention for Reliable Sub-Agent Modeling in Large-scale Multi-Agent Systems
Shuang Xie, Yunan Lu, Han Li et al.
arXiv · 2026-06-17
EARS (Explanatory Abstention for Reliable Sub-Agent Modeling) is a framework designed to improve the reliability of large-scale multi-agent systems (MAS) used in enterprise settings, where a coordinator delegates tasks to specialized sub-agents. The core problem is that sub-agents built on smaller fine-tuned models tend to over-answer ambiguous or misrouted requests, producing hallucinated outputs rather than useful feedback. EARS reframes abstention as a structured inter-agent communication protocol: sub-agents are fine-tuned to detect failure conditions and return actionable rationales to the coordinator for clarification, rerouting, or fallback. Evaluated in a production e-commerce business intelligence assistant, EARS improved the overall response pass rate from 68.5% to 78.9%, demonstrating measurable reliability gains in enterprise AI deployments.
- Enterprise
- Quality assurance
Research
Gender Bias in LLM Hiring Decisions: Evidence from a Japanese Context and Evaluation of Mitigation Strategies
Serena A. Hoffstedde, Machiko Hirota, Akshara Nadayanur Sathis Kanna et al.
arXiv · 2026-06-17
This study investigates gender bias in LLM-assisted hiring decisions within a Japanese corporate context, using 60 rirekisho-format resumes, 12 name pairs, and five leading LLMs across 43,200 API calls. The results confirm a significant pro-female bias across all five models, replicating patterns previously found in Western research and extending them to a non-Western setting. A prompt-level gender-neutrality instruction failed to meaningfully reduce the bias, but removing candidate names from prompts nearly eliminated the female effect, identifying the name as the primary channel of gender information. The study also surfaces a practical deployment challenge: a 42% refusal rate when combining name anonymization with GPT-4o's content safety filter, complicating the use of privacy filters in real hiring pipelines.
- Workforce
- AI policy
Research
AI-Driven Assessment of Human Tutors: Linking Training Performance to Real-Life Practice
Danielle R. Thomas, Marie Cynthia Abijuru Kamikazi, Clara Brandt et al.
arXiv · 2026-06-17
This paper presents an AI-driven system using Gemini-2.5-pro to assess both training performance and real-life tutoring sessions for human math tutors, bridging the gap between simulation-based training and authentic classroom practice. Across 86 tutors and 405 session-to-lesson pairs, training performance significantly predicted real-life tutoring quality with an effect size of 0.25 SD, with open-response scores being more predictive than multiple-choice. Tutors showed a 7.4% average learning gain during training and demonstrated measurable improvements in both encountering and executing pedagogical opportunities in real sessions over time. The system contributes open datasets, AI prompts, and scoring rubrics to support reproducibility in AI-driven tutor evaluation.
- Workforce
- Quality assurance
Research
Green AI adoption for circular and agile supply chains: ESG-driven pathways in Vietnam's emerging technological economy
Bang Nguyen‐Viet, Luu Chi-Luong, Ngan Nguyen-Khanh
International Journal of Productivity and Performance Management · 2026-06-17
This study surveys 780 Vietnamese medium- and large-sized enterprises to examine how green AI adoption drives sustainable business performance. Using structural equation modeling, the findings show that green AI positively enhances green circular capacity, supply chain agility, and green innovation, all of which improve business performance. Green digital orientation amplifies AI's effect on supply chain agility, and ESG compliance strengthens the link between green innovation and performance. The results demonstrate measurable pathways through which AI and sustainability capabilities translate into accountable outcomes in emerging markets.
- Enterprise
- AI policy
- Quality assurance
Research
The Impact of Government Green Procurement on Corporate Carbon Emission Reduction: A Dual Mediation Perspective of Artificial Intelligence and Green Finance
Z K Zhang, Jianmin Wu
Sustainability · 2026-06-17
This study examines how government green procurement policies affect corporate carbon emission reduction among Chinese A-share listed companies from 2020 to 2024. Using two-way fixed effects models and Bootstrap methods, the authors find that green public procurement significantly improves firms' carbon reduction performance, with AI adoption and government green subsidies acting as mediating mechanisms that amplify this effect. The impact is strongest for state-owned enterprises, high-tech firms, and companies in regions with more advanced digital economies. The findings offer actionable guidance for policymakers seeking to align procurement policy with digital technology and green finance to accelerate decarbonization.
- AI policy
- Enterprise
Research
Competency Based Education in Information Technology: Designing AI-Driven, Outcome-Focused Learning Pathways
Salmon Oliech Owidi, Kelvin Kabeti Omieno
Zenodo (CERN European Organization for Nuclear Research) · 2026-06-17
This systematic review of 124 studies explores how AI-driven Competency-Based Education (CBE) frameworks can replace time-based IT curricula with skill-mastery pathways. The paper finds that adaptive learning improves outcomes by 0.35–0.65 standard deviations, explainable AI boosts complex problem-solving with effect sizes up to 0.58, and automated assessment cuts feedback latency from days to seconds while raising student satisfaction 25–30%. The authors propose a framework integrating explainable AI and PEARL principles, and recommend aligning competency taxonomies with CC2020 and industry certifications, piloting AI assessments in low-stakes settings, and funding longitudinal research on graduate employment outcomes.
- Workforce
- Certifications
- Quality assurance
- Enterprise
- AI policy
Research
Human Capital Ratio (HCR): A Fiscal Framework for the Age of Artificial Intelligence
Nandeep Nagarkar
Zenodo (CERN European Organization for Nuclear Research) · 2026-06-17
This paper proposes the Human Capital Ratio (HCR), a fiscal policy framework designed to counteract AI-driven labor displacement in the enterprise economy. Modeled on healthcare's Medical Loss Ratio and 401(k) non-discrimination testing, HCR creates a dual-trigger levy system that activates when firms' AI spending outpaces human labor spending or when compensation becomes skewed toward highly paid employees. Levy proceeds fund a cost-of-living-indexed social minimum income for displaced workers, with eligibility tied to trigger events rather than documented terminations, and surplus funds directed to workforce retraining. The framework also mandates forward-looking efficiency disclosures at the point of AI capital commitment to close avoidance vectors that the spending ratio alone cannot address.
- Workforce
- Enterprise
- AI policy
Research
Environmental Competitiveness and Green Entrepreneurship: The Mediating Role of Generative Artificial Intelligence in Pakistan’s Manufacturing Sector
Muhammad Ahsan Iqbal, Asifa Muhammad Sabir, Hufza Iqbal
Inclusive Society and Sustainability Studies · 2026-06-17
This study examines how environmental competitiveness drives green entrepreneurship in Pakistan's manufacturing sector, with generative AI (Gen-AI) adoption acting as a partial mediator. Using PLS-SEM with survey data from 250 employees, the authors find that environmental competitiveness significantly predicts both Gen-AI adoption (β=0.257) and green entrepreneurship (β=0.178), while Gen-AI adoption in turn predicts green entrepreneurship (β=0.216) with a confirmed indirect mediation effect (β=0.055). The findings suggest that AI adoption can serve as a strategic mechanism linking competitive environmental pressures to sustainable business outcomes, with practical implications for managers and policymakers in emerging economies.
- Enterprise
- AI policy
- Workforce
Research
Policy, Regulation, and the Public Good of Artificial Intelligence
Robert Ipiin Gnankob, Jayanta Kumar Mohapatra, Samuel Etse Dzakpasu
Advances in computational intelligence and robotics book series · 2026-06-17
This chapter develops a 'Policy–Regulation–Governance Compass' framework grounded in public value theory, responsible innovation, and stakeholder governance to guide ethical AI development. Through comparative analysis of global regulatory approaches—including the EU AI Act, U.S. AI Bill of Rights, and emerging African and Asian strategies—it identifies asymmetries in regulatory capacity and persistent challenges such as regulatory lag, power concentration, and global inequality. The authors argue that governance functions as an ethical enabler rather than a constraint on innovation, and propose actionable pathways emphasizing adaptive frameworks, institutional capacity, inclusivity, and international cooperation. The work matters because it offers policymakers a structured lens for aligning AI development with equitable and sustainable public good outcomes.
- AI policy
Research
Artificial Intelligence as a Strategic Driver of Environmental Sustainability: Unpacking the Mediating Role of Green Governance in GCC Industrial Firms
Ruaa Binsaddig, Amina Toumi, Reem Khamis et al.
Sustainability · 2026-06-17
This study examines how AI adoption drives environmental sustainability among 75 publicly listed industrial firms across six GCC countries from 2018 to 2025, using fixed-effects and bootstrapped mediation analyses. Results show AI adoption is positively and significantly associated with environmental sustainability, with green governance (board-level ESG structures) partially mediating this relationship. The findings suggest AI's environmental benefits are most fully realized when embedded within sound corporate governance frameworks, offering actionable guidance for policymakers and managers in resource-intensive industries pursuing digital transformation.
- Enterprise
- AI policy
- Quality assurance
Research
AI-Augmented Compliance Auditing for Cloud Systems: A Hybrid ML–LLM Approach
Moise Iradukunda Ingabire, Jema David Ndibwile
Future Internet · 2026-06-17
This paper presents a hybrid AI compliance auditing system combining XGBoost multi-label classification and GPT-4o-mini large language model analysis, evaluated against Rwanda's National Cyber Security Authority standards (169 controls across 14 families). The system achieves 85.1% F1 on real-world logs with a 6.4% false-positive rate, and 92.8% macro detection across adversarial MITRE ATT\_CK scenarios with 0.0% false-positive rate on compliant logs. A key contribution is identifying and correcting an 86.3% data-leakage flaw that had artificially inflated prior results to 99.99%, improving result credibility. The system runs on $50/month cloud infrastructure and generates audit reports in 2–5 seconds, demonstrating that effectiveness-based compliance auditing is viable without enterprise-grade resources.
- Enterprise
- Quality assurance
- Certifications
- AI policy
Research
Competency Based Education in Information Technology: Designing AI-Driven, Outcome-Focused Learning Pathways
Salmon Oliech Owidi, Kelvin Kabeti Omieno
Zenodo (CERN European Organization for Nuclear Research) · 2026-06-17
This systematic review of 124 studies proposes a framework for integrating AI-driven Competency-Based Education (CBE) into IT programs, combining explainable AI tools with adaptive learning to create transparent, personalized learning pathways. The paper reports that adaptive learning improves outcomes by 0.35 to 0.65 standard deviations, automated assessment reduces feedback latency from days to seconds, and increases student satisfaction by 25 to 30 percent, while cutting per-learner assessment costs from $45–$60 (human) to $12–$18 (AI). Key recommendations include aligning competency taxonomies with CC2020 and industry certifications, piloting AI assessment in low-stakes contexts, and investing in faculty AI literacy to build workforce-ready graduates. The findings are directly relevant to educators, certifying bodies, and policymakers seeking to modernize IT credentialing and improve graduate employment outcomes.
- Workforce
- Certifications
- Quality assurance
- AI policy
Research
An Adaptive Governance-Centric MLOps Framework for Risk-Tiered Control and Continuous Assurance of Responsible AI in High-Stakes Domains
Sunilkumar Reddy Eraganeni
International Journal of Computational and Experimental Science and Engineering · 2026-06-17
This paper proposes an Adaptive Governance-Centric MLOps Framework that embeds responsible AI principles—including explainability, fairness checking, and audit logging—directly into the machine learning lifecycle for high-stakes domains like finance and healthcare. Validated on the German Credit Dataset using six ML and deep learning models, the framework shows that governance enforcement (regulatory compliance, bias detection, drift monitoring) can be integrated without sacrificing predictive performance, with XGBoost reaching 94.57% accuracy and the MLP model meeting strict fairness thresholds. The study demonstrates that continuous compliance monitoring, including automated retraining triggers upon distributional shift, is achievable within a unified MLOps architecture. These findings are directly relevant to organizations and regulators seeking to operationalize requirements under the EU AI Act and GDPR in production AI systems.
- Quality assurance
- Certifications
- AI policy
- Enterprise
Research
Pedagogical Audit Framework: The Multimodal Video Evaluator (MVE) Knowledgebase Version 1.1
Koichi SATO
arXiv · 2026-06-17
This paper presents the Multimodal Video Evaluator (MVE) Version 1.1, an AI-assisted system designed to automate pedagogical auditing of educational video content in higher education. The system uses a knowledgebase grounded in the Cognitive Theory of Multimedia Learning, video engagement research, and visual scaffolding frameworks to evaluate multimodal attributes like visual transitions, text density, audio-visual synchronization, and instructional pacing. Version 1.1 refines the prior architecture by eliminating semantic inconsistencies in its rule set to improve deterministic reasoning and reduce LLM processing conflicts. The system addresses the scaling challenge of auditing large institutional media libraries, moving beyond manual low-percentage sampling to programmatic identification of cognitive bottlenecks and prediction of student engagement patterns.
- Quality assurance
- Certifications
- Enterprise
Research
FPGA-Based Reconfigurable SoCs for Safety-Critical AI Inference: A Systematic Literature Review
Yasmeen M. Hussein, Raaed F. Hassan, Raad Farhood Chisab
Electronics · 2026-06-17
This systematic review of 36 studies on FPGA-based reconfigurable SoC platforms for safety-critical AI inference finds that 92% of surveyed work ignores safety certification entirely. While these platforms demonstrate strong performance—including energy efficiencies up to 60 GOPS/W and 2–5× throughput improvements via dynamic partial reconfiguration—none of the AI accelerator studies provide worst-case execution time bounds or formal verification required for safety standards like ISO 26262, ISO 21448, and ISO/PAS 8800. The authors identify conformal prediction as a promising hardware-compatible framework for uncertainty quantification on resource-constrained FPGAs and propose a research agenda to close the gap between performance optimization and certified deployment in transportation and industrial automation contexts.
- Certifications
- Quality assurance
- AI policy
Research
Human Capital Ratio (HCR): A Fiscal Framework for the Age of Artificial Intelligence
Nandeep Nagarkar
Zenodo (CERN European Organization for Nuclear Research) · 2026-06-17
This paper proposes the Human Capital Ratio (HCR), a fiscal policy framework designed to counteract AI-driven labor displacement in the enterprise economy. Modeled on the Medical Loss Ratio in healthcare and 401(k) non-discrimination testing, HCR uses a dual trigger — one based on human-vs-AI spend ratios and one based on compensation skew toward highly paid employees — to levy funds from organizations whose AI adoption displaces rank-and-file workers. Levy proceeds flow into a pooled national fund providing cost-of-living-indexed income support and retraining resources for displaced workers, with eligibility anchored to the levy trigger rather than documented termination reasons. The framework also includes forward-looking efficiency disclosure requirements and a chain-of-custody model covering contractors, gig platforms, and open-source deployments to close key avoidance vectors.
- Workforce
- Enterprise
- AI policy
Research
Integrating Artificial Intelligence Into Financial and Strategic Frameworks: A Path To Entrepreneurial Success In Emerging Markets
Oktaria Ardika Putri
KASTA Jurnal Ilmu Sosial Agama Budaya dan Terapan · 2026-06-17
This study investigates how AI integration in financial management and strategic planning affects entrepreneurial success among 200 Indonesian startups, using a mixed-methods design combining PLS-SEM analysis and qualitative interviews. Results show that AI-driven financial management (β=0.312, p<0.001) and AI-driven strategic planning (β=0.287, p<0.001) both significantly predict entrepreneurial success, together explaining 52.4% of the variance, with operational efficiency mediating the strategy-success relationship. The findings extend Dynamic Capabilities Theory into AI-driven startup ecosystems in emerging economies, offering practical guidance for startups seeking sustainable growth through AI adoption. This matters for enterprise decision-makers and policymakers in emerging markets looking to understand how AI tools can be strategically embedded to drive measurable business outcomes.
- Enterprise
- Workforce
- AI policy
Research
AI Sandboxes: A Threat Model, Taxonomy, and Measurement Framework
Inderjeet Singh, Haitham Mahmoud, Andrés Murillo
arXiv · 2026-06-16
This paper develops a formal framework for AI sandboxes—controlled environments used to test and validate AI systems, including physical AI, AIoT, and cyber-physical systems. It introduces a threat model that accounts for attacks on the assurance apparatus itself, a taxonomy of sandbox archetypes, and a measurement framework covering fidelity, controllability, observability, containment, reproducibility, and governance artifacts. Three real-world case studies are used to instantiate the framework. The work clarifies what a sandbox can validly test, which risks it can contain, and what evidence it can provide to support safety, security, and regulatory assurance—making it directly relevant to certification and quality-assurance processes for AI systems.
- Quality assurance
- Certifications
- AI policy
Research
Evaluating Prompting-Based Defenses Against Domain-Camouflaged Injection Attacks
Aaditya Pai
arXiv · 2026-06-16
This paper evaluates five prompting-based defenses against domain-camouflaged prompt injection attacks, where malicious instructions are hidden in retrieved content using domain-appropriate vocabulary to evade standard detectors. Across 3,510 trials spanning three AI model families (Claude Haiku, Llama 3.1 8B, Gemini 2.0 Flash) and three deployment domains (financial, legal, general), the authors find that paraphrasing retrieved content before agent processing is the most consistently effective defense, reducing camouflage attack success rates by 55–84% depending on the model. Defense effectiveness is strongly model-dependent—spotlighting halves attack success on Claude Haiku but provides no benefit on Llama 3.1 8B—and financial domain deployments face the highest residual risk at 26–33% baseline attack success rate, with no prompting-based defense fully eliminating the threat on weaker models. These findings provide the first systematic benchmark of prompting-based defenses against camouflage-class injection attacks and offer practical recommendations for enterprise AI deployments.
- Enterprise
- Quality assurance
Research
Towards Scalable Customization and Deployment of Multi-Agent Systems for Enterprise Applications
Paresh Dashore, Shreyas Kulkarni, Uttam Gurram et al.
arXiv · 2026-06-16
This paper presents a unified framework for adapting and deploying large language model (LLM)-based multi-agent systems in enterprise settings. The first stage combines continual pretraining, supervised fine-tuning, and preference optimization to tailor compact models to specialized domains while preserving agentic capabilities. The second stage applies speculative decoding and FP8 quantization to reduce inference costs, achieving a 4.48x speedup in throughput with minimal quality loss. The framework enables rapid domain adaptation and improved robustness on long-tail scenarios across enterprise workloads.
- Enterprise
Research
Ghost Vectors: Soft-Deleted Embeddings Remain Reconstructible in HNSW Vector Databases
Chandranil Chakraborttii, Jackeline García Alvarado, Sitora Abdulofizova et al.
arXiv · 2026-06-16
This paper exposes a critical privacy vulnerability in HNSW vector databases commonly used in RAG (retrieval-augmented generation) pipelines: when records are 'soft-deleted,' their embeddings remain physically on disk and can be reconstructed using the Vec2Text inversion model without any domain-specific fine-tuning. Experiments across multiple datasets demonstrate alarming recovery rates—including 100% recovery of patient age and gender from medical data, 99% top-1 identity recovery from facial embeddings, and 100% tissue classification from histopathology images—raising serious compliance concerns under GDPR Article 17 and HIPAA. The authors propose 'Epoch Key Rotation,' an encryption-based deletion mechanism that reduces PII recovery to 0% and completes in approximately 0.005 ms per record, while generating an ECDSA-signed cryptographic proof as an auditable deletion record. These findings have direct implications for enterprises deploying AI systems that handle sensitive personal data and for regulators assessing compliance of AI data pipelines.
- AI policy
- Enterprise