News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
Governance and Security-by-Design: Embedding Safety and Alignment into Agentic AI Systems
Himanshu Joshi, Shivani Shukla
arXiv · 2026-07-15
This paper proposes a governance and security-by-design framework that embeds safety and alignment mechanisms directly into agentic AI architectures rather than relying on external oversight. Across 800 experiments, the authors identify three critical failure modes: AI-generated code introduces memory safety vulnerabilities in 42.7% of efficiency-focused cases and cryptographic flaws in 21.1% of security-focused cases, security issues worsen by 37.6% after five feedback iterations, and domain expertise degrades by 47% as irrelevant context accumulates. Their multi-agent governance architecture, validated through industry partnerships, achieves a 40% reduction in post-deployment safety incidents. The findings argue that autonomous AI systems require governance to be built into their architecture to remain safe and aligned as complexity scales.
- AI policy
- Quality assurance
Research
Caught in Between Protection and Autonomy: A Scoping Review of Youth’s Right to Participation in Artificial Intelligence
Lauriane Lalande, Sarah Bouhouita-Guermech, Hazar Haidar
Youth · 2026-07-15
This scoping review examines how scientific literature conceptualizes young people's right to participate in AI governance and design. Analyzing 15 sources from 2010–2024, the study finds that while frameworks like the OECD AI Principles and UNESCO's AI Ethics Recommendation exist, they rarely incorporate youth perspectives despite children's widespread exposure to AI in educational and social contexts. The literature frames youth participation as both a children's rights requirement under UN General Comment No. 25 and a condition for legitimate digital governance, yet a major gap remains between youth exposure to AI systems and their actual influence over decisions. The authors argue that integrating youth participation into AI governance and education could lead to more inclusive and developmentally informed technological ecosystems.
- AI policy
Research
Artificial Intelligence and the Future of the Labour Economy:A Multi-Criteria Expert Evaluation of Institutional Models of Adaptation
Jurica Bosna, Adis Puška, Darko Božanić
Journal of Soft Computing and Decision Analytics · 2026-07-15
This study uses a hybrid fuzzy-rough multi-criteria decision-making framework (combining SiWeC and WASPAS methods) to evaluate six institutional models for adapting labour markets to AI and automation. Expert judgments across eight criteria—most notably innovation/productivity incentives and institutional feasibility—were applied in a Croatian case study, finding that mass retraining and participatory AI capital models are the most suitable responses. The research supports evidence-based policymaking for workforce adaptation in the AI era while also contributing a methodological advance by explicitly handling uncertainty and imprecision in expert evaluations.
- Workforce
- AI policy
Research
The Missing Link to Safe AI Homologation
Mark Locherer, Michael Kordovan, Bernd Buxbaum et al.
arXiv · 2026-07-15
This paper examines the challenges of certifying AI-based safety-related systems under the EU AI Act and Machinery Regulation, arguing that traditional safety assurance approaches are insufficient. The authors introduce the concept of an 'efficient assurance argument' as a structured means to demonstrate regulatory compliance throughout a system's lifetime, covering risk management, design, verification, and validation. They conclude that this assurance argument acts as the critical missing link in AI homologation, serving both engineering teams and notified bodies responsible for type approval.
- Certifications
- AI policy
Research
Addressing benchmarking gaps in large language models for health and medicine with dynamic red-teaming
Jiazhen Pan, Bailiang Jian, Paul Hager et al.
Nature Health · 2026-07-15
This paper introduces DAS (Dynamic, Automatic and Systematic), a red-teaming audit framework that continuously stress-tests large language models across four safety-critical dimensions: robustness, privacy, bias, and hallucination. Applied to 15 state-of-the-art LLMs, the framework revealed a severe 'benchmarking gap': despite median MedQA accuracy above 80%, 94% of correct answers failed under dynamic robustness testing, privacy leaks were elicited in 86% of scenarios, cognitive bias altered recommendations in 81% of fairness tests, and hallucination rates exceeded 74%. The findings suggest that high static benchmark scores may reflect superficial memorization rather than genuine safety, and that continuous adversarial auditing is necessary before LLMs are deployed in consumer health or clinical settings.
- Quality assurance
- Certifications
Research
ЭКОНОМИЧЕСКАЯ ОЦЕНКА ВНЕДРЕНИЯ ИНСТРУМЕНТОВ ИСКУССТВЕННОГО ИНТЕЛЛЕКТА В БИЗНЕС-ПРОЦЕССЫ МАЛОГО ПРЕДПРИЯТИЯ
Егоров Д. Б.
CyberLeninK (CyberLeninka) · 2026-07-15
This paper develops a methodological framework for economically evaluating AI tool adoption in small business processes. It argues that simply equating saved labor time with monetary savings overstates AI project returns, since freed-up capacity does not automatically translate into cost reductions or additional output. The proposed model integrates total cost of ownership with effects from labor time savings, marginal revenue growth, error reduction, and loss prevention, alongside correction coefficients for actual utilization, quality, and result attribution. A pilot algorithm and model example in a small business services firm demonstrate that investment attractiveness depends heavily on how freed time is actually used and whether commercial effects are sustained.
- Enterprise
Research
Rethinking Penetration Testing for AI-Enabled Systems: From Resource Compromise to Behavioral Objective Violation
Mohammad Allahbakhsh, Mohammad Hassan Bahari, Moslem Attar-Raouf
arXiv (Cornell University) · 2026-07-15
This paper argues that traditional penetration testing—focused on infrastructure and resource compromise—is insufficient for AI-enabled systems, where adversaries can alter system behavior through prompt injection, data poisoning, sensor manipulation, and other influence pathways without ever breaching underlying infrastructure. The authors reframe penetration testing as objective-driven behavioral evaluation, defining success as the feasible induction of AI-governed behavior that violates operational objectives under an explicit threat model. They propose a structured testing workflow and illustrate the approach with an AI-enabled security operations center assistant. The framework provides practical guidance for evaluating adversarial risk in deployed AI systems.
- Quality assurance
- Certifications
Research
CAVA: Canonical Action Verification and Attestation for Runtime Governance of Agentic AI Systems
Zexun Wang
arXiv (Cornell University) · 2026-07-15
CAVA introduces a runtime-semantics layer that converts heterogeneous agentic AI activity—spanning coding hooks, browser automation, API gateways, and workflow engines—into canonical runtime action objects, enabling consistent governance across incompatible runtime records. The system formalizes how an approved action is identified, how approval is bound to execution evidence, and how an independent verifier can later reproduce the same action identity. A 384-variant benchmark covering semantic equivalence, wrapper bypass, tamper detection, and runtime portability validates the reference implementation. The work addresses a foundational gap in deployer-side AI governance: ensuring that what was approved is verifiably what was executed.
- AI policy
- Enterprise
Research
From promise to practice: artificial intelligence in mental health care in the MENA region
Sara El Hajj, Ahmad Nsouli, Mohamad Wehbe et al.
Frontiers in Psychiatry · 2026-07-15
This narrative review examines the use of AI-driven conversational tools—including large language models and psychotherapy chatbots—for mental health care in the Middle East and North Africa (MENA) region, where treatment gaps reach 80–95% due to provider shortages, stigma, and cultural barriers. The review finds that while AI tools offer high accessibility and user engagement for low-intensity support, their effectiveness is constrained by linguistic mismatches such as Arabic diglossia and poor alignment with locally grounded expressions of distress. A paradox emerges in which stigma and privacy concerns drive users toward anonymous AI tools yet simultaneously limit trust in their clinical reliability, reinforcing preference for hybrid human-oversight models. The authors conclude that current systems are insufficiently adapted to the MENA context and call for culturally grounded, dialect-sensitive, and clinically supervised approaches.
- AI policy
- Workforce
Research
Human-like conversational agents as social partners: a scoping review of socioaffective mechanisms, well-being outcomes, risks and governance in the post-Turing era
Qian Li, Han Geng, Xin Hu et al.
Frontiers in Artificial Intelligence · 2026-07-15
This scoping review synthesizes 58 sources on how large-language-model-based companion agents produce human-like social behavior and what psychosocial consequences follow. Therapeutic chatbot studies showed the strongest evidence for short-term symptom reduction, while evidence for sustained loneliness reduction in open-domain companions remained preliminary. The review identifies risks including dependency, displacement of human relationships, sycophancy, manipulation, and harms to vulnerable users, and proposes a 'relational safety stack' as an evaluation and governance agenda for managing psychosocial impact as these systems scale.
- AI policy
- Quality assurance
Research
Human Capital Readiness for Cloud–AI–Data Center Ecosystems: A Bibliometric Review
Rinaldi Noor, Agus Rahayu, Ratih Hurriyati et al.
Human Resources Management and Services · 2026-07-15
This bibliometric review examines 783 Scopus-indexed journal articles (2010–2025) to map the intellectual landscape of human capital readiness for cloud, AI, and data center ecosystems. The study finds that publication activity accelerated after 2019 and again after 2023, reflecting growing scholarly attention to digital workforce readiness. Analysis reveals a persistent divide between technology-oriented research (AI, cloud computing, Industry 4.0) and human-centered research (digital skills, competencies, workforce readiness), and identifies gaps in Asia-Pacific scholarship despite the region's expanding digital infrastructure role. The authors argue for repositioning human capital readiness as a strategic HR capability integrating digital talent development, HR analytics, and workforce planning.
- Workforce
- Enterprise
Research
Algorithmic Management and Employee Outcomes: Dual Effects of Control and Capability on Autonomy and Career Security
Stefan Milojević, Miroslav Knežević, Aleksandra Vujko
Administrative Sciences · 2026-07-15
This study of 441 hotel employees in Bavaria, Germany finds that algorithmic management has two distinct effects on workers: algorithmic control reduces perceived autonomy, while AI-supported skill development increases both autonomy and career security. Using structural equation modeling and mediation analysis, the research shows that the control-oriented and developmental dimensions of AI have meaningfully different consequences for employee work experiences. The findings highlight the importance of distinguishing between these two dimensions when designing or studying AI-enabled management systems in labor-intensive service organizations.
- Workforce
Research
Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)
Olivier Martinez
arXiv (Cornell University) · 2026-07-15
This critical survey reviews 45 studies on Generative Engine Optimization (GEO)—the practice of making content more likely to be cited or used by AI-powered generative search engines. The authors find that while some content-level tactics (such as topical relevance and position in context) reliably affect citation within a fixed retrieval setting, no reviewed technique demonstrates a stable, long-term causal effect on organic discoverability or user behavior across platforms. The survey identifies key methodological weaknesses across the field, including inconsistent terminology, low reproducibility, run-to-run variability, and fidelity gaps, and proposes a formal multistage model and evidence standards to guide future work. The findings matter for enterprises and content creators seeking to maintain visibility in AI-mediated information environments, as the evidence base for most GEO tactics remains narrow and context-dependent.
- Enterprise
- Quality assurance
Research
From disinformation to Foreign Information Manipulation and Interference: policy and legal challenges for European security
Carlos Galán-Cordero, Álvaro Cremades-Guisado
Frontiers in Political Science · 2026-07-15
This article analyzes the European legal and policy framework for addressing disinformation and Foreign Information Manipulation and Interference (FIMI), examining instruments such as the Digital Services Act, the Code of Practice on Disinformation, and the European Democracy Action Plan. It finds that classical disinformation paradigms are insufficient to counter modern information operations characterized by inauthentic coordination, algorithmic amplification, generative AI use, and attribution difficulties. The study identifies strengths in the European model—transparency, due diligence, and systemic risk assessment—while flagging persistent gaps in institutional coordination, attribution capabilities, and technological adaptation. It argues for a preventive, rights-based, multilevel architecture that strengthens democratic resilience without indefinitely securitizing the digital public sphere.
- AI policy
Research
AI-accelerated End-to-End Framework for Rapid Professional Upskilling
Tam Nguyen, H S. Nguyen, Robert Ogburn
arXiv (Cornell University) · 2026-07-15
This paper presents an end-to-end AI-accelerated framework for professional upskilling spanning five stages: knowledge acquisition, content development, content review, teaching, and assessment development. The framework was validated externally when NASBA approved a program built on it for continuing professional education credits, three learners passed the NVIDIA Certified Professional in Agentic AI exam in a notably short time, and the framework's knowledge base generated a 1,267-item risk dataset for managing multi-agent AI systems. The authors argue this addresses a growing enterprise skills gap, where the average time to close such gaps grew from roughly 3 days in 2014 to 36 days in 2018. The work is notable for covering the full upskilling pipeline rather than accelerating only individual stages.
- Workforce
- Certifications
Research
Persona Migration and Expectation Recalibration in Generative AI Adoption: A Longitudinal Study at a State Department of Transportation
Omidreza Shoghli, Fatemeh Banani Ardecani, Amin Mohamadi Hezaveh
arXiv (Cornell University) · 2026-07-15
This longitudinal study tracked 124 state Department of Transportation employees over an eight-week Microsoft 365 Copilot pilot, finding that perceived usefulness declined significantly after hands-on use, suggesting initial expectations were overstated. Using persona-based clustering, the study found substantial individual shifts: 40% of initial Skeptics became more positive, while 68% of initial Champions became less enthusiastic. Job and skills concerns increased over the pilot period, even as accuracy and privacy concerns fell. The authors recommend dynamic monitoring of AI adoption and persona-specific training, workflow examples, and trust-calibration safeguards for public-sector deployments.
- Workforce
- Enterprise
- AI policy
Research
Messy Research, Certification and the Monetization of Science
J. Fourie
arXiv (Cornell University) · 2026-07-15
This theoretical economics paper examines how AI tools that lower the cost of producing polished research manuscripts reshape the institutions that certify scientific quality. Because AI reduces production costs faster than it reduces the cost of evaluating genuine contribution, polish becomes a less reliable signal and the average quality of uncertified submissions can decline. This deterioration of the outside option increases willingness to pay for credible certification, giving certifiers with market power the ability to capture a premium, while fixed review capacity can lead to certification dilution instead. The core finding is that cheaper AI-assisted research production shifts scarcity from making work look credible to verifying which work actually is credible.
- Certifications
- Quality assurance
- AI policy
Research
Thinking about the Impact of Artificial Intelligence on U.S. Health Care Costs and Spending Growth
Bob Kocher, Brian Zhao, Erin Duffy
NEJM Catalyst · 2026-07-15
This article examines how AI is likely to affect U.S. health care costs and spending across areas such as drug innovation, remote patient monitoring, chronic care management, clinical decision support, and administrative labor. The authors argue that under the current fee-for-service payment model and consolidated hospital and insurance markets, AI is more likely to increase total costs and spending growth in the short to medium term, even while improving patient access and clinical quality. They contend that the cost-reducing potential of AI depends heavily on payment model structure—value-based versus fee-for-service—and that without policy and reimbursement reforms, AI will not slow cost growth. Regulators are urged to introduce policy and reimbursement levers to ensure AI delivers on its cost-bending potential.
- AI policy
- Enterprise
- Workforce
Research
From the EU AI Act to Audit Practice: A Governance-to-Controls Framework for Quality Management and Evidence
János Kálmán
Accounting and Auditing · 2026-07-15
This conceptual paper translates the EU AI Act's (Regulation 2024/1689) governance requirements into practical audit-quality controls and evidence-evaluation criteria for audit firms using AI tools such as machine-learning models and generative AI. It maps AI Act obligations—covering risk management, data governance, transparency, human oversight, and logging—onto firm-level and engagement-level quality-management frameworks, producing artefacts including a crosswalk with IAASB standards, an evidence-risk typology, a documentation checklist, and a maturity model. The framework specifies three modes of AI Act relevance (direct legal, indirect, and benchmark) and clarifies the conditions under which AI outputs may serve as triage, corroborative, or substantive audit evidence. The work is significant for audit practice and standard-setting because it provides a scalable, inspection-ready governance structure for defensible reliance on AI-enabled audit tools without overstating the Act's direct legal applicability.
- Quality assurance
- Certifications
- AI policy
- Enterprise
Research
Privacy Preserving Recommender Systems Balancing Personalization with Privacy
Ranjeet K Jha, Venkata Suresh Gummadilli
arXiv · 2026-07-14
This paper proposes and evaluates a privacy-preserving recommendation framework for e-commerce platforms that combines federated learning, differential privacy, cohort-level modeling, and privacy-aware agents to keep raw user data decentralized while adding mathematically bounded noise to model updates. Experiments on synthetic retail datasets show that the framework maintains competitive recommendation quality—measured via CTR, Precision@K, Recall@K, and NDCG@K—at moderate privacy budgets (approximately ε≈5), demonstrating that strong privacy guarantees can be achieved with limited impact on recommendation effectiveness. The work directly addresses compliance requirements under GDPR, CCPA, and CPRA, offering a scalable approach for AI-driven retail platforms that must balance personalization, regulatory compliance, and business objectives.
- Enterprise
- AI policy
- Quality assurance
Research
Where Does the Noise Come From? A Variance-Components Decomposition of Non-Determinism in LLM Brand Answers
Dmitrij Żatuchin
arXiv · 2026-07-14
This paper investigates why brand scores produced by large language models (LLMs) are unstable, decomposing the total variance of LLM brand-recommendation answers into four sources: within-prompt resampling, prompt paraphrase, model identity, and query language. Using nearly 13,000 LLM responses across 20 brands, 8 languages, and 3 models, the study finds that query language accounts for 26.5% of response variance while brand identity accounts for only 1.5%, meaning a single AI answer carries almost no brand-discriminating signal. The research shows that reliability is best improved by diversifying across languages and models rather than simply repeating prompts, with each additional repeat past the fifth contributing negligibly to reducing error variance. These findings matter for enterprises and quality-assurance practitioners who use LLM outputs to measure brand perception, as they reveal that common measurement practices based on a handful of prompt repetitions are insufficient for reliable brand tracking.
- Enterprise
- Quality assurance
Research
Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners
Haseeb Shah, Lingwei Zhu, Adam White et al.
arXiv · 2026-07-14
This paper conducts a large-scale empirical study of over 33,000 experiments to evaluate how key design components in actor-critic reinforcement learning algorithms—such as action distribution type, gradient estimators, and update scheduling—affect reliability and hyperparameter sensitivity. Using a control task derived from a real water treatment plant, the authors find that commonly used defaults like Gaussian action distributions with pathwise gradient estimators are among the least reliable configurations, while bounded distributions with adaptive update schedules prove more robust. The findings provide concrete, component-level guidance for practitioners deploying reinforcement learning in high-stakes real-world systems where reliability is critical and tuning resources are limited.
- Enterprise
- Quality assurance
Research
Cost-Optimal Foundation Model Deployment Portfolio for Transportation Management
Xi Cheng, Ke Liu, Siyuan Feng et al.
arXiv · 2026-07-14
This paper addresses the challenge of deploying multiple AI foundation models—including large language models and vision-language models—across transportation management center (TMC) functions such as anomaly detection and incident reporting. The authors formulate the Foundation Model Deployment Portfolio (FMDP) problem as a mixed-integer program that minimizes total cost of ownership while satisfying quality, latency, and safety constraints under a shared GPU budget, and prove the problem is NP-hard. A greedy heuristic is proposed and applied in a case study with five TMC functions and 19 candidate model-deployment pairs, yielding a mixed portfolio costing $34/month—97% below the cheapest all-closed-API baseline—by routing most functions to open-source APIs and only one quality-constrained function to a closed API. The work provides actionable guidance for transportation agencies on model selection, deployment mode, and when on-premise GPU investment becomes cost-justified (above approximately 309 vision queries/hour or if API prices double).
- Enterprise
- AI policy
Research
Continuously Evolving Deepfake Detection: An Architecture and Public-Benchmark Evaluation of a Dynamic Detection System
Ken Jon Miyachi, Dylan Uys
arXiv · 2026-07-14
This paper presents BitMind Forensics (BMF), a deepfake detection system designed to continuously update its training distribution through an open adversarial competition called Bittensor SN34, directly addressing the well-documented problem of static detectors collapsing on real-world content—where state-of-the-art models have shown AUC drops of 45–50% in in-the-wild evaluations. Evaluated across nineteen public datasets spanning face-swap, AI-generated images, and video benchmarks, BMF achieves strong results including 0.936 AUC on Sumsub original images, 0.915 on Deepfake-Eval-2024 images (matching the best commercial detector), and 0.947 on DFDC—outperforming the FF++-trained frontier (0.843). A temporal study further shows successive model exports improve detection of generators unseen during static baseline training, demonstrating measurable gains from continual retraining. The work is relevant to quality assurance and policy by providing a publicly reproducible evaluation harness and a verifiable production API snapshot, offering practitioners and regulators a credible benchmark for assessing deepfake detection robustness.
- Quality assurance
- AI policy
- Certifications
Research
AI-Native Insurance for Agentic AI: Pricing, Underwriting, and End-to-End Automation
Quanyan Zhu
arXiv · 2026-07-14
This paper develops a mathematical framework for insuring agentic AI systems—autonomous AI that can make decisions, use tools, and interact with external services—addressing the novel risks these systems introduce. The framework models deployments via a 'risk state' capturing autonomy level, operational authority, permission exposure, governance maturity, and dependency concentration, then maps these factors to event probabilities, loss severities, premiums, deductibles, and policy covenants. The authors establish structural properties of insurability, including an insurability region, how feasibility deteriorates with increasing exposure, and governance certification thresholds, and illustrate the framework through a healthcare case study with contract optimization and automated claims processing. The work is relevant to enterprise AI deployment, certification and governance standards, and emerging policy frameworks for AI liability.
- Enterprise
- Certifications
- AI policy