News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
5526 items
Research
Bounded Agents: Delegation Security for Multi-Agent AI Systems
Xabier Muruaga
arXiv · 2026-08-16
This paper introduces the Agentic Principal Chain (APC), an authorization architecture for multi-agent AI systems that tracks and restricts delegated authority across agent sessions. Rather than treating prompt injection purely as a model-behavior problem, APC enforces six authorization checks against accumulated session state, carries forward scope and budget limits, and uses composition closure to block prohibited combinations of individually permitted actions. In benchmark evaluations across 3,154 instances—including InjecAgent, AgentDojo, and ASB—APC reduced exfiltration to 0% across all four AgentDojo domains, blocked all 544 InjecAgent data-stealing cases, and cut manipulation success from 90.5% to 12.1%, with authorization latency of just 0.24 ms at the 99th percentile. The work matters because it demonstrates that agentic security risks are fundamentally an authorization architecture problem, and provides a formally proven, publicly available implementation to address them.
- Enterprise
- AI policy
Research
Propaganda Forensics: Recovering the Generation Pipeline of an AI-Driven Influence Campaign
Benjamin Icard, Elouan Vuichard, Louis Lefebvre et al.
arXiv · 2026-08-16
This paper presents a forensic analysis of an AI-driven influence campaign, introducing PROPAGIA, a corpus of 2,646 propagandist French articles from the Storm-1516/CopyCop campaign. By comparing these articles to human-written French press (SIPA), the authors identify propaganda techniques including higher vagueness, subjectivity, negativity, and fewer source citations in the AI-generated content. The researchers also recover prompt instruction leaks on 50 of 84 campaign websites, revealing a ten-point editorial specification, and use rewriting-based detection to attribute the content to Llama 3 and possibly Mistral-family models. These findings matter for policy and quality-assurance efforts, offering concrete methods to detect and trace AI-generated disinformation campaigns.
- AI policy
- Quality assurance
News
The CPU Comeback Is Upon Us
spectrum.ieee.org · 2026-08-16
IEEE Spectrum reports that the rapid growth of agentic AI systems is triggering an unexpected surge in CPU demand, catching major cloud providers like Amazon Web Services off-guard. Unlike traditional AI inference workloads that rely heavily on GPUs, agentic AI pipelines — where models autonomously spawn sub-agents, make API calls, and invoke software tools — are largely CPU-bound tasks, with AMD researchers finding that seven of eight stages in typical agentic pipelines run entirely on the CPU. Researchers from Intel and the Georgia Institute of Technology also found that tokenization bottlenecks worsen significantly as sequence lengths grow, and that insufficient CPU core counts can cause GPU stalls, with more CPU cores reducing time-to-first-token latency by 1.5x to 7x in tests. Analysts warn that this CPU crunch could deepen, with Intel already sold out of server CPUs through year-end and AMD doubling its server CPU forecast, potentially driving broader shortages and price increases similar to what occurred with GPUs.
- Enterprise
Research
PL-Guard: Probabilistic Logic Reasoning for LLM Guardrails
Satchit Chatterji, Shihan Wang, Giovanni Sileno et al.
arXiv · 2026-08-16
PL-Guard is a neurosymbolic guardrail architecture for large language models that separates semantic grounding from policy reasoning by using a local LLM to convert prompt-response pairs into predicate probabilities and then applying ProbLog probabilistic rules for explicit policy inference. On the XSTest benchmark, PL-Guard with a hand-curated policy reduces unsafe compliance from 22.0% for the base model to 0.5%, outperforming an LLM-as-a-judge baseline at 6.0%, though at the cost of higher over-refusal (14.4% vs. 5.2%). The approach makes guardrail reasoning steps explicit and auditable, exposing the safety-helpfulness tradeoff inherent in LLM content moderation.
- Quality assurance
- AI policy
Research
Do Assessment Instruments Measure the Same Thing for Humans and LLMs? A Latent Structure Analysis
Alona Strugatski, Licol Zeinfeld, Giora Alexandron
arXiv · 2026-08-16
This paper investigates whether standardized educational assessments—originally designed to measure human skills—can validly be applied to evaluate large language models (LLMs). Using exploratory factor analysis, factor congruence, and resampling techniques, the researchers compared response patterns from human learners and six multimodal LLMs on two instruments: a high-school chemistry exam and a university entrance exam's quantitative reasoning section. They find systematic differences in the latent factor structures between humans and LLMs, suggesting these assessments do not measure the same underlying constructs in both groups. The findings challenge the common practice of using human-normed assessments as evidence for generalizable claims about AI capabilities.
- Quality assurance
- Certifications
Research
Turning AI Capability into Performance: How AI Understanding and AI Skills Shape Employee Productivity through Employee–AI Collaboration in a Chinese Smart Hospital
Tang Song, Nor ‘Ain Bt Abdullah, Zunirah Mohd Talib et al.
International Journal of Computer Information Systems and Industrial Management Applications · 2026-08-16
This study of 694 employees at a Chinese tertiary public hospital finds that AI understanding and AI skills improve worker productivity primarily by enabling day-to-day collaboration with AI systems, rather than through direct knowledge alone. Path analysis shows that employee–AI collaboration is the strongest predictor of productivity (β=0.426), and that collaboration mediates 52–64% of the total effect of AI capabilities on output. The results suggest that hospital training and system design should be judged by whether they change how employees actually work with AI, not just what they know about it. This shifts the focus from AI acceptance attitudes to enacted collaborative work practice as the key driver of productivity gains.
- Workforce
- Enterprise
Research
RoboSafe: A Quantitative Character Safety Certification Framework for Social Robot Deployments in Public-Facing Environments
Chang Xiong
Zenodo (CERN European Organization for Nuclear Research) · 2026-08-16
RoboSafe Standard v1.0 is a normative certification framework designed to establish measurable safety requirements for the 'character safety layer' of physical AI systems—robots and AI-driven hardware deployed in public-facing environments. The framework defines three certification levels tied to deployment risk (retail/corporate, hospitality/elder care, and clinical/pediatric settings), each with specific KPI thresholds such as hard block accuracy, false positive rates, response latency, and PHI redaction coverage. A four-stage certification process (Configure, Simulate, Validate KPIs, Maintain) provides a repeatable compliance path, and the standard is explicitly designed to be citable in procurement documents, RFP responses, enterprise contracts, and regulatory filings. This matters because it addresses the current absence of a shared safety standard for embodied AI platforms, reducing procurement ambiguity and accountability gaps in high-stakes public deployments.
- Certifications
- Enterprise
- AI policy
Research
Governing generative AI in organizations: a design theory and quasi-experimental field study of sociotechnical guardrails
Maikel Leon
The Journal of Supercomputing · 2026-08-16
This paper develops a design theory for governing generative AI in organizations through 'sociotechnical guardrails'—mechanisms combining policy, technical, and workflow components to embed organizational norms into deployed AI systems. A quasi-experiment at a Fortune 500 firm across 20 teams and 28 weeks found that guardrails reduced interaction entropy by 35%, cut hallucinations in half, narrowed a fairness gap from 0.18 to 0.05, and raised audit-trail completeness from 53% to 96%, though at a 12% task-time cost and with at least 35% of teams circumventing guardrails they found opaque or disproportionate. The findings highlight perceived legitimacy as a critical factor in governance effectiveness and offer design implications drawn from analysis of nine US executive orders on AI (2019–2025).
- Enterprise
- AI policy
- Quality assurance
Research
AI-Generated Evidence And Judicial Decision-Making In India: Constitutional Limits Of Admissibility, Reliability, Human Oversight, And The Role Of Artificial Intelligence In Judicial Discretion
Dr. Prashant Yadav
Zenodo (CERN European Organization for Nuclear Research) · 2026-08-16
This article examines the constitutional and procedural challenges posed by AI-generated evidence—such as facial-recognition outputs, algorithmic analytics, and synthetic media—in Indian courts. It argues that while the Bharatiya Sakshya Adhiniyam 2023 treats such material as admissible electronic records upon certification, the statute lacks reliability standards or explainability requirements adequate to address risks of bias, hallucination, and deepfake manipulation. Grounding the analysis in Articles 14 and 21 of the Indian Constitution and the non-delegable nature of judicial discretion, the article proposes a framework of mandatory disclosure, independent expert validation, and a human-in-the-loop requirement to protect due process. The paper concludes that AI may assist but cannot replace judicial decision-making in Indian adjudication.
- AI policy
- Quality assurance
Research
RoboSafe: A Quantitative Character Safety Certification Framework for Social Robot Deployments in Public-Facing Environments
Chang Xiong
Open MIND · 2026-08-16
RoboSafe Standard v1.0 is a normative certification framework designed to assess and certify the character safety of AI-driven social robots deployed in public-facing physical environments. It defines three certification levels tied to deployment risk — retail/corporate, hospitality/elder care, and clinical/pediatric — each with measurable KPI thresholds such as hard block accuracy, false positive rates, response latency, and PHI redaction coverage. A four-stage process (Configure, Simulate, Validate KPIs, Maintain) provides a repeatable compliance path. The framework is intended to be citable in procurement documents, RFP responses, enterprise contracts, and regulatory filings, directly addressing procurement ambiguity and accountability gaps in embodied AI deployments.
- Certifications
- Enterprise
- AI policy
Research
AI-Generated Evidence And Judicial Decision-Making In India: Constitutional Limits Of Admissibility, Reliability, Human Oversight, And The Role Of Artificial Intelligence In Judicial Discretion
Dr. Prashant Yadav
Zenodo (CERN European Organization for Nuclear Research) · 2026-08-16
This article examines how AI-generated evidence—such as facial-recognition outputs, algorithmic analytics, and synthetic media—is being introduced into Indian courts under the Bharatiya Sakshya Adhiniyam, 2023, which permits such material as electronic records upon certification but provides no reliability standards or explainability requirements. The authors argue that the 'black-box' nature of many AI systems, combined with risks of bias, hallucination, and deepfake manipulation, poses serious threats to due process and constitutional guarantees under Articles 14 and 21. The article proposes a framework of mandatory disclosure, independent expert validation, and an explicit human-in-the-loop requirement to ensure AI assists rather than displaces judicial discretion. This matters because it directly addresses the legal and constitutional boundaries courts must observe as AI becomes embedded in forensic and investigative processes.
- AI policy
- Certifications
Research
Bridging the gap between vocational AI curricula and industry skill demand: Evidence from China
Huixiang Xiao, Hoi Leong Lee, Kaige Zheng et al.
Industry and Higher Education · 2026-08-16
This study analyzes the gap between vocational AI curricula and industry skill demand in China by comparing nearly 500,000 job advertisements (2020–2024) with 46 institutional training plans. Using a bilingual taxonomy of 198 skill keywords and a demand-weighted coverage index, the researchers find a selective technology lag: programming and AI practicum courses align reasonably well with employer needs, while cloud computing and big-data skills are underrepresented. Soft skills are also unevenly covered, often implicit rather than systematically taught. Work-integrated learning formats—practicums, internships, and capstone projects—show the strongest alignment and are identified as key levers for curriculum renewal.
- Workforce
- Certifications
Research
Quantifying Systemic Risk from Correlated Models: A Methodological Review
Yuanyuan Li, Michael von Gablenz
Zenodo (CERN European Organization for Nuclear Research) · 2026-08-16
This paper examines how deploying multiple AI systems that share training data, architectures, or foundation models can produce correlated errors and synchronized failures that standard single-model evaluations miss. The authors survey methods for measuring behavioral similarity across models, propose a risk-oriented evaluation framework with desirable statistical properties, and identify common drivers of correlated behavior. They argue that effective AI governance requires shifting from isolated model validation to portfolio-level auditing and dependency-aware risk management aligned with emerging regulatory frameworks.
- Quality assurance
- AI policy
Research
AI-Generated Evidence And Judicial Decision-Making In India: Constitutional Limits Of Admissibility, Reliability, Human Oversight, And The Role Of Artificial Intelligence In Judicial Discretion
Dr. Prashant Yadav
Zenodo (CERN European Organization for Nuclear Research) · 2026-08-16
This article examines how AI-generated evidence—including facial-recognition outputs, algorithmic analytics, and synthetic media—is being introduced into Indian courts under the Bharatiya Sakshya Adhiniyam, 2023, which admits such material as electronic records upon certification but sets no reliability or explainability standards. The authors argue that the black-box nature of AI systems, combined with risks of bias, hallucination, and deepfake manipulation, threatens constitutional guarantees under Articles 14 and 21 and the right to a fair trial. The paper proposes a framework of mandatory disclosure, independent expert validation, and a human-in-the-loop requirement to ensure that AI may assist but never replace judicial discretion in Indian adjudication.
- AI policy
- Quality assurance
Research
GALENA: A Governance-Aware LLM Enterprise Navigation Architecture for Autonomous Multi-Agent Workflow Automation with Compliance Enforcement
Narasimha Rao Boinapalli
arXiv · 2026-08-16
GALENA is a multi-agent LLM orchestration framework that embeds regulatory compliance (GDPR, HIPAA, and domain-specific rules) as a formal architectural constraint evaluated before every agent action, rather than as a post-hoc filter. Across three enterprise task categories—Regulated Document Processing, Financial Workflow Automation, and IT Service Management—the system achieves 97.0% task completion accuracy, outperforming the strongest baseline by 18.3%, while reducing governance violation rates by 72.7% to a median of 0.03 with a median latency of 164 ms. The paper argues that treating governance as a first-class invariant, including drift-resilient lifecycle management and role-aware agent routing, is both feasible and necessary for production-grade enterprise automation. These results matter for enterprises seeking to deploy autonomous AI workflows without sacrificing regulatory compliance.
- Enterprise
- AI policy
Research
AI-Generated Evidence And Judicial Decision-Making In India: Constitutional Limits Of Admissibility, Reliability, Human Oversight, And The Role Of Artificial Intelligence In Judicial Discretion
Dr. Prashant Yadav
Zenodo (CERN European Organization for Nuclear Research) · 2026-08-16
This article examines the constitutional and procedural challenges posed by AI-generated evidence—including facial-recognition outputs, algorithmic analytics, and synthetic media—in Indian courts. It argues that while the Bharatiya Sakshya Adhiniyam, 2023 treats such material as admissible electronic records upon certification, the statute lacks reliability standards, explainability requirements, or scrutiny of generative processes, threatening due process under Articles 14 and 21. The authors contend that the black-box nature of AI, combined with risks of bias, hallucination, and deepfake manipulation, makes rigorous human oversight essential, and that AI may assist but cannot replace judicial discretion. The article proposes a framework of mandatory disclosure, independent expert validation, and an explicit human-in-the-loop requirement to safeguard the integrity of Indian adjudication.
- AI policy
- Quality assurance
Research
Nurse educators' experiences and perceptions using generative artificial intelligence: a systematic review
Ani Henttonen, Maria Christidis, Helena Kullenberg et al.
BMC Medical Education · 2026-08-15
This systematic review of 13 studies (3,082 participants) examines nurse educators' experiences with generative AI in teaching, finding a tension between optimism about pedagogical efficiency and concerns over academic integrity, critical thinking erosion, and professional role loss. Educators' confidence and use were shaped by institutional position, organizational policy, and prior GenAI experience. The review concludes that effective integration requires structured training, competency development, and clear governance frameworks. These findings highlight the urgent need for policy and capacity-building support within nursing education institutions.
- Workforce
- AI policy
Research
Insurance as AI Risk Infrastructure: A Generative-Agent Simulation of AI Adoption
Yixuan Yuan, Dedai Wei, Chudong Qian et al.
arXiv (Cornell University) · 2026-08-15
This paper proposes using insurance as a financial risk-transfer mechanism to address enterprise hesitancy around AI adoption, arguing that technical safeguards alone cannot protect firms from residual financial losses caused by AI system failures. The authors build an LLM-driven agent-based social simulation to model how insurance frameworks affect firm-level behavior, finding that insurance reduces financial exposure, accelerates aggregate AI adoption, and improves firm solvency and capital. The work frames insurance as complementary infrastructure to existing AI safety measures, with implications for how businesses and policymakers approach AI integration risk.
- Enterprise
- AI policy
Research
IA generativa e LGPD: o desafio da transparência algorítmica
Marise Miglioli Lorusso
Derecho y cambio social. · 2026-08-15
This article examines the legal and technical obstacles to implementing Brazil's General Data Protection Law (LGPD), specifically the right to review automated decisions (Article 20), in the context of Generative AI systems that operate as 'black boxes.' Drawing on case studies involving Meta, X (Grok), the Smart Sampa surveillance program, AI use in the judiciary, and the platformization of public education, the study finds a deep structural conflict between the legal requirement for transparency and the inherent opacity of Large Language Models. The authors conclude that reactive enforcement models or direct transplantation of Global North regulations are insufficient for Brazil's socioeconomic context, and recommend regulatory sandboxes, Algorithmic Impact Assessments, and mandatory Privacy by Design as more appropriate governance approaches.
- AI policy
- Enterprise
Research
Working with semantic machines: The role of semantic reconciliation in public sector data governance
Jonny Holmström
Government Information Quarterly · 2026-08-15
This qualitative case study of a large Swedish municipality examines an AI interviewing system used in early-stage recruitment, which translated candidates' open-ended responses into competency indicators and narrative summaries. The authors develop a 'semantic reconciliation' process model identifying five recurring forms of organizational work needed to make AI outputs meaningful across a decentralized public sector hierarchy—including mapping vendor categories to local vocabularies and integrating AI reports with human judgment. The study shows that while AI tools can make evaluative information portable, continuous organizational work is required to restore contextual adequacy and preserve human decision authority. The findings offer practical guidance for municipalities seeking to adopt common AI tools without letting standardized outputs displace role-specific knowledge.
- Workforce
- AI policy
Research
Artificial Intelligence in Nursing Education: A Scoping Review of Academic Perspectives
Natasha Hawkins, Anthea Fagan, Yumiko Coffey et al.
Journal of Advanced Nursing · 2026-08-15
This scoping review of 15 studies across 8 countries (2,004 nursing academics) finds an 'adoption paradox' in AI integration in nursing education: while most academics believe AI will revolutionize the field, implementation remains conservative, with two-thirds of applications operating only at an augmentation level and none achieving transformative redefinition. Key barriers include knowledge gaps, institutional policy vacuums, and global access inequities, with academics expressing concern about critical thinking erosion and threats to professional identity. The findings point to a need for faculty development, institutional policy frameworks, and curriculum strategies that balance technological advancement with person-centred nursing values.
- Workforce
- AI policy
Research
The Complementarity Paradox: Human-AI Collaboration, Overreliance, and Skill Decline in Entrepreneurial Contexts
Srinivas Subramanya, Satish Krishnan, Nasreen Azad
Information Systems Frontiers · 2026-08-15
This paper examines how knowledge workers with strong entrepreneurial orientations adopt generative AI tools deeply, which can lead to over-reliance and perceived erosion of professional skills—a pattern the authors call 'hustle-to-handicap.' Drawing on cognitive offloading theory and skill decay literature, the researchers theorize and find support (via survey data from 143 knowledge workers across five countries) for a sequential mechanism: entrepreneurial orientation drives deeper AI collaboration, which fosters AI dependence, which in turn correlates with perceived skill decline. A long-term orientation amplifies this effect because future-focused individuals reframe routine AI use as strategic investment rather than a threat to skill maintenance. The findings contribute to debates on human-AI collaboration by showing how individual mindset and temporal outlook shape whether AI augments or erodes capability.
- Workforce
- Enterprise
Research
Digital Development and Participation in EU Digital Governance: Evidence from Central and Eastern Europe
Xiaoqing Wang
Journal of International Relations and Foreign Policy · 2026-08-15
This article investigates how levels of digital development among Central and Eastern European (CEE) EU member states affect their ability to participate in EU digital governance processes, using the Digital Economy and Society Index 2020 as a baseline. Examining the Digital Markets Act and the Artificial Intelligence Act, it finds that CEE countries remain relatively peripheral during agenda-setting and legislative negotiation, while implementation capacity varies across the region. The study argues that common EU digital rules do not produce equal participation, meaning digital development is both a policy outcome and a structural condition shaping governance influence. This matters for policy because it highlights systemic inequalities in how EU AI and digital regulations are shaped and adopted across member states.
- AI policy
Research
From data to decisions: how neural network input structures propagate to air quality policies
Laura Zecchi, Michele Francesco Arrighini, Claudio Marchesi et al.
npj Clean Air · 2026-08-15
This study examines how different artificial neural network (ANN) architectures used as surrogate models in air-quality planning affect both predictive accuracy and the policy recommendations that emerge from optimization. The researchers tested multiple ANN designs varying in spatial aggregation for the Po Valley basin and found that regionalized ANNs achieve relative mean absolute errors of 2.4–3.3%, outperforming a single basin-wide ANN at 6.4%. At higher cost levels, these architectural differences propagate into divergent optimal emission-reduction policy portfolios, affecting how resources are allocated across sectors like domestic heating, transport, and agriculture. The findings demonstrate that surrogate model design choices meaningfully shape air quality policy outcomes, not just prediction quality.
- AI policy
- Quality assurance
Research
Brain Capital Management: A Firm-Level Theory of Cognitive Capability, Its Three Constraints, and an Agenda for Its Measurement in the Age of AI
Naoki Kadowaki
Zenodo (CERN European Organization for Nuclear Research) · 2026-08-15
This working paper introduces a firm-level theory of 'enterprise brain capital,' defining it as the cognitive capability a company can actually deploy, shaped by three constraints: Belonging (whether cognitive capacity can be expressed organizationally), Base (whether strain consumes available capacity), and Build (how capability accumulates or depreciates over time). The framework develops fourteen falsifiable propositions and a three-tier measurement architecture, while distinguishing genuine capability accumulation from AI-assisted productivity gains—arguing the two do not necessarily coincide. A key finding is a measurement asymmetry in existing human-capital disclosure regimes, which report inputs and costs but largely omit validated measures of workforce cognitive and psychological states. The paper positions Brain Capital Management as a diagnostic and research agenda requiring further empirical validation rather than a finished framework.
- Enterprise
- Workforce