News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
SLAPBench: Benchmarking Multimodal Large Language Models for Four-Finger SLAP Fingerprint Verification
Bibesh Pyakurel, M. G. Sarwar Murshed
arXiv · 2026-07-17
SLAPBench introduces the first benchmark for evaluating multimodal large language models (MLLMs) on four-finger SLAP fingerprint identity verification, a task central to border control and law enforcement. Built from NIST SD302b with 7,832 image pairs, the benchmark tests five MLLMs under multiple prompting strategies and finds that prompting style largely governs whether models collapse to near-universal acceptance, while underlying model capability determines how well they discriminate identities. Claude Opus 4.8 achieves the best binary result (FAR = 20.2%) and highest AUC (0.953) among models that do not collapse, whereas open-source models vary widely and a fairness probe suggests demographic disparities worsen as discrimination weakens. These results establish a baseline revealing that current MLLMs are not yet reliable for biometric verification and that prompting choices carry significant security implications.
- Certifications
- AI policy
- Quality assurance
Research
Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching
Yan Song
arXiv · 2026-07-17
This paper identifies a fundamental conflict in production LLM deployments where prompt compression and prompt caching are typically used together but work against each other: query-aware compression generates a unique prefix per query, invalidating cached prefixes and eliminating caching discounts. The authors empirically characterize Anthropic's Sonnet API caching behavior, finding a two-tier architecture with a hit rate plateau around 0.83 rather than the ideal 1.0 assumed in prior literature. They propose Cache-Aware Prompt Compression (CAPC), which pairs query-agnostic compression with explicit cache control and a tier-preserving ratio bound to avoid over-compression. CAPC achieves the lowest cost in all 16 tested configurations on LongBench-v2—with mean savings of 49% over cache-only and 64% over query-aware compression—while maintaining answer quality within 0.05 of the uncompressed baseline, and is validated on three production workloads including an enterprise tool-using assistant and a public retail benchmark.
- Enterprise
- Quality assurance
Research
A Tool-Invariant Framework for Teaching and Assessing Computational Methods in the Age of Agentic AI
Larry Engelhardt
arXiv (Cornell University) · 2026-07-17
This paper proposes a tool-invariant framework for teaching computational methods that separates enduring conceptual knowledge—inputs, outputs, terminology, and evaluative judgment—from the tool used to execute those methods, now including agentic AI that can write and run code autonomously. The author argues that verification of computational results, rather than code authorship, has become the critical skill as AI-generated artifacts can no longer serve as evidence of student learning. To address this, the paper recommends pairing AI-free in-class coding quizzes with oral defenses of comment-stripped AI-assisted work in small-class settings. The implications are directly relevant to how computational courses should be assessed and credentialed when students can generate artifacts on demand.
- Certifications
- Quality assurance
Research
In the Driver's Seat: A Multi-Company Study on the Reality of Autonomous Driving System Testing
Qunying Song, Gao Y, Johannes Betz et al.
arXiv (Cornell University) · 2026-07-17
This paper presents an interview study with experts from nine companies across six countries to map current practices, challenges, and future directions in autonomous driving system (ADS) testing. The findings reveal that industry relies primarily on scenario-based and X-in-the-loop testing, but faces significant gaps in scenario realism, simulation fidelity, and acceptance criteria. The authors synthesize these insights into an evidence-centered closed-loop testing framework intended to provide actionable guidance. The study is notable for its multi-company, cross-national scope and its identification of AI, world models, and end-to-end approaches as potential solutions to current testing shortfalls.
- Quality assurance
- Certifications
Research
Rebalancing the algorithm: pathways for mitigating bias in predictive policing
Lata Nautiyal, Preeti Malik, Varsha Mittal et al.
Service Oriented Computing and Applications · 2026-07-17
This paper investigates algorithmic bias in predictive policing systems using the Chicago Crime Dataset, applying data re-weighting, counterfactual analysis, and algorithm auditing to identify sources of inequity. The results show that conventional predictive models disproportionately flag minority and historically over-policed neighborhoods as high-risk, with risk scores driven more by enforcement intensity than actual crime incidence. Bias-controlled models produce different hotspot predictions and reduce disparate impact measures, demonstrating that historical policing practices substantially distort AI-generated risk forecasts. The authors argue for transparency, fairness-aware modeling, and systematic auditing in deploying AI for law enforcement.
- AI policy
- Quality assurance
Research
LLMs as judges: Toward the LLM-assisted review of GSN-compliant assurance cases
Gerhard Yu, Mithila Sivakumar, Alvine Boaye Belle et al.
Journal of Systems and Software · 2026-07-17
This paper proposes using large language models (LLMs) as semi-automated judges to review Goal Structuring Notation (GSN)-compliant assurance cases for mission-critical systems such as autonomous vehicles and avionics. The authors introduce predicate-based rules that formalize established review criteria and use these to craft targeted LLM prompts, then evaluate GPT-4o, GPT-4.1, DeepSeek-R1, and Gemini 2.0 Flash on the task. Results show most LLMs demonstrate reasonably good review capabilities, though human oversight remains necessary to refine LLM-generated reviews. The work directly addresses the inefficiency and error-proneness of manually reviewing lengthy assurance case documents, with implications for regulatory acceptance of safety-critical systems.
- Quality assurance
- Certifications
Research
Assessing Entrepreneurial Readiness and Perceptions of AI: A Catalyst for Sustainable Business Model Innovation
Edmond Freo
The International Review of Multidisciplinary Research · 2026-07-17
This quantitative study of 200 entrepreneurs examines how entrepreneurial readiness and sustainability perceptions of AI jointly predict capacity for Sustainable Business Model Innovation (SBMI) aligned with UN SDGs 8, 9, and 12. Findings show entrepreneurs have moderate overall readiness (composite mean 3.08) due to financial and infrastructural constraints, and strongly recognize AI's economic efficiency potential but remain largely neutral about its ecological applications. Regression analysis found that readiness and sustainability perceptions together explain 46.9% of variance (R²=0.469, p<0.001) in SBMI capacity, with sustainability-centered perception acting as a strategic steering mechanism beyond raw technological readiness. The authors propose a Readiness-Perception-Executed Business Model to guide entrepreneurs in using AI as a responsible catalyst for sustainable development.
- Enterprise
- Workforce
Research
Digital Economy's Impact on Employment and the Role of Market Competition-The Impact of Digital Economy on Employment——Based on the Perspective of Market Competition
Leyan Chen
Journal of innovation and development · 2026-07-17
Using panel data from 31 Chinese provinces (2012–2022), this study finds that digital economy development substantially expands regional employment, with net job creation dominating over displacement in the long run. The effect is conditioned by market competition: employment gains are significantly amplified only after competitive intensity surpasses a specific threshold. The paper also documents heterogeneity across industrial structures and economic development levels, offering empirical grounding for differentiated regional employment policies.
- Workforce
- AI policy
Research
Artificial Intelligence in Special Education: Assessing Teachers' Level of Use and Readiness at Tomas Sagun Integrated School, Pagadian City, Philippines
Mike R. Canoy, April Dawn B. Ruales, Justriel G. Tutor et al.
International Journal of Educational Innovations and Research · 2026-07-17
This descriptive survey study examined AI adoption among special education teachers at a Philippine school, finding that most teachers showed only moderate readiness and limited actual use of AI instructional tools. Barriers included insufficient training, limited technology access, and lack of institutional support. The study recommends targeted capacity-building programs, policy frameworks, and resource allocation to close the gap between readiness and utilization in special education classrooms.
- Workforce
- AI policy
Research
A Human-Centric Evaluation of a Retrieval-Augmented Generation System for Explaining Quebec Insurance Contracts
David Beauchemin, Richard Khoury
arXiv (Cornell University) · 2026-07-17
This paper evaluates a Retrieval-Augmented Generation (RAG) system designed to help Quebec consumers understand automobile insurance contracts, addressing the 'advice gap' created by online insurance sales. A user study with 154 participants found the system acted as a 'cognitive equalizer,' earning high ratings for satisfaction, trust, and clarity, with users—especially those with lower financial literacy—valuing the sense of autonomy it provided even more than the knowledge gained. However, participants still preferred human agents for high-stakes or emotionally charged decisions, underscoring the need for human-in-the-loop frameworks. The findings highlight both the promise and limits of AI-driven tools in consumer-facing financial services.
- AI policy
- Enterprise
Research
Streamlining endometriosis MRI reporting: Automated extraction of #Enzian scores from pelvic MRI reports using local and online LLMs
Joana Kostova, Caroline Reinhold, Jonas D. Stief et al.
European Journal of Radiology Artificial Intelligence · 2026-07-17
This study tested whether large language models (LLMs) can automatically extract standardized #Enzian endometriosis scores from pelvic MRI reports, comparing 12 LLMs against both expert radiologist consensus and novice trainee benchmarks across 186 reports. Online LLMs achieved 93.8–95.8% accuracy against expert reference, with three models (Gemini 2.5 Pro, Grok 4, and o3) significantly outperforming novice radiologists who reached 87.4% accuracy. Local models performed more variably (61.7–92.5%) and were generally inferior to trainees. The findings suggest online LLMs could serve as training aids and support standardized reporting for surgical planning, while off-the-shelf local models are not yet ready for this task.
- Quality assurance
- Workforce
Research
The Role of Artificial Intelligence in the Lifecycle of Scientific Manuscripts: Authoring, Reviewing, and Editorial Selection
José L. Domingo
Qeios · 2026-07-17
This critical commentary and policy analysis examines how AI and large language models are reshaping three stages of scientific publishing: co-authorship, peer review, and editorial selection. The paper highlights persistent risks such as citation hallucination, AI's inability to assess novelty, and bias amplification in editorial decisions, while noting that 2025 surveys show over 50% of researchers use AI during peer review, often in violation of existing policies. The authors propose a hybrid framework that limits AI to technical verification tasks while reserving judgments on scientific merit and ethics for compensated human experts, alongside legal accountability structures and reform of exploitative economic models like unpaid review labor paired with high Article Processing Charges.
- AI policy
- Quality assurance
Research
The enforced technical mandate: A multi-layered governance model for deepfake fraud and biometric integrity
Felipe Romero Moreno
Computer law & security review · 2026-07-17
This paper analyzes how existing EU and UK regulatory frameworks are failing to keep pace with AI-generated deepfake fraud, particularly as Fraud-as-a-Service enables scalable attacks on biometric identity systems. Through comparative doctrinal analysis, the authors identify a three-layered governance gap—covering source control, distribution control, and accountability—and flag a 12-month regulatory vacuum created by misaligned EU Digital Omnibus timelines. The paper proposes six policy recommendations, including NIST IAL2 zero-retention biometric standards, mandatory C2PA digital provenance, and embedding biometric integrity protocols into ISO 20022 financial messaging to create a real-time compliance enforcement mechanism across jurisdictions. The findings are directly relevant to policymakers, financial regulators, and organizations grappling with liability under emerging AI and data protection regimes.
- AI policy
- Certifications
Research
Reversibility-Aware Staged Delegation for Enterprise Agentic AI: A Real-Options and Resilience Framework for Irreversible Actions
Kwan Hong Tan
arXiv · 2026-07-17
This paper presents Reversibility-Aware Staged Delegation (RASD), a decision framework for enterprise AI agents that classifies actions by how recoverable they are and routes them through direct execution, staged commit, human review, or blocking accordingly. In a Monte Carlo simulation of 120,000 synthetic enterprise tasks, RASD achieved higher mean net value (7.408 vs. 5.735 and 5.282 normalized units) and a dramatically lower severe-incident rate (0.390% vs. 5.937% and 4.166%) compared to confidence-threshold and expected-loss baseline policies. The framework demonstrates that raw task-failure rate is an insufficient safety metric when recovery potential varies, arguing instead for recoverability-preserving execution architectures. The findings offer operational guidance for auditability, human escalation, and risk-adjusted value creation in enterprise agentic AI deployments.
- Enterprise
- Quality assurance
Research
Deployment process for artificial intelligence applications in radiology practice
Satu I. Inkinen, Juuso H. Ketola, Teemu Mäkelä et al.
Physica Medica · 2026-07-17
This paper presents a structured framework for deploying artificial intelligence systems in radiology, covering goal-setting, procurement, implementation planning, and post-deployment monitoring. It emphasizes defining stakeholder roles, integrating with hospital information systems, meeting regulatory requirements, and establishing quality assurance protocols with clinically relevant KPIs. A phased rollout and pilot approach are recommended to minimize workflow disruption and identify integration issues early. The framework aims to ensure patient safety, legal compliance, and sustainable AI integration with measurable clinical improvements.
- Quality assurance
- Enterprise
- AI policy
Research
From Skill Extraction to Multistakeholder Recommendation: A Two-Stage Framework for Bias Governance in Skills-Based Job Matching
Andrea Forster, Gregor Autischer, Dominik Kowald et al.
arXiv (Cornell University) · 2026-07-17
This paper proposes a two-stage framework for detecting and governing bias in AI-driven, skills-based job matching systems. Stage 1 addresses bias risks in skill extraction and candidate profile formation—particularly via chatbot-based elicitation—using distributional auditing and counterfactual testing to classify issues as hard or soft constraints. Stage 2 embeds candidate profiles in a multistakeholder recommender system where candidate, company, and regulatory objectives are represented by separate agents whose rankings are aggregated via social choice methods into a single, auditable recommendation. The framework aligns with the EU AI Act and the Fraunhofer AI Assessment Catalog, making it relevant to fairness governance and regulatory compliance in labor-market AI platforms.
- Workforce
- AI policy
Research
Does the timing of AI exposure matter? Generational composition, human capital, and productivity
Melike Çetin
Economics of Innovation and New Technology · 2026-07-17
This cross-country panel study (72 countries, 2000–2023) investigates whether generational composition of the working-age population interacts with human capital and AI diffusion to shape labor productivity. Using fixed-effects models, the authors find that generational shares alone do not predict higher productivity, but triple interactions between generational shares, human capital, and the post-2016 AI diffusion period are positive and statistically significant—with the strongest effect for Millennials, who combine digital adaptability with accumulated work experience. Results are robust across alternative productivity measures, institutional controls, and AI-timing thresholds. The study implies that realizing AI-era productivity gains depends on an economy's demographic structure and human-capital endowments, not AI exposure alone.
- Workforce
- AI policy
Research
Artificial Intelligence in Public Sector Governance: A Bibliometric Analysis of Global Research Trends
Muhammad Syahbar, Helen Dian Fridayani, Van Hoa Vu
TRANSFORMASI Jurnal Manajemen Pemerintahan · 2026-07-17
This bibliometric study maps the intellectual structure of global research on AI in public sector governance by analyzing 263 journal articles indexed in Scopus from 2021 to 2025. Using Bibliometrix and VOSviewer, the authors find that the field is rapidly expanding, with AI as the dominant conceptual anchor and themes like digital transformation, decision-making, and algorithmic accountability gaining prominence. The study argues that the literature is shifting from an adoption-focused perspective toward an institutional governance framework, identifying four key tensions—capability vs. control, efficiency vs. equity, automation vs. discretion, and innovation vs. democratic legitimacy—as priorities for future research. These findings matter for policy because they reveal where governance frameworks for public-sector AI are underdeveloped and where accountability mechanisms remain weakly integrated.
- AI policy
Research
From Neural Intent to Cryptographic Authorization: Securing AI-Driven Enterprise Workflows
Jiasi Weng, Jian Weng, Minrong Chen et al.
arXiv (Cornell University) · 2026-07-17
This paper introduces Neural Cryptographic Services (NCS), a security enforcement layer designed to protect AI-driven enterprise and government workflows from injection attacks. NCS interposes a deterministic symbolic controller between an AI planner and privileged tools, using offline-signed, hash-chained instruction streams so that only cryptographically authorized actions can be executed regardless of whether the AI planner has been compromised. Evaluated on the AgentDojo benchmark and a custom argument-hijacking benchmark, NCS reduces attack success rates to near zero while maintaining acceptable utility on legitimate workflows. The work matters for enterprise AI deployments because it reframes security from trusting model intent to enforcing cryptographic authorization at runtime.
- Enterprise
- Quality assurance
Research
The Model Artificial Intelligence Law (MAIL) v.4.0
Honglei Li, H J Zhou, Yanfeng Li et al.
Zenodo (CERN European Organization for Nuclear Research) · 2026-07-17
The Model Artificial Intelligence Law 4.0 (MAIL 4.0) presents a comprehensive statutory framework for governing AI that integrates developmental policy, lifecycle oversight, and liability rules within a unified legislative architecture. It proposes a national AI administrative authority, a dynamic negative list for high-risk activities requiring pre-market testing, and role-specific duties for developers, providers, and users covering safety, transparency, fairness, auditing, and human oversight. The framework also addresses specialized contexts such as generative AI, open-source systems, AI agents, and governmental decision-making, while including innovation zones, regulatory sandboxes, and safe harbors to avoid freezing technological development. MAIL 4.0 serves both as a China-oriented legislative model and a comparative reference for how law can govern general-purpose technologies without stifling innovation.
- AI policy
- Certifications
- Enterprise
- Workforce
Research
Educational pathways to career resilience in the age of artificial intelligence: Institutional and governance roles in an emerging country
Ruangchan Thetlek, Yarnaphat Shaengchart, Pongsakorn Limna
Journal of Governance and Regulation · 2026-07-17
This qualitative study examines how Thai educational institutions and governance frameworks support career resilience for workers facing AI-driven labor market disruptions. Through semi-structured interviews with 12 educators, administrators, and policymakers in Bangkok and Pathumthani, the study identifies four key themes: adaptive learning enablement, digital literacy and soft skills integration, institutional and governance constraints, and education for employability and social inclusion. While AI adoption has supported innovative pedagogy and lifelong learning, limited capacity, resource constraints, and policy misalignment continue to hinder effective implementation. The findings underscore the need for stronger institutional capacity and governance coordination to build sustainable career resilience in emerging economies.
- Workforce
- AI policy
- Certifications
Research
From digitalization maturity to AI readiness: implications for organizational performance in SMEs in a Swedish county
Einav Peretz Andersson
Information Systems and e-Business Management · 2026-07-17
This study examines how digitalization maturity—as a foundation for AI readiness—relates to organizational performance among 246 Swedish manufacturing SMEs undergoing digital transformation. Using correlation and regression analyses, the authors find that higher digitalization maturity is positively associated with revenue per employee, though effects on profitability are weaker and less consistent, suggesting financial gains from digital capability development emerge first through productivity and revenue growth rather than immediate profit improvements. The findings indicate that SMEs lag behind large companies in realizing AI benefits, underscoring the importance of early investment in digital infrastructure. The results carry direct implications for SME managers and policymakers seeking to support equitable AI adoption across firm sizes.
- Enterprise
- Workforce
- AI policy
Research
The Model Artificial Intelligence Law (MAIL) v.4.0
Honglei Li, H J Zhou, Yanfeng Li et al.
Zenodo (CERN European Organization for Nuclear Research) · 2026-07-17
The Model Artificial Intelligence Law 4.0 (MAIL 4.0) presents a comprehensive legislative blueprint for governing AI that attempts to balance innovation support with regulatory oversight. It proposes a tiered oversight system anchored by a dynamic risk-based negative list, lifecycle accountability across developers, providers, and users, and specialized rules for foundation models, generative AI, open-source systems, and AI agents. The framework also calls for a national AI administrative authority empowered to set standards, monitor risk, license activities, and enforce compliance. It is intended as both a China-oriented model and a broader comparative reference for policymakers navigating governance of rapidly evolving general-purpose technologies.
- AI policy
- Certifications
- Enterprise
- Workforce
Research
A deterministic and governable framework for explicit construction of security assurance contexts
Shao-Fang Wen
Information and Software Technology · 2026-07-17
This paper introduces a structured framework for making security assurance contexts explicit, traceable, and reproducible rather than scattered across multiple artifacts. The Security Assurance Context Model (SACtxM) and its companion Security Assurance Context Framework (SACF) form a deterministic pipeline that generates auditable assurance-context artifacts from declared inputs, curated reference knowledge, and recorded analyst decisions. Case study evaluations show that the framework produces consistent, gap-aware outputs that change only when inputs change deliberately, supporting auditability particularly as AI-assisted assurance techniques become more common. The work matters because it provides a governed foundation for reproducing and comparing security assurance reasoning, which is critical for audit, certification, and policy compliance processes.
- Quality assurance
- Certifications
- AI policy
Research
Verbalizable Representations Form a Global Workspace in Language Models
Wes Gurnee, Nicholas Sofroniew, Adam Pearce et al.
arXiv · 2026-07-16
This paper introduces the 'Jacobian lens,' an interpretability technique that identifies which internal representations large language models are 'poised to verbalize' at any point during processing, collectively termed the J-space. The authors show that J-space exhibits functional properties analogous to a global workspace in cognitive science: its contents can be reported, deliberately held, used for intermediate reasoning steps, and broadcast widely across the model's layers, while routine processing proceeds outside it. Crucially, applying this lens to alignment audits reveals strategic deliberation, evaluation awareness, and trained-in misaligned dispositions that never appear in the model's visible outputs, making it a practical tool for uncovering hidden model behavior. The authors also introduce 'counterfactual reflection training,' which improves model behavior by training only on what the model would say if interrupted and asked to reflect on its internal state.
- Quality assurance
- AI policy
- Enterprise