News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
5608 items
Research
Privacy-Aware Synthetic Video Benchmarking and Relational Evaluation for Worker-Under-Suspended-Load Detection
Anshu Singh, Alejandro Seif
arXiv · 2026-07-17
This paper introduces SynthSite, a synthetic video benchmark of 55 clips designed to evaluate detection of workers positioned under suspended loads on construction sites — a safety-critical, relational hazard that depends on spatial geometry and temporal persistence rather than simple object detection. The authors develop a privacy-aware hybrid generation workflow and test five whole-body privacy obfuscation conditions, finding that structure-preserving methods retain more downstream hazard-recognition utility than appearance-smoothing approaches. Notably, retaining a raw visual reference alone does not guarantee the best alignment with human hazard labels. The work argues that privacy evaluation in construction safety analytics must account for preservation of geometric cues, not just suppression of worker appearance.
- Workforce
- Quality assurance
Research
The Information Shadow: Measuring Structural Limits on What Language Models Can Learn
Priyansh Srivastava, Romit Chatterjee
arXiv · 2026-07-17
This paper introduces the 'information shadow,' a framework identifying three structural categories of knowledge that language models cannot acquire from text-based training regardless of scale: (I) phenomena that text cannot express, (II) functions statistically non-identifiable from the training distribution, and (III) functions representable but unreachable by gradient-based optimization. The authors design targeted probes for each type, demonstrating, for example, that text learners hit an expressibility ceiling that does not close with 300x more data, that counterfactual behavior is governed by inductive bias rather than data volume, and that some functions achievable by hand construction are never reached by standard training. These findings matter for enterprise and quality-assurance contexts because they suggest fundamental, provable limits on what auditing or scaling alone can guarantee about model capabilities. The paper also discusses implications for benchmark design and capability auditing, offering a released probe suite for shadow-aware uncertainty estimation.
- Quality assurance
- Certifications
- AI policy
Research
Neuro-Symbolic AI for LEED compliance: Document-Centric Benchmarking, Deterministic Numeric Checking, and When Multimodal Hurts
Aritro De, Juliana Felkner
arXiv (Cornell University) · 2026-07-17
This paper investigates whether small, locally deployed language models combined with deterministic symbolic components can screen LEED v4.1 BD+C certification documents, a process that normally requires reviewers to manually read hundreds of pages of project evidence. The authors introduce a neuro-symbolic pipeline that aligns project PDFs to LEED credit sections, retrieves evidence using credit-aware keyword signatures, and applies a deterministic numeric checker to quantitative thresholds alongside a 4-billion-parameter language model. Experiments on four university buildings (484 PDFs, 153 credit-level decisions) show that the 4-billion-parameter model achieves 67.3% accuracy as a text-only verifier, while the deterministic numeric checker improves accuracy on specific quantitative credits (e.g., moving EA-p2 from 50% to 100%), though the full neuro-symbolic pipeline trails the best text-only baseline at 61.6% due to extraction failures and conservative behavior. The findings matter for certification and quality-assurance workflows, offering an initial reproducible reference point for AI-assisted compliance verification and highlighting failure modes such as low-resolution images consistently reducing accuracy.
- Certifications
- Quality assurance
- Enterprise
Research
Think at 5 Hz, Act at 20 Hz: Asynchronous Fast-Slow Vision-Language-Action Inference for Closed-Loop Driving
Yun Li, Jiachen Gong, Simon Thompson et al.
arXiv · 2026-07-17
This paper presents a fast-slow architecture for autonomous driving that decouples a large 7B vision-language model (running at low frequency) from a lightweight action expert (running at every 50 ms simulation tick), allowing fresh control outputs without waiting for the slow model's inference. On LangAuto-Short routes in the CARLA simulator, the system raises route completion from 37.0 to 94.0 compared to a frame-skipping baseline, cuts red-light violations by a third, and reduces open-loop waypoint error by nearly a factor of four versus the backbone's own action head. The approach also generalizes zero-shot to unseen towns, achieving 84–94% route completion where the baseline reaches only 31–41%. This architecture matters for enterprise and workforce contexts where deploying large AI models in real-time safety-critical systems requires balancing reasoning capability with strict latency constraints.
- Enterprise
- Quality assurance
Research
AEGIS: Assay-Aware Protocol Validation and Runtime Monitoring for Open-Source Liquid Handling Robots
Priyanka V. Setty, Arvind Ramanathan, Ian Foster et al.
arXiv · 2026-07-17
AEGIS is a two-layer AI system designed to catch failures in open-source liquid handling robots (specifically the Opentrons OT-2) that lack built-in monitoring. The first layer combines a machine-readable assay rule database with a large language model to validate lab protocols before execution, achieving an adjusted F1 of 0.97 on a 24-protocol benchmark across five assay families. The second layer uses computer vision (YOLO-cropped trajectories and PCA modeling) to detect physical failures like partial dispenses and missing tips at runtime, reaching an AUROC of 0.80 under deployment-faithful evaluation, with live tests catching planted failures deterministically. AEGIS is open source and, per the authors, the first system to unify pre-flight protocol validation with runtime visual monitoring for an open-source liquid handler, reducing VLM costs to roughly $1.63 per plate versus $10.33 for an always-on baseline.
- Quality assurance
- Enterprise
Research
Scalable LLM Agent Tool Access in the Cloud
Mingxin Li, Enge Song, Yueshang Zuo et al.
arXiv · 2026-07-17
This paper presents a cloud-scale gateway system for the Model Context Protocol (MCP), which has become the standard interface for LLM agents to call external tools. The gateway addresses key challenges in operating MCP at scale: integrating legacy services, consolidating incompatible protocol variants, managing access control, and handling session-aware routing. Using hybrid retrieval, the system achieves 98% Top-15 recall while scaling agent tool access to over 3,000 tools, reducing tool selection time by 8.9× and token usage by 23.8×. These results matter for enterprise deployments where large-scale, low-latency tool access is critical for reliable AI agent performance.
- Enterprise
- Workforce
Research
Hard Rules, Soft Preferences: Bridging Reasoning, Learning, and Optimization for Personalized Packing Checklist Generation
Himel Dev, Madhusudan Basak, Tanmoy Sen et al.
arXiv · 2026-07-17
This paper presents a three-stage framework for generating personalized air travel packing checklists that balance hard regulatory constraints with individual user preferences. A symbolic engine first builds a regulation-aware seed checklist (achieving 99.7% recall and 0.96 rubric validity, outperforming frontier LLMs at 0.78–0.81), a preference learner then estimates item utilities from user add/remove actions, and a CP-SAT optimizer selects a compliant final list with 100% constraint satisfaction versus 28% for greedy baselines. Evaluated on 604 labeled trip scenarios with 29K inclusion labels and 343K pairwise comparisons, the system achieved an AUC-ROC of 0.943 and NDCG@5 of 0.923. Deployed in the FlyEnJoy iOS app, the approach doubled checklist completions and reduced both editing and completion time, demonstrating practical value for constrained personalization problems.
- Enterprise
- Workforce
Research
From Feasibility to Desirability: Plan, Learn, Adapt (PLA) Framework for Personalized On-Device Itinerary Generation
Himel Dev, Tanmoy Sen, Madhusudan Basak et al.
arXiv · 2026-07-17
The PLA (Plan, Learn, Adapt) framework addresses the challenge of generating personalized trip itineraries on mobile devices by combining feasibility-constrained planning with learned user preference modeling. A heterogeneous ensemble of lightweight planners produces structurally diverse feasible candidates, while a compact Bradley-Terry reward model trained on 2,519 pairwise human comparisons captures schedule quality properties like pacing and geographic coherence. In head-to-head evaluations across more than 100 U.S. cities, PLA achieved a 67.8% win rate and 100% feasibility, outperforming frontier LLMs (GPT-5, Claude Opus 4.5, Gemini 3 Pro) which achieved 0% feasibility under the same constraints. Deployed in the FlyEnJoy app, PLA increased itinerary completion rates by 91% with an average on-device latency of 109.9 ms, demonstrating meaningful enterprise and user-facing impact.
- Enterprise
- Quality assurance
Research
SeerGuard: A Safety Framework for Mobile GUI Agents via World Model Prediction
Xue Yu, Bo Yuan, Pengshuai Yang et al.
arXiv · 2026-07-17
SeerGuard is a safety framework for mobile GUI agents that assesses risks before actions are executed, rather than reacting after the fact. It combines instruction-level screening with action-level risk assessment using a safety-augmented world model (SAWM) built via multi-task learning, which predicts likely next GUI states and evaluates associated risks. Experiments show substantial improvements: on Qwen3-VL-8B-Instruct, the safety-utility score rises from 0.191 to 0.596 and the risk-cost score drops from 0.347 to 0.130, demonstrating that proactive, consequence-aware safety checks can meaningfully reduce the chance of irreversible errors in automated mobile tasks.
- Quality assurance
- Enterprise
Research
Boundary-Seeking GAN-Augmented TabTransformer for Adversarially Robust Intrusion Detection
Raihan Sultan Pasha Basuki, Aliyah Kurniasih
arXiv · 2026-07-17
This paper proposes combining a Boundary-Seeking Generative Adversarial Network (BGAN) with a TabTransformer model to improve machine learning-based intrusion detection. BGAN serves dual roles: generating synthetic minority-class samples to address class imbalance, and producing adversarial samples to test and strengthen model robustness. Tested on the CICIDS2017 dataset, BGAN augmentation raised the TabTransformer's Macro-F1 score from 82.96% to 86.50%, with non-augmented models suffering a 100% Performance Drop Rate under adversarial testing while BGAN-augmented models achieved negative PDR values indicating improved resilience. These results suggest the framework offers a more robust and adaptive intrusion detection solution for adversarial network environments.
- Quality assurance
- Enterprise
- AI policy
Research
SLAPBench: Benchmarking Multimodal Large Language Models for Four-Finger SLAP Fingerprint Verification
Bibesh Pyakurel, M. G. Sarwar Murshed
arXiv · 2026-07-17
SLAPBench introduces the first benchmark for evaluating multimodal large language models (MLLMs) on four-finger SLAP fingerprint identity verification, a task central to border control and law enforcement. Built from NIST SD302b with 7,832 image pairs, the benchmark tests five MLLMs under multiple prompting strategies and finds that prompting style largely governs whether models collapse to near-universal acceptance, while underlying model capability determines how well they discriminate identities. Claude Opus 4.8 achieves the best binary result (FAR = 20.2%) and highest AUC (0.953) among models that do not collapse, whereas open-source models vary widely and a fairness probe suggests demographic disparities worsen as discrimination weakens. These results establish a baseline revealing that current MLLMs are not yet reliable for biometric verification and that prompting choices carry significant security implications.
- Certifications
- AI policy
- Quality assurance
Research
Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching
Yan Song
arXiv · 2026-07-17
This paper identifies a fundamental conflict in production LLM deployments where prompt compression and prompt caching are typically used together but work against each other: query-aware compression generates a unique prefix per query, invalidating cached prefixes and eliminating caching discounts. The authors empirically characterize Anthropic's Sonnet API caching behavior, finding a two-tier architecture with a hit rate plateau around 0.83 rather than the ideal 1.0 assumed in prior literature. They propose Cache-Aware Prompt Compression (CAPC), which pairs query-agnostic compression with explicit cache control and a tier-preserving ratio bound to avoid over-compression. CAPC achieves the lowest cost in all 16 tested configurations on LongBench-v2—with mean savings of 49% over cache-only and 64% over query-aware compression—while maintaining answer quality within 0.05 of the uncompressed baseline, and is validated on three production workloads including an enterprise tool-using assistant and a public retail benchmark.
- Enterprise
- Quality assurance
Research
A Tool-Invariant Framework for Teaching and Assessing Computational Methods in the Age of Agentic AI
Larry Engelhardt
arXiv (Cornell University) · 2026-07-17
This paper proposes a tool-invariant framework for teaching computational methods that separates enduring conceptual knowledge—inputs, outputs, terminology, and evaluative judgment—from the tool used to execute those methods, now including agentic AI that can write and run code autonomously. The author argues that verification of computational results, rather than code authorship, has become the critical skill as AI-generated artifacts can no longer serve as evidence of student learning. To address this, the paper recommends pairing AI-free in-class coding quizzes with oral defenses of comment-stripped AI-assisted work in small-class settings. The implications are directly relevant to how computational courses should be assessed and credentialed when students can generate artifacts on demand.
- Certifications
- Quality assurance
Research
In the Driver's Seat: A Multi-Company Study on the Reality of Autonomous Driving System Testing
Qunying Song, Gao Y, Johannes Betz et al.
arXiv (Cornell University) · 2026-07-17
This paper presents an interview study with experts from nine companies across six countries to map current practices, challenges, and future directions in autonomous driving system (ADS) testing. The findings reveal that industry relies primarily on scenario-based and X-in-the-loop testing, but faces significant gaps in scenario realism, simulation fidelity, and acceptance criteria. The authors synthesize these insights into an evidence-centered closed-loop testing framework intended to provide actionable guidance. The study is notable for its multi-company, cross-national scope and its identification of AI, world models, and end-to-end approaches as potential solutions to current testing shortfalls.
- Quality assurance
- Certifications
Research
Rebalancing the algorithm: pathways for mitigating bias in predictive policing
Lata Nautiyal, Preeti Malik, Varsha Mittal et al.
Service Oriented Computing and Applications · 2026-07-17
This paper investigates algorithmic bias in predictive policing systems using the Chicago Crime Dataset, applying data re-weighting, counterfactual analysis, and algorithm auditing to identify sources of inequity. The results show that conventional predictive models disproportionately flag minority and historically over-policed neighborhoods as high-risk, with risk scores driven more by enforcement intensity than actual crime incidence. Bias-controlled models produce different hotspot predictions and reduce disparate impact measures, demonstrating that historical policing practices substantially distort AI-generated risk forecasts. The authors argue for transparency, fairness-aware modeling, and systematic auditing in deploying AI for law enforcement.
- AI policy
- Quality assurance
Research
LLMs as judges: Toward the LLM-assisted review of GSN-compliant assurance cases
Gerhard Yu, Mithila Sivakumar, Alvine Boaye Belle et al.
Journal of Systems and Software · 2026-07-17
This paper proposes using large language models (LLMs) as semi-automated judges to review Goal Structuring Notation (GSN)-compliant assurance cases for mission-critical systems such as autonomous vehicles and avionics. The authors introduce predicate-based rules that formalize established review criteria and use these to craft targeted LLM prompts, then evaluate GPT-4o, GPT-4.1, DeepSeek-R1, and Gemini 2.0 Flash on the task. Results show most LLMs demonstrate reasonably good review capabilities, though human oversight remains necessary to refine LLM-generated reviews. The work directly addresses the inefficiency and error-proneness of manually reviewing lengthy assurance case documents, with implications for regulatory acceptance of safety-critical systems.
- Quality assurance
- Certifications
Research
Assessing Entrepreneurial Readiness and Perceptions of AI: A Catalyst for Sustainable Business Model Innovation
Edmond Freo
The International Review of Multidisciplinary Research · 2026-07-17
This quantitative study of 200 entrepreneurs examines how entrepreneurial readiness and sustainability perceptions of AI jointly predict capacity for Sustainable Business Model Innovation (SBMI) aligned with UN SDGs 8, 9, and 12. Findings show entrepreneurs have moderate overall readiness (composite mean 3.08) due to financial and infrastructural constraints, and strongly recognize AI's economic efficiency potential but remain largely neutral about its ecological applications. Regression analysis found that readiness and sustainability perceptions together explain 46.9% of variance (R²=0.469, p<0.001) in SBMI capacity, with sustainability-centered perception acting as a strategic steering mechanism beyond raw technological readiness. The authors propose a Readiness-Perception-Executed Business Model to guide entrepreneurs in using AI as a responsible catalyst for sustainable development.
- Enterprise
- Workforce
Research
Digital Economy's Impact on Employment and the Role of Market Competition-The Impact of Digital Economy on Employment——Based on the Perspective of Market Competition
Leyan Chen
Journal of innovation and development · 2026-07-17
Using panel data from 31 Chinese provinces (2012–2022), this study finds that digital economy development substantially expands regional employment, with net job creation dominating over displacement in the long run. The effect is conditioned by market competition: employment gains are significantly amplified only after competitive intensity surpasses a specific threshold. The paper also documents heterogeneity across industrial structures and economic development levels, offering empirical grounding for differentiated regional employment policies.
- Workforce
- AI policy
Research
Artificial Intelligence in Special Education: Assessing Teachers' Level of Use and Readiness at Tomas Sagun Integrated School, Pagadian City, Philippines
Mike R. Canoy, April Dawn B. Ruales, Justriel G. Tutor et al.
International Journal of Educational Innovations and Research · 2026-07-17
This descriptive survey study examined AI adoption among special education teachers at a Philippine school, finding that most teachers showed only moderate readiness and limited actual use of AI instructional tools. Barriers included insufficient training, limited technology access, and lack of institutional support. The study recommends targeted capacity-building programs, policy frameworks, and resource allocation to close the gap between readiness and utilization in special education classrooms.
- Workforce
- AI policy
Research
A Human-Centric Evaluation of a Retrieval-Augmented Generation System for Explaining Quebec Insurance Contracts
David Beauchemin, Richard Khoury
arXiv (Cornell University) · 2026-07-17
This paper evaluates a Retrieval-Augmented Generation (RAG) system designed to help Quebec consumers understand automobile insurance contracts, addressing the 'advice gap' created by online insurance sales. A user study with 154 participants found the system acted as a 'cognitive equalizer,' earning high ratings for satisfaction, trust, and clarity, with users—especially those with lower financial literacy—valuing the sense of autonomy it provided even more than the knowledge gained. However, participants still preferred human agents for high-stakes or emotionally charged decisions, underscoring the need for human-in-the-loop frameworks. The findings highlight both the promise and limits of AI-driven tools in consumer-facing financial services.
- AI policy
- Enterprise
Research
Streamlining endometriosis MRI reporting: Automated extraction of #Enzian scores from pelvic MRI reports using local and online LLMs
Joana Kostova, Caroline Reinhold, Jonas D. Stief et al.
European Journal of Radiology Artificial Intelligence · 2026-07-17
This study tested whether large language models (LLMs) can automatically extract standardized #Enzian endometriosis scores from pelvic MRI reports, comparing 12 LLMs against both expert radiologist consensus and novice trainee benchmarks across 186 reports. Online LLMs achieved 93.8–95.8% accuracy against expert reference, with three models (Gemini 2.5 Pro, Grok 4, and o3) significantly outperforming novice radiologists who reached 87.4% accuracy. Local models performed more variably (61.7–92.5%) and were generally inferior to trainees. The findings suggest online LLMs could serve as training aids and support standardized reporting for surgical planning, while off-the-shelf local models are not yet ready for this task.
- Quality assurance
- Workforce
Research
The Role of Artificial Intelligence in the Lifecycle of Scientific Manuscripts: Authoring, Reviewing, and Editorial Selection
José L. Domingo
Qeios · 2026-07-17
This critical commentary and policy analysis examines how AI and large language models are reshaping three stages of scientific publishing: co-authorship, peer review, and editorial selection. The paper highlights persistent risks such as citation hallucination, AI's inability to assess novelty, and bias amplification in editorial decisions, while noting that 2025 surveys show over 50% of researchers use AI during peer review, often in violation of existing policies. The authors propose a hybrid framework that limits AI to technical verification tasks while reserving judgments on scientific merit and ethics for compensated human experts, alongside legal accountability structures and reform of exploitative economic models like unpaid review labor paired with high Article Processing Charges.
- AI policy
- Quality assurance
Research
The enforced technical mandate: A multi-layered governance model for deepfake fraud and biometric integrity
Felipe Romero Moreno
Computer law & security review · 2026-07-17
This paper analyzes how existing EU and UK regulatory frameworks are failing to keep pace with AI-generated deepfake fraud, particularly as Fraud-as-a-Service enables scalable attacks on biometric identity systems. Through comparative doctrinal analysis, the authors identify a three-layered governance gap—covering source control, distribution control, and accountability—and flag a 12-month regulatory vacuum created by misaligned EU Digital Omnibus timelines. The paper proposes six policy recommendations, including NIST IAL2 zero-retention biometric standards, mandatory C2PA digital provenance, and embedding biometric integrity protocols into ISO 20022 financial messaging to create a real-time compliance enforcement mechanism across jurisdictions. The findings are directly relevant to policymakers, financial regulators, and organizations grappling with liability under emerging AI and data protection regimes.
- AI policy
- Certifications
Research
Reversibility-Aware Staged Delegation for Enterprise Agentic AI: A Real-Options and Resilience Framework for Irreversible Actions
Kwan Hong Tan
arXiv · 2026-07-17
This paper presents Reversibility-Aware Staged Delegation (RASD), a decision framework for enterprise AI agents that classifies actions by how recoverable they are and routes them through direct execution, staged commit, human review, or blocking accordingly. In a Monte Carlo simulation of 120,000 synthetic enterprise tasks, RASD achieved higher mean net value (7.408 vs. 5.735 and 5.282 normalized units) and a dramatically lower severe-incident rate (0.390% vs. 5.937% and 4.166%) compared to confidence-threshold and expected-loss baseline policies. The framework demonstrates that raw task-failure rate is an insufficient safety metric when recovery potential varies, arguing instead for recoverability-preserving execution architectures. The findings offer operational guidance for auditability, human escalation, and risk-adjusted value creation in enterprise agentic AI deployments.
- Enterprise
- Quality assurance
Research
Deployment process for artificial intelligence applications in radiology practice
Satu I. Inkinen, Juuso H. Ketola, Teemu Mäkelä et al.
Physica Medica · 2026-07-17
This paper presents a structured framework for deploying artificial intelligence systems in radiology, covering goal-setting, procurement, implementation planning, and post-deployment monitoring. It emphasizes defining stakeholder roles, integrating with hospital information systems, meeting regulatory requirements, and establishing quality assurance protocols with clinically relevant KPIs. A phased rollout and pilot approach are recommended to minimize workflow disruption and identify integration issues early. The framework aims to ensure patient safety, legal compliance, and sustainable AI integration with measurable clinical improvements.
- Quality assurance
- Enterprise
- AI policy