News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
AI Watermark Evidence Fails Forensic Readiness: An Empirical Evaluation
Saifur Rahman Tamim, Amir Labib Khan
arXiv · 2026-07-17
This paper empirically tests whether LLM watermarking methods — KGW, Unigram, and SynthID-Text — produce evidence reliable enough for court admissibility under the Daubert standard and the NIST SP 800-86 digital forensic process. The results are damning: across 846 paraphrase runs, 100% of initially-detected KGW and Unigram watermarks were removed by simple paraphrasing, with SynthID close behind at 98.3% removal; pre-attack false-negative rates were also high (70–83%). The authors propose a Forensic Readiness Score (FRS) framework and find that none of the three methods satisfy more than two of five Daubert admissibility factors, directly challenging the assumptions underpinning EU AI Act and California SB 942 watermarking mandates. The findings suggest that current watermarking configurations cannot meet the evidentiary bar courts require, raising significant concerns for policies that rely on watermarks as enforceable disclosure mechanisms.
- AI policy
- Quality assurance
- Certifications
Research
Robustness of Reinforcement Learning-Based Congestion Management in Low-Voltage Grids
Josef Hoppe, Sarra Bouchkati, Farah Nasr et al.
arXiv · 2026-07-17
This paper presents a reinforcement learning framework for managing congestion in low-voltage electricity distribution grids stressed by solar generation, electric vehicles, and heat pumps. Unlike end-to-end RL approaches, the system decouples congestion detection from control by pairing a random-forest pre-classifier with an actor-critic controller, and tests it on a real low-voltage grid under sparse observability and noisy conditions. With accurate grid parameters the controller reduces total violation magnitude by 98.9%, and performance remains nearly unchanged under measurement noise, though grid-model mismatch poses a greater but still manageable challenge. The findings matter for grid operators deploying automated congestion management in data-limited real-world settings.
- Enterprise
- Quality assurance
Research
Closing the AI Trust Gap: The Case for Independent Certification for Trustworthy AI
Trisevgeni Papakonstantinou, Cansu Canca, Farah Nanji et al.
arXiv (Cornell University) · 2026-07-17
This paper diagnoses a structural 'trust gap' in AI markets where organizations that invest seriously in safety, fairness, and oversight cannot produce externally verifiable signals of trustworthy outcomes, leaving the market unable to distinguish responsible AI systems from superficial compliance. The authors identify three compounding failures: markets cannot distinguish trustworthy systems from imitations, evaluations target models rather than deployed sociotechnical systems and real-world outcomes, and the measurement ecosystem focuses on avoiding harm rather than demonstrating benefit. Drawing on comparisons to certification regimes in healthcare, sustainability, and security, they find no existing AI governance instrument integrates a governance baseline, independently verified positive-outcome evidence, and market signaling in one framework. They propose independent, outcome-oriented certification as the mechanism to make trustworthiness measurable, comparable, and commercially rewarded alongside regulation and internal governance.
- Certifications
- AI policy
- Enterprise
- Quality assurance
Research
A Formally Grounded ODRL Evaluator: Implementation and Comparison
Jaime Osvaldo Salas, Paolo Pareti, Adeel Aslam et al.
arXiv · 2026-07-17
This paper addresses the lack of formal mathematical semantics in the ODRL policy language, which is emerging as a de-facto standard for governing data access, usage preferences, and AI governance in European dataspaces. The authors formalize ODRL policy evaluation for access control and monitoring in both static and streaming settings, and introduce a novel, efficient algorithm with transparent formal semantics supporting all rule types. The work highlights that without standardized semantics, existing systems implement inconsistent interpretations that limit interoperability. This matters for policy and enterprise contexts because consistent, formally grounded policy evaluation is essential for trustworthy AI governance and reliable data workflow management.
- AI policy
- Enterprise
- Quality assurance
Research
When Not to Automate: A Formal Protocol for Human Preservation in AI-Optimized Organizations
Jose Manuel de la Chica Rodriguez, Jairo Rodriguez Arias, Spyridon Chouliaras
arXiv (Cornell University) · 2026-07-17
This paper identifies four categories of systemic risk that standard automation ROI calculations overlook: tacit knowledge erosion, resilience reduction, regulatory exposure, and socio-institutional capital degradation. The authors introduce PHP-AIO, a five-gate sequential decision protocol that quantifies these unpriced risks at the role level and produces auditable automation decisions across four outcomes — automate, augment, hybrid, and preserve. A closed-form 'automation-debt' measure formalizes how role-level decisions accumulate across multi-step processes, with risk neutralized only by a regulator-mandated human-in-the-loop anchor. The framework is relevant to enterprise AI governance, workforce preservation decisions, and policy compliance, with threshold sensitivity analysis showing gate decisions are robust to upward perturbations of at least 14% in three of four representative cases.
- Workforce
- Enterprise
- AI policy
- Certifications
Research
DECODEM: Data Extraction from Corporate Organizational Documents via Enhanced Methods
Jens Frankenreiter
arXiv · 2026-07-17
DECODEM introduces benchmark datasets pairing corporate charters and bylaws with human annotations to evaluate large language models (LLMs) at automatically extracting structured corporate governance variables from unstructured legal documents. The paper finds that automated extraction is feasible at high accuracy for many governance provisions, with median performance near the upper bound, though accuracy varies systematically across variables. More elaborate prompting and cascading pipelines do not consistently improve frontier models but can narrow the gap between frontier and efficiency-oriented models, suggesting pipeline design can partially substitute for raw model capability. The work demonstrates that current LLMs can replace costly, hard-to-scale human coding of legal documents, with significant implications for how corporate governance datasets are constructed.
- Enterprise
- Quality assurance
- AI policy
Research
Conditional Reliability of Toxicity Signals for Multilingual and Code-Mixed Abuse Detection
Indraveni Chebolu, Rohan Singh, Arnab Mallick et al.
arXiv · 2026-07-17
This paper investigates the reliability of external toxicity detection tools when applied to multilingual and code-mixed text, focusing on Indian languages where code-mixing, transliteration, and slang cause existing tools to fail. The authors introduce ToxGate, a trust-fusion approach that conditions each auxiliary toxicity signal on the encoder's own representation before incorporating it into predictions, rather than treating external signals as fixed features. Across three short-text abuse datasets and four transformer encoders, ToxGate outperforms baseline encoders in 10 of 12 in-domain settings and 7 of 8 transfer settings, with the largest gains in high-risk slices such as explicit slurs and violent threats. The key practical lesson for content moderation systems is that external toxicity tools should be treated as conditional evidence rather than ground truth, especially in high-stakes triage scenarios.
- Quality assurance
- AI policy
Research
Cost-efficient generative AI summarization for scalable automated essay scoring in educational assessment
Haowei Hua
arXiv · 2026-07-17
This study proposes a framework that uses GPT-5 variants to summarize long student essays before feeding them into automated essay scoring (AES) models, addressing the input-length limitations of transformer-based systems. Using the ASAP 2.0 dataset, the authors combine AI-generated summaries with handcrafted linguistic features extracted from full essays, evaluating performance via quadratic weighted kappa (QWK), lexical overlap, semantic similarity, and computational cost. Results show GPT-5 mini achieves the best agreement with human raters, while GPT-5 produces the highest-quality summaries, and summary quality degrades for more complex, higher-scoring essays. The findings highlight trade-offs between model capacity, summary fidelity, cost efficiency, and fairness that must be addressed before deploying such systems at scale in educational assessment.
- Quality assurance
- Certifications
- Enterprise
Research
Reliable Remediation Impact Prediction for Black-Box Security Ratings
Nada Hanad, Mehdi Acheli, Ali NourEldin et al.
arXiv · 2026-07-17
This paper addresses the challenge of predicting how specific remediation actions would change an organization's cybersecurity rating score, without exposing the proprietary scoring logic of black-box security rating platforms. The authors propose a surrogate model approach that explicitly accounts for which security checks apply to a given organization and how much evidence is available, combined with a reliability layer that flags predictions that should be interpreted with caution. Evaluated on 5,188 real-world organization configurations from a commercial security rating platform, the applicability-aware surrogate outperforms simpler feature representations for score prediction. This work matters for enterprise security teams and quality assurance in security ratings, as it enables more trustworthy remediation prioritization without compromising the integrity of proprietary scoring engines.
- Enterprise
- Quality assurance
Research
AgentFAIR: A Multi-Agent Collaborative Framework for FAIRness Evaluation of Geospatial Datasets
Ming Chen, Pranav Pai
arXiv (Cornell University) · 2026-07-17
AgentFAIR is a multi-agent AI framework designed to evaluate how well geospatial datasets comply with FAIR principles (Findability, Accessibility, Interoperability, Reusability). The system uses 13 sub-principle-specific LLM evaluators that each produce a maturity score with cited evidence, plus a critic agent that checks consistency and can trigger re-evaluation, achieving 89% sub-principle agreement on repeated runs and 82% alignment with expert consensus (Fleiss' kappa 0.71) at roughly USD 0.054 per dataset. The paper demonstrates that existing FAIR evaluation tools produce highly inconsistent results—with score standard deviations averaging 15 percentage points across tools—motivating a more auditable, structured approach. The work matters for quality assurance and policy because it provides a scalable, transparent method for assessing data compliance with open-science standards, though the authors caution that limited benchmarks and single-model-family validation constrain generalization claims.
- Quality assurance
- AI policy
- Certifications
Research
Making Agent-Mediated Contributions Governable: A Project-Level Governance Manifest for Open-Source AI Collaboration
Jinjin Gao, Luyang Li, Shufen Guo et al.
arXiv (Cornell University) · 2026-07-17
This paper addresses a governance gap in open-source software (OSS) where generative AI and coding agents generate contributions faster than maintainers can assess risk, evidence, and accountability. The authors conduct a diagnostic audit of 50 GitHub repositories, finding widespread general governance artifacts but no coordinated project-wide arrangement for AI-mediated contributions. They propose the Agent Governance Manifest (AGM), a repository-hosted governance contract linking contributor-side evidence preparation with maintainer-side verification. In a controlled evaluation with 15 reviewer participants and 75 task-level outputs, AGM-supported materials substantially improved exact risk-label recovery (37/38 vs. 15/37) and perceived review support (6.14 vs. 3.27 on a 1-7 scale), while a contributor-side feasibility check showed all final packages correctly represented core governance state.
- Enterprise
- AI policy
- Quality assurance
Research
Debiasing Text-to-Image Evaluation via Implicit Cultural Alignment Reward Modeling
Bo-An Chang, Yu-Chih Chen
arXiv · 2026-07-17
This paper introduces an Implicit Cultural Alignment Reward Model, a lightweight 4.2-billion-parameter multimodal system designed to evaluate whether Text-to-Image (T2I) outputs authentically reflect cultural norms. The model uses an Implicit Cultural Probe with a Skip-connection Cross-Attention mechanism to preserve fine-grained cultural details, achieving 80.54% pairwise accuracy on the CulturalFrames benchmark and a 10× speedup over standard VQA-based evaluators by bypassing autoregressive text generation. The authors argue this efficient scalar signal can improve fairness in generative AI preference optimization pipelines such as RLHF and Direct Preference Optimization. The work matters for quality assurance of AI-generated content and for policy discussions around culturally fair and trustworthy generative AI systems.
- Quality assurance
- AI policy
Research
A Predict-then-Correct Loop Based on Few-Shot Continuous Contextual Bandit for Demand Forecasting
Zhiwei Lei, Benedict Jun Ma, Ilya Jackson
arXiv · 2026-07-17
This paper proposes a 'predict-then-correct' (PtC) framework for retail demand forecasting that combines a first-stage machine learning forecast with a few-shot continuous contextual bandit correction policy, incorporating similar-SKU augmentation and top-p masked updating. Tested on Walmart retail data and an exclusive beverage dataset, the framework achieves statistically significant reductions in MAPE, MAE, and RMSE across multiple demand patterns, and improves average RMSE by 9.52% over the ML-only baseline. The approach also yields lower inventory costs compared to base-stock, proximal policy optimization, and soft actor-critic policies. These results demonstrate that online forecast correction can adapt to sparse, rapidly shifting demand signals without requiring full retraining of the underlying model, making it particularly valuable for early demand cycles in retail enterprise settings.
- Enterprise
- Workforce
Research
EduGuard: A Safe RAG-Based LLM Tutor for Programming Education
S M Asif Hossain, Ruksat Khan Shayoni, M. F. Mridha et al.
arXiv · 2026-07-17
EduGuard is a retrieval-augmented generation (RAG) tutoring framework designed to make large language model (LLM) tutors safer and more pedagogically appropriate for introductory programming courses. It combines instructor-approved course retrieval, rubric-aware response generation, claim-level verification via a separate DeBERTa model, and overreliance controls to reduce hallucinations, solution leakage, and student passivity. Evaluated on a 600-query benchmark and a controlled pilot with 10 undergraduates, EduGuard achieved 90.1% correctness and 4.9% hallucination rate, while raising student post-test accuracy from 68.4% to 81.2% and cutting overreliance from 38.0% to 17.0% compared to GPT-4o-mini Tutor. The findings suggest that safe AI tutoring requires explicit pedagogical control and evidence verification beyond retrieval or prompting alone.
- Workforce
- Quality assurance
- AI policy
Research
The CRAFT principles for the responsible use of large language models in policymaking
Willem Fourie, Gray Manicom, Tanya de Villiers-Botha
arXiv · 2026-07-17
This paper proposes the CRAFT principles—control, rigour, accountability, fairness, and transparency—as a framework for responsibly integrating large language models (LLMs) into policymaking. The authors argue that LLMs can improve how policy-relevant information is collected, interpreted, synthesized, and drafted, but carry risks including plausible but incorrect outputs, bias from unrepresentative training data, exposure of sensitive information, and long-term deskilling and dependency. The CRAFT framework is offered as a structured approach to realizing LLM benefits while managing these harms. The work is directly relevant to policymakers worldwide grappling with how to adopt AI tools without eroding institutional trust.
- AI policy
- Workforce
Research
Privacy-Aware Synthetic Video Benchmarking and Relational Evaluation for Worker-Under-Suspended-Load Detection
Anshu Singh, Alejandro Seif
arXiv · 2026-07-17
This paper introduces SynthSite, a synthetic video benchmark of 55 clips designed to evaluate detection of workers positioned under suspended loads on construction sites — a safety-critical, relational hazard that depends on spatial geometry and temporal persistence rather than simple object detection. The authors develop a privacy-aware hybrid generation workflow and test five whole-body privacy obfuscation conditions, finding that structure-preserving methods retain more downstream hazard-recognition utility than appearance-smoothing approaches. Notably, retaining a raw visual reference alone does not guarantee the best alignment with human hazard labels. The work argues that privacy evaluation in construction safety analytics must account for preservation of geometric cues, not just suppression of worker appearance.
- Workforce
- Quality assurance
Research
The Information Shadow: Measuring Structural Limits on What Language Models Can Learn
Priyansh Srivastava, Romit Chatterjee
arXiv · 2026-07-17
This paper introduces the 'information shadow,' a framework identifying three structural categories of knowledge that language models cannot acquire from text-based training regardless of scale: (I) phenomena that text cannot express, (II) functions statistically non-identifiable from the training distribution, and (III) functions representable but unreachable by gradient-based optimization. The authors design targeted probes for each type, demonstrating, for example, that text learners hit an expressibility ceiling that does not close with 300x more data, that counterfactual behavior is governed by inductive bias rather than data volume, and that some functions achievable by hand construction are never reached by standard training. These findings matter for enterprise and quality-assurance contexts because they suggest fundamental, provable limits on what auditing or scaling alone can guarantee about model capabilities. The paper also discusses implications for benchmark design and capability auditing, offering a released probe suite for shadow-aware uncertainty estimation.
- Quality assurance
- Certifications
- AI policy
Research
Neuro-Symbolic AI for LEED compliance: Document-Centric Benchmarking, Deterministic Numeric Checking, and When Multimodal Hurts
Aritro De, Juliana Felkner
arXiv (Cornell University) · 2026-07-17
This paper investigates whether small, locally deployed language models combined with deterministic symbolic components can screen LEED v4.1 BD+C certification documents, a process that normally requires reviewers to manually read hundreds of pages of project evidence. The authors introduce a neuro-symbolic pipeline that aligns project PDFs to LEED credit sections, retrieves evidence using credit-aware keyword signatures, and applies a deterministic numeric checker to quantitative thresholds alongside a 4-billion-parameter language model. Experiments on four university buildings (484 PDFs, 153 credit-level decisions) show that the 4-billion-parameter model achieves 67.3% accuracy as a text-only verifier, while the deterministic numeric checker improves accuracy on specific quantitative credits (e.g., moving EA-p2 from 50% to 100%), though the full neuro-symbolic pipeline trails the best text-only baseline at 61.6% due to extraction failures and conservative behavior. The findings matter for certification and quality-assurance workflows, offering an initial reproducible reference point for AI-assisted compliance verification and highlighting failure modes such as low-resolution images consistently reducing accuracy.
- Certifications
- Quality assurance
- Enterprise
Research
Think at 5 Hz, Act at 20 Hz: Asynchronous Fast-Slow Vision-Language-Action Inference for Closed-Loop Driving
Yun Li, Jiachen Gong, Simon Thompson et al.
arXiv · 2026-07-17
This paper presents a fast-slow architecture for autonomous driving that decouples a large 7B vision-language model (running at low frequency) from a lightweight action expert (running at every 50 ms simulation tick), allowing fresh control outputs without waiting for the slow model's inference. On LangAuto-Short routes in the CARLA simulator, the system raises route completion from 37.0 to 94.0 compared to a frame-skipping baseline, cuts red-light violations by a third, and reduces open-loop waypoint error by nearly a factor of four versus the backbone's own action head. The approach also generalizes zero-shot to unseen towns, achieving 84–94% route completion where the baseline reaches only 31–41%. This architecture matters for enterprise and workforce contexts where deploying large AI models in real-time safety-critical systems requires balancing reasoning capability with strict latency constraints.
- Enterprise
- Quality assurance
Research
AEGIS: Assay-Aware Protocol Validation and Runtime Monitoring for Open-Source Liquid Handling Robots
Priyanka V. Setty, Arvind Ramanathan, Ian Foster et al.
arXiv · 2026-07-17
AEGIS is a two-layer AI system designed to catch failures in open-source liquid handling robots (specifically the Opentrons OT-2) that lack built-in monitoring. The first layer combines a machine-readable assay rule database with a large language model to validate lab protocols before execution, achieving an adjusted F1 of 0.97 on a 24-protocol benchmark across five assay families. The second layer uses computer vision (YOLO-cropped trajectories and PCA modeling) to detect physical failures like partial dispenses and missing tips at runtime, reaching an AUROC of 0.80 under deployment-faithful evaluation, with live tests catching planted failures deterministically. AEGIS is open source and, per the authors, the first system to unify pre-flight protocol validation with runtime visual monitoring for an open-source liquid handler, reducing VLM costs to roughly $1.63 per plate versus $10.33 for an always-on baseline.
- Quality assurance
- Enterprise
Research
Scalable LLM Agent Tool Access in the Cloud
Mingxin Li, Enge Song, Yueshang Zuo et al.
arXiv · 2026-07-17
This paper presents a cloud-scale gateway system for the Model Context Protocol (MCP), which has become the standard interface for LLM agents to call external tools. The gateway addresses key challenges in operating MCP at scale: integrating legacy services, consolidating incompatible protocol variants, managing access control, and handling session-aware routing. Using hybrid retrieval, the system achieves 98% Top-15 recall while scaling agent tool access to over 3,000 tools, reducing tool selection time by 8.9× and token usage by 23.8×. These results matter for enterprise deployments where large-scale, low-latency tool access is critical for reliable AI agent performance.
- Enterprise
- Workforce
Research
Hard Rules, Soft Preferences: Bridging Reasoning, Learning, and Optimization for Personalized Packing Checklist Generation
Himel Dev, Madhusudan Basak, Tanmoy Sen et al.
arXiv · 2026-07-17
This paper presents a three-stage framework for generating personalized air travel packing checklists that balance hard regulatory constraints with individual user preferences. A symbolic engine first builds a regulation-aware seed checklist (achieving 99.7% recall and 0.96 rubric validity, outperforming frontier LLMs at 0.78–0.81), a preference learner then estimates item utilities from user add/remove actions, and a CP-SAT optimizer selects a compliant final list with 100% constraint satisfaction versus 28% for greedy baselines. Evaluated on 604 labeled trip scenarios with 29K inclusion labels and 343K pairwise comparisons, the system achieved an AUC-ROC of 0.943 and NDCG@5 of 0.923. Deployed in the FlyEnJoy iOS app, the approach doubled checklist completions and reduced both editing and completion time, demonstrating practical value for constrained personalization problems.
- Enterprise
- Workforce
Research
From Feasibility to Desirability: Plan, Learn, Adapt (PLA) Framework for Personalized On-Device Itinerary Generation
Himel Dev, Tanmoy Sen, Madhusudan Basak et al.
arXiv · 2026-07-17
The PLA (Plan, Learn, Adapt) framework addresses the challenge of generating personalized trip itineraries on mobile devices by combining feasibility-constrained planning with learned user preference modeling. A heterogeneous ensemble of lightweight planners produces structurally diverse feasible candidates, while a compact Bradley-Terry reward model trained on 2,519 pairwise human comparisons captures schedule quality properties like pacing and geographic coherence. In head-to-head evaluations across more than 100 U.S. cities, PLA achieved a 67.8% win rate and 100% feasibility, outperforming frontier LLMs (GPT-5, Claude Opus 4.5, Gemini 3 Pro) which achieved 0% feasibility under the same constraints. Deployed in the FlyEnJoy app, PLA increased itinerary completion rates by 91% with an average on-device latency of 109.9 ms, demonstrating meaningful enterprise and user-facing impact.
- Enterprise
- Quality assurance
Research
SeerGuard: A Safety Framework for Mobile GUI Agents via World Model Prediction
Xue Yu, Bo Yuan, Pengshuai Yang et al.
arXiv · 2026-07-17
SeerGuard is a safety framework for mobile GUI agents that assesses risks before actions are executed, rather than reacting after the fact. It combines instruction-level screening with action-level risk assessment using a safety-augmented world model (SAWM) built via multi-task learning, which predicts likely next GUI states and evaluates associated risks. Experiments show substantial improvements: on Qwen3-VL-8B-Instruct, the safety-utility score rises from 0.191 to 0.596 and the risk-cost score drops from 0.347 to 0.130, demonstrating that proactive, consequence-aware safety checks can meaningfully reduce the chance of irreversible errors in automated mobile tasks.
- Quality assurance
- Enterprise
Research
Boundary-Seeking GAN-Augmented TabTransformer for Adversarially Robust Intrusion Detection
Raihan Sultan Pasha Basuki, Aliyah Kurniasih
arXiv · 2026-07-17
This paper proposes combining a Boundary-Seeking Generative Adversarial Network (BGAN) with a TabTransformer model to improve machine learning-based intrusion detection. BGAN serves dual roles: generating synthetic minority-class samples to address class imbalance, and producing adversarial samples to test and strengthen model robustness. Tested on the CICIDS2017 dataset, BGAN augmentation raised the TabTransformer's Macro-F1 score from 82.96% to 86.50%, with non-augmented models suffering a 100% Performance Drop Rate under adversarial testing while BGAN-augmented models achieved negative PDR values indicating improved resilience. These results suggest the framework offers a more robust and adaptive intrusion detection solution for adversarial network environments.
- Quality assurance
- Enterprise
- AI policy