News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
5608 items
Research
Automated Hardware Validation Test Plan Generation for Large Scale AI Datacenter Platforms Using a Generative AI Multi-Agents Architecture
Mohammed-Khalil Ghali, Saurabh Kulkarni, Prathamesh Kulkarni et al.
arXiv (Cornell University) · 2026-07-17
This paper presents a generative AI multi-agent system that automates the creation of hardware validation test plans for large-scale AI datacenters, replacing a manual, labor-intensive process. The system uses three specialized agents—ingestion, classification, and generation—to convert hardware documents and bills of materials into structured, traceable test cases. Evaluated on two production platforms, the framework achieved test coverage expansions of 74.2% and 51.4%, reduced authoring time from days to hours, and demonstrated 100% extraction fidelity according to automated and expert evaluations. The approach matters because it reduces dependence on institutional knowledge, eliminates coverage gaps, and provides full traceability from test cases back to source specifications.
- Quality assurance
- Enterprise
- Workforce
Research
Fantastic Adaptive Taxonomies and How to Use Them
Mert Cemri, Andrei Cojocaru, Melissa Pan et al.
arXiv · 2026-07-17
This paper introduces AdaMAST, a system that automatically builds structured 'failure taxonomies' from AI agent execution traces — compact vocabularies of named failure codes organized along system, role, and domain axes — without any human annotation. These taxonomies serve as a shared feedback interface that improves agent systems in three ways: guiding agent-system search, providing runtime feedback (raising SWE-agent's resolution on SWE-bench Verified Mini from 60% to 70%, and Claude Code from 64.0% to 70.7%), and enabling a verifier (AdaMAST-Judge) that improves best-of-5 accuracy on Terminal-Bench 2.0 by 8–15 points over Pass@1. The induced vocabulary is an order-of-magnitude compression of raw traces, matches expert failure annotations more closely than hand-crafted vocabularies, and adapts across domains. This matters for quality assurance and enterprise AI deployment, as it offers a principled, scalable way to diagnose and correct recurring agent failures without changing model weights or requiring human labeling.
- Quality assurance
- Enterprise
Research
A Methodology for Auditable Trustworthiness Levels in AI Lifecycle Governance
Andrea Ferrario
arXiv (Cornell University) · 2026-07-17
This paper proposes a methodology for assigning and tracking auditable trustworthiness levels to AI systems throughout their lifecycle, addressing a gap between high-level governance principles and narrow metric-driven approaches. The methodology combines a formal framework—using interpretable decision trees to learn trustworthiness levels from measurable dimensions—with a governance procedure covering design-time labeling, post-deployment monitoring, reassessment, and reporting. Key diagnostics include boundary margins and profile drift, which help detect when an AI system's trustworthiness has changed in governance-relevant ways. The approach is intended to support conformity documentation and provide an evidential basis for human decision-makers without replacing legal or expert judgment.
- Certifications
- AI policy
- Quality assurance
Research
CRAFT: Clustering Rubrics to Diagnose Weak LLM Capabilities and Generate Targeted Fine-Tuning Data
Vipul Gupta, Zihao Wang, Razvan-Gabriel Dumitru et al.
arXiv · 2026-07-17
CRAFT is a method that converts rubric-based evaluation datasets into model-specific diagnoses of weak capabilities by extracting capability descriptions from prompt-rubric pairs, clustering them into a hierarchical tree, and identifying low-performing nodes to guide targeted fine-tuning data generation. Rather than simply identifying where a model fails, CRAFT explains why it fails at the level of individual grading criteria. Tested across four open-source models in finance and legal domains against 13 held-out benchmarks, CRAFT achieves the strongest domain average performance for most models compared to prompt-level clustering and untargeted random data generation. This matters for enterprise AI development and quality assurance, as it provides a principled pipeline for diagnosing and systematically improving LLM weaknesses before deployment.
- Enterprise
- Quality assurance
Research
Harmonizing AI Safety Thresholds
Wilber Sean Anterola, Matthew Ball, Luis F. Lafuerza et al.
arXiv (Cornell University) · 2026-07-17
This paper addresses the fragmented landscape of AI safety thresholds published by frontier AI companies, which vary substantially and hinder third-party verification and cross-company comparison. The authors develop a methodology for harmonizing thresholds across three risk domains: cyber misuse, biological misuse, and automated AI R&D. For misuse risks, they use expected harm as the key primitive with an explicit risk-modeling approach accounting for risk channels and model release conditions; for automated AI R&D, they base their proposed threshold on observed rates of AI progress. The work highlights empirical gaps and limitations in current approaches and warns that without common minimum thresholds, inconsistent risk mitigation could create a race to the bottom in safety standards.
- AI policy
- Certifications
- Quality assurance
Research
Controlling Implicit Shortcut Reliance in L2 Spoken English Auto-markers
Shilin Gao, Mark J. F. Gales, Kate M. Knill
arXiv · 2026-07-17
This paper addresses a critical flaw in automated language proficiency assessment systems: transformer-based auto-markers can learn 'shortcuts'—over-relying on specific input features—allowing learners to game the system without genuinely improving their language ability. The authors introduce a novel training criterion designed to reduce this shortcut reliance in both audio-based and speech-recognition-text-based assessment systems. Results show that without the fix, both systems exhibit higher correlations with exploitable features than human raters do, and that the new training criterion brings these correlations closer to the human reference. This matters for the integrity and fairness of automated English proficiency certification and assessment tools.
- Certifications
- Quality assurance
Research
Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning
Ajay Patel, Kartik Hosanagar, Ramayya Krishnan et al.
arXiv (Cornell University) · 2026-07-17
This paper introduces BusinessCaseBench, a benchmark built from hundreds of business school case study questions spanning eighteen disciplines, each graded against expert-written instructor rubrics. The study finds that frontier large language models already score highly on these evaluations and that capability within at least one model family has improved substantially over two years. The results suggest AI performance on complex analytical knowledge work—such as synthesizing information, exercising judgment under uncertainty, and multi-stakeholder reasoning—is already strong and rapidly advancing. The authors argue this has direct implications for business school education and for entry-level professional roles historically anchored by these analytical skills.
- Workforce
- Enterprise
- Certifications
- AI policy
Research
Revisiting data-driven dynamic security assessment with a tabular foundation model
Olayiwola Arowolo, Maosheng Yang, Jochen Cremer
arXiv · 2026-07-17
This paper applies a tabular foundation model (TFM) to power system dynamic security assessment (DSA), which evaluates the risk of electrical faults using machine learning. Unlike existing approaches that require large labeled datasets and a separate trained model for each contingency scenario, the TFM uses in-context learning so a single model handles many contingencies without retraining or hyperparameter tuning. Case studies on the IEEE 68-bus system show the TFM achieves an average Macro F1 score of about 90% with only 120 labeled samples per contingency—roughly two orders of magnitude fewer than conventional methods—and can generalize to unseen contingencies using as few as 10 labeled samples with electrical distance coordinate encoding. This work matters for enterprise power system operations by dramatically reducing the data and maintenance burden of deploying AI-based grid security tools.
- Enterprise
- Quality assurance
News
Introducing Gemini 3.5 Flash Cyber
deepmind.google · 2026-07-17
Google DeepMind Blog announces Gemini 3.5 Flash Cyber, a lightweight AI model fine-tuned specifically for cybersecurity tasks such as finding, validating, and patching software vulnerabilities. Built on top of the existing Gemini 3.5 Flash architecture, the model is designed to be cost-efficient and scalable, enabling it to be invoked multiple times within an agent pipeline to scan larger codebases and discover more unique vulnerabilities than larger, more expensive models. In internal tests, the model uncovered remote code execution vulnerabilities and memory-corruption flaws in Google's own production systems within two hours, and outperformed both mainline Gemini Flash models and Claude Opus 4.6 on several benchmarks. Due to the dual-use risks of the technology, Google is initially limiting access to governments and trusted partners via its CodeMender platform, with plans to expand availability over time.
- Enterprise
- Quality assurance
- AI policy
Research
AI Watermark Evidence Fails Forensic Readiness: An Empirical Evaluation
Saifur Rahman Tamim, Amir Labib Khan
arXiv · 2026-07-17
This paper empirically tests whether LLM watermarking methods — KGW, Unigram, and SynthID-Text — produce evidence reliable enough for court admissibility under the Daubert standard and the NIST SP 800-86 digital forensic process. The results are damning: across 846 paraphrase runs, 100% of initially-detected KGW and Unigram watermarks were removed by simple paraphrasing, with SynthID close behind at 98.3% removal; pre-attack false-negative rates were also high (70–83%). The authors propose a Forensic Readiness Score (FRS) framework and find that none of the three methods satisfy more than two of five Daubert admissibility factors, directly challenging the assumptions underpinning EU AI Act and California SB 942 watermarking mandates. The findings suggest that current watermarking configurations cannot meet the evidentiary bar courts require, raising significant concerns for policies that rely on watermarks as enforceable disclosure mechanisms.
- AI policy
- Quality assurance
- Certifications
Research
Robustness of Reinforcement Learning-Based Congestion Management in Low-Voltage Grids
Josef Hoppe, Sarra Bouchkati, Farah Nasr et al.
arXiv · 2026-07-17
This paper presents a reinforcement learning framework for managing congestion in low-voltage electricity distribution grids stressed by solar generation, electric vehicles, and heat pumps. Unlike end-to-end RL approaches, the system decouples congestion detection from control by pairing a random-forest pre-classifier with an actor-critic controller, and tests it on a real low-voltage grid under sparse observability and noisy conditions. With accurate grid parameters the controller reduces total violation magnitude by 98.9%, and performance remains nearly unchanged under measurement noise, though grid-model mismatch poses a greater but still manageable challenge. The findings matter for grid operators deploying automated congestion management in data-limited real-world settings.
- Enterprise
- Quality assurance
Research
Closing the AI Trust Gap: The Case for Independent Certification for Trustworthy AI
Trisevgeni Papakonstantinou, Cansu Canca, Farah Nanji et al.
arXiv (Cornell University) · 2026-07-17
This paper diagnoses a structural 'trust gap' in AI markets where organizations that invest seriously in safety, fairness, and oversight cannot produce externally verifiable signals of trustworthy outcomes, leaving the market unable to distinguish responsible AI systems from superficial compliance. The authors identify three compounding failures: markets cannot distinguish trustworthy systems from imitations, evaluations target models rather than deployed sociotechnical systems and real-world outcomes, and the measurement ecosystem focuses on avoiding harm rather than demonstrating benefit. Drawing on comparisons to certification regimes in healthcare, sustainability, and security, they find no existing AI governance instrument integrates a governance baseline, independently verified positive-outcome evidence, and market signaling in one framework. They propose independent, outcome-oriented certification as the mechanism to make trustworthiness measurable, comparable, and commercially rewarded alongside regulation and internal governance.
- Certifications
- AI policy
- Enterprise
- Quality assurance
Research
A Formally Grounded ODRL Evaluator: Implementation and Comparison
Jaime Osvaldo Salas, Paolo Pareti, Adeel Aslam et al.
arXiv · 2026-07-17
This paper addresses the lack of formal mathematical semantics in the ODRL policy language, which is emerging as a de-facto standard for governing data access, usage preferences, and AI governance in European dataspaces. The authors formalize ODRL policy evaluation for access control and monitoring in both static and streaming settings, and introduce a novel, efficient algorithm with transparent formal semantics supporting all rule types. The work highlights that without standardized semantics, existing systems implement inconsistent interpretations that limit interoperability. This matters for policy and enterprise contexts because consistent, formally grounded policy evaluation is essential for trustworthy AI governance and reliable data workflow management.
- AI policy
- Enterprise
- Quality assurance
Research
When Not to Automate: A Formal Protocol for Human Preservation in AI-Optimized Organizations
Jose Manuel de la Chica Rodriguez, Jairo Rodriguez Arias, Spyridon Chouliaras
arXiv (Cornell University) · 2026-07-17
This paper identifies four categories of systemic risk that standard automation ROI calculations overlook: tacit knowledge erosion, resilience reduction, regulatory exposure, and socio-institutional capital degradation. The authors introduce PHP-AIO, a five-gate sequential decision protocol that quantifies these unpriced risks at the role level and produces auditable automation decisions across four outcomes — automate, augment, hybrid, and preserve. A closed-form 'automation-debt' measure formalizes how role-level decisions accumulate across multi-step processes, with risk neutralized only by a regulator-mandated human-in-the-loop anchor. The framework is relevant to enterprise AI governance, workforce preservation decisions, and policy compliance, with threshold sensitivity analysis showing gate decisions are robust to upward perturbations of at least 14% in three of four representative cases.
- Workforce
- Enterprise
- AI policy
- Certifications
Research
DECODEM: Data Extraction from Corporate Organizational Documents via Enhanced Methods
Jens Frankenreiter
arXiv · 2026-07-17
DECODEM introduces benchmark datasets pairing corporate charters and bylaws with human annotations to evaluate large language models (LLMs) at automatically extracting structured corporate governance variables from unstructured legal documents. The paper finds that automated extraction is feasible at high accuracy for many governance provisions, with median performance near the upper bound, though accuracy varies systematically across variables. More elaborate prompting and cascading pipelines do not consistently improve frontier models but can narrow the gap between frontier and efficiency-oriented models, suggesting pipeline design can partially substitute for raw model capability. The work demonstrates that current LLMs can replace costly, hard-to-scale human coding of legal documents, with significant implications for how corporate governance datasets are constructed.
- Enterprise
- Quality assurance
- AI policy
Research
Conditional Reliability of Toxicity Signals for Multilingual and Code-Mixed Abuse Detection
Indraveni Chebolu, Rohan Singh, Arnab Mallick et al.
arXiv · 2026-07-17
This paper investigates the reliability of external toxicity detection tools when applied to multilingual and code-mixed text, focusing on Indian languages where code-mixing, transliteration, and slang cause existing tools to fail. The authors introduce ToxGate, a trust-fusion approach that conditions each auxiliary toxicity signal on the encoder's own representation before incorporating it into predictions, rather than treating external signals as fixed features. Across three short-text abuse datasets and four transformer encoders, ToxGate outperforms baseline encoders in 10 of 12 in-domain settings and 7 of 8 transfer settings, with the largest gains in high-risk slices such as explicit slurs and violent threats. The key practical lesson for content moderation systems is that external toxicity tools should be treated as conditional evidence rather than ground truth, especially in high-stakes triage scenarios.
- Quality assurance
- AI policy
Research
Cost-efficient generative AI summarization for scalable automated essay scoring in educational assessment
Haowei Hua
arXiv · 2026-07-17
This study proposes a framework that uses GPT-5 variants to summarize long student essays before feeding them into automated essay scoring (AES) models, addressing the input-length limitations of transformer-based systems. Using the ASAP 2.0 dataset, the authors combine AI-generated summaries with handcrafted linguistic features extracted from full essays, evaluating performance via quadratic weighted kappa (QWK), lexical overlap, semantic similarity, and computational cost. Results show GPT-5 mini achieves the best agreement with human raters, while GPT-5 produces the highest-quality summaries, and summary quality degrades for more complex, higher-scoring essays. The findings highlight trade-offs between model capacity, summary fidelity, cost efficiency, and fairness that must be addressed before deploying such systems at scale in educational assessment.
- Quality assurance
- Certifications
- Enterprise
Research
Reliable Remediation Impact Prediction for Black-Box Security Ratings
Nada Hanad, Mehdi Acheli, Ali NourEldin et al.
arXiv · 2026-07-17
This paper addresses the challenge of predicting how specific remediation actions would change an organization's cybersecurity rating score, without exposing the proprietary scoring logic of black-box security rating platforms. The authors propose a surrogate model approach that explicitly accounts for which security checks apply to a given organization and how much evidence is available, combined with a reliability layer that flags predictions that should be interpreted with caution. Evaluated on 5,188 real-world organization configurations from a commercial security rating platform, the applicability-aware surrogate outperforms simpler feature representations for score prediction. This work matters for enterprise security teams and quality assurance in security ratings, as it enables more trustworthy remediation prioritization without compromising the integrity of proprietary scoring engines.
- Enterprise
- Quality assurance
Research
AgentFAIR: A Multi-Agent Collaborative Framework for FAIRness Evaluation of Geospatial Datasets
Ming Chen, Pranav Pai
arXiv (Cornell University) · 2026-07-17
AgentFAIR is a multi-agent AI framework designed to evaluate how well geospatial datasets comply with FAIR principles (Findability, Accessibility, Interoperability, Reusability). The system uses 13 sub-principle-specific LLM evaluators that each produce a maturity score with cited evidence, plus a critic agent that checks consistency and can trigger re-evaluation, achieving 89% sub-principle agreement on repeated runs and 82% alignment with expert consensus (Fleiss' kappa 0.71) at roughly USD 0.054 per dataset. The paper demonstrates that existing FAIR evaluation tools produce highly inconsistent results—with score standard deviations averaging 15 percentage points across tools—motivating a more auditable, structured approach. The work matters for quality assurance and policy because it provides a scalable, transparent method for assessing data compliance with open-science standards, though the authors caution that limited benchmarks and single-model-family validation constrain generalization claims.
- Quality assurance
- AI policy
- Certifications
Research
Making Agent-Mediated Contributions Governable: A Project-Level Governance Manifest for Open-Source AI Collaboration
Jinjin Gao, Luyang Li, Shufen Guo et al.
arXiv (Cornell University) · 2026-07-17
This paper addresses a governance gap in open-source software (OSS) where generative AI and coding agents generate contributions faster than maintainers can assess risk, evidence, and accountability. The authors conduct a diagnostic audit of 50 GitHub repositories, finding widespread general governance artifacts but no coordinated project-wide arrangement for AI-mediated contributions. They propose the Agent Governance Manifest (AGM), a repository-hosted governance contract linking contributor-side evidence preparation with maintainer-side verification. In a controlled evaluation with 15 reviewer participants and 75 task-level outputs, AGM-supported materials substantially improved exact risk-label recovery (37/38 vs. 15/37) and perceived review support (6.14 vs. 3.27 on a 1-7 scale), while a contributor-side feasibility check showed all final packages correctly represented core governance state.
- Enterprise
- AI policy
- Quality assurance
News
The risk of weather data sabotage is rising
technologyreview.com · 2026-07-17
MIT Technology Review reports that the growing use of AI in weather forecasting is increasing the risk posed by manipulation of weather station data, as illustrated by a real incident at Paris Charles de Gaulle Airport where temperature readings were artificially spiked in early 2026, resulting in a $20,000 payout to a prediction-market gambler. The authors, a group of climate and AI scientists, warn that while traditional forecasting systems include safeguards like data assimilation and human oversight that caught the CDG case, AI-driven 'data-driven models' are even more dependent on raw observational accuracy and may skip those filtering steps entirely. The piece outlines a risk spectrum ranging from individual fraud to coordinated market manipulation to potential national security threats, arguing that adversaries will continue probing for weaknesses as financial incentives grow. The authors recommend continuous station monitoring, embedding adversarial robustness tools throughout AI pipelines, and strengthening accountability across the full chain of data custodians.
- Quality assurance
- AI policy
- Enterprise
Research
Debiasing Text-to-Image Evaluation via Implicit Cultural Alignment Reward Modeling
Bo-An Chang, Yu-Chih Chen
arXiv · 2026-07-17
This paper introduces an Implicit Cultural Alignment Reward Model, a lightweight 4.2-billion-parameter multimodal system designed to evaluate whether Text-to-Image (T2I) outputs authentically reflect cultural norms. The model uses an Implicit Cultural Probe with a Skip-connection Cross-Attention mechanism to preserve fine-grained cultural details, achieving 80.54% pairwise accuracy on the CulturalFrames benchmark and a 10× speedup over standard VQA-based evaluators by bypassing autoregressive text generation. The authors argue this efficient scalar signal can improve fairness in generative AI preference optimization pipelines such as RLHF and Direct Preference Optimization. The work matters for quality assurance of AI-generated content and for policy discussions around culturally fair and trustworthy generative AI systems.
- Quality assurance
- AI policy
Research
A Predict-then-Correct Loop Based on Few-Shot Continuous Contextual Bandit for Demand Forecasting
Zhiwei Lei, Benedict Jun Ma, Ilya Jackson
arXiv · 2026-07-17
This paper proposes a 'predict-then-correct' (PtC) framework for retail demand forecasting that combines a first-stage machine learning forecast with a few-shot continuous contextual bandit correction policy, incorporating similar-SKU augmentation and top-p masked updating. Tested on Walmart retail data and an exclusive beverage dataset, the framework achieves statistically significant reductions in MAPE, MAE, and RMSE across multiple demand patterns, and improves average RMSE by 9.52% over the ML-only baseline. The approach also yields lower inventory costs compared to base-stock, proximal policy optimization, and soft actor-critic policies. These results demonstrate that online forecast correction can adapt to sparse, rapidly shifting demand signals without requiring full retraining of the underlying model, making it particularly valuable for early demand cycles in retail enterprise settings.
- Enterprise
- Workforce
Research
EduGuard: A Safe RAG-Based LLM Tutor for Programming Education
S M Asif Hossain, Ruksat Khan Shayoni, M. F. Mridha et al.
arXiv · 2026-07-17
EduGuard is a retrieval-augmented generation (RAG) tutoring framework designed to make large language model (LLM) tutors safer and more pedagogically appropriate for introductory programming courses. It combines instructor-approved course retrieval, rubric-aware response generation, claim-level verification via a separate DeBERTa model, and overreliance controls to reduce hallucinations, solution leakage, and student passivity. Evaluated on a 600-query benchmark and a controlled pilot with 10 undergraduates, EduGuard achieved 90.1% correctness and 4.9% hallucination rate, while raising student post-test accuracy from 68.4% to 81.2% and cutting overreliance from 38.0% to 17.0% compared to GPT-4o-mini Tutor. The findings suggest that safe AI tutoring requires explicit pedagogical control and evidence verification beyond retrieval or prompting alone.
- Workforce
- Quality assurance
- AI policy
Research
The CRAFT principles for the responsible use of large language models in policymaking
Willem Fourie, Gray Manicom, Tanya de Villiers-Botha
arXiv · 2026-07-17
This paper proposes the CRAFT principles—control, rigour, accountability, fairness, and transparency—as a framework for responsibly integrating large language models (LLMs) into policymaking. The authors argue that LLMs can improve how policy-relevant information is collected, interpreted, synthesized, and drafted, but carry risks including plausible but incorrect outputs, bias from unrepresentative training data, exposure of sensitive information, and long-term deskilling and dependency. The CRAFT framework is offered as a structured approach to realizing LLM benefits while managing these harms. The work is directly relevant to policymakers worldwide grappling with how to adopt AI tools without eroding institutional trust.
- AI policy
- Workforce