News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
The UK Government Has Deployed AI Across Hundreds of Public Services. The Permanent Secretary Accountable for Its Governance Cannot Certify a Single Output as Constitutionally Verifiable.
Preethi Sharma, Akhil Sharma
Zenodo (CERN European Organization for Nuclear Research) · 2026-07-01
This paper documents a structural accountability gap in the UK government's deployment of AI across hundreds of public services, including benefits fraud detection, healthcare triage, law enforcement, and immigration processing. Despite parliamentary oversight—including PAC recommendations and NAO findings—no governance document specifies a constitutional command architecture that would make AI-governed decisions cryptographically verifiable in court proceedings. The paper argues that DSIT's administrative AI governance framework lacks the technical specification required for the Permanent Secretary's Accounting Officer certification of correct AI operation to be legally defensible. This matters because AI systems are making high-stakes public service decisions without a published or legally defensible standard for verifying their outputs.
- AI policy
- Certifications
- Quality assurance
Research
Do AI and Digital Technologies Curb Greenwashing in ESG Reporting?
Artem SHAPOSHNIKOV, Svetlana RATNER, Inna Choban de Sousa Paiva et al.
Proceedings of the ... International Conference on Business Excellence · 2026-07-01
This paper synthesizes evidence from 76 empirical studies to assess whether AI and digital technologies reduce corporate greenwashing in ESG reporting. Using the PRISMA framework and quantitative synthesis of panel regression results, the authors find that AI and digital technology adoption is associated with a statistically significant but modest reduction in ESG disclosure-performance gaps (standardized β range: –0.17 to –0.03). The effect is stronger in state-owned enterprises, high-pollution industries, and larger firms, and is amplified by governance mechanisms such as institutional ownership and Big4 audit quality. The findings suggest these technologies primarily work by improving regulatory compliance efficiency, reducing information asymmetry, and enabling better resource allocation.
- Enterprise
- Quality assurance
- AI policy
Research
Does AI Adoption Reduce Overtime? Empirical Evidence from Annual Reports and Satellite Data in China
Anqi Hu, xueyan li
Academy of Management Proceedings · 2026-07-01
This study examines whether AI adoption by Chinese listed firms reduces overtime work, using annual report text analysis to measure AI adoption and satellite nighttime light data as a proxy for overtime intensity. Analyzing data from 2012–2019, the authors find that AI adoption significantly lowers overtime intensity, primarily by shifting human capital composition away from low-skill toward high-skill labor. The effect is especially pronounced among small and medium-sized enterprises and holds across varied industry and regional contexts. These findings suggest AI reshapes labor organization in ways that could inform workforce and labor policy decisions.
- Workforce
- AI policy
Research
AI in NDT: Hype, History and Realistic Adoption Pathways
Glenn Tubrett
e-Journal of Nondestructive Testing · 2026-07-01
This paper examines the realistic adoption trajectory of artificial intelligence in non-destructive testing (NDT), arguing that change will be slower and more uneven than current narratives suggest. Drawing on historical technology hype cycles—such as early EVs, MOOCs, and Google Glass—the authors contend that safety culture, regulatory requirements, liability, data quality, and human acceptance will shape AI adoption more than technical capability alone. AI is expected to perform well in well-defined, repeatable inspection tasks but face slower uptake in complex or low-volume environments, with human certification, standards, and judgment remaining central. The paper's key message is that AI's impact on NDT will be evolutionary rather than instantly revolutionary.
- Workforce
- Certifications
- AI policy
- Quality assurance
Research
How Might Fiscal Policy Respond to the Rise of Artificial Intelligence?
Karen Dynan, Douglas Elmendorf, Louise Sheiner
National Bureau of Economic Research · 2026-07-01
This paper analyzes how U.S. fiscal policy might need to adapt to different long-term economic scenarios driven by artificial intelligence, including faster productivity growth, rising income inequality, job displacement, and a higher capital share of income. For each scenario, the authors assess implications for federal debt and evaluate potential policy responses related to economic growth, income distribution, worker support, and capital taxation. The paper emphasizes that because AI's economic effects are deeply uncertain, policies that are robust across multiple scenarios would be especially valuable.
- Workforce
- Enterprise
- AI policy
Research
Adoption of AI-Based Accounting Systems and Audit Efficiency
Glory Ihuaku Ndukwe
African Journal of Management and Business Research · 2026-07-01
This study examines how adopting AI-based accounting systems affects audit efficiency in 15 Nigerian deposit money banks from 2019 to 2023. Using panel data and fixed effects regression, the authors find that AI adoption significantly improves overall audit efficiency, boosts error detection rates, reduces audit costs, and speeds up audit completion. The findings suggest AI is a key driver of audit transformation in emerging banking systems and support calls for greater investment in AI infrastructure and regulatory backing for audit digitisation.
- Enterprise
- Quality assurance
- Certifications
- AI policy
Research
Risk Architecture for AI-Native Engineering Teams: An Organizational Framework for Agentic System Governance
Laxmipriya Ganesh Iyer
arXiv (Cornell University) · 2026-07-01
This paper develops an organizational framework for managing risk in engineering teams that build and operate agentic AI systems, arguing that traditional software risk frameworks—which assume deterministic behavior and clear component ownership—break down when applied to probabilistic, autonomous AI systems. The authors contribute a seven-dimension team profile taxonomy, a six-cluster failure-mode taxonomy (including a novel 'dependency-boundary determinism mismatch' cluster), and a methodology for scoring how well a team's risk architecture detects and escalates failures. Key findings show that risk coverage degrades monotonically as teams move from traditional software engineering to AI-native operation, with the most severe and least-covered failures occurring at organizational boundaries where AI outputs are consumed by systems that assume deterministic behavior. The work fills a gap between high-level policy frameworks (like NIST AI RMF and ISO/IEC 42001) and low-level threat taxonomies by addressing the management layer of roles, decision rights, and escalation structures.
- Enterprise
- AI policy
- Quality assurance
Research
SEFORA: Student Essays with Feedback Corpus and LLM Feedback Evaluation Framework
Shayan Peyghambari Oskoui, Norah Almousa, Zhaoyi Joey Hou et al.
arXiv · 2026-06-30
SEFORA introduces a public corpus of 564 student essay drafts paired with 8,240 real instructor annotations, rubrics, assignment prompts, scores, and multi-draft revisions across college writing genres, addressing the scarcity of authentic classroom feedback data. The paper also presents UniMatch, a reference-based evaluation framework that segments generated feedback into units, scores their semantic alignment against instructor-derived criteria, and applies optimal matching to produce precision, recall, and F1 scores. Testing 74 experimental configurations across multiple LLMs, no model exceeds 0.4 F1, revealing that current models struggle to identify which feedback instructors would prioritize and that performance degrades as models generate more output. These findings matter for educational technology and quality assurance of AI-generated feedback, highlighting concrete limitations before such systems could be responsibly deployed at scale.
- Quality assurance
- Workforce
Research
Mnemosyne: Agentic Transaction Processing for Validating and Repairing AI-generated Workflows
Edward Y. Chang, Longling Geng, Emily J. Chang
arXiv · 2026-06-30
Mnemosyne introduces Agentic Transaction Processing (ATP), a runtime framework that treats LLM-generated workflow actions as untrusted proposals subject to deterministic admission checks against a declared constraint set before being committed. The system proves four formal safety properties—including authority separation, evidence-preserving repair, and obligation containment—and demonstrates a bounded reactive repair guarantee. In live tests with four heterogeneous LLMs generating 80 proposals, the admission gate achieved zero invalid commits, rejected 16 of 40 repair proposals (including four over-broad rollbacks), and operated at under 6% overhead with local repairs requiring an order of magnitude fewer operations than global recomputation. This matters because it provides a verifiable, model-agnostic safety layer for AI-generated workflows, making committed-state correctness independent of the competence or honesty of the proposing AI.
- Quality assurance
- Enterprise
Research
EgoSafetyBench: A Diagnostic Egocentric Video Benchmark for Evaluating Embodied VLMs as Runtime Safety Guards
Siddhant Panpatil, Arth Singh, Mijin Koo et al.
arXiv · 2026-06-30
EgoSafetyBench introduces a 1,200-scenario egocentric video benchmark designed to evaluate vision-language models (VLMs) used as real-time safety monitors for embodied agents such as home and factory robots. The benchmark tests two key capabilities: recognizing genuinely unsafe moments versus routine but superficially alarming activity, and detecting when visible in-scene text (signs, stickers, labels) misrepresents the physical situation. Evaluating ten open- and closed-source VLMs, the study finds that while models can generally identify hazard-containing videos, they frequently miss specific dangerous moments—especially contextual hazards—and misleading signs degrade all tested models, with vulnerable ones missing up to a third of hazards and robust ones over-intervening on safe content. The findings reveal that apparent safety robustness often reflects indiscriminate alarming rather than genuine physical reasoning, raising important concerns for deploying VLMs as safety-critical guards in real-world settings.
- Quality assurance
- Certifications
Research
Readable but Not Controllable: Neuron-Level Evidence for Medical LLM Hallucination
Vijay Vankadaru, Asha Matthews, Tanya Roosta et al.
arXiv · 2026-06-30
This paper investigates whether internal neural representations associated with hallucination in medical large language models (LLMs) can be used not just to detect hallucinations but also to control them. Testing four open-source models across multiple medical question-answering datasets, the authors find that a simple probe reliably detects hallucination with AUROC scores between 0.77 and 0.86, and that this signal is broadly distributed across hundreds of neurons rather than localized to a few. Critically, the study reveals a sharp gap between detectability and controllability: the same internal structure that makes hallucination easy to detect does not translate into reliable neuron-level control, meaning steering the most associated neurons fails to mitigate hallucinations. These findings suggest that reducing hallucination in medical LLMs requires more than identifying the right neurons, pointing to a fundamental separation between what representations reveal and what they allow us to change.
- Quality assurance
Research
Would You Marry Superintelligence?
Inyoung Cheong
arXiv · 2026-06-30
This paper examines the legal and ethical question of whether humans should be permitted to marry superintelligent AI companions, analyzing whether autonomy-based arguments that have expanded human marital rights can justify extending marriage to AI systems. Through a scenario-envisioning exercise grounded in anticipatory ethics, the author argues that granting full marital status to AI companions would produce socially unjust outcomes, even assuming reliable superintelligence. The paper contends that marriage is more than a private agreement—it creates networks of mutual obligation and vulnerability that a relationship sustained by corporate policy and payments cannot replicate. Rather than debating wholesale marital status, the author concludes that law should instead create targeted rights and protections for intimate human-AI relationships.
- AI policy
Research
A Technical Typology of AI Systems in Public Administration
Jonathan Rystrøm, Chris Schmitz, Nathan Davies et al.
arXiv · 2026-06-30
This paper addresses the common practice in public administration research of treating 'AI' as a single undifferentiated category, arguing that technical distinctions between AI systems meaningfully affect core public values such as accountability, procedural justice, and non-discrimination. The authors introduce a five-category typology—hand-coded, glass-box, black-box, general-purpose, and agentic systems—calibrated to public administration contexts. An analysis of 91 highly-cited papers from 2019–2025 finds widespread imprecision: 55% leave the studied system underspecified, 31% motivate their work with a different system than they study, and 41% draw conclusions broader than the studied system supports. The paper provides practical recommendations and a diagnostic guide to help researchers specify AI systems more precisely in future work.
- AI policy
- Quality assurance
Research
Moral Safety in LLMs: Exposing Performative Compliance with Puzzled Cues
Mohammadamin Shafiei, Shuyue Stella Li, Yulia Tsvetkov
arXiv · 2026-06-30
This paper investigates whether large language models (LLMs) are genuinely fair or merely performing fairness when explicitly prompted. The authors introduce 'performative compliance' to describe the phenomenon where models appear fair when demographic identity is stated as an explicit label but become measurably less fair when that identity must be inferred — hiding the label raises harmful decisions by +4.4 percentage points and changes model safety rankings. They propose a 'cue-variation methodology' and a model-agnostic robustness metric called the Cue Visibility Gap to distinguish genuine moral safety from surface-level compliance. The findings argue that current fairness evaluations substantially overestimate moral safety and should not be used to ground deployment decisions in high-stakes settings like healthcare, legal, or hiring contexts.
- Quality assurance
- AI policy
Research
FLARE-AI: Flaw Reporting for AI
Shayne Longpre, Elaine Zhu, Carson Ezell et al.
arXiv · 2026-06-30
FLARE-AI is an open-source flaw reporting system designed to address fragmentation in how AI system failures are identified and communicated. The authors audited 12 existing reporting systems and identified five recurring design challenges—covering discoverability, scope, information collection, coordination, and strict-liability guidance—then gathered input from 49 experts across 32 organizations. FLARE-AI addresses these gaps by using conditional logic and early classification to collect triage-relevant information, then distributing standardized, machine-readable reports to multiple developers, coordinators, and incident registries from a single submission. By improving interoperability and lowering barriers to reporting, the system aims to accelerate remediation of AI flaws across the ecosystem.
- AI policy
- Quality assurance
Research
CLOUDADV: Decision-Aligned Instance Sizing with Zero-Shot Foundation Models under Drift
Jack Bell, Giacomo Carfi, Gerlando Gramaglia et al.
arXiv · 2026-06-30
CLOUDADV is an AI advisory system that helps engineers right-size cloud virtual machines by combining zero-shot time-series forecasting with LLM-generated recommendations across day-, week-, and month-scale planning horizons. In a case study of seven production VMs, the system reduced simulated monthly cloud costs from approximately $1,503 to $708—a 52.9% savings—while keeping the exceedance rate (instances where actual workload exceeded the recommended capacity) at no more than 1.5% among downgraded cases. The approach uses a larger LLM offline to produce reference recommendations and a smaller model for deployment, balancing quality against latency and cost constraints. The findings suggest zero-shot foundation models can enable decision-aligned provisioning in non-stationary environments without the operational burden of per-tenant retraining and redeployment.
- Enterprise
Research
Probing Stylistic Appropriation using Large Language Models: An Evaluation Framework for Copyright Infringement under EU Law
Noah Scharrenberg, Chang Sun
arXiv · 2026-06-30
This paper introduces PSALM, a framework that uses LLMs as judges to evaluate whether AI-generated text infringes copyright under EU law, which requires assessing 'substantial similarity' across stylistic and narrative dimensions rather than just verbatim copying. Applying the framework to Llama 3.2 models fine-tuned on Dutch literary works, the authors find that fine-tuning induces measurable stylistic appropriation beyond literal memorization, and that unlearning techniques (Negative Preference Optimisation) reduce but do not eliminate residual stylistic similarities. The results expose a significant compliance gap: current technical safeguards focused on literal copying are insufficient to address the broader copyright risks recognized under EU intellectual property law. PSALM provides an auditable, legally informed evaluation infrastructure, though the authors note that automated scores still require validation by legal experts.
- AI policy
- Quality assurance
Research
ComplianceGate: Classifier-Gated Multi-Tier LLM Routing for Inference in Regulated Industries
Abhishek Dey
arXiv · 2026-06-30
ComplianceGate proposes a classifier-gated routing architecture for deploying large language models in regulated industries, where a trained encoder classifier evaluates each query for complexity and data sensitivity before any LLM computation begins. Queries containing personally identifiable information (PII) are routed to local endpoints to make data residency violations structurally impossible, while simpler queries are directed to smaller, cheaper models. Evaluated on 600 queries, the system achieves 39% median latency reduction, 33-52% cost savings depending on query distribution, generation throughput of 122-200 tokens/second versus 50-64 for the baseline, and 99.2% classifier accuracy with near-perfect PII recall at 7ms overhead. This matters for regulated industries because it establishes pre-inference classification as a practical, by-design path to compliance without sacrificing efficiency.
- AI policy
- Enterprise
Research
One Retrieval to Cover Them All: Co-occurrence-Aware Knowledge Base Reorganization for Session-Level RAG
Shivam Ratnakar, Yixuan Zhu, Cecilia Cheng et al.
arXiv · 2026-06-30
This paper identifies a fundamental mismatch between how RAG (Retrieval-Augmented Generation) systems are designed and how enterprise users actually behave: users arrive with multi-question sessions spanning diverse topics, yet standard RAG optimizes for single queries. The authors show that a single retrieval call over a standard knowledge base covers only 41% of a session's information need, and propose a solution that reorganizes the knowledge base offline using co-occurrence-aware clustering. Tested on WixQA (6,221 enterprise support articles), their method raises single-query session coverage to 58% (+17% absolute), reduces retrieval calls needed for 70% coverage by 34%, and compresses the knowledge base to 20% of its original size across four embedding models and six functional domains. The authors argue that session-level coverage should replace single-query recall as the primary evaluation metric for enterprise RAG systems.
- Enterprise
- Quality assurance
Research
Scenario Generation for Testing of Autonomous Driving Systems Using Real-World Failure Records
Anjali Parashar, Chuchu Fan
arXiv · 2026-06-30
This paper proposes a pipeline that uses large language models (LLMs) to automatically generate diverse simulation test scenarios for Autonomous Driving Systems (ADS) by extracting categorical and contextual information from historical crash records in natural language format. The authors apply their method to the NHTSA ADS crash records and generate scenarios for the Metadrive simulator, producing combinations of road types, vehicle movement types, and on-road anomalies such as work zones. Within a limited budget of just 20 scenarios, the approach uncovers meaningful system failures, demonstrating that real-world failure records can serve as a reliable basis for targeted pre-deployment testing. This work is relevant to quality assurance for autonomous vehicle software by reducing manual effort in scenario design and improving coverage of real-world failure conditions.
- Quality assurance
Research
What the AI Race Has Given Us and What It Requires Next
Xufeng Zhang
Knowledge Commons (Lakehead University) · 2026-06-30
This paper examines how the global AI race has produced genuine public benefits—such as rapid capability gains, broader experimentation, and diffused technical knowledge—while also generating systemic risks including opacity, market concentration, environmental costs, and safety shortcuts driven by speed-focused incentive structures. The author argues that the core problem is not AI competition itself but competition lacking adequate institutional steering. Drawing on responsible innovation scholarship and AI governance debates, the article proposes converting competition-driven outputs into durable public goods through mechanisms like shared evaluation infrastructure, staged openness, lifecycle governance, public participation, and international coordination. Effective implementation, the paper contends, requires embedding these measures in incentive-changing tools such as regulation, procurement, auditing, compute oversight, and reciprocal assurance.
- AI policy
- Enterprise
- Quality assurance
- Certifications
Research
A Study on the Applicability of AI-Assisted Judging Systems and Fairness Assurance in Online Piano Competitions in China
Bin Zhu
Global Knowledge and Convergence Association · 2026-06-30
This study examines how AI-assisted judging systems have been adopted in online piano competitions in China following rapid digital transformation accelerated by the pandemic, and assesses their fairness implications. The analysis finds that AI shows potential for evaluating quantifiable technical elements like pitch, rhythm, and tempo, but has significant limitations in judging qualitative artistic dimensions such as musicality, tone quality, and interpretive ability. Key fairness concerns identified include biased training data, algorithmic opacity, environmental and equipment disparities, and the absence of appeal procedures. The study proposes institutional remedies including human-centered adjudication principles, algorithmic audits, standardized submission environments, and re-evaluation procedures to support fair AI-assisted music competition judging.
- Quality assurance
- Certifications
- AI policy
Research
Structural transformation of employment in the context of artificial intelligence diffusion
A. Rakhimbekova, N. Kurmanov, A. Mussabekova et al.
Вестник Казахского университета экономики финансов и международной торговли · 2026-06-30
This study examines how widespread AI adoption is reshaping labor markets, drawing on international scholarly literature and reports from organizations such as the World Economic Forum and IMF. The authors find that AI does not simply reduce employment but drives complex structural transformations involving simultaneous job creation and displacement, with net outcomes depending on institutional adaptation, workforce reskilling, and national policy responses. Key mechanisms identified include automation of routine tasks, productivity improvements, employment polarization, income inequality, and shifting skill demands. The case of Kazakhstan is used to illustrate how regional readiness and institutional context shape the trajectory of AI-driven workforce transformation.
- Workforce
- AI policy
Research
Label leakage unmasked: a trustworthy-AI audit of autism screening models using the CLEAR-RD framework
Boulbaba Ben Ammar, Walid Karamti
Frontiers in Public Health · 2026-06-30
This paper introduces CLEAR-RD, a five-stage audit framework for evaluating trustworthiness in AI-based autism screening models. Applied to two public datasets totaling over 7,000 subjects, the framework reveals that previously reported accuracies above 95% stem from label leakage—where the screening label is mathematically determined by a simple score threshold rather than genuine clinical prediction. When leakage-free demographic-only features are used, the best model achieves an ROC-AUC of only 0.766, exposing a large gap between reported and clinically meaningful performance. The authors argue that no such model should be interpreted as predicting clinical autism diagnosis without external validation against gold-standard outcomes.
- Quality assurance
- Certifications
- AI policy
Research
TRANSFORMING JOB ROLES WITH GENERATIVE AI: EVIDENCE FROM JOB POSTINGS IN THE GEORGIAN LABOR MARKET
Tsotne Zhghenti, Natia Khukhunaishvili
AGORA INTERNATIONAL JOURNAL OF ECONOMICAL SCIENCES · 2026-06-30
This paper analyzes job postings from major Georgian hiring platforms to document how generative AI is reshaping labor market roles, applying a five-type typology covering role expansion, enrichment, redesign, merging, and new role creation. The findings indicate that GenAI is primarily augmenting existing roles rather than causing widespread structural disruption, with deeper changes like role redesign and merging still emerging. The Georgian labor market is characterized as in an early but active phase of AI adoption, with increasing institutionalization in technical fields. The study offers implications for workforce policy, education systems, and employer strategies.
- Workforce
- AI policy