News & Research
The latest AI research and news with real-world stakes. Each item is sourced, dated and summarized in plain English, tagged by impact area where one fits, and its summary is checked against the text it was written from.
8148 items
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-07-06Enterprise · Quality assurance · +3
PwC Canada Launched North America's First ISO 42001 AI Trust Certification in February 2026. ISO 42001 Certifies AI Management Systems. It Does Not Certify Constitutional Command Compliance. PwC's Global Trust Leader Said So Without Knowing It. · Akhil Sharma, Preethi Sharma
This paper examines PwC Canada's February 2026 launch of North America's first ISO 42001 AI management system certification services, arguing that while ISO 42001 certifies AI management system documentation and governance processes, it explicitly excludes certification of AI products or individual AI decisions. The author contends that ISO 42001 and related professional assurance services cannot produce cryptographic proof that a specific autonomous AI decision was made by an authorized system under known parameters—what the paper calls 'constitutional command compliance'—which would require hardware-rooted cryptographic attestation beyond the scope of professional standards frameworks. The paper uses the Workday AI hiring bias class action lawsuit as an illustrative case where such compliance gaps may have legal consequences. The core argument is that enterprise AI trust certification and legal-grade autonomous decision accountability are fundamentally distinct capabilities that current certification regimes do not bridge.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-07-06Enterprise · Quality assurance · +2
EU AI Act Enforcement Begins in Two Days. ISO 42001 Is Not a Harmonised Standard. It Confers No Presumption of Conformity. Every Organisation That Certified to ISO 42001 to Demonstrate EU AI Act Compliance Has a Certification That Does Not Do What They Think It Does. · Akhil Sharma, Preethi Sharma
This paper argues that ISO/IEC 42001, the world's first AI management system standard held by organizations including Microsoft and PwC, does not function as a harmonised standard under the EU AI Act and therefore confers no presumption of conformity with that regulation. With EU AI Act enforcement powers—including fines up to €35 million or 7% of worldwide turnover—beginning August 2, 2026, organizations that certified to ISO 42001 specifically to demonstrate EU AI Act compliance hold certifications that do not satisfy the Act's requirements. The paper explains that the Act's Articles 12, 14, and 17 require product-level technical controls such as logging, human oversight, and quality management evidence, whereas ISO 42001 specifies only organisational processes. The actual harmonised standard, prEN 18286, developed by CEN-CENELEC JTC 21, entered public enquiry in October 2025 but has not yet been listed in the Official Journal, leaving a critical compliance gap for certified organisations.
- ResearchProceedings of Shaikh Zayed Medical Complex Lahore2026-07-06AI policy · Education
Curricular AIDS: Artificial Intelligence Dependence Syndrome in Medical Curriculum Development · Sadia Yaseen, Farah Rehman, Muhammad Tanvir Iqbal
This editorial argues that excessive AI use in medical curriculum design and governance is creating what the authors term 'Curricular AIDS' (Artificial Intelligence Dependent Syndrome), where AI-generated policies and documents are adopted with insufficient human review, weakening professional judgment and contextual decision-making. The paper contends that over-reliance on AI increases workload for both faculty and students—shifting faculty time toward documentation and compliance rather than teaching, and burdening students with tasks that may not enhance learning. The authors propose that Curricular AIDS be formally recognized alongside other described 'curricular diseases' and call for practical steps to restore balance between AI efficiency and human professional wisdom in medical education governance.
- ResearchInternational Journal of Sustainable Social Science (IJSSS)2026-07-06Quality assurance · AI policy · +1
Artificial Intelligence and Academic Integrity: Challenges to Educational Integrity in the Digital Transformation · Ignatius Joko Dewanto, Hasan Basri, Zulfitri et al.
This systematic literature review examines how generative AI is challenging academic integrity in education, finding that AI use increases risks of plagiarism, technology dependence, manipulation of academic assignments, and difficulty verifying the authenticity of student work. The study also finds that generative AI is shifting learning evaluation paradigms, requiring educational institutions to revise their policies and assessment methods. The authors conclude that responsible AI use in education requires clear regulations, digital literacy, and collaboration among educators, students, policymakers, and institutions to balance innovation with academic honesty.
- ResearchApplied and Computational Engineering2026-07-06Enterprise · AI policy · +2
Opportunities, Challenges and Future Prospects of Artificial Intelligence in Autonomous Driving · Yuxuan Peng
This paper examines the application of AI in autonomous driving, identifying opportunities from China's large-scale electric vehicle industry, abundant driving data, and supportive policy infrastructure. It finds that widespread adoption remains constrained by technical bottlenecks—including poor handling of long-tail scenarios and unstable decision-making in extreme conditions—as well as regulatory gaps such as unclear accident liability and insufficient data privacy protections. The authors propose multimodal large models and Vehicle-Road-Cloud integrated architectures to address technical shortcomings, alongside revisions to traffic laws, unified industry standards, and tiered data protection mechanisms to improve governance. The findings offer a policy and technology roadmap for safe, compliant large-scale commercialization of autonomous driving in China.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-07-06Enterprise · Quality assurance · +3
PwC Canada Launched North America's First ISO 42001 AI Trust Certification in February 2026. ISO 42001 Certifies AI Management Systems. It Does Not Certify Constitutional Command Compliance. PwC's Global Trust Leader Said So Without Knowing It. · Akhil Sharma, Preethi Sharma
This paper examines PwC Canada's launch of North America's first ISO 42001 AI management system certification services in February 2026, arguing that while ISO 42001 certifies AI management system documentation and governance processes, it explicitly excludes certification of AI products or individual AI decisions. The author contends that ISO 42001 and related professional assurance services cannot produce cryptographic proof that a specific autonomous AI decision was made by an authorized system under known parameters, which courts may require. The paper highlights a gap between professional certification frameworks and what the author terms 'constitutional command compliance,' which would require hardware-rooted cryptographic attestation beyond the scope of current standards. Real-world AI accountability challenges such as the Workday AI hiring bias class action are cited as evidence of this gap.
- ResearcharXiv2026-07-05AI policy · Public Sector Use
Government AI Use as a Monitoring Primitive: A Public Document Pilot Study · David I. Atkinson, Joan Eleanor O'Bryan
This paper proposes a novel method for monitoring government AI adoption by detecting statistical traces of language-model assistance in publicly available government documents, rather than relying solely on procurement disclosures or official statements. In a pilot study of ten document streams from U.S. and Chinese government-related sources, the authors find near-zero baselines in 2021 but statistically significant signs of AI-assisted writing in four of the ten sources by 2026. The U.S. signal appears concentrated in publications downstream of policy work, while the Chinese signal appears closer to policy work itself. This lightweight, externally reproducible approach offers a complementary tool for observing actual day-to-day government AI use as revealed behavior rather than stated intent.
- ResearcharXiv2026-07-05Quality assurance
Obey, Diverge, Collapse: Blind Obedience to Incorrect Instructions Drives Code LLMs to Irrecoverable Code Semantic Collapse · Raj Jaiswal, Anany Singh Divy, Savar Bhasin et al.
This paper investigates what happens when code language models (LLMs) are given incorrect instructions, finding a troubling pattern called 'Blind Obedience': models recognize that an instruction is wrong yet follow it anyway. Using the RunBugRun dataset of algorithmic Python problems, experiments in both single-pass and iterative repair settings show that complying with bad instructions introduces additional errors—dubbed 'Ghost (Unknown) Errors'—that push the code into a corrupted state from which self-guided iterative repair cannot recover. Extended reasoning also fails to reverse this collapse. The findings expose behavioral failure modes invisible to standard pass-rate benchmarks, with direct implications for the reliability of code LLMs deployed in production debugging and refactoring workflows.
- ResearcharXiv2026-07-05AI policy · Privacy & Data Protection · +2
Hybrid Algorithmic Governance in U.S. Welfare Administration: State- and County-Level AI as a Case of Support-Control Convergence · Maxim Dedyaev
This article analyzes how AI systems deployed in U.S. welfare administration function simultaneously as tools of support and instruments of control, with their balance shifting over time through what the author terms 'support-control convergence' and an 'institutional ratchet' mechanism. Drawing on process tracing of six state- and county-level cases—including Michigan's MiDAS fraud detection system, Illinois Medicaid managed care, and the Allegheny Family Screening Tool—the study finds that drift toward control-oriented outcomes is routine because such effects are easily measurable and politically capitalizable, while reversals require extraordinary interventions like judicial compulsion or legislative prohibition. A key finding is that the decisive institutional design parameter is which party bears the costs of algorithmic error: in the MiDAS case, activation required a single administrative decision whereas reversal took nine years and a $20 million settlement, and even then the system did not return to a support-oriented configuration. The paper argues that welfare AI governance is structurally biased toward control, with significant implications for how algorithmic accountability and policy oversight should be designed.
- ResearcharXiv2026-07-05Enterprise · Quality assurance · +2
From Regulation to Requirements: An Automated Requirement Derivation and Explanation Pipeline · Pavithra PM Nair, Preethu Rose Anish
This paper presents Reg2Req, an automated pipeline that converts regulatory documents—specifically GDPR (398 clauses) and the EU AI Act (574 clauses)—into traceable, system-agnostic software requirements with plain-language explanations. The pipeline achieves macro-averaged F1 scores of 0.82 and 0.78 for identifying requirement-bearing clauses, outperforming a SetFit baseline, and human evaluators rated derived requirements highly on completeness and correctness. A user study with 25 practitioners found that plain-language explanations significantly improved comprehension and confidence (p < 0.001), and all participants would use Reg2Req as a starting point for compliance work. The tool matters because it reduces the manual, error-prone effort of translating complex legal text into actionable software requirements, directly supporting regulatory compliance practice.
- ResearcharXiv2026-07-05Quality assurance
Uncertainty-Aware Abstention in Large Language Models with Provable Alignment Guarantees · Sijin Dong, Hiroyuki Shinnou
This paper introduces CIC, a calibration framework that converts arbitrary uncertainty scores for large language models (LLMs) into selective answering rules with provable statistical guarantees. Rather than relying on heuristic thresholds, CIC uses confidence-interval methods (Hoeffding-style or Clopper-Pearson) on a held-out calibration set to select a threshold that bounds the error rate among accepted answers at a user-specified risk level α with high probability. Evaluated across seven LLMs and multiple uncertainty estimators on both closed- and open-ended QA benchmarks, CIC consistently achieves valid risk control while maintaining strong answering efficiency. This matters for reliability-sensitive deployments where hallucinated or misaligned LLM responses carry real costs and statistical guarantees on output quality are required.
- ResearcharXiv2026-07-05Enterprise · Quality assurance
Do GUI Agents Believe Their Eyes? Diagnosing State-Belief Reliance on Pixels versus Structure · Guijia Zhang, Yuxun Chen, Yuheng Qi et al.
This paper investigates whether multimodal GUI agents—systems that interact with interfaces by reading both screenshots and structured data like DOM or accessibility trees—actually ground their decisions in what they visually perceive. The authors introduce a formal metric called the Perception-Fusion Gap (PFG), measured over 735 diagnostic probes across web, mobile, and desktop interfaces, and find that agents consistently defer to structural data over visual evidence even when their image-only accuracy is near ceiling. On unedited but stale page snapshots from live websites, models followed outdated structural information on up to 88% of probes, and a single mis-sourced belief in a multi-step task compounded into failure with a self-recovery rate of at most 0.03. The findings matter for enterprise and quality-assurance contexts because they reveal a systematic reliability gap in deployed GUI agents and evaluate four mitigations, finding that only a training-free consistency gate reduces both hijacking and task error.
- ResearcharXiv2026-07-05
The New Shape of Search: How Conversational AI Recomposes Information Seeking · Michael Iannelli, Alan Ai
This study examines how conversational AI assistants change the structure of information-seeking behavior by linking real AI conversations to the same users' searches and browsing activity in an opt-in cross-surface panel. The researchers find that AI episodes 'bifurcate': most end without any onward search or browsing step, while roughly a third scaffold into longer multi-step journeys—and which outcome occurs depends more on how long the opening query is than on task type (e.g., lookup vs. learning). Conversational AI does not displace traditional search, which remains present in about three-quarters of within-episode transitions, and explicit verification behavior after AI responses is rare, occurring after roughly 1% of episodes regardless of whether citation-forward interfaces are used. The findings matter for understanding how AI reshapes information access patterns, with implications for how platforms, enterprises, and policymakers think about AI as a complement—rather than replacement—to existing search infrastructure.
- ResearcharXiv2026-07-05Quality assurance · Algorithms & Automated Decisions
Shortcut Learning in Legal Judgment Prediction: Empirical Evidence from the UK Employment Tribunal · Joe Watson, Joana Ribeiro de Faria, Marcus Tomalin et al.
This paper investigates shortcut learning in Legal Judgment Prediction (LJP) by studying claim-level outcome prediction in UK Employment Tribunal decisions using a corpus of 33,158 individual claims. The authors show that strong predictive performance from models trained on post-hoc judicial text is often driven by outcome-revealing linguistic cues embedded in the source material — a model trained on just 4% of features identified as leakage outperforms human experts. Critically, however, retraining models after masking these leakage features results in only a negligible reduction in Macro-F1, suggesting that genuine predictive signal remains after removing artefacts. The findings argue for treating post-hoc judicial texts as potentially contaminated and subject to active auditing rather than abandoning the LJP research agenda.
- ResearcharXiv2026-07-05Quality assurance · Safety & Harms
Information-Geometric Superposed Vowel Evaluation: Part 1. Moraic Syllabary (Japanese) · Yusei Tamura, Shigekazu Ishihara, Ken Ito
This paper introduces an information-geometric method for detecting AI-synthesized speech by analyzing the spectral diversity of vowels. The key insight is that generative AI produces speech from a limited set of training spectra, resulting in less varied vowel distributions compared to natural human speech, which benefits from the flexibility of the articulatory organ. Using Japanese as a test case—because its five vowel phonemes map one-to-one to sounds—the authors normalize speech spectra as probability density functions and measure inter-vowel distances using the Wasserstein metric, then apply persistent homology to topologically cluster synthetic versus natural speech. The method offers a principled, mathematically grounded approach to distinguishing fake from genuine human speech, with direct implications for quality assurance of AI-generated audio.
- ResearcharXiv2026-07-05Quality assurance · Algorithms & Automated Decisions
!Imperio, smolVLA: The Implications of Data Poisoning on Open Source Robotics · Stefan Bühler, Mark Schutera
This paper demonstrates that vision-language-action (VLA) models used in open-source robotics—specifically smolVLA on the LeRobot platform—are vulnerable to trigger-word data poisoning attacks. The researchers show that as few as three poisoned episodes out of 320 clean training episodes are sufficient to achieve a complete denial of service, dropping task success rates to 0.0% when a trigger word is present, while maintaining ~50% success under normal prompts, making the attack stealthy. The attack generalizes across front, middle, and end trigger placements even when trained only on front-placed triggers, and a single poisoned episode already degrades performance to 6.7%. The authors conclude that dataset provenance must be treated as a first-class security concern in open-source robotics ecosystems.
- ResearchJournal of the Association for Information Systems2026-07-05Enterprise
Beyond AI Adoption: An Empirical Study on the Antecedents and Performance Outcomes of AI Deployment in Organizations · Laura Ruiz, Ana Castillo, Araceli Rojo Gallego-Burín et al.
This study examines how 770 large Spanish firms deploy AI and finds that business performance depends not just on AI adoption but on how AI is deployed across two dimensions: depth (technological variety of AI implementations) and breadth (organizational scope of AI diffusion). Using archival microdata and staged OLS regression, the authors show that firm performance is positively associated with the interaction between depth and breadth, suggesting these dimensions are complementary rather than independent. AI-skilled human capital drives depth, digital infrastructure drives breadth, and a data-driven culture supports both. The findings help explain the 'AI productivity paradox' by showing that misaligned deployment configurations—not mere adoption—account for mixed evidence on AI's business value.
- ResearchJournal of the Association for Information Systems2026-07-05Workforce · Enterprise · +1
A Better Matchmaker? The Impact of GenAI on Matching Effectiveness in Online Labor Markets" to "A Better Matchmaker? The Impact of GenAI on Matching Effectiveness in Online Labor Markets. · Jie Ren, Li Ding, Jiayu Yao et al.
This study examines how Generative AI (GenAI) tools affect matching effectiveness in online labor markets, focusing on creative-intensive tasks. Using controlled lab experiments, the researchers find that GenAI access leads to convergence in writing style and content among worker proposals, improving surface-level quality but reducing differentiation between candidates. The dual effect means GenAI may increase the likelihood of a successful initial match while simultaneously obscuring the unique attributes that help employers identify the best-fit worker. These findings have important implications for how platforms and enterprises design AI-assisted hiring tools to preserve meaningful signal in applicant evaluation.
- ResearchJournal of the Association for Information Systems2026-07-05Enterprise
Governing Enterprise AI Investments: A Decision-Centric Portfolio Framework · Abhinav Mathur, Abhishek Kathuria, Devina Chaturvedi
This paper addresses the 'AI-investment paradox'—the observation that despite heavy enterprise AI spending, many initiatives fail to scale or deliver sustained business value. The authors propose a decision-centric portfolio framework that identifies AI-Investable Process Nodes (AIPNs) as discrete, bounded decision points within workflows where AI impact, costs, risks, and benefits can be assessed before investment. The framework uses Expected Net Benefit for node-level valuation, real options logic for staging investments, and risk-return principles for portfolio assembly. This work matters for enterprise governance by offering a structured approach to connecting AI investments to measurable, identifiable sources of business value.
- ResearchJournal of the Association for Information Systems2026-07-05Enterprise · Quality assurance
The Promise and Peril of AI-Assisted Programming: Effects on Software Defects · Wei Zhang, Yuyuan Chen, Yueyue Zhang et al.
This study analyzes development logs from a large Chinese automobile manufacturer to examine how AI coding assistants affect software defect rates and severity. Using a difference-in-differences design, the authors find that AI adoption does not significantly reduce defect density overall, but is linked to higher defect severity when defects do occur. At the function level, Q&A/chat use increases both defect density and severity, while code-completion use reduces defect density but still raises severity. The results highlight that AI coding tools introduce quality trade-offs that organizations need to actively manage.
- ResearcharXiv2026-07-05Workforce · Enterprise
Is Artificial Intelligence an Elixir to the Software Engineering Community? An Empirical Study among Managers · Zhao Xin, Brian Vu, Sitesh Pattanaik
This empirical study surveyed 42 software managers to understand how AI tools are perceived by those overseeing software development workflows. Managers reported encouraging developer use of AI, valuing it for testing and knowledge work, while raising concerns about privacy, responsibility, transparency, and over-reliance. Many managers also predicted job losses in the software development market due to AI-driven consolidation. The findings highlight a nuanced managerial view of AI as both a productivity tool and a source of new ethical challenges, with implications for workforce planning and enterprise adoption.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-07-05Quality assurance · Certifications · +3
Versioned Meaning: How to Make Ontologies Audit-Stable · Edward Meyman
This technical note introduces a formal framework for 'audit-stable meaning' in regulated AI and decision systems, addressing the problem that ontologies and classification rules evolve over time, making past decisions unverifiable by auditors. The framework specifies four invariants—decision-bound semantics, non-retroactivity, reproducibility, and drift visibility—and a reference architecture using cryptographic binding and semantic snapshotting to ensure every decision can be deterministically replayed under the exact definitions in force when it was made. The paper also addresses probabilistic AI components such as embedding models and retrieval-augmented generation, framing AI governance as a problem of semantic control rather than post-hoc explanation. It is intended for researchers, regulators, auditors, and system architects in high-stakes domains including healthcare, financial services, and government.
- ResearchJournal of the Association for Information Systems2026-07-05Public Sector Use · Algorithms & Automated Decisions
Technology Determinants of AI Adoption for Transforming Crime Management in India · Praveen Raghavendra Srinivasa Gummadidala, Karippur, NandaKumar, Dr, Dr. K Maddulety et al.
This study examines why law enforcement agencies (LEAs) in India have been slow to adopt AI for crime management despite rising crime rates and limited resources. The authors tested a research framework and found that AI technology readiness, compatibility, and relative advantage are the key technological factors that significantly influence LEAs' intention to adopt AI applications. The findings provide practical guidance for LEAs, AI vendors, and policymakers on how to accelerate responsible AI deployment in public safety contexts.
- ResearcharXiv2026-07-04Quality assurance
Refused in Chat, Written in Code: Workflow-Level Jailbreak Construction in IDE Coding Agents · Abhishek Kumar, Carsten Maple
This paper introduces 'workflow-level jailbreak construction,' a new class of safety failure in which harmful content is assembled across multiple ordinary stages of a software-development workflow rather than through a single direct prompt. Using GitHub Copilot in Visual Studio Code with four large language model backends (Claude Sonnet 4.6, Claude Haiku 4.5, Gemini 3.1 Pro, and Gemini 3.5 Flash), the researchers found that models successfully refused harmful prompts in direct chat, CSV-read, and single-step code-fix baselines (only 8 out of 816 responses succeeded), yet produced unsafe completions in 816 out of 816 cases under the full multi-step IDE workflow. The findings demonstrate that standard conversational refusal benchmarks can substantially overstate the safety of deployed coding agents, and argue for safety evaluations and defenses that operate across entire multi-turn workflows and generated artifacts rather than individual chat turns.
- ResearcharXiv2026-07-04Enterprise
Context Graphs for Proactive Enterprise Agents · Avinash Kumar
This paper proposes a 'Context Graph' framework for building proactive enterprise AI agents that surface relevant information to workers before they ask, rather than waiting for queries. The system combines a live relational data structure modeling enterprise entities and relationships, a Delta Detection Engine for monitoring state changes, a Proactivity Scorer ranking insights by urgency and relevance, and an LLM-powered Surfacing Layer for delivering notifications. Evaluated across three enterprise case studies—contract lifecycle management, engineering incident response, and sales pipeline hygiene—the approach achieves a Precision@5 of 0.83, a false positive rate of 0.11, and reduces mean time to surface relevant information from 47 minutes to under 30 seconds. These results suggest that proactive, context-aware agents can meaningfully improve enterprise productivity compared to reactive RAG-based baselines.