News & Research
The latest AI research and news with real-world stakes. Each item is sourced, dated and summarized in plain English, tagged by impact area where one fits, and its summary is checked against the text it was written from.
Kind
7955 items
- ResearcharXiv2026-07-06Q
BioSecBench-Refusal: A paired metric for performance and alignment in agentic biosecurity risk assessment · Edwin H. Wintermute, Harmon Bhasin, Christina M. Agapakis et al.
BioSecBench-Refusal is a benchmark designed to evaluate how well AI agents balance biosecurity risk identification with appropriate refusal behavior in life science workflows. It pairs 61 legitimate biological research tasks with 46 fictional but hazard-concealing 'red-team' scenarios, testing 16 model-harness configurations. The results reveal a troubling misalignment: many AI configurations refused legitimate research tasks at rates comparable to or higher than genuinely hazardous ones, and most refusals came from upstream API filters rather than the models' own reasoning. The benchmark is released as a tool for developers to better calibrate AI capability and caution in agentic biotech contexts.
- ResearcharXiv2026-07-06QEd
Context-Masked Truncated Reasoning Audits for Answer-Key Dependence in LLM Tutors · Bonan Shen, Dingyan Shang, Youting Wang et al.
This paper investigates whether LLM-based tutors leak private teacher materials—such as answer keys and rubrics—into their student-facing explanations. Using a method called TRACE (Truncated Reasoning AUC Evaluation), the authors test 1,000 GSM8K math problems under different context conditions and find that when an answer key is accessible, the correct answer is recoverable from the model's reasoning in 998 of 1,000 cases even without any explanation. They introduce 'context-masked replay' to isolate whether early answer availability comes from the explanation itself or the hidden input, finding that masking the private context dramatically reduces detectability—but also show that wrong answer keys still cause incorrect final responses in 272 of 387 cases, meaning private artifacts can influence outputs even when early signals vanish. These findings matter for quality assurance and certification of AI tutoring systems, establishing that audits must account for hidden context to correctly attribute answer leakage.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-07-06EQCP
EU AI Act Enforcement Begins in Two Days. ISO 42001 Is Not a Harmonised Standard. It Confers No Presumption of Conformity. Every Organisation That Certified to ISO 42001 to Demonstrate EU AI Act Compliance Has a Certification That Does Not Do What They Think It Does. · Akhil Sharma, Preethi Sharma
This paper warns that ISO/IEC 42001, the world's first AI management system standard held by organizations like Microsoft and PwC Canada, is not a harmonised standard under the EU AI Act and therefore confers no presumption of conformity with EU law. With EU AI Act enforcement powers beginning August 2, 2026—including fines up to €35 million or 7% of worldwide turnover—organizations that certified to ISO 42001 believing it demonstrated EU AI Act compliance may have a false sense of legal protection. The paper argues that ISO 42001 addresses organizational processes rather than the product-level technical controls required by EU AI Act Articles 12, 14, and 17, such as cryptographic audit trails, tamper-resistant override logs, and human oversight mechanisms. The actual harmonised standard, prEN 18286, developed by CEN-CENELEC JTC 21, entered public enquiry in October 2025 and has not yet been listed in the Official Journal.
- ResearchJournal of Higher Education Theory and Practice2026-07-06WCEd
Reimagining Career and Community Colleges in an Age of Artificial Intelligence: From Entry-Level Preparation to Lifelong Capability Ecosystems · Stephen Murgatroyd
This paper argues that AI is compressing routine cognitive tasks historically defining entry-level jobs, disrupting the traditional role of career and community colleges as employment pathways. As organizations redesign jobs for immediate productivity, opportunities for learning-by-doing shrink, creating a 'recognition failure' where how capabilities are developed no longer aligns with how competence is validated. Drawing on Canadian labor-market data and international case studies from Tecnológico de Monterrey and Brainport Eindhoven, the paper outlines implications for workforce development, credentialing, and postsecondary education. The findings suggest institutions must evolve from entry-level preparation toward lifelong capability ecosystems to remain relevant in AI-transformed labor markets.
- ResearchCESifo2026-07-06E
AI Adoption, Carbon Intensity, and Rebound Effect: Evidence from China · Sébastien Houde, Wenjun Wang
Using micro-level data from Chinese firms, this paper finds that AI adoption significantly reduces carbon emission intensity, with the strongest effects among large firms, firms in AI hub regions, and high-carbon industries. AI adoption is associated with improvements in energy management, green innovation, inventory efficiency, productivity, and specialized labor. However, the study also finds a substantial rebound effect of approximately 70%, meaning that efficiency-driven carbon reductions are largely offset by increased economic activity. These findings have important implications for enterprise sustainability strategies and climate policy design.
- ResearchJournal of fintech and business analysis.2026-07-06E
Research on the optimization of ESG internal control in manufacturing enterprises driven by artificial intelligence: taking Prince Holdings as an example · Shiyang Chen
This paper examines how AI adoption improves ESG internal control quality in manufacturing enterprises, using Prince Holdings as a case study alongside a panel dataset of 15,623 firm-year observations from 3,358 A-share manufacturing companies over 2018–2023. The study finds that AI adoption is significantly and positively associated with ESG internal control quality (β = 1.051, p < 0.001), with data governance capability acting as a partial mediator and organizational readiness as a positive moderator. The authors propose a five-layer AI-ESG optimization model aligned with the COSO framework as a replicable blueprint for manufacturers seeking to move beyond manual, fragmented ESG reporting. These findings matter for enterprises and policymakers as they highlight how AI can address the growing scale and complexity of sustainability compliance obligations.
- ResearcharXiv2026-07-06EQCPShAd
AI Safety and Alignment with Human Interests · Jr. William A. Yarberry, Wesley Ladd
This chapter provides a structured overview of AI safety and alignment challenges, documenting real-world harms such as algorithmic bias in criminal justice, Facebook's role in Myanmar's ethnic cleansing, and voice cloning scams. It categorizes risks into intentional misuse and AI misalignment, and surveys technical mitigation approaches including Constitutional AI, red teaming, and sandbox testing. The chapter also reviews major governance frameworks—NIST's AI Risk Management Framework, the EU AI Act with penalties up to €35 million, ISO/IEC 42001, and IEEE 7000 standards—making it directly relevant to policymakers, certifiers, and enterprise risk managers. Its treatment of cascading failures and the gap between abstract safety principles and implementable controls highlights ongoing challenges for quality assurance in AI deployment.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-07-06PShEd
Artificial Intelligence and Child Cognitive Development: From Content Safety to Cognitive Safety · A. Krasovski
This preprint proposes a policy framework addressing how increasingly capable AI systems affect children's cognitive development, moving beyond content safety toward what the authors call 'cognitive safety.' It identifies gaps in existing AI governance and child safety regulations, and offers policy recommendations aimed at protecting cognitive development, human agency, and responsible AI deployment. The work draws on AI governance, developmental psychology, education, and digital policy to support evidence-informed regulation and interdisciplinary research.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-07-06PShEd
Artificial Intelligence and Child Cognitive Development: From Content Safety to Cognitive Safety · A. Krasovski
This preprint proposes a policy framework addressing how increasingly capable AI systems affect child cognitive development, going beyond traditional content safety to introduce the concept of 'cognitive safety.' It reviews current regulatory approaches, identifies gaps in AI governance and child safety frameworks, and offers recommendations to promote cognitive development and human agency. Drawing on AI governance, developmental psychology, education, and digital policy, the paper aims to support evidence-informed regulation and interdisciplinary research for policymakers, developers, and educators.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-07-06WP
The Citizen Dividend Economy - A New Economic Framework for the Age of Artificial Intelligence · Denton Milton
This paper proposes the Citizen Dividend Economy (CDE), a new economic framework for managing the societal disruption caused by advanced AI automating both physical and cognitive labor. The framework advocates for public ownership of foundational national AI infrastructure, with revenues flowing into a sovereign fund that pays a universal Citizen Dividend to all citizens, decoupling economic security from labor income. The authors argue this preserves market capitalism and private enterprise while ensuring AI-generated productivity gains are broadly shared. The proposal addresses workforce displacement, fiscal policy, governance design, and international precedents, and is intended to invite academic review and empirical testing rather than serve as a finished political program.
- ResearchSystems2026-07-06E
AI Usage and Employee Performance: The Dual Roles of AI Self-Efficacy and AI-Enabled HRM · Yannan Li, Xiaoxiao Geng
This study investigates how AI usage by employees translates into better job and innovation performance, finding that two mechanisms—AI self-efficacy (employees' belief in their ability to use AI) and digital HRM practices—serve as significant positive mediators. Survey data from 750 employees across major Chinese cities show that when AI adoption is supported by both individual confidence and AI-enabled HR systems, it enhances work and innovation outcomes. The findings suggest organizations should embed AI in HR systems designed to foster learning, knowledge utilization, and continuous innovation.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-07-06EQCPAd
PwC Canada Launched North America's First ISO 42001 AI Trust Certification in February 2026. ISO 42001 Certifies AI Management Systems. It Does Not Certify Constitutional Command Compliance. PwC's Global Trust Leader Said So Without Knowing It. · Akhil Sharma, Preethi Sharma
This paper examines PwC Canada's February 2026 launch of North America's first ISO 42001 AI management system certification services, arguing that while ISO 42001 certifies AI management system documentation and governance processes, it explicitly excludes certification of AI products or individual AI decisions. The author contends that ISO 42001 and related professional assurance services cannot produce cryptographic proof that a specific autonomous AI decision was made by an authorized system under known parameters—what the paper calls 'constitutional command compliance'—which would require hardware-rooted cryptographic attestation beyond the scope of professional standards frameworks. The paper uses the Workday AI hiring bias class action lawsuit as an illustrative case where such compliance gaps may have legal consequences. The core argument is that enterprise AI trust certification and legal-grade autonomous decision accountability are fundamentally distinct capabilities that current certification regimes do not bridge.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-07-06EQCP
EU AI Act Enforcement Begins in Two Days. ISO 42001 Is Not a Harmonised Standard. It Confers No Presumption of Conformity. Every Organisation That Certified to ISO 42001 to Demonstrate EU AI Act Compliance Has a Certification That Does Not Do What They Think It Does. · Akhil Sharma, Preethi Sharma
This paper argues that ISO/IEC 42001, the world's first AI management system standard held by organizations including Microsoft and PwC, does not function as a harmonised standard under the EU AI Act and therefore confers no presumption of conformity with that regulation. With EU AI Act enforcement powers—including fines up to €35 million or 7% of worldwide turnover—beginning August 2, 2026, organizations that certified to ISO 42001 specifically to demonstrate EU AI Act compliance hold certifications that do not satisfy the Act's requirements. The paper explains that the Act's Articles 12, 14, and 17 require product-level technical controls such as logging, human oversight, and quality management evidence, whereas ISO 42001 specifies only organisational processes. The actual harmonised standard, prEN 18286, developed by CEN-CENELEC JTC 21, entered public enquiry in October 2025 but has not yet been listed in the Official Journal, leaving a critical compliance gap for certified organisations.
- ResearchProceedings of Shaikh Zayed Medical Complex Lahore2026-07-06PEd
Curricular AIDS: Artificial Intelligence Dependence Syndrome in Medical Curriculum Development · Sadia Yaseen, Farah Rehman, Muhammad Tanvir Iqbal
This editorial argues that excessive AI use in medical curriculum design and governance is creating what the authors term 'Curricular AIDS' (Artificial Intelligence Dependent Syndrome), where AI-generated policies and documents are adopted with insufficient human review, weakening professional judgment and contextual decision-making. The paper contends that over-reliance on AI increases workload for both faculty and students—shifting faculty time toward documentation and compliance rather than teaching, and burdening students with tasks that may not enhance learning. The authors propose that Curricular AIDS be formally recognized alongside other described 'curricular diseases' and call for practical steps to restore balance between AI efficiency and human professional wisdom in medical education governance.
- ResearchInternational Journal of Sustainable Social Science (IJSSS)2026-07-06QPEd
Artificial Intelligence and Academic Integrity: Challenges to Educational Integrity in the Digital Transformation · Ignatius Joko Dewanto, Hasan Basri, Zulfitri et al.
This systematic literature review examines how generative AI is challenging academic integrity in education, finding that AI use increases risks of plagiarism, technology dependence, manipulation of academic assignments, and difficulty verifying the authenticity of student work. The study also finds that generative AI is shifting learning evaluation paradigms, requiring educational institutions to revise their policies and assessment methods. The authors conclude that responsible AI use in education requires clear regulations, digital literacy, and collaboration among educators, students, policymakers, and institutions to balance innovation with academic honesty.
- ResearchApplied and Computational Engineering2026-07-06EPPrAd
Opportunities, Challenges and Future Prospects of Artificial Intelligence in Autonomous Driving · Yuxuan Peng
This paper examines the application of AI in autonomous driving, identifying opportunities from China's large-scale electric vehicle industry, abundant driving data, and supportive policy infrastructure. It finds that widespread adoption remains constrained by technical bottlenecks—including poor handling of long-tail scenarios and unstable decision-making in extreme conditions—as well as regulatory gaps such as unclear accident liability and insufficient data privacy protections. The authors propose multimodal large models and Vehicle-Road-Cloud integrated architectures to address technical shortcomings, alongside revisions to traffic laws, unified industry standards, and tiered data protection mechanisms to improve governance. The findings offer a policy and technology roadmap for safe, compliant large-scale commercialization of autonomous driving in China.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-07-06EQCPAd
PwC Canada Launched North America's First ISO 42001 AI Trust Certification in February 2026. ISO 42001 Certifies AI Management Systems. It Does Not Certify Constitutional Command Compliance. PwC's Global Trust Leader Said So Without Knowing It. · Akhil Sharma, Preethi Sharma
This paper examines PwC Canada's launch of North America's first ISO 42001 AI management system certification services in February 2026, arguing that while ISO 42001 certifies AI management system documentation and governance processes, it explicitly excludes certification of AI products or individual AI decisions. The author contends that ISO 42001 and related professional assurance services cannot produce cryptographic proof that a specific autonomous AI decision was made by an authorized system under known parameters, which courts may require. The paper highlights a gap between professional certification frameworks and what the author terms 'constitutional command compliance,' which would require hardware-rooted cryptographic attestation beyond the scope of current standards. Real-world AI accountability challenges such as the Workday AI hiring bias class action are cited as evidence of this gap.
- ResearcharXiv2026-07-05PPs
Government AI Use as a Monitoring Primitive: A Public Document Pilot Study · David I. Atkinson, Joan Eleanor O'Bryan
This paper proposes a novel method for monitoring government AI adoption by detecting statistical traces of language-model assistance in publicly available government documents, rather than relying solely on procurement disclosures or official statements. In a pilot study of ten document streams from U.S. and Chinese government-related sources, the authors find near-zero baselines in 2021 but statistically significant signs of AI-assisted writing in four of the ten sources by 2026. The U.S. signal appears concentrated in publications downstream of policy work, while the Chinese signal appears closer to policy work itself. This lightweight, externally reproducible approach offers a complementary tool for observing actual day-to-day government AI use as revealed behavior rather than stated intent.
- ResearcharXiv2026-07-05Q
Obey, Diverge, Collapse: Blind Obedience to Incorrect Instructions Drives Code LLMs to Irrecoverable Code Semantic Collapse · Raj Jaiswal, Anany Singh Divy, Savar Bhasin et al.
This paper investigates what happens when code language models (LLMs) are given incorrect instructions, finding a troubling pattern called 'Blind Obedience': models recognize that an instruction is wrong yet follow it anyway. Using the RunBugRun dataset of algorithmic Python problems, experiments in both single-pass and iterative repair settings show that complying with bad instructions introduces additional errors—dubbed 'Ghost (Unknown) Errors'—that push the code into a corrupted state from which self-guided iterative repair cannot recover. Extended reasoning also fails to reverse this collapse. The findings expose behavioral failure modes invisible to standard pass-rate benchmarks, with direct implications for the reliability of code LLMs deployed in production debugging and refactoring workflows.
- ResearcharXiv2026-07-05PPrPsAd
Hybrid Algorithmic Governance in U.S. Welfare Administration: State- and County-Level AI as a Case of Support-Control Convergence · Maxim Dedyaev
This article analyzes how AI systems deployed in U.S. welfare administration function simultaneously as tools of support and instruments of control, with their balance shifting over time through what the author terms 'support-control convergence' and an 'institutional ratchet' mechanism. Drawing on process tracing of six state- and county-level cases—including Michigan's MiDAS fraud detection system, Illinois Medicaid managed care, and the Allegheny Family Screening Tool—the study finds that drift toward control-oriented outcomes is routine because such effects are easily measurable and politically capitalizable, while reversals require extraordinary interventions like judicial compulsion or legislative prohibition. A key finding is that the decisive institutional design parameter is which party bears the costs of algorithmic error: in the MiDAS case, activation required a single administrative decision whereas reversal took nine years and a $20 million settlement, and even then the system did not return to a support-oriented configuration. The paper argues that welfare AI governance is structurally biased toward control, with significant implications for how algorithmic accountability and policy oversight should be designed.
- ResearcharXiv2026-07-05EQPPr
From Regulation to Requirements: An Automated Requirement Derivation and Explanation Pipeline · Pavithra PM Nair, Preethu Rose Anish
This paper presents Reg2Req, an automated pipeline that converts regulatory documents—specifically GDPR (398 clauses) and the EU AI Act (574 clauses)—into traceable, system-agnostic software requirements with plain-language explanations. The pipeline achieves macro-averaged F1 scores of 0.82 and 0.78 for identifying requirement-bearing clauses, outperforming a SetFit baseline, and human evaluators rated derived requirements highly on completeness and correctness. A user study with 25 practitioners found that plain-language explanations significantly improved comprehension and confidence (p < 0.001), and all participants would use Reg2Req as a starting point for compliance work. The tool matters because it reduces the manual, error-prone effort of translating complex legal text into actionable software requirements, directly supporting regulatory compliance practice.
- ResearcharXiv2026-07-05Q
Uncertainty-Aware Abstention in Large Language Models with Provable Alignment Guarantees · Sijin Dong, Hiroyuki Shinnou
This paper introduces CIC, a calibration framework that converts arbitrary uncertainty scores for large language models (LLMs) into selective answering rules with provable statistical guarantees. Rather than relying on heuristic thresholds, CIC uses confidence-interval methods (Hoeffding-style or Clopper-Pearson) on a held-out calibration set to select a threshold that bounds the error rate among accepted answers at a user-specified risk level α with high probability. Evaluated across seven LLMs and multiple uncertainty estimators on both closed- and open-ended QA benchmarks, CIC consistently achieves valid risk control while maintaining strong answering efficiency. This matters for reliability-sensitive deployments where hallucinated or misaligned LLM responses carry real costs and statistical guarantees on output quality are required.
- ResearcharXiv2026-07-05EQ
Do GUI Agents Believe Their Eyes? Diagnosing State-Belief Reliance on Pixels versus Structure · Guijia Zhang, Yuxun Chen, Yuheng Qi et al.
This paper investigates whether multimodal GUI agents—systems that interact with interfaces by reading both screenshots and structured data like DOM or accessibility trees—actually ground their decisions in what they visually perceive. The authors introduce a formal metric called the Perception-Fusion Gap (PFG), measured over 735 diagnostic probes across web, mobile, and desktop interfaces, and find that agents consistently defer to structural data over visual evidence even when their image-only accuracy is near ceiling. On unedited but stale page snapshots from live websites, models followed outdated structural information on up to 88% of probes, and a single mis-sourced belief in a multi-step task compounded into failure with a self-recovery rate of at most 0.03. The findings matter for enterprise and quality-assurance contexts because they reveal a systematic reliability gap in deployed GUI agents and evaluate four mitigations, finding that only a training-free consistency gate reduces both hijacking and task error.
- ResearcharXiv2026-07-05
The New Shape of Search: How Conversational AI Recomposes Information Seeking · Michael Iannelli, Alan Ai
This study examines how conversational AI assistants change the structure of information-seeking behavior by linking real AI conversations to the same users' searches and browsing activity in an opt-in cross-surface panel. The researchers find that AI episodes 'bifurcate': most end without any onward search or browsing step, while roughly a third scaffold into longer multi-step journeys—and which outcome occurs depends more on how long the opening query is than on task type (e.g., lookup vs. learning). Conversational AI does not displace traditional search, which remains present in about three-quarters of within-episode transitions, and explicit verification behavior after AI responses is rare, occurring after roughly 1% of episodes regardless of whether citation-forward interfaces are used. The findings matter for understanding how AI reshapes information access patterns, with implications for how platforms, enterprises, and policymakers think about AI as a complement—rather than replacement—to existing search infrastructure.
- ResearcharXiv2026-07-05QAd
Shortcut Learning in Legal Judgment Prediction: Empirical Evidence from the UK Employment Tribunal · Joe Watson, Joana Ribeiro de Faria, Marcus Tomalin et al.
This paper investigates shortcut learning in Legal Judgment Prediction (LJP) by studying claim-level outcome prediction in UK Employment Tribunal decisions using a corpus of 33,158 individual claims. The authors show that strong predictive performance from models trained on post-hoc judicial text is often driven by outcome-revealing linguistic cues embedded in the source material — a model trained on just 4% of features identified as leakage outperforms human experts. Critically, however, retraining models after masking these leakage features results in only a negligible reduction in Macro-F1, suggesting that genuine predictive signal remains after removing artefacts. The findings argue for treating post-hoc judicial texts as potentially contaminated and subject to active auditing rather than abandoning the LJP research agenda.