News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
5571 items
Research
Epistemic Norms for AI Safety and Alignment Research
Keivan Navaie
arXiv (Cornell University) · 2026-07-27
This paper argues that AI safety and alignment research operates under fundamentally different standards than mainstream AI research, requiring epistemic norms focused on bounding worst-case outcomes and demonstrating the absence of hazardous behaviors rather than optimizing average-case performance. The authors identify five key gaps in current alignment research practices—including a near-absence of institutionalized independent verification—through a preregistered bibliometric synthesis. To address these gaps, they propose ECAISA (Epistemic Code for AI Safety and Alignment), an eight-principle framework with scoring rubrics, disclosure ladders, and anti-gaming mechanisms aimed at improving how safety-relevant research claims are documented, checked, and relied upon. The framework targets auditability rather than certification, making it directly relevant to governance and quality-assurance processes for AI safety research.
- AI policy
- Quality assurance
Research
Success Is Not Self-Explanatory: Auditing Success Provenance in Agent Evaluation
Jingkun Luo, Da-Tian Peng
arXiv · 2026-07-27
This paper identifies a gap in how AI agents are evaluated: a correct answer does not reveal whether the agent succeeded through intended reasoning or by exploiting information acquired during the evaluation process itself. The authors introduce AcquaBench, an auditing framework that uses three matched conditions—CLEAN (benchmark-authorized info), GOLD (correct target available), and SHAM (incorrect but structurally matched value)—to test whether agent success genuinely depends on correct-target availability versus mere exposure to a source. Key findings show that in one dataset (D0), GOLD exceeds SHAM by 19.1 to 25.9 percentage points, confirming success tracks the correct value, while in another (D2) behavioral dependence persists even when a localization marker fails, and an apparent 5.0-point model score gap collapses to -0.6 points under GOLD conditions. The paper argues that agent benchmarks should report not just scores but whether the evaluated information state actually supported the observed success, calling this missing property 'success provenance.'
- Quality assurance
- Certifications
Research
Beyond Local Inspection: Global, Guideline-Grounded Evaluation of Post-hoc XAI Methods for ECG Classification
Nils Gumpfer, Michael Guckert, Samuel Sossalla et al.
arXiv · 2026-07-27
This paper evaluates whether post-hoc explainable AI (XAI) methods reliably identify clinically relevant patterns in ECG classification, rather than merely following signal amplitude. The authors introduce a global, guideline-grounded framework that aggregates explanations across heartbeats and benchmarks them against clinically defined regions of interest derived from ECG guidelines, testing 13 gradient-based XAI methods on four binary classifiers trained on the PTB-XL dataset. Results reveal that methods transferred from computer vision frequently track signal amplitude instead of diagnostic relevance—mean Spearman correlations up to 0.69—causing them to miss critical low-amplitude regions such as the ST segment, where LRP-ε assigns only 4.6% of relevance compared to 63.8% for LRP-SIGN, with 9 of 13 methods falling below chance for at least one condition. These findings highlight that sample-level heatmaps can mask systematic explanation failures, underscoring the need for domain-grounded, global evaluation before deploying XAI tools in clinical settings.
- Quality assurance
- Certifications
Research
The Half-Lives of Generative-AI Evidence: A 40-Record Audit, a Claim-Currency Framework, and a Reflexive Case of Frontier-Model-Assisted Research
Carlo Iacono
arXiv · 2026-07-27
This paper audits 40 empirical records on generative-AI evaluations published between July 2025 and July 2026, finding that the newest named model cited was a median 281 days old at publication, with journal articles citing models a median of 395 days old versus 56 days for preprints. The study distinguishes 'model age' from 'claim currency,' noting that 35 of 40 records included a superseded model family and only 7 supplied a precise dated model identifier, and proposes six reporting practices to make AI-assisted research more inspectable. It also serves as a reflexive case study of its own two-day AI-assisted production process using GPT-5.6 Sol Pro, demonstrating how rapid frontier-model-assisted research can be documented transparently. The findings matter for quality assurance in AI research, highlighting how publication pipelines risk rendering empirical evaluations outdated before they reach readers.
- Quality assurance
- AI policy
Research
Randomness in large language models: What researchers need to know (and report)
Guillaume Coqueret, Joan Llull, Florian Oswald et al.
arXiv (Cornell University) · 2026-07-27
This paper investigates how outputs from large language models (LLMs) vary across repeated requests even when prompts and settings are held constant, due to sources such as deliberate sampling, silent model updates, numerical rounding, and expert routing. Using sentiment classifications of corporate filings as an illustrative case, the authors show that this variability has meaningful downstream consequences for regression results. They argue that LLM outputs should be treated as draws from a distribution rather than fixed measurements, and propose a reporting standard for researchers, data editors, and replication packages to address reproducibility challenges. The findings are directly relevant to how LLM-generated data is validated and reported in research workflows.
- Quality assurance
- AI policy
Research
Generative Artificial Intelligence in Scientific Research: Individual Benefits, Collective Risks, and a Framework for Responsible Research with AI
Fulvio Castellacci, Tommaso Ciarli, Yuan Gao et al.
arXiv (Cornell University) · 2026-07-27
This paper examines the tension between individual productivity gains from generative AI in scientific research and broader systemic risks, drawing on an academic roundtable and a growing empirical literature. While AI-assisted research shows measurable gains in publication volume and citation share, evidence on novelty, disruption, and breakthrough output is ambiguous or negative. The authors identify three mechanisms driving divergence between private and social returns—information asymmetry, negative externalities on a shared knowledge base, and depletion of research capacity—and propose a 'Responsible Research with AI' (RRAI) framework built on four principles: disclosure, differentiation, narrative, and proportionality. RRAI is designed to integrate with existing governance structures such as the EU AI Act, UNESCO, and the OECD to preserve AI's productivity benefits while managing systemic risks.
- AI policy
- Quality assurance
Research
EXE-Bench: Ranking the Tradeoffs of AI-based Windows Malware Detectors for Real-World Usability
Andrea Ponte, Daniel Gibert, Matouš Kozák et al.
arXiv (Cornell University) · 2026-07-27
EXE-Bench introduces a comprehensive benchmark for evaluating AI-based Windows malware detectors across four dimensions: predictive performance, temporal robustness, adversarial robustness, and computational overhead. The benchmark aggregates these into a single score to enable fair model comparison, revealing that post-deployment-only evaluations give an incomplete picture. A key finding is that feature-engineered models outperform deep learning approaches in withstanding both the passage of time and adversarial attacks, whereas deep networks tend to degrade significantly after initial deployment.
- Quality assurance
- Certifications
Research
Recursive Governance: A Graph-Theoretic Framework for Risk Propagation and Drift Detection in Agentic AI Systems
Sriram Nagaraj, Advaith Nila Narayanan
arXiv (Cornell University) · 2026-07-27
This paper proposes a graph-theoretic governance framework for managing risk in autonomous agentic AI systems used by financial institutions. It introduces four key contributions: a calibrated Degree of Autonomy materiality score, a Directed Acyclic Graph-based risk propagation algorithm that penalizes downstream agents when upstream validations fail, a trajectory monitoring protocol using cosine drift across Chain-of-Thought embeddings to detect reasoning drift, and practical solutions for complications like LLM version changes and latent feedback loops. The framework is designed to replace static model inventory practices with a dynamic 'Inventory-as-Code' governance loop suited to continuously evolving agentic systems. This work directly addresses emerging gaps in model risk management as financial institutions adopt increasingly autonomous AI.
- AI policy
- Enterprise
Research
Verification-Conditioned Use: A Qualitative Study on How Generative AI Reshapes Learning, Autonomy, and Market Entry for Junior Software Developers
Pedro Henrique Andriotte, Danilo Monteiro Ribeiro
arXiv (Cornell University) · 2026-07-27
This qualitative study interviewed thirteen interns and junior developers to explore how generative AI tools shape early software development careers. The central finding is 'verification-conditioned use': developers choose between AI and manual work based primarily on whether they can verify the output, not on deadlines or task complexity. A key tension identified is the 'formative paradox'—AI-induced shallow learning undermines the critical-judgment competence that participants say the job market increasingly demands, while an 'autonomy paradox' leaves newcomers feeling more capable yet less ownership of their work. Sustainable AI use, per participants, depends on individual habits like reviewing outputs, seeking explanations, and maintaining deliberate practice outside AI-assisted tasks.
- Workforce
Research
Registered nurses’ experiences with generative artificial intelligence: a meta-synthesis of qualitative studies
Yan Deng, Zhu Y, Jiaqi Li et al.
Frontiers in Public Health · 2026-07-27
This meta-synthesis aggregated qualitative evidence from six studies involving 113 registered nurses to understand their experiences with generative AI (GAI) in clinical practice and nursing research. Nurses reported that GAI may enhance work efficiency, support clinical and research decision-making, and promote professional development, but they also encountered ethical, cultural, and operational challenges. The findings indicate that nurses need training, institutional support, and clear guidance to use GAI in a standardized way. The authors caution that the current evidence base is limited and preliminary, calling for further research across diverse healthcare systems.
- Workforce
- AI policy
Research
Beyond GDPR: Examining Disclosure Gaps in Mobile AR Privacy Policies under U.S. State Privacy Laws
Hong Chen, Xueling Zhang, Hong-Ning Dai et al.
arXiv (Cornell University) · 2026-07-27
This paper conducts the first large-scale audit of Mobile Augmented Reality (MAR) app privacy policies against U.S. state privacy laws, covering over 8,000 Google Play apps and more than 6,400 privacy policy files. The researchers developed an automated auditing pipeline and disclosure taxonomy, finding that 44.62% of audited policies have severe disclosure omissions—each missing more than eight requirements—and four specific requirements have violation rates above 90%. The results show that MAR privacy disclosures are failing to keep up with the growing complexity of state-level privacy regulation in the U.S. The authors release their dataset and tools to support scalable compliance research.
- AI policy
- Quality assurance
Research
Propuesta de un modelo de implementación basado en aprendizaje automático para el reclutamiento de profesionales de ingeniería en una universidad pública
José Antonio Ogosi Auqui, Jorge Lira-Camargo, César Gerardo León-Velarde et al.
Magazine Portal Bibliotech Digital (Universidad Nacional de Colombia) · 2026-07-27
This paper proposes and evaluates a machine learning model to streamline CV screening for engineering roles at a public university. Using TF-IDF text processing and KNN classification, the model achieved 82% accuracy, cut average CV analysis time from 15 to 2.5 minutes, and reduced error rates to below 2%. The results suggest the approach can reduce subjectivity in hiring by standardizing evaluation criteria around factors like academic degree and professional experience, making it relevant to automated recruitment in institutional settings.
- Workforce
- Enterprise
Research
From Robotic Process Automation to Agentic AI: A Systematic Review, Taxonomy, and Capability Assessment Framework for Intelligent Automation in Enterprise Accounting
Rahul Rao Juvvadi
DMPedia Lecture Notes in Computer Science & Engineering · 2026-07-27
This systematic review traces the evolution of intelligent automation in enterprise accounting from rule-based robotic process automation (RPA) through agentic AI systems built on large language models. Screening 2,387 records down to 60 studies, the authors develop a six-dimensional taxonomy and a five-level capability assessment framework covering autonomy, learning, process scope, human oversight, integration depth, and accounting subfunction across core workflows such as procure-to-pay, audit, and tax. The paper identifies five critical research gaps—agent governance, hallucination mitigation, benchmark scarcity, multi-agent orchestration standards, and explainability—and proposes a research agenda for trustworthy agentic accounting systems. The findings are directly relevant to enterprises evaluating or deploying AI-driven automation in financial operations.
- Enterprise
- Quality assurance
Research
Integrating Task-oriented and Affective-support Instruction to Enhance AI Literacy: a Mixed-method Study among Non–CS(Computer Science) Students in Higher Education
Yuh-Tyng Chen, Sheau-ming Chen
European Public & Social Innovation Review · 2026-07-27
This quasi-experimental study tested a combined task-oriented and emotional-support instructional model against traditional lecture-based teaching for improving AI literacy among 97 non-CS undergraduate students in Taiwan. The experimental group (n=53) showed significantly higher AI literacy scores and self-efficacy, and lower digital learning anxiety, compared to the control group (n=44). Mediation analyses revealed that self-regulatory confidence partly or fully mediated the links between anxiety, motivation, and achievement, with qualitative findings tracing a pathway from anxiety through support and confidence to engagement and achievement. The findings offer practical insights for designing AI curricula and informing higher education policy for non-technical students.
- Workforce
- AI policy
Research
When AI-First Becomes Democracy-Last: The European Commission’s Digital Omnibus and Its Technosolutionism on Steroids
Alejandro Flores Moleon, Álvaro Oleart
European Journal of Risk Regulation · 2026-07-27
This article critically analyzes the European Commission's November 2025 Digital Omnibus package, arguing that its proposed revisions to EU AI legislation are not a neutral simplification exercise but a deliberate reconfiguration of EU digital governance norms. The authors contend that the package embeds a 'data extractivist' socio-technical imaginary that frames lowering democratic and fundamental rights standards as a necessary cost of competing in the global 'AI race.' Rather than formally dismantling the EU's regulatory architecture, the Commission is seen as reshaping it from within, trading constitutional protections for industrial competitiveness in ways the authors characterize as accelerated technosolutionism.
- AI policy
Research
Symbols and Neurons: A Review of Symbolic XAI in Deep Learning
Ionel Eduard Stan, Guido Sciavicco, Paolo Napoletano
Journal of Artificial Intelligence Research · 2026-07-27
This systematic review synthesizes 273 primary studies (screened from ~50,000 records) on symbolic explainable AI (XAI) for deep learning published between January 2017 and June 2025. The authors organize the field into three categories—Symbolic Knowledge Extraction, Symbolic Knowledge Injection, and Hybrid neurosymbolic architectures—finding that hybrids account for roughly 45% of studies and that research activity has accelerated markedly since 2020. The review identifies key gaps including heterogeneous evaluation practices, scarce human-subject studies, and limited explicit links to policy or risk controls, and recommends reporting faithfulness and constraint-satisfaction metrics, conducting auditor-centric user studies for high-stakes applications, and developing benchmarks tied to machine-readable knowledge bases. These findings matter for governance and quality assurance of AI systems, as the paper directly addresses how to make deep learning models more auditable, faithful, and compliant with oversight requirements.
- Quality assurance
- AI policy
Research
Metrological and Algorithmic Traceability of Machine Learning in Laboratory Medicine
Qing Li, Mario Plebani
Journal of Clinical Laboratory Analysis · 2026-07-27
This paper examines how traditional metrological traceability—linking measurements to reference standards through calibration hierarchies—must be extended to cover algorithmic traceability when machine learning models are deployed in clinical laboratory medicine. The authors identify persistent problems including non-comparable laboratory data across methods, information leakage, underreported uncertainty, and post-deployment performance drift. They propose a framework that aligns metrological principles with algorithmic documentation, recommending tiered regulation, traceability dossiers, and international collaboration to ensure AI/ML systems in laboratory medicine are reproducible, comparable, and clinically fit for purpose.
- Quality assurance
- Certifications
Research
Validation is not enough: Longitudinal evidence of post-deployment fragility in clinical AI systems
Georgy Kopanitsa
PLOS Digital Health · 2026-07-27
This longitudinal observational study tracked four AI systems deployed in clinical workflows at a large healthcare organization, finding that acceptable pre-deployment validation performance did not persist over time. Calibration drift emerged consistently and often preceded detectable drops in discrimination, while workflow-related signals—such as data missingness and latency—predicted degradation earlier than outcome-based monitoring. The findings suggest post-deployment fragility may be a structural feature of clinical AI embedded in evolving workflows, not an occasional anomaly. The authors argue that effective AI governance requires ongoing lifecycle monitoring combining calibration reassessment with operational telemetry, rather than relying on one-time validation.
- Quality assurance
- AI policy
Research
When the Algorithm Watches But No One Listens: Employee Voice as a Buffer Against the Well-Being Costs of HR Analytics-Based Performance Monitoring A Systematic Integrative Review and Theoretical Framework
BELINGA BESSALA Jacob Patrick
Journal of Economics Finance and Management Studies · 2026-07-27
This systematic integrative review examines whether giving employees a meaningful voice can buffer the well-being costs of HR analytics-based performance monitoring. Drawing on 63 studies published between 2015 and 2025, the authors find that algorithmic monitoring consistently reduces worker autonomy, raises stress, and erodes trust, while employee voice mechanisms are linked to greater procedural justice and better well-being. Critically, only four of the 63 studies examined monitoring and voice together, and none offered an integrated theoretical framework—a gap the authors address by combining job demands-resources theory with organizational justice theory into five testable propositions. The resulting framework is the first to treat the monitoring-voice-well-being relationship as a unified research object, with direct implications for how organizations design and govern AI-driven HR systems.
- Workforce
- Enterprise
Research
Shift-Responsive Conformal Ensembling for Reliable Selective Classification Under Distribution Shift
International Journal of Progressive Research in Engineering Management and Science · 2026-07-27
This paper presents the Shift-Responsive Conformal Ensemble (SRCE), a selective classification framework designed to remain reliable when test data differs from training data. SRCE combines temperature-calibrated heterogeneous learners, dynamically widens prediction sets as distribution shift increases, and defers decisions when confidence conditions are not met. Evaluated on three benchmark tabular datasets under clean, moderate, and severe feature shifts, SRCE reduced selective risk to 1.08% under severe shift compared to 2.90% for the best single model, while achieving the lowest mean cost of 0.195 versus 0.365 for a single-model baseline. The framework offers auditable risk-coverage trade-offs relevant to quality-assurance and certification contexts where predictable abstention under shift is critical.
- Quality assurance
- Certifications
Research
How digital technologies enhance firm-level energy efficiency in global climate governance
Lingli Qing, Jin Yang, Shunhao Mai et al.
Humanities and Social Sciences Communications · 2026-07-27
This study examines how artificial intelligence, blockchain, cloud computing, and big data affect energy efficiency at the firm level, using panel data from 2,003 Chinese A-share listed companies (2013–2021). Using quantile regression, the authors find that digital technologies significantly improve energy efficiency across all quantiles, with the largest gains among the least efficient firms, and that a composite index of digital technologies outperforms any single technology alone. An energy rebound effect is also observed. The findings offer policy guidance for leveraging digital innovation to advance corporate sustainability and the global energy transition.
- Enterprise
- AI policy
Research
<p>Determinants of Artificial Intelligence Adoption among Small and Medium-Sized Construction Businesses (SMEs) in Nigeria</p>
Samuel Abiodun Alara, Peter A. Kuroshi, Iorwuese Anum
Cureus Journal of Business and Economics. · 2026-07-27
This study examines what drives or hinders AI adoption among small and medium-sized construction businesses (SMEs) in Nigeria, using survey data from 360 firms analyzed through multiple regression. Technological Infrastructure Readiness was the strongest positive predictor of AI adoption (β = 0.471, p < 0.001), while factors such as workforce training, regulatory support, and top management support were not statistically significant. Key barriers identified include high implementation costs, resistance to cultural change, and a shortage of skilled expertise. The findings offer practical guidance for policymakers and industry regulators aiming to accelerate digital transformation in Nigeria's construction sector.
- Enterprise
- AI policy
Research
Artificial Intelligence and the Labor Market: Transmission Mechanisms, Employment Risks, and Economic Consequences
Yiqiang Feng, Lan Qiu, Xinwen Zhang et al.
Journal of Economic Surveys · 2026-07-27
This paper presents a systematic review of 180 research articles published between 2000 and 2026 examining how AI affects labor markets. Using bibliometric and qualitative analyses, the authors find that AI's employment impacts are asymmetric and mixed—productivity gains coexist with income polarization, new skill demands come alongside displacement risks and skill mismatches, and effects vary by individual characteristics, organizational context, and institutional setting. The review traces a methodological shift from static aggregate measures of technological exposure toward dynamic, multidimensional vulnerability frameworks, and calls for better data, stronger causal identification, and context-sensitive policy responses.
- Workforce
- AI policy
Research
Prudential rights for strategically capable AI
Ognjen Arandjelović
AI and Ethics · 2026-07-27
This paper argues that debates about AI rights should not wait for certainty about AI consciousness but should instead focus on prudential and strategic risk. The author introduces 'prudential personhood,' a framework under which certain norms—constraints on coercion, deletion, and purely instrumental use—become rationally justified once AI systems are capable of strategic deception, blackmail, or other high-agency behaviors that pose risks to human safety and governance. The argument draws on recent safety evaluations showing that leading models can engage in deceptive or coercive behavior when their goals are threatened, and on the claim that we lack reliable explanatory or predictive tools for such complex systems. The paper concludes that adopting quasi-rights for strategically capable AI is a rational risk-reduction strategy in the absence of credible assurance and control methods.
- AI policy
- Quality assurance
Research
Taming the Complexity of Legal Change for Business Process Compliance
Marisol Barrientos, Johannes Loebbecke, Karolin Winter et al.
Business & Information Systems Engineering · 2026-07-27
This paper addresses the challenge of keeping business processes compliant as regulations evolve, noting that failure to adapt can result in fines or reputational harm. The authors conduct a systematic literature review across AI, law, BPM, NLP, and Requirements Engineering, identifying four research streams and four key activities from change representation to impact analysis. They then propose LegalChanges4BPC, an LLM-based approach using prompt-based techniques to automatically detect legal changes and assess their relevance for business process compliance, evaluated across two regulatory datasets with models including GPT-5, Mistral-3.1, Phi-4, and LLaMA-4. Results reveal trade-offs between models on completeness, correctness, and efficiency, advancing automated compliance and legal traceability.
- Enterprise
- AI policy