News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
Safe to Hire: Predicting Recidivism Risk for Job Candidates with Criminal Records
Elizabeth C. Chase, Shawn D. Bushway, Bethany Saunders-Medina et al.
Statistics and Public Policy · 2026-07-20
This paper investigates whether statistical models can improve employment decision-making for job candidates with criminal records, which currently suffers from opacity, inconsistency, and racial disparities. Using multi-state data from the Criminal Justice Administrative Records System (CJARS) spanning 1992–2021, the authors build a Cox proportional hazards model to predict recidivism risk and evaluate its predictive performance, fairness, and generalizability. They find their model outperforms some existing approaches, but note persistent challenges in fairness and practical usability. The work is relevant to hiring policy and equity for Black and Hispanic Americans disproportionately affected by current practices.
- Workforce
- AI policy
Research
AI-Driven Innovation and Optimization of Packaging Design
Y Zhang
Information Resources Management Journal · 2026-07-20
This study proposes and empirically tests a four-component AI-assisted packaging design model across 120 comparative projects. Using stratified sampling and regression analysis, the authors find that AI integration reduced development cycle times by up to 68.3%, cut costs by 36.9%, boosted ROI by 88%, doubled creative concepts, and halved revision rounds, with small and medium-sized enterprises (SMEs) benefiting the most. The findings highlight AI's scalable value when combined with human collaboration, while also identifying gaps in cultural modeling, sustainability tooling, and IP governance frameworks. The results matter for enterprises—especially smaller firms—looking to modernize design workflows and remain competitive in personalization and e-commerce contexts.
- Enterprise
Research
Code, capital, and clusters: understanding firm performance in the UK AI economy
Waqar Muhammad Ashraf, Diane Coyle, Ramit Debnath
npj Artificial Intelligence · 2026-07-20
This study analyzes a comprehensive dataset of UK AI firms from 2000 to 2024, finding that 41.3% of entities are concentrated in London and that firm size and AI specialisation intensity are the primary drivers of revenue performance. Local socio-economic factors—including qualification rates, population density, and employment levels—also contribute meaningfully, underscoring AI growth's dependence on regional ecosystems. Forecasting models project approximately 4,651 total entities by 2030 alongside a rising dissolution ratio, signaling sector consolidation. The authors argue these findings justify place-sensitive policy interventions to cultivate regional AI capabilities beyond London and to balance scaling support with deeper technical specialisation.
- Enterprise
- AI policy
Research
(Over)Reliance on Test Agents in AI-Assisted Software Testing
Eduard Paul Enoiu
arXiv (Cornell University) · 2026-07-20
This paper argues that AI-based test agents introduce two interrelated risks in software testing: an agency problem, where engineers may cede cognitive control over test design decisions, and an assurance problem, where generated testing artifacts may be accepted as valid evidence without sufficient scrutiny. Drawing on three theoretical lenses—software testing as cognitive problem-solving, test agents as adaptively autonomous entities, and test design argumentation—the authors develop a framework for identifying and measuring overreliance in test agent workflows. The work aims to help organizations capture the speed and scalability benefits of AI-assisted testing without eroding engineer judgment or the evidentiary value of test outputs.
- Quality assurance
- Workforce
Research
Is Archaeological Data Labor Sustainable? Only through Ethical and Policy Reforms, Open Data Emphases on Public Engagement, and Indigenous Data Sovereignty
Joshua Wells, Neha Gupta, Kelsey Noack Myers
Advances in Archaeological Practice · 2026-07-20
This paper argues that the rapid adoption of information and communication technologies in US archaeology has made archaeological data labor unsustainable, outpacing the development of professional and ethical frameworks. The authors critically review how traditional Western legal and intellectual property models shape data standards in ways that alienate the public, disregard Indigenous Data Sovereignty, and enable extractive commodification of heritage by AI enterprises. They call for reforms centered on public goods and Indigenous rights, drawing on governance frameworks such as CARE, FAIR, and the US OPEN Government Data Act to guide sustainable and accountable data labor practices across the discipline.
- Workforce
- AI policy
Research
Gratified use of AI moderating the pathway from sense of empowerment towards infusion use
Yingnan Shi, Yifan Zhong, Chenxiao Wang
Humanities and Social Sciences Communications · 2026-07-20
This study of 515 employees using generative AI tools (ChatGPT, Midjourney, DeepSeek, DALL·E 3) examines how psychological empowerment from AI translates into deeper workplace use. Using a moderated-mediation structural equation model, the authors find that AI-enabled empowerment positively drives integrative and extended use, which in turn promote emergent use, but that utilitarian motives amplify these pathways while hedonic (enjoyment-oriented) motives can weaken them. The findings suggest organizations should frame AI empowerment initiatives around utility-focused goals and watch for 'hedonic drift' that may undermine systematic feature adoption.
- Workforce
- Enterprise
Research
Algorithmic management and the future of work
Ivan Žegarac
Repository of the University of Primorsk (University of Primorska) · 2026-07-20
This thesis examines how algorithmic management affects working conditions across the EU, with a focus on Slovenia, using Eurofound's 2024 European Working Conditions Survey microdata covering 36,644 respondents from 35 countries. Workers exposed to algorithmic management reported significantly higher stress and lower wellbeing than those not exposed, with stronger effects observed in Slovenia than the EU average. Notably, the study found no reduction in worker autonomy, contrasting with prior empirical literature on the topic. The findings highlight measurable workforce wellbeing costs associated with algorithmic management in contemporary labor markets.
- Workforce
- AI policy
Research
Evolving Universal Requirements for AI-based Medical Devices: An International Comparison of Regulatory Approaches
Hirokazu Arima, Shingo Kano
IntechOpen eBooks · 2026-07-20
This study compares how Japan, the USA, and the EU regulate AI-based medical devices across multiple time points, using an integrated analytical framework drawn from regulatory guidance and ethical, legal, and social literature. It finds that coverage of evaluation requirements has improved in all three regions, but each has developed distinct policy priorities: Japan emphasizes postmarket surveillance and data governance; the USA focuses on performance changes and risk-based change management including predetermined change control plans; and the EU stresses governance, conformity assessment, and fundamental rights protection. Cross-cutting issues such as human–machine teaming and fairness remain unevenly addressed across regions. The paper offers practical guidance for developers on combining localization strategies with universal, total product lifecycle-based approaches to meet diverging international requirements.
- Certifications
- AI policy
Research
Legal Regulation of the Application of Formative Assessment and the Use of Artificial Intelligence in Schools in Lithuania: Insights from Thematic Analysis
Rūta Gedminienė, Julija Melnikova
Acta Paedagogica Vilnensia · 2026-07-20
This study analyzes 15 Lithuanian legal documents governing formative assessment and AI use in general education schools, finding that while formative assessment is well-established in law as a continuous, feedback-driven process, its regulation lacks full coherence. AI use in schools is governed in a fragmented manner, with no clear definitions, implementation guidelines, or explicit links to formative assessment practices. The authors conclude that existing legal frameworks inadequately address the challenges of integrating AI into education and call for targeted regulatory improvements.
- AI policy
Research
Artificial Intelligence Adoption and Ethical Governance in Australian Insurance: Evidence from Web-Based Content Analysis
Matias A. Morales Armijo, Jinhui Zhang, Yanlin Shi
Risks · 2026-07-20
Analyzing 156 AI-related web pages from Australian insurers, this study finds that 58% of companies make no reference to any AI ethics principles, revealing a significant gap between AI adoption and ethical governance. Operational themes dominate public-facing communications, while accountability and contestability principles are notably underrepresented compared to wellbeing and privacy. The authors argue that stronger, more consistent public commitments to responsible AI could enhance stakeholder trust and support sustainable innovation in the insurance sector. This is described as one of the first empirical assessments of publicly articulated ethical AI governance in Australian insurance.
- AI policy
- Enterprise
Research
Digital Transformation and Structural Challenges in the Moroccan Construction Sector: A Qualitative Study of Construction Professionals’ Perspectives on Building Information Modeling and Artificial Intelligence Adoption
Yasser Tajmout, Aniss Moumen
Buildings · 2026-07-20
This qualitative study examines why building information modeling (BIM) and AI adoption remains limited in Morocco's construction sector, drawing on semi-structured interviews with eight construction professionals. Findings show BIM implementation is still predominantly at Level 1, held back not by technology gaps but by economic, institutional, legal, and capacity barriers, even as practitioners report measurable benefits like reduced design conflicts through clash detection. AI is perceived as augmenting rather than replacing professional expertise. The study recommends a national BIM mandate, a national BIM competence center, and stronger regulatory frameworks to accelerate digital transition.
- Workforce
- AI policy
Research
A Diagnostic Framework for Staged AI Adoption in Batik Motif Recognition: Integrating CNN Evidence and Implementation Readiness
Irwan Sembriring
Journal of Applied Data Sciences · 2026-07-20
This study introduces a dual-layer diagnostic framework for deciding when and how to deploy AI in batik motif recognition, combining technical model performance with organizational readiness assessments. Three CNN transfer-learning models (VGGNet-16, ResNet50, MobileNetV2) were tested on 983 batik images; the best performer, ResNet50, achieved only 45% accuracy and a macro F1-score of 0.40, indicating low technical readiness. Meanwhile, 173 IT practitioners reported high perceived implementation readiness at 78.7% agreement, revealing a 'readiness asymmetry' where organizational support exists even though the AI model is technically immature. The framework provides a structured, risk-aware basis for staged AI adoption decisions relevant to cultural heritage preservation, enterprise governance, and quality assurance.
- Enterprise
- Quality assurance
- AI policy
Research
Guest editorial: The next wave of innovation. Human, enterprise, and artificial intelligence united for impactful change
Nicola Cucari, Francesco Schiavone, Biagio Palese
European Journal of Innovation Management · 2026-07-20
This guest editorial introduces a special issue of 11 empirical studies that collectively argue the next wave of AI-driven innovation depends on the quality of human-AI partnership rather than machine autonomy alone. Spanning individual, enterprise, and technology levels of analysis, the studies find that AI benefits are neither automatic nor uniform—they require calibrated human engagement, organizational readiness, strategic alliances, ethical governance, and national context sensitivity. The editorial synthesizes findings across manufacturing, healthcare, financial services, public administration, startups, and nonprofits to advance an 'Innovation 5.0' thesis in which humans, extended by AI, collaborate to produce change that is both efficient and socially meaningful. Managers are advised to dose AI assistance, invest in training, pursue alliances over isolated capability announcements, and govern AI with attention to ethics and shifting organizational identity.
- Enterprise
- Workforce
- AI policy
- Quality assurance
Research
Designing a safe generative AI-powered chat assistance system for mental health using a bounded generative framework
Udochukwu John-bright Okike, Amarachi Tyndale Eche, Chimeremeze Victor Amaechi et al.
Discover Artificial Intelligence · 2026-07-20
This paper proposes the Bounded Generative Framework (BGF), an architectural pattern for deploying large language models in mental health chat applications while maintaining protocol fidelity, crisis safety, and auditability. Tested in a web-based Emotional Freedom Techniques app called Tapaway, the framework uses a two-layer pipeline with a 10-state therapeutic protocol controller and structured tool calls to constrain LLM behavior. Across a synthetic benchmark, GPT-4o achieved the strongest protocol fidelity and crisis-intent detection, with no false negatives across crisis-positive test cases, and a convenience sample of university students showed large within-session distress reductions (Cohen's d = 1.85), though the authors caution these are feasibility signals only. The work argues that safety assurance in therapeutic AI should shift from model-only alignment toward application-level constraints, while noting controlled trials are needed to establish clinical equivalence to human-delivered therapy.
- Quality assurance
- AI policy
- Certifications
Research
Controlled, Not Correct A Computer Software Assurance framework for domain-specific regulatory AI agents
Rudolf Wagner
arXiv · 2026-07-20
This paper argues that AI accuracy is the wrong metric for evaluating regulatory AI tools, and that what matters instead is whether the tool's output is subject to adequate control proportionate to its intended use and risk. Drawing on existing regulatory instruments (FDA CSA draft guidance 2022, ISO 13485:2016, MDCG 2019-11, EU AI Act Article 10), the authors derive a three-layer human-in-the-loop control architecture and quantify it using a defect-escape model. Their key finding is that a model operating at 93.3% accuracy (66,800 DPMO) placed under three mature control layers achieves a residual defect rate of 401 DPMO—equivalent to a 99.960% process yield—demonstrating that the number of control layers, not model accuracy, is the dominant factor in regulatory AI performance.
- Quality assurance
- Certifications
- AI policy
- Enterprise
Research
Reconfiguring Value Capture: The Impact of Artificial Intelligence on Income Distribution in Global Value Chains
Yu Zheng, Shibao Xu, Yixin Dai
Emerging Markets Finance and Trade · 2026-07-20
Using panel data from 35 economies over 2000–2014, this paper finds that AI development significantly increases a country's share of global value chain (GVC) income, with effects being strongest in developed economies, larger domestic markets, and countries positioned higher in GVCs. AI primarily boosts value created by capital and high-skilled labor, confirming its capital-augmenting and skill-biased nature. The study identifies three transmission channels: factor structure optimization, skill-biased technological progress, and intermediate goods inward orientation. These findings have direct implications for how countries design AI-related industrial and trade policies to capture greater gains from global production networks.
- Workforce
- Enterprise
- AI policy
Research
Governing AI reasoning to mitigate assumption injection in assurance workflows
Shao-Fang Wen
Discover Artificial Intelligence · 2026-07-20
This paper identifies a risk called 'assumption injection,' where AI tools silently introduce or modify unverified contextual assumptions when generating security assurance artifacts from incomplete system descriptions. To mitigate this, the authors propose a governed context evolution mechanism that treats unresolved assumptions as explicit 'context gaps,' allowing AI to suggest advisory refinements but restricting authoritative changes to human-controlled, provenance-tracked decision records with replay-based validation. A proof-of-concept implementation and case study demonstrate that this approach supports auditability, traceability, and human oversight in AI-assisted assurance workflows. The work is relevant to ensuring AI outputs in high-stakes assurance processes remain trustworthy and controllable.
- Quality assurance
- Certifications
- AI policy
Research
Empirical Grounding Improves the Realism of LLM Agents Simulating Human Behavior During Disruptions
Chen Xia, Zexi Kuang, Yuqing Hu
arXiv · 2026-07-19
This paper tests whether embedding real demographic and behavioral data into large language model (LLM) agents improves how realistically those agents simulate human behavior during disasters. The researchers built an empirically grounded framework using U.S. Census demographic profiles, national time-use survey data, and urban spatial context, then validated it against an independent household survey collected during the July 2024 Philadelphia heatwave. The grounded model dramatically outperformed an ungrounded baseline, raising mean correlation with observed activity profiles from 0.349 to 0.836 under heatwave conditions and capturing 46.4% of observed behavioral response amplitude versus 20.6% for the baseline. The findings matter for disaster and infrastructure-disruption planning, showing that empirical grounding is a concrete method for making LLM-based population simulations more statistically credible while also exposing remaining gaps in modeling human adaptation.
- AI policy
- Enterprise
Research
Abliteration Is Not a Scalpel: Off-Target Effects of Refusal Removal on Decision Disposition Across Model Families
Aleksander Fafuła
arXiv · 2026-07-19
This paper investigates 'abliteration,' a technique that removes refusal behavior from open-weight AI models by deleting a refusal direction from model weights, and finds it produces unintended side effects on decision-making beyond just removing content restrictions. Using 21,600 investment decisions on Warsaw Stock Exchange equities as a refusal-free probe task, the authors show that abliterated versions of two large language model families (Gemma-4-26B and Qwen3-30B) become systematically more optimistic, produce longer justifications, and use fewer uncertainty expressions compared to their base counterparts. Critically, the effect on expressed confidence reverses direction between model families, demonstrating that the same weight surgery can produce opposite behavioral shifts depending on the model. The findings warn that deploying 'uncensored' community-modified models as autonomous agents means deploying a measurably different decision-maker, not simply the base model with safety filters removed.
- Enterprise
- Quality assurance
- AI policy
Research
The Librarian Who Refused to Code: Model-Dependent Identity Enactment in LLM Code Generation
Shayell Aharon Salomon, Noam Israel, Ido Safruti et al.
arXiv · 2026-07-19
This study evaluates how biographical personas in system prompts affect large language model (LLM) code generation across controlled, pre-registered conditions using 480 completions from two frontier models (Claude Opus and GPT-5.5). Results show that persona effects are strongly model-dependent: on Claude Opus, a minimalist engineer persona reduced output length by 30% without improving correctness, while a research-librarian persona triggered in-character refusals in 55 of 60 responses and 12 genuine no-code outputs, dropping mean correctness from 0.92 to 0.67. GPT-5.5 showed neither behavior. The findings suggest that personas function as model-dependent behavioral-policy biases rather than universal quality improvements, with significant implications for how prompt engineering is evaluated and deployed in software development workflows.
- Enterprise
- Quality assurance
Research
Team DACTYL at PAN 2026: Bayesian Data Mixing and Empirical X-risk Minimization for AI-text Detection
Shantanu Thorat
arXiv · 2026-07-19
This paper presents Team DACTYL's approach to detecting AI-generated text at the PAN 2026 competition, tackling the common problem that AI-text classifiers perform well on familiar data but poorly on out-of-distribution (OOD) text. The team addresses this by using BERT-tiny models with Bayesian classification heads to curate a consolidated training set from three datasets, then training DeBERTa-V3-large and ModernBERT-large classifiers via empirical X-risk minimization, with an additional MCGrad model for calibration. The MCGrad model achieves a mean score of 0.974 across five metrics (AUROC, F1, C@1, Brier score, and F0.5u) on the PAN 2026 test set, ranking second on the leaderboard, demonstrating that careful dataset curation can significantly improve OOD generalization. These findings are relevant to quality assurance and policy contexts where reliable, generalizable AI-text detection is needed to authenticate or flag machine-generated content.
- Quality assurance
- AI policy
Research
Auditing Differential Visibility of Political Content on TikTok
Hazem Ibrahim, Tewoflos Girmay
arXiv · 2026-07-19
This study audits TikTok's alleged political content shadow-banning by analyzing over 556,000 hourly view observations across 2,753 videos from 67 accounts spanning pro and anti positions on U.S. immigration, Trump coverage, and Israel/Palestine. While pooled video-hour data appeared to show large, statistically significant reach gaps, account-level analysis—the appropriate independent unit—found no evidence of moderate-to-large suppression on any topic after multiple-comparison correction. The apparent gap was explained by methodological artifacts: pseudoreplication from autocorrelated video-hours treated as independent observations, and confounding factors such as account size and language differences. The study concludes that what has been interpreted as shadow-banning is better explained by audience engagement patterns, and offers guidance on what credible platform visibility audits require.
- AI policy
- Enterprise
Research
Agentic ERP: Multi-Agent Large Language Model Architecture for Autonomous Enterprise Resource Planning
Zhihao Liu, Tianyu Wang, Xi Vincent Wang et al.
arXiv · 2026-07-19
Agentic ERP proposes a multi-agent large language model architecture that transforms Enterprise Resource Planning systems from passive transaction recorders into active decision-makers capable of executing end-to-end business workflows autonomously. The architecture uses role-aligned LLM agents, a graph-based Planner–Executor–Reflector–Responder orchestrator, and a risk-tiered human-in-the-loop harness to handle cross-functional operational decisions that classical rule-based automation cannot manage. Evaluation across scenario-based tasks, a comparison of six orchestration paradigms, and a 365-day simulation shows the proposed system significantly outperforms rule-based baselines—sustaining a simulated year of operation with zero stockouts while the rule-based baseline accumulated hundreds under the same demand conditions. This work provides a reference architecture and evaluation protocol for autonomous ERP operation, with direct implications for how enterprises automate complex operational decision-making with human oversight.
- Enterprise
- Workforce
- Quality assurance
Research
DRNOISE: Benchmarking Deep Research Agents in Misleading Evidence Environments
Jun Nie, Zhiqin Yang, Zhenheng Tang et al.
arXiv · 2026-07-19
DRNOISE introduces a 100-task benchmark designed to test how well deep research agents handle misleading evidence on the open web. Each task includes a gold answer supported by indirect evidence chains, plus one plausible but false document offering a conflicting shortcut answer. Across agents that perform well under clean conditions, adding this single misleading document causes accuracy drops of 66–88 percentage points, with 'verification inertia'—stopping before fully reconciling evidence—identified as the dominant failure mode. The findings highlight that reliable AI-driven research requires active reconciliation of direct claims against record-level evidence, not just retrieval and citation.
- Quality assurance
- Enterprise
Research
Lookahead Branching for Neural Network Verification
Liam Davis, Duo Zhou, Huan Zhang et al.
arXiv · 2026-07-19
This paper investigates lookahead branching strategies for neural network verification using branch-and-bound methods. The authors present a general framework for integrating lookahead into any branch-and-bound verifier, showing that the existing FSB heuristic is a special case of this approach. By also generating additional lemmas during lookahead, the method improves branching decisions and accelerates verification. Experiments with two representative verifiers (Marabou and α-β-CROWN) show consistent speedups and up to 57% more solved verification instances, which matters for ensuring the reliability and safety of neural network systems.
- Quality assurance
- Certifications