News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
How does AI adoption in construction organisations influence employee well-being? The mediating role of job crafting
Sulafa Badi, Abubakr Suliman, Alex Torku et al.
Construction Management and Economics · 2026-05-20
This study investigates how organizational AI adoption shapes the job crafting behaviors of 393 construction project managers and, in turn, affects their work engagement and emotional exhaustion. Using structural equation modeling grounded in the Job Crafting and Job Demands–Resources frameworks, the authors find that AI adoption boosts all dimensions of job crafting, with the strongest effects on structural and challenging job resources. Notably, reducing hindering demands through AI unexpectedly decreased engagement and increased exhaustion, revealing a dual-edged effect of AI in construction project management. The findings suggest that technology adoption must be paired with social support and skill development to genuinely improve employee well-being.
- Workforce
- Enterprise
Research
Toward Cybersecurity Testing and Monitoring of IoT Ecosystems
Stephen Taylor, Martin Gile Jaatun, Aida Omerovic et al.
SN Computer Science · 2026-05-20
This paper presents an extensible cybersecurity framework for IoT ecosystems that unifies testing, runtime monitoring, risk modelling, secure update mechanisms, and auditable evidence management across the full device lifecycle. The architecture integrates component-level techniques—including SBOM generation, network fuzzing, machine learning-based anomaly detection, and access control risk evaluation—with system-level knowledge-based risk modelling to capture threat propagation across interconnected assets. Validated through three industrial use cases in aviation cargo monitoring, smart manufacturing, and telecommunications, the results demonstrate feasibility of combining static analysis, runtime indicators, and dynamic risk assessment to prioritise vulnerabilities contextually and support secure patch deployment in resource-constrained environments. The work is relevant to quality assurance, certification, and policy by advancing lifecycle-integrated, interoperable cybersecurity assurance for complex IoT systems.
- Quality assurance
- Certifications
- AI policy
Research
Predicting affinity and potency of new psychoactive substances at cannabinoid 1 receptor with explainable artificial intelligence
Verena Schöning, Gaia Alluisetti, Katharina Elisabeth Grafinger et al.
Frontiers in Pharmacology · 2026-05-20
This study developed machine learning models to predict how strongly new psychoactive substances (NPS)—specifically synthetic cannabinoid receptor agonists—bind to and activate the cannabinoid 1 receptor (CB1). Using XGBoost and Random Forest algorithms with molecular descriptors and fingerprints, the models achieved recall, precision, and F1 scores above 90% for both binding affinity and functional potency. Explainable AI (SHAP values) revealed that affinity is driven mainly by lipophilicity, while potency depends on a broader set of molecular features. These findings could support drug regulators in identifying and scheduling dangerous NPS before they reach the market.
- AI policy
- Quality assurance
Research
AI-driven innovations in higher learning institutions: a review of pre-service teachers’ pedagogical competence through technology integration
Patrick Clement Mwananyama, Ayubu Ismail Ngao
Cogent Education · 2026-05-20
This literature review examines how AI-driven tools—including adaptive learning systems, VR/AR simulations, generative AI (ChatGPT, DeepSeek), and learning analytics—are being integrated into pre-service teacher education at higher learning institutions. Drawing on studies published between 2015 and 2025 across five major databases, the authors identify both opportunities and challenges, noting gaps such as limited longitudinal research, insufficient cross-cultural studies, and inadequate AI literacy among pre-service teachers. The study proposes a TPACK-AI competency framework with three developmental tiers—foundational literacy, applied pedagogical integration, and professional AI fluency—to guide curriculum design for AI-enhanced teacher preparation, though empirical validation of this framework is still needed. The findings matter for workforce development and certification because they highlight the growing need to systematically build AI competence in future educators before they enter the classroom.
- Workforce
- Certifications
- AI policy
Research
Navigating the data protection landscape in Saudi Arabia: policy effectiveness, barriers, and a strategic roadmap
Alia Mohammed AlSulaimi
Frontiers in Computer Science · 2026-05-20
This study examines data protection legislative frameworks and organizational policies in Saudi Arabia using a mixed-methods approach combining analysis of six public portals and a survey of 200 professionals. It finds foundational policy frameworks exist but significant gaps remain in procedural clarity, contact channel disclosure, and third-party risk management, with statistical evidence of regulatory inconsistencies and technological limitations as systemic barriers. The authors propose a six-pillar strategic roadmap to improve regulatory compliance and build digital trust in alignment with Saudi Vision 2030. The findings are directly relevant to how policy effectiveness and organizational compliance posture can be strengthened in a rapidly digitalizing environment.
- AI policy
- Enterprise
- Quality assurance
Research
Gender Differences in AI Literacy Workshop Outcomes and Deepfake Engagement
Jake Renzella, Christian Bergh, Natasha Banks et al.
arXiv · 2026-05-19
This study examines how gender shapes AI literacy outcomes among Australian secondary students (Years 7, 8, and 10; N=199 pre-workshop, n=136 post-workshop) who attended a one-day AI literacy workshop. Before the workshop, male students reported higher STEM career interest across AI, computer science, and engineering, while female students were more likely to use AI for schoolwork and seek advice from AI tools; males were also significantly more likely to have created or shared deepfake content. After the intervention, both genders improved in AI knowledge, but females showed broader gains including wider conceptual understanding, greater confidence, and meaningful increases in AI and CS career interest that partially narrowed the STEM gender gap. The findings argue for gender-responsive AI curricula — particularly deepfake safety education targeting male students — and show that even single-day workshops can reduce gender gaps in STEM aspirations and AI confidence.
- Workforce
- AI policy
Research
Privacy-by-Design Adaptive Group Assignment for Digital Lifestyle Coaching at Scale
Nariman Mani, Salma Attaranasl
arXiv · 2026-05-19
PRISM-Coach is a privacy-by-design architecture for digital lifestyle coaching that separates each user's data into four bounded views (Identity, Operational, Learning, and Coaching) to prevent PII and health data from leaking into AI pipelines while still enabling personalized peer-group assignment. The system uses a privacy-constrained contextual bandit for adaptive group matching and a human-in-the-loop coaching assistant that generates de-identified summaries without sending raw personal data to external AI services. Evaluated on roughly 2,800 users over three years, the AI-enabled workflow raised daily check-in adherence from 0.35 to 0.74 (versus 0.48 under static grouping) and produced greater average weight loss (5.2 kg versus 3.1 kg) over a matched 19-week window, while 92% of surveyed users reported increased privacy confidence after transparency disclosures. The paper offers a practical blueprint for building adaptive wellness platforms that meet privacy-engineering requirements without sacrificing personalization effectiveness.
- Enterprise
- AI policy
Research
Stage-Audit: Auditable Source-Frontier Discovery for Cross-Wiki Tables
Chen Shen
arXiv · 2026-05-19
Stage-Audit addresses a specific reliability hazard in LLM-curated tables: the model may draw on parametric memory and attach page-level citations post-hoc rather than grounding each row in its actual source. The paper introduces a system with disjoint curator-auditor write rights, a row-level source-citation gate, and a 12-check audit taxonomy to enforce genuine source traceability in cross-wiki table construction. On a 51-instance Seed2Frontier benchmark spanning 15 domains, Stage-Audit raises source-frontier precision from 0.356 to 0.505 (+42% relative) and F1 from 0.334 to 0.451 (+35%) over a vanilla LLM curator, while maintaining explicit per-row provenance. These results demonstrate that auditing policy design—not just LLM capability—is a meaningful lever for improving the verifiability of AI-generated structured content.
- Quality assurance
Research
Does Code Cleanliness Affect Coding Agents? A Controlled Minimal-Pair Study
Priyansh Trivedi, Olivier Schmitt
arXiv · 2026-05-19
This study investigates whether code cleanliness—structural and stylistic quality as measured by static-analysis rule violations and cognitive complexity—affects the performance of autonomous coding agents. Using a controlled minimal-pair protocol, the authors constructed matched repository pairs that differ only in cleanliness and authored 33 tasks across six pairs, evaluated over 660 trials with Claude Code. Results show that code cleanliness does not change task completion (pass) rates, but agents working on cleaner code use 7–8% fewer tokens and revisit files 34% less often. The findings suggest that traditional software maintainability principles remain relevant in AI-driven development, materially shaping computational cost and navigational efficiency even if not raw task success.
- Enterprise
- Quality assurance
Research
Detecting Fluent Optimization-Based Adversarial Prompts via Sequential Entropy Changes
Mohammed Alshaalan, Miguel R. D. Rodrigues
arXiv · 2026-05-19
This paper introduces CPD Online (CPD), a training-free, model-agnostic detector for adversarial prompt attacks against large language models (LLMs). It reframes adversarial suffix detection as an online change-point detection problem, monitoring token-level entropy shifts via a one-sided CUSUM statistic. Tested on 1,012 optimization-based attacks (GCG, AutoDAN, AdvPrompter, BEAST, AutoDAN-HGA) and 1,012 benign prompts across six open-weight models, CPD outperforms windowed-perplexity baselines in F1 on all models, reaches AUROC 0.88 and F1 0.82 on LLaMA-2-7B, and reduces safety-guard API calls by 17–22% in benign-dominated deployments. This matters for enterprise LLM deployments where detecting and filtering adversarial jailbreak attempts efficiently is critical to maintaining system safety without prohibitive computational overhead.
- Enterprise
- Quality assurance
Research
What Are LLMs Doing to Scientific Communication? Measuring Changes in Writing Practices and Reading Experience
Filip Miletić, Neele Falk
arXiv · 2026-05-19
This paper investigates how large language models are changing scientific writing in the NLP domain by analyzing a corpus of over 37,000 ACL Anthology papers (2020–2024) and a synthetic dataset of 3,000 human-written passages paired with LLM-generated revisions. Diachronic lexical analyses show significant shifts in word frequency and usage contexts over time, while stylistic modeling reveals that LLM-modified texts tend to feature specific syntactic constructions, longer and more complex words, and lower lexical diversity. A pilot annotation study with 20 domain experts found that LLM-improved texts are generally rated as more understandable and exciting, yet experts also expressed negative qualitative attitudes toward LLM involvement in writing. The findings highlight the measurable and subjective impacts of AI-assisted writing on scientific communication quality and reader experience.
- Quality assurance
Research
Explainable Wastewater Digital Twins: Adaptive Context-Conditioned Structured Simulators with Self-Falsifying Decision Support
Gary Simethy, Daniel Ortiz Arroyo, Petar Durdevic
arXiv · 2026-05-19
This paper presents CCSS-IX, an explainable digital twin for wastewater treatment plant aeration and dosing setpoints, designed to help operators navigate the daily trade-off between energy efficiency and safety compliance. The simulator uses interpretable locally linear state-space models mixed by a context-aware gating network, and a runtime decision layer applies conformal risk control to either certify, reject, or return a falsifying temporal witness for any proposed operator action. Tested on two full-scale Danish plants and the BSM2 international benchmark, the system achieves within 1.08% RMSE of an unconstrained black-box reference, cuts aggregate two-plant regret by 43.6%, and prevents 93 of 187 false-safe nitrous-oxide approvals—about 4.65x the dyadic baseline (paired McNemar p < 1e-21). The approach demonstrates that safety-critical industrial simulators can combine operator-readable dynamics with finite-sample statistical guarantees, directly supporting safer automated decision-making in wastewater infrastructure.
- Quality assurance
- Enterprise
Research
OriginBlame: Record- and Token-Level Data Provenance for AI Training Datasets
Haolin Xue
arXiv · 2026-05-19
OriginBlame (ob) is a record- and token-level data provenance system that tracks author identity through AI training data pipelines, enabling precise identification of which training records belong to a contributor who has requested removal. Evaluated on over 219,000 Wikipedia pages, the system reduces dataset-level over-deletion from 101x down to 1.3x compared to existing file- or dataset-level approaches, while adding only modest throughput overhead (1.3–19.0% depending on the pipeline). On a 1.7B parameter model, provenance-based forget sets improved machine unlearning performance by 42% over random baselines, making data removal requests far more tractable for model trainers.
- AI policy
- Enterprise
Research
Generative-Evaluative Agreement: A Necessary Validity Criterion for LLM-Enabled Adaptive Assessment
Grandee Lee, Yue Wang, Che Yee Lye et al.
arXiv · 2026-05-19
This paper introduces Generative-Evaluative Agreement (GEA), a validity criterion for detecting a self-referential flaw that arises when a single large language model generates assessment items, simulates student responses, and scores them. In the first direct measurement of GEA on a two-stage adaptive assessment, the authors find the model recovers roughly half the intended variance (r = 0.698) with systematic positive bias, and that GEA is strong (r > 0.7) for syntactically verifiable skills but near zero for design-level skills. Low-skill overestimation near routing thresholds further inflates scores where accurate measurement matters most. The authors argue that granular, skill-decomposed rubrics are the primary mechanism for improving GEA, with complementary mitigations also outlined, making this directly relevant to the reliability and fairness of AI-driven assessments.
- Quality assurance
- Certifications
Research
Base Models Look Human To AI Detectors
Yixuan Even Xu, Ziqian Zhong, Aditi Raghunathan et al.
arXiv · 2026-05-19
This paper reveals that AI-text detectors like GPTZero and Pangram systematically misclassify outputs from base (non-instruction-tuned) language models as human-written, while correctly flagging instruction-tuned model outputs. Building on this, the authors propose Humanization by Iterative Paraphrasing (HIP), a pipeline that fine-tunes a base model as a paraphraser to evade detectors while preserving meaning, tested across Llama-3 and Qwen-3 model families ranging from 0.6B to 70B parameters. The findings suggest current detectors are tracking artifacts of instruction tuning rather than any fundamental property of machine-generated text. This has direct implications for academic integrity and education workflows that rely on these commercial tools, exposing a significant vulnerability in their deployment.
- Quality assurance
- AI policy
Research
Generative Auto-Bidding with Unified Modeling and Exploration
Mingming Zhang, Feiqing Zhuang, Na Li et al.
arXiv · 2026-05-19
GUIDE is a generative auto-bidding framework for digital advertising that combines a Decision Transformer with a Q-value exploration module and an Inverse Dynamics Module to balance exploration and safety in automated bidding. Unlike prior reinforcement learning and generative approaches, it uses an 'explore-safeguard-select' pipeline that prevents unsafe financial decisions while still pursuing performance gains. Deployed at scale on Taobao, GUIDE achieved measurable real-world improvements of +4.10% ad GMV, +1.40% ad clicks, +1.66% ad cost, and +3.52% ad ROI compared to existing baselines. These results demonstrate the system's practical value for advertising platforms seeking more efficient and safer automated bidding.
- Enterprise
Research
Do Vision Models Truly Forget? New Findings from Representation-Level Certification of Visual Unlearning in Vertical Federated Learning
Zhenyu Yu, Yangchen Zeng, Chunlei Meng et al.
arXiv · 2026-05-19
This paper introduces Mirage, a representation-level auditing framework for evaluating machine unlearning in Vertical Federated Learning (VFL). Using four diagnostics—linear probe recovery, centered kernel alignment, feature separability scoring, and layer-wise recovery analysis—the authors test seven baseline unlearning methods across seven datasets and find that methods passing output-level certification still retain substantial class structure in their internal representations, with linear probe recovery exceeding the retrained baseline by up to 15.4 points. A key finding is the 'unlearning trilemma': no existing method simultaneously achieves high utility, output-level forgetting, and representation-level forgetting. These results argue that current output-level certification standards are insufficient and that representation-aware evaluation should become standard practice in federated unlearning research.
- Certifications
- Quality assurance
Research
Agentic Trading: When LLM Agents Meet Financial Markets
Yihan Xia, Panpan You, Taotao Wang et al.
arXiv · 2026-05-19
This paper systematically reviews 77 studies on Large Language Model (LLM)-based trading agents, framing them as expert-system decision pipelines and conducting a reproducibility audit. Of 19 studies meeting minimum empirical criteria, only 2 report time-consistent evaluation protocols, 1 includes a transaction-cost model, and no study achieves the paper's highest reproducibility tier (R3). The central finding is 'protocol incomparability': architectural experimentation is expanding rapidly, but comparable evaluation standards, execution semantics, and reproducible artifacts remain critical bottlenecks. The paper's main contributions are an evidence ledger, reproducibility audit, and reporting checklist intended to raise methodological standards in the field.
- Quality assurance
- Enterprise
Research
Lost in Interpretation: The Plausibility-Faithfulness Trade-off in Cross-Lingual Explanations
Somnath Banerjee, Pranav Jha, Rima Hazra et al.
arXiv · 2026-05-19
This paper investigates a systematic trade-off in how multilingual large language models are audited: when explanations for non-English inputs are generated in English (English-pivot), those explanations can appear plausible—achieving higher agreement with human-annotated rationale spans—yet become less causally faithful to the model's actual predictions. Tested across 3 tasks, 5 languages, and 2 multilingual LLM families, the authors find that comprehensiveness (a faithfulness metric) degrades by up to 5.7x relative to native-language explanations, even when task accuracy stays stable. For socially nuanced classification tasks, English pivots also fail to preserve pragmatic cues, further reducing faithfulness. The authors recommend auditing AI systems in the input language, using multi-faceted faithfulness metrics beyond lexical overlap, and treating English rationales as communication summaries rather than reliable decision traces.
- Quality assurance
- AI policy
Research
DECOR: Auditing LLM Deception via Information Manipulation Theory
Linyue Cai, Samuel Yeh, Jwala Dhamala et al.
arXiv · 2026-05-19
DECOR is a multi-agent auditing framework that detects deception in large language model (LLM) outputs by grounding its analysis in Information Manipulation Theory. It breaks input contexts into atomic informational units and scores each against the model's response across four dimensions of manipulation—such as omission, focus-shifting, or meaning obscuration—producing interpretable 'manipulation profiles' and a global deception index. Evaluated on single-turn and multi-turn benchmarks across real-world domains and 15 frontier models, DECOR achieves state-of-the-art deception detection performance while outperforming competitive baselines. This matters because it offers a fine-grained, theory-grounded approach to auditing AI honesty, moving beyond coarse black-box judgments toward pinpointing exactly which facts were distorted and how.
- Quality assurance
- AI policy
Research
Can Large Language Models Revolutionize Survey Research? Experiments with Disaster Preparedness Responses
Yan Wang, Ziyi Guo, Christopher McCarty
arXiv · 2026-05-19
This paper evaluates whether large language models (LLMs) can improve survey research workflows, using a 2024 Hurricane Milton preparedness survey of 946 Florida residents as a testbed. The authors propose a five-stage LLM integration framework—covering questionnaire design, sample selection, pilot testing, missing-data imputation, and post-collection analysis—and develop seven LLM configurations, including a novel Anchored Marginal Theory-Informed LLM (A-TLM) grounded in Protection Motivation Theory. A-TLM outperforms classical imputation methods (IPW/MI, MICE+PMM, missForest) on RMSE under block-wise missing-not-at-random conditions and achieves near-zero signed bias compared to the largest absolute bias produced by the random-forest imputer. The study also highlights that aggregate bias metrics can mask opposing subgroup errors, proposing subgroup-stratified bias auditing as a reporting standard—a finding with direct implications for data quality assurance in high-stakes survey contexts.
- Quality assurance
Research
Position: Uncertainty Quantification in LLMs is Just Unsupervised Clustering
Tiejin Chen, Longchao Da, Xiaoou Liu et al.
arXiv · 2026-05-19
This position paper argues that mainstream uncertainty quantification (UQ) methods for large language models (LLMs) are fundamentally flawed because they function as unsupervised clustering algorithms that measure internal consistency of model outputs rather than factual correctness. The authors identify three critical pathologies: hyperparameter sensitivity that makes deployment unsafe, an internal evaluation cycle that confuses stability with truth, and reliance on unstable proxy metrics in the absence of ground truth. As a result, current UQ methods cannot detect 'confident hallucinations'—cases where a model is consistently wrong—and may create a false sense of safety in high-stakes deployments. The paper calls for a paradigm shift toward UQ approaches grounded in objective truth, with better evaluation metrics and mechanisms for native uncertainty estimation.
- Quality assurance
- AI policy
Research
SimGym: A Framework for A/B Test Simulation in E-Commerce with Traffic-Grounded VLM Agents
Han Li, Vibhor Malik, Zahra Zanjani Foumani et al.
arXiv · 2026-05-19
SimGym is a framework that simulates A/B tests on e-commerce storefronts using vision-language model (VLM) agents operating in a live browser, eliminating the need to divert real user traffic. The system generates buyer personas from production clickstream data, runs coherent shopping sessions across control and treatment storefronts, and compares simulated outcomes against real buyer behavior. Validated on UI theme change experiments from a major e-commerce platform, SimGym achieves 77% directional alignment with observed add-to-cart shifts, reducing experimental cycles from weeks to under an hour. This matters for enterprises seeking faster, lower-risk experimentation and for quality assurance teams needing reliable pre-deployment evaluation of storefront changes.
- Enterprise
- Quality assurance
Research
LQS v3.1: A Procurement-Grade Quality Standard for AI Training Data with Cryptographically Verifiable Certificates
Alex Adrion
Open MIND · 2026-05-19
LQS v3.1 is a 19-dimension quality standard for AI training datasets designed to meet model-risk audit requirements in regulated industries such as financial services, healthcare, and legal sectors. It addresses weaknesses in existing quality measurement approaches by using a 7-oracle consensus across 5 algorithm families, statistical uncertainty intervals, and conformal prediction to produce verifiable performance forecasts with coverage guarantees. Every quality score is cryptographically signed with an Ed25519 keypair, generating an offline-auditable certificate suitable for procurement and compliance workflows. The specification is presented as a candidate reference methodology for IEEE P2841, NIST AI RMF, and ISO/IEC JTC 1 SC 42 standards efforts.
- Certifications
- Quality assurance
- AI policy
- Enterprise
Research
Impact of EU Laws and Regulations on the Adoption of Artificial Intelligence in Cyber–Physical Systems: A Review of Regulatory Barriers, Technological Challenges, and Cross-Sector Implications
Bo Nørregaard Jørgensen, Zheng Grace Ma
Electronics · 2026-05-19
This scoping review examines how EU regulations—covering data processing, cybersecurity, accountability, and safety—shape the adoption of AI in cyber-physical systems across energy, smart buildings, mobility, and industrial sectors. The authors find that regulation simultaneously acts as a constraint, increasing compliance burden and slowing deployment in high-risk settings, and as an enabler, promoting trustworthy AI, stronger cybersecurity, and interoperable digital infrastructures. The central finding is that regulation actively shapes the design space for AI-enabled cyber-physical systems rather than sitting external to it, meaning future progress requires regulation-aware systems engineering and cross-sector reference architectures. The paper has direct implications for policy, enterprise adoption, and certification frameworks governing AI in critical infrastructure.
- AI policy
- Enterprise
- Certifications
- Quality assurance