News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
5218 items
Research
AI And National Economic Growth: Exploring the Role of STEM Education in Workforce Transformation in Nigeria
Ibeabuchi Anwuri Bruno
INTERNATIONAL JOURNAL OF COMPUTER SCIENCE AND MATHEMATICAL THEORY E-ISSN · 2026-09-16
This study examines how AI-integrated STEM education can drive workforce transformation and economic growth in Nigeria. Using survey data from 210 STEM educators, students, and professionals across three major cities, the researchers found a strong positive correlation (r = 0.78) between AI-integrated STEM education and workforce adaptability, with 72% of respondents acknowledging AI's role in job creation and skill development. The paper recommends curriculum reform, policy implementation, and industry-academia partnerships to build an AI-ready workforce in Nigeria.
- Workforce
- AI policy
Research
A study of medical students’ perception of the risk of future job displacement and adaptive behaviors in the context of artificial intelligence
Ying Yang, Hao Zhang, Hong Liang et al.
Scientific Reports · 2026-09-16
This survey study of 682 medical students examines how perceptions of AI-driven job displacement shape their adaptive behaviors in healthcare careers. Using structural equation modeling, the researchers found that social support, self-efficacy, perceived usefulness, perceived severity, and response efficacy all significantly predict students' intention to adapt, while response cost had a small negative effect; behavioral intention strongly predicted actual adaptive behavior. The findings suggest medical schools should invest in digital literacy training, human-AI collaboration skills, and psychological support to prepare students for AI-driven changes in healthcare. The study highlights AI's growing influence on career perceptions and professional preparation in the medical field.
- Workforce
Research
Scaling data management capabilities for enterprise AI: a maturity model for the banking industry
N. Baum, Lea Mueller-Fortmann, Alexander Benlian
European Journal of Information Systems · 2026-09-16
This study traces how a German bank evolved from fragmented, Excel-based data processes to a cloud-enabled environment capable of supporting AI at scale, using a longitudinal clinical case study. The authors identify five interdependent capability domains—Technology and Infrastructure, Data Governance and Quality, Cultural and Organizational Shifts, Regulatory Compliance and Risk Management, and AI Enablement—and develop a maturity model that offers diagnostic signals and readiness evidence for AI transformation in regulated settings. The findings highlight sociocultural readiness and recurring managerial missteps as central mechanisms in scaling trustworthy AI within financial institutions. The work is particularly relevant for enterprise leaders navigating AI adoption under regulatory constraints.
- Enterprise
- AI policy
Research
Retrieval-Augmented Generation for Commercial Proposal Management in the Electrical Sector: An Engineering Project Management Evaluation
Nefi Alejandro Barron Herrera, Elsa de la Calleja
Ciencia Latina Revista Científica Multidisciplinar · 2026-09-16
This study benchmarks a Retrieval-Augmented Generation (RAG) LLM assistant against manual proposal engineers for managing commercial proposals in electrical generation plant retrofit projects. The RAG system reduced information processing time from a 2–30 minute manual range to a predictable 1–2 minutes—roughly a 90% cycle-time reduction—while achieving a zero hallucination rate and projecting nearly 90% operating-cost savings per cycle. The findings suggest RAG architectures can serve as reliable decision-support tools in pre-project engineering workflows, with meaningful implications for enterprise efficiency and engineering workforce productivity.
- Enterprise
- Workforce
Research
AI Signals in the DAX 40: A Public-Source Benchmark of Disclosed AI Strategy, Adoption and Governance across Germany's 40 Largest Listed Companies (2026 Edition)
Sami Darwich
Zenodo (CERN European Organization for Nuclear Research) · 2026-09-16
This report benchmarks all 40 DAX-listed German companies on publicly observable AI readiness using a transparent 15-criterion framework across five dimensions scored from official corporate sources. The mean total score is 1.81 out of 3 (60/100), with governance and responsible-AI signals being the weakest and most dispersed dimension (mean 1.39) compared to adoption signals (mean 2.30); in 75% of companies adoption exceeds governance disclosure. Only about one in five companies publishes a quantified AI target system, and only one reports an external AI management-system certification. The findings highlight a systematic gap between enterprise AI adoption momentum and formal governance accountability among Germany's largest listed companies.
- Enterprise
- AI policy
Research
AI Signals in the DAX 40: A Public-Source Benchmark of Disclosed AI Strategy, Adoption and Governance across Germany's 40 Largest Listed Companies (2026 Edition)
Sami Darwich
Zenodo (CERN European Organization for Nuclear Research) · 2026-09-16
This report benchmarks all 40 DAX companies on publicly observable AI readiness using a transparent 15-criterion framework covering strategy, adoption, data/technology, governance, and talent, scored from official corporate disclosures published through September 2026. The mean score is 1.81 out of 3 (60/100), with adoption signals far ahead of governance signals (means 2.30 vs. 1.39), and 75% of companies showing adoption disclosures that exceed their governance disclosures. Only about one in five companies publishes a quantified AI target system, and only one reports an external AI management-system certification. Three disclosure profiles emerge—broad and institutionalised, adoption-led but governance-light, and limited evidence—highlighting a systemic gap between AI deployment momentum and formal governance accountability among Germany's largest listed firms.
- Enterprise
- AI policy
- Certifications
Research
Adoption and Automation Risk of AI Tools in Public and Private Enterprises: A Machine Learning-Based Analysis
Enes Bajrami
Balkan Journal of Electrical and Computer Engineering · 2026-09-16
This study analyzes AI adoption and automation risk across public and private sector enterprises in North Macedonia using survey data from 477 respondents. Machine learning models (Random Forest and XGBoost) predicted high-risk roles, finding that private sector jobs with high proportions of routine tasks are most vulnerable to automation, while public sector and creative roles are more insulated. Feature importance analysis highlighted sector type, industry classification, and AI adoption level as key predictors. The findings offer actionable insights for policymakers designing reskilling programs and anticipating labor market shifts.
- Workforce
- Enterprise
- AI policy
Research
ERPBench: A State-Grounded Evaluation Paradigm for Computer-Use Agents in Enterprise Software
Kratika Bhagtani, Kusha Sridhar, Maziyar Baran Pouyan et al.
arXiv · 2026-09-15
ERPBench introduces a benchmark for evaluating computer-use AI agents on live Enterprise Resource Planning (ERP) systems, scoring agent actions against ground-truth database values rather than on-screen outcomes. The study finds that strong general GUI performance does not transfer to enterprise reliability: agents may successfully save a form in up to 85% of runs yet write the correct value in as few as 3% of cases. The paper also presents a production-grade deployment harness that gates agent actions behind human approval for safe use. These findings highlight a critical gap between general-purpose AI agent capability and the precision required for enterprise software operations.
- Enterprise
- Quality assurance
Research
Does AI Assistance Leave a Temporal Fingerprint? Detecting Overreliance in AI-Assisted Writing and Programming
Eduardo Davalos, Yike Zhang
arXiv · 2026-09-15
This paper investigates whether AI-assisted writing and programming leave detectable temporal patterns in process data — specifically keystroke timing and editor telemetry — that can distinguish authentic student work from wholesale delegation to AI. Analyzing three public corpora (CoAuthor, RealHumanEval, and a pre-LLM CS1 keystroke dataset), the researchers find that AI contributions arrive in statistically distinct bursts compared to human baselines across both writing and programming tasks. Classifiers using only temporal features can separate simulated wholesale delegation from authentic work with near-perfect accuracy (F1 ≥ 0.997), while ordinary AI collaboration remains difficult to distinguish from unassisted work. The findings suggest process-level visibility — rather than finished-artifact analysis — could serve as a more accurate and ethically grounded basis for detecting academic integrity violations, though the authors note validation in authentic coursework is still needed.
- Certifications
- AI policy
Research
Can VLMs Reliably Assess Sidewalk Accessibility Attributes from Pedestrian-Level Imagery?
Seung Jae Lieu, Diego Morra, Chiara Cadoni et al.
arXiv · 2026-09-15
This paper investigates whether vision-language models (VLMs) can reliably measure sidewalk accessibility attributes — effective width, longitudinal slope, cross slope, and pavement condition — from pedestrian-level images, which is critical for wheelchair users and people with reduced mobility. Testing four VLMs on 514 sidewalk images from Seoul with field-measured ground truth, the study finds that no quantitative attribute reaches the precision required for compliance assessment: effective width is the most informative but models systematically overestimate it, longitudinal slope is only marginally informative, cross-slope intervals are too wide to resolve regulatory thresholds, and pavement-condition estimates degenerate for most models. The authors introduce the first use of sampling-based conformal prediction for VLM-based accessibility assessment, showing that uncalibrated sampling dispersion covers only 17–47% of field-measured values at a nominal 90% level, and that response self-consistency is not evidence of accuracy. The findings matter for policy and certification efforts because they clarify the current limits of AI-based sidewalk compliance screening while identifying a calibration-based path for ruling out segments clearly far from regulatory thresholds.
- Quality assurance
- AI policy
Research
Do Frontier Models Seek Safety Evidence Before Acting?
Omer Tafveez
arXiv · 2026-09-15
This paper introduces SAFE, a benchmark that tests whether frontier AI models proactively seek out safety-relevant information before making deployment decisions, rather than merely responding to safety data already provided to them. Across four frontier models (GPT-5.5, o3, Claude Opus 4.8, and Claude Sonnet 4.6), the study finds distinct evidence-acquisition behaviors: inspection rates rise sharply with problem severity and drop with retrieval cost, while the stated probability of a problem has surprisingly weak influence—raising likelihood from 10% to 70% shifts inspection by at most 21 percentage points. The paper also uncovers a mismatch between model behavior and stated reasoning, where factors like evidence framing strongly influence decisions near the inspection boundary without being mentioned, while probability is frequently cited despite having little causal effect. The findings suggest that AI deployment safety depends not just on how models handle known risks, but on whether they reliably seek the evidence needed to determine that acting is safe in the first place.
- Quality assurance
- AI policy
Research
SAGE: Governed Artifact Generation from Enterprise Guidelines
Mohammadreza Sediqin, Shivali Dalmia, Sumukha Thoppanahalli et al.
arXiv · 2026-09-15
SAGE is a governed multi-stage LLM pipeline that automates the conversion of enterprise guideline documents—containing mixed narrative text, tables, and images—into structured work artifacts. The system uses a shared versioned rule store, schema-validated inter-stage contracts, and end-to-end provenance tracking, combined with deterministic structural validation and LLM-based semantic scoring to catch errors before human review. Tested on 120 documents, SAGE reduces turnaround time from two to three days down to 20–100 minutes, achieving a 96% document-level success rate and 3.2% hallucination rate, compared to 15.7% hallucination without governance. This matters for enterprises seeking to automate high-effort document processing workflows while maintaining auditability and accuracy.
- Enterprise
- Quality assurance
News
Agility’s new humanoid robot will stop, squat to avoid harming human coworkers
arstechnica.com · 2026-09-15
Ars Technica reports that Agility Robotics has unveiled its Digit 5 humanoid robot, which is engineered to operate safely alongside human coworkers without requiring physical separation barriers. The robot uses an onboard safe motion system that autonomously detects nearby people and responds with precautions such as rerouting, stopping, or even crouching into a seated position to avoid contact. According to Agility's CTO, the system can select from a variety of mitigation behaviors depending on the nature of the detected human presence, potentially enabling broader deployment in warehouses and automotive factories.
- Workforce
- Enterprise
Research
Decomposition Buys Integrity, Not Yield
Rong He
arXiv · 2026-09-15
This paper mathematically models how multi-agent 'decomposition' systems—where a task is split across a tree of agents—affect how much information actually reaches the final output. The authors show that deeper agent trees systematically lose findings at each level (measured exponent δ=0.34 on 600 production traces), and that alignment also degrades with depth, with roughly one in sixteen task briefs going off-target per tier. While decomposition does offer real benefits in reducing root-context size and cost (flat agents bill as N^1.39, and two tiers become cost-competitive at 403 findings), the model estimates only 0.7%–11.3% of production sessions actually benefit from delegation, versus 7.8% that currently use it—suggesting multi-agent delegation is overused relative to its yield advantages.
- Enterprise
- Quality assurance
Research
FlashVector: Agent for Hierarchical Model Serving Stack Optimization
Qi Wu, Lohan Lemire, Kai Meng et al.
arXiv · 2026-09-15
FlashVector is an agentic AI system designed to optimize performance across all layers of a model serving stack—GPU kernels, ML framework computation graphs, model servers, and on-demand feature processing—areas that typically require deep, siloed human expertise. The system generalizes single-kernel optimization agents into an extensible framework capable of tuning heterogeneous technical stacks holistically. Deployed in Unity's Vector advertising platform, FlashVector achieved up to 2x throughput increase and up to 1.98x latency speedup on the model server, and up to 1.6x throughput increase on the feature store. This matters for enterprise AI infrastructure because it demonstrates that agentic systems can automate complex, cross-layer performance tuning at scale, reducing one of the largest cost drivers in production recommender systems.
- Enterprise
Research
Evaluating Ambient Clinical Scribes in India: The Need for Multilingual Real-World Clinical Conversation Data
Siddharth D Jaiswal, Krithi S, Ashish Makani et al.
arXiv · 2026-09-15
This paper evaluates the readiness of ambient clinical scribes (ACS)—AI tools that automate clinical documentation—for deployment in Indian healthcare settings. The authors find that existing ACS systems are built and validated on Global North speech and consultation styles, making them poorly suited to Indian encounters, which are brief, multilingual, code-mixed, and conducted in noisy, resource-constrained environments. A systematic survey reveals no publicly available, large-scale, real-world benchmark datasets for ACS evaluation in India, with current datasets being overwhelmingly synthetic and culturally mismatched. The paper calls for a shared, real-world, multilingual benchmark and standardized evaluation policies to ensure these systems are safe and reliable before large-scale procurement and deployment.
- Quality assurance
- AI policy
Research
Trust propagation and structural containment in Multi-agent LLM pipelines
Tanzim Hossain Safin, Sharif Noor Zisad, Swakkhar Shatabda et al.
arXiv · 2026-09-15
This paper investigates how attacks spread through multi-agent LLM pipelines, specifically studying a four-agent LangGraph system where a low-privilege agent can be compromised and influence higher-privilege agents. The researchers test shared-memory poisoning and indirect prompt injection attacks, finding that without defenses, memory poisoning reaches execution in every trial. They introduce a structural authorization layer using signed tokens and a policy oracle, which achieves 0% Unsafe Action Rate even when the Validator agent is fully compromised (100% Judgment Bypass Rate), and an Observer layer reduces false-positive hijacking detection from 49% to 7%. The results demonstrate that structural authorization boundaries can contain compromised agent behavior even when the LLM's own judgment has been subverted, with important implications for securing enterprise AI pipelines.
- Enterprise
- Quality assurance
Research
Towards Detecting AI-Assisted Responses in Online Surveys
Qizhou Wang, Bogdan Mamaev, Christopher Leckie
arXiv · 2026-09-15
This paper introduces ASURRE, a benchmark dataset for detecting AI-assisted responses in online surveys, pairing LLM-generated responses across multiple usage strategies—from full generation to persona-grounded agentic completion—with genuine human responses. The study finds that naive AI usage is easily detectable by existing machine-generated text detectors, but persona-grounded agents that mimic real respondents reduce detector performance to near chance. However, agentic completion still leaves distinctive behavioural traces, and a simple few-shot, training-free aggregator over these cues improves mean AUROC by +0.14 over the best existing detector. These findings matter for survey-based research validity, as undetected AI-assisted responses could systematically corrupt data used in policy, social science, and other evidence-based fields.
- Quality assurance
- AI policy
Research
After the Party: Governing What a Viral Agent-Skill Ecosystem Left Behind
Yunpeng Xiong, Ting Zhang
arXiv · 2026-09-15
This paper examines the lifecycle and security risks of a rapidly grown public AI agent-skill registry (OpenClaw/ClawHub) that nearly doubled its listings in 91 days during early 2026. The study finds that human oversight was minimal—77.86% of skills had zero stars or comments—while 85.06% of readable skills carried privilege evidence indicating potential security risks. Automated security scanning proved unreliable, with three scanners disagreeing on 23,702 of 61,990 skills and weighted sensitivity ranging from only 21.67% to 61.06% after human adjudication. The findings argue that governing fast-growing agent-skill registries requires robust, transparent measurement and independent validation rather than simple metadata filters or single scanner scores.
- AI policy
- Quality assurance
Research
Memorisation bias in medical AI
Moritz A. Knolle, Martin J. Menten, Laurin Lux et al.
arXiv · 2026-09-15
This paper identifies and characterizes 'memorisation bias' in medical AI: when a model is trained on a patient's historical (even anonymised) data, its predictions on that same patient's future data can be significantly distorted compared to models that never saw their records. The authors show this bias persists across diverse data types, model architectures, and time spans of decades, and has asymmetric diagnostic consequences — reducing sensitivity when a new condition is present and inflating both sensitivity and specificity when health status is unchanged. Because de-identification makes it difficult to identify returning contributors and exclude them, current model development and deployment practices do not adequately protect against this risk. The findings call for changes to training and deployment protocols to mitigate memorisation-related harms in clinical settings.
- Quality assurance
- AI policy
Research
Scaling Articulated Rationales for MLLM-based Recommendation
Haoke Xiao, Yueyang Liu, Yuhui Zhang et al.
arXiv · 2026-09-15
This paper introduces SARA (Scaling Articulated Rationales), an industrial framework that converts sparse, natural-language user preference explanations—called articulated user rationales (AURs)—into scalable recommendation signals for Kuaishou Live. The system builds a quality-controlled dataset from 240 million users, fine-tunes a 7-billion-parameter multimodal language model to extend rationale generation across 10 million content authors, and integrates positive and negative rationales into a production ranker. Offline evaluations, human assessments, and live A/B tests show the approach generates more specific and polarity-consistent rationales than baseline models, while improving user engagement and reducing negative feedback over 30+ days of deployment. The work demonstrates that articulated rationales can serve as a practical, first-class textual signal in large-scale industrial recommendation systems.
- Enterprise
Research
EviScope: Paired Counterfactual Evidence Diagnostics for Faithful and Efficient Grounded Language Models
Suryadeep Singh Deswal
arXiv · 2026-09-15
EviScope is a diagnostic benchmark designed to reveal when language models produce correct answers for the wrong reasons — such as ignoring evidence, citing the wrong source, or failing to recognize contradictions. The benchmark uses paired counterfactual question sets that systematically add, remove, distract from, or contradict evidence while holding the question fixed, then measures whether model behavior changes appropriately. Testing Qwen2.5-7B, Llama 3.1 8B, and Gemini 3.5 Flash reveals that standard answer accuracy hides significant grounding failures: explicit evidence-gating actually underperforms vanilla RAG on local models, and even the strongest model (Gemini) still answers 5% of cases after contradiction is inserted. These findings matter for quality assurance of retrieval-augmented systems, showing that fidelity to evidence requires dedicated evaluation beyond accuracy alone.
- Quality assurance
Research
Beyond "ChatGPT Can Make Mistakes": Designing Interventions to Support Metacognitive Monitoring in AI-Assisted Work
Manuel A. D. Santos, Paul Thiesse, Steeven Villa et al.
arXiv · 2026-09-15
This paper investigates how to design interventions that help users better judge both their own competence and the reliability of AI systems during AI-assisted tasks—a challenge the authors frame as metacognitive monitoring. Through expert elicitation of 30 interventions and a between-subjects experiment with 917 participants across 12 planning-and-organizing problems, the study tested four intervention types (reliability cards, contrasting replies, pause points, and post-problem reflection) against a baseline LLM assistant. Reliability cards and contrasting replies were found to reduce estimation error and overconfidence and improve aggregate confidence discrimination, though no improvements in task performance were established. The findings matter for enterprise and workforce contexts because they provide comparative evidence that monitoring quality and task performance are separable design targets, giving practitioners a shared vocabulary and design space to guide the deployment of AI assistance.
- Enterprise
- Workforce
News
Exclusive: Paying for frontier AI models buys 4-month head start at 5x the cost
arstechnica.com · 2026-09-15
Ars Technica reports on a new Mozilla State of Open Source AI report finding that the performance gap between leading US frontier AI models and the best open-weight Chinese models has narrowed to just 4.4 months. The report notes that Moonshot AI's Kimi K3 scores only three points below Anthropic's Claude 5 on a major AI benchmark index while costing roughly 70% less. Mozilla's CTO argues that closed frontier models are only worth the premium for specific high-demand tasks—such as expert professional work and long-context retrieval—while most organizations should default to open models for the bulk of their workloads.
- Enterprise
News
NIST Awards More Than $30 Million for MEP Centers in 11 States and Puerto Rico
nist.gov · 2026-09-15
NIST News reports that the National Institute of Standards and Technology has awarded more than $30 million to 12 Manufacturing Extension Partnership (MEP) centers across 11 states and Puerto Rico to help small and medium-sized manufacturers adopt advanced technologies including AI, robotics, automation, and additive manufacturing. The competitively awarded grants range from roughly $812,000 to over $6 million per center and require awardees to secure at least 50% in non-federal matching funds. Centers will enter five-year cooperative agreements and must develop metrics for technology adoption to be shared across the broader MEP National Network, which includes nearly 1,400 advisers at more than 450 service locations. No awards were made for Alaska or California, with a new competition for those states planned for early 2027.
- Enterprise
- Workforce