News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
5571 items
Research
Comparing the text-based diagnostic reasoning performance of emergency medicine physicians and large language models in both definitive and differential diagnoses using standardized clinical vignettes: a preliminary study
Mehdi Arzani Shamsabadi, Roya Vatankhah, Hasan Jalilvand et al.
International Journal of Emergency Medicine · 2026-07-31
This preliminary study compared the text-based diagnostic accuracy of 10 emergency medicine physicians against four large language models (ChatGPT GPT-5.2, Gemini 3, Microsoft Copilot GPT-4, and Claude Opus 4.1) using 10 standardized clinical vignettes from emergency department presentations. AI models achieved significantly higher overall diagnostic accuracy (73.75%) than physicians (57.00%, p=0.014), and crucially maintained consistent performance across both definitive and differential diagnoses, while physicians showed a marked drop in accuracy for differential diagnoses (45.0% vs. 69.0%). The authors conclude that LLMs demonstrate robust pattern-recognition and reasoning capabilities that could make them reliable clinical decision support tools, especially for complex differential diagnostic reasoning, though they caution that the small sample of cases and evaluators limits generalizability.
- Quality assurance
- Enterprise
Research
Avoiding De‐Skilling and Dependency: A Practical Guide to Using Generative <scp>AI</scp> in <scp>TESOL</scp> Teacher Education
Lucas Kohnke, Benjamin Luke Moorhouse
TESOL Quarterly · 2026-07-31
This paper argues that generative AI tools in TESOL teacher education risk causing de-skilling and professional dependency if teachers offload core pedagogical work to AI before developing the judgment to use it responsibly. Drawing on a five-dimension framework of professional GenAI competence, the authors propose design principles and practical activities—such as low-stakes tool engagement, scenario-based microteaching, and structured reflection—for both pre-service and in-service teacher education programs. The paper emphasizes teacher agency, transparency, and ethical responsibility, including concerns about digital exclusion, labor conditions, and environmental impacts. It concludes that effective GenAI integration must be grounded in teacher competence and professional judgment rather than tool reliance.
- Workforce
Research
La evaluación educativa ecuatoriana ante la inteligencia artificial generativa: Desafíos para la autenticidad del aprendizaje
Henry Fabricio Carrillo Chacón, Fátima Lourdes Vega Barberán, Diego Alexander Cevallos Torres et al.
REDSA Revista Ecuatoriana de Desarrollo Social y Ambiental · 2026-07-31
This integrative review examines how generative AI threatens the authenticity of student assessment in Ecuador's educational system, finding that tasks focused only on final products are especially vulnerable to cognitive delegation and that automated plagiarism detectors cannot reliably establish authorship. The study analyzed 30 academic, institutional, and regulatory documents and coded evidence into six thematic categories, concluding that the most effective pedagogical responses combine process traceability, oral dialogue, situated performance, and transparent disclosure of AI use. The authors argue that Ecuadorian regulations provide formative foundations for redesigning assessment, but that learning authenticity depends less on prohibiting AI than on aligning permitted use with learning outcomes and gathering complementary evidence of student reasoning.
- Quality assurance
- AI policy
Research
Bridging the skills gap in recruitment: A RAG-based LLM framework for cybersecurity job advertisement analysis
Abdeslam Rehaimi, Yassine Sadqi, Abdessamad Elboushaki et al.
Information Processing & Management · 2026-07-31
This paper presents a retrieval-augmented generation (RAG) framework using large language models (GPT-3.5, GPT-4.1, and Meta Llama 3) to automatically extract and structure information from cybersecurity job advertisements at scale. Applied to 1,681 cybersecurity postings drawn from LinkedIn, Indeed, and Rekrute—with a focus on Morocco's market—the system achieves near-perfect scores across six evaluation metrics, with GPT-4.1 reaching 96.3% correctness and 100% completeness. Key findings reveal that most postings target mid-level candidates with advanced degrees and three or more years of experience, and that CISSP is the most in-demand certification. By automating labor-market analysis that previously relied on manual or semi-automated NLP approaches, the framework offers a scalable tool for understanding cybersecurity workforce demand and skills gaps.
- Workforce
- Certifications
Research
WHEN FAIR AI BECOMES UNFAIR: A COUNTERFACTUAL AUDIT OF POSITIONAL BIAS IN LARGE LANGUAGE MODELS FOR HIRING DECISIONS
Arthur Mesquita Camargo, Rafaela Silva Figueiredo Camargo
Seven Editora eBooks · 2026-07-31
This study audits four leading large language models (GPT-4.1-mini, GPT-5.2, Claude Sonnet 4.6, Gemini 2.5 Pro) and one lower-capacity model for gender, racial, and positional bias in a simulated CEO-selection task using 60 functionally identical candidate profiles. Frontier models showed no statistically significant demographic bias, but a lower-capacity model produced extreme rank segregation driven not by gender bias but by primacy bias—favoring candidates listed earlier in the prompt—yielding a perfect effect size (Cliff's δ = 1.0). The findings demonstrate that even demographically neutral AI systems can generate discriminatory hiring outcomes through structural artifacts like presentation order, and that bias auditing must encompass these interaction effects alongside traditional demographic parity checks. An ecosystem-level analysis of 2,653 model listings further highlights systemic risks beyond individual model behavior.
- Workforce
- AI policy
Research
Impacts of the popularization of artificial intelligence on the three principles of information integrity
Carlos Alberto Ávila Araújo
adComunica revista científica de estrategias tendencias e innovación en comunicación · 2026-07-31
This study analyzes how the rise of generative AI (GAI) tools challenges the three core principles of information integrity—accuracy, consistency, and reliability—as defined in multilateral frameworks from organizations like the UN and G20. Using Habermas's theory of communicative action to critically examine institutional documents and recent GAI research, the authors identify threats such as increased disinterest in truth, compromised cognitive authorities, algorithmic discrimination, and greater difficulty distinguishing truth from falsehood. The paper concludes that GAI poses risks to science, democracy, public health, and environmental protection, while also offering benefits like efficiency and accessibility. The authors argue that GAI-related risks must be incorporated into international policy documents and actions promoting information integrity.
- AI policy
Research
The Ontological Attack Surface: Measured Distortion Channels as Adversarial Primitives in Clinical AI
Florian O. Stummer
arXiv · 2026-07-31
This paper reframes the security threat model for clinical AI systems, arguing that the real attack surface lies in measurable ontological distortion channels—systematic ways that coded administrative data already diverge from clinical reality—rather than in conventional adversarial input perturbation or data poisoning. Using three empirical datasets (synthetic Synthea simulations, real MIMIC-IV EHR data, and aggregate primary-care data from 23 practices), the authors demonstrate that distortion primitives such as coding drift, set-membership rescue, and salient-code overshadowing are real, quantifiable, and exploitable by adversaries who amplify existing feedback channels rather than injecting high-magnitude noise. The study proposes a six-class threat taxonomy mapping each distortion channel to an adversarial primitive, and crucially shows that because these channels are detectable in non-invertible aggregate data, the attack surface can be monitored without requiring patient-level access. This matters for clinical AI quality assurance and policy because it identifies a tractable, observable vulnerability class specific to administrative healthcare pipelines that current security frameworks largely overlook.
- Quality assurance
- AI policy
Research
Artificial Intelligence Use and Cognitive Resource Allocation: Nonlinear Associations with Mental Workload and Perceived Performance
Şahin Danışman, Filiz Evran Acar
Journal of Intelligence · 2026-07-31
This study of 464 pre-service teachers finds that how often someone uses AI tools is associated with their mental workload and perceived performance on lesson-planning tasks in a nonlinear way. Infrequent AI users reported higher mental demand, effort, and frustration and lower perceived task performance, while more frequent users tended to report lower frustration and higher perceived performance. Polynomial trend analyses show the relationship is not simply linear across all NASA-TLX workload dimensions, with some dimensions lowest among occasional users. The findings suggest that workforce training and onboarding strategies for AI tools should account for the complexity of how usage frequency shapes cognitive burden and self-assessed output quality.
- Workforce
Research
Curriculum innovation through artificial intelligence and its influence on pupils’ independent learning in Azerbaijan
Galandarov Sahil, Yanping Li, Baghirova Aytan et al.
Frontiers in Education · 2026-07-31
This quantitative study of 1,250 students and 250 teachers in Azerbaijani secondary schools finds that AI integration in classrooms strongly predicts perceived curriculum innovation (β=.55), which in turn significantly predicts students' independent learning skills (β=.51), with curriculum innovation mediating much of AI's effect on student autonomy. Teacher AI literacy moderates this relationship, amplifying the positive impact of AI on curriculum redesign. The authors conclude that effective AI adoption requires both thoughtful curriculum redesign and substantial investment in teacher professional development, and call for a comprehensive national policy framework in Azerbaijan to coordinate these efforts.
- Workforce
- AI policy
Research
From learners to contributors: how an AI-infused STEM program shaped youth identity and initiated them to an AI-future
Mark Weckel, Preeti Gupta, Katherine S. Moore et al.
Frontiers in Education · 2026-07-31
This study evaluated SRMPmachine, a 150-hour out-of-school STEM program that embedded machine learning literacy into scientific research internships for 42 high school students. Using a mixed-methods time-series design, researchers found significant gains in ML knowledge and skills—especially among youth from underrepresented groups—alongside emerging self-efficacy and a sense of belonging in AI communities. Qualitative interviews revealed nuanced shifts toward 'informed ambivalence,' reflecting greater ethical awareness rather than simple attitude change. The findings suggest that integrating ML into authentic science mentorship can strengthen the AI readiness of the next generation of workers and contributors.
- Workforce
Research
Can the fourth industrial revolution solve the productivity problem?
Gerbrand Tholen, Andrew Westwood
The Economic and Labour Relations Review · 2026-07-31
This paper critically examines whether generative AI and the fourth industrial revolution will actually improve workers' economic wellbeing, challenging the assumption that AI-driven productivity gains will be broadly shared. The authors distinguish between 'zero-sum productivity' (gains captured by capital owners) and 'positive-sum productivity' (gains shared with workers), arguing that three factors undermine equitable outcomes: corporations' historical tendency to retain productivity gains, AI's threat to knowledge workers with specialized expertise, and uneven AI adoption across organizations. Drawing on evidence of wage-productivity decoupling since the 1970s and rent-seeking behavior, the paper concludes that AI may deepen inequality rather than resolve it, and that equitable outcomes require active state intervention through industrial policy, job creation incentives, and work design reforms.
- Workforce
- AI policy
Research
Policy analysis of artificial intelligence in social science research at higher education institutions: problems and possibilities
Rashmi Gopi, Smita Agarwal
Frontiers in Education · 2026-07-31
This policy analysis examines how artificial intelligence is being integrated into social science research and education at Indian higher education institutions, evaluating key policy documents including India's National Education Policy 2020 and the 2025–26 budget commitment of ₹500 crore for Centres of Excellence. Using human coding of textual documents and a digital ethics framework, the study finds that major problems include AI's 'invisibility,' algorithmic bias reflecting caste, class, gender, and regional disparities, unreliable detection tools, and erosion of critical thinking among Gen-Z learners. The authors conclude that current regulatory mechanisms are fragmented and indirect, and argue that establishing a robust India-centric regulatory framework for AI in higher education is now essential rather than optional.
- AI policy
- Quality assurance
Research
Strengthening Vietnam’s legal framework for personal data protection and social responsibility: Lessons from Japan’s experience
Thi Phuong Cham Nguyen
Knowledge and Performance Management · 2026-07-31
This study evaluates Vietnam's legal framework for personal data protection—including the new Personal Data Protection Law 2025—against Japan's Personal Information Protection Law, using theoretical legal analysis and comparative law methods. The findings reveal that Vietnam's regime is hampered by fragmented regulations, an overly consent-based approach, and weak enforcement mechanisms that render its core principles largely symbolic in practice. Japan's model, featuring a centralized supervisory body, risk-based governance, and regular legal review cycles, is identified as a more effective approach for promoting socially responsible data use. The authors recommend that Vietnam establish a synchronized legal framework, adopt risk-based data governance, and create an independent agency for nationwide oversight and implementation.
- AI policy
Research
Navigating academic integrity in the age of on-demand artificial intelligence: implications for globalizing higher education at the University of Ibadan
Solomon O. Ojedeji
Frontiers in Education · 2026-07-31
This study of 205 undergraduate students at the University of Ibadan, Nigeria finds a strong positive correlation (r = .611, p < .05) between AI tool use and academic dishonesty, indicating that unrestricted use of on-demand AI significantly undermines academic integrity. While AI tools improve learning effectiveness and knowledge access, the authors conclude that Nigerian universities must redesign their academic integrity frameworks to address these risks. The study recommends context-sensitive institutional policies that embed AI ethics into curricula and establish clear standards for identifying AI-generated work, arguing these steps are essential for maintaining the legitimacy and competitiveness of African higher education systems.
- AI policy
- Quality assurance
Research
Beyond Code Generation: AI Across the Product Development Lifecycle
Iuliia Mineeva
Zenodo (CERN European Organization for Nuclear Research) · 2026-07-31
This comparative case study examines AI adoption across the full product development lifecycle in a lean technology startup, comparing two similar projects with differing levels of AI use. The AI-assisted project achieved a 35.8% reduction in overall labor effort, with the largest gains in research, requirements preparation, documentation, and design, alongside improvements in software development and testing. The findings demonstrate that AI can support activities spanning market research and hypothesis validation through to post-release improvement, while still leaving key decisions to human judgment. The results are relevant to enterprise teams and workforce planning, showing that AI's productivity benefits extend well beyond code generation.
- Enterprise
- Workforce
Research
Responsible Artificial Intelligence for Managing Vocational Certificate Education in Thailand
Chaimongkhol Pugsuwan
Zenodo (CERN European Organization for Nuclear Research) · 2026-07-31
This paper develops a context-specific framework for responsible AI adoption in Thailand's Vocational Certificate (PWC) education system, addressing the unique challenges of serving upper-secondary learners—many of them minors—across safety-sensitive and workplace-connected occupational fields. Through an integrative review of Thai legal and policy documents, international standards, and peer-reviewed research, the authors identify seven management domains where AI can add value and six cross-cutting risk clusters, including child rights, bias, and cybersecurity. The resulting Responsible AI Management Framework proposes seven principles, four risk tiers, lifecycle governance gates, and a phased implementation roadmap, arguing that AI should augment rather than displace professional judgment in consequential decisions. The framework provides actionable guidance for vocational education authorities, colleges, quality-assurance bodies, and employers in Thailand.
- Certifications
- Quality assurance
- AI policy
Research
An explainable AI-based workforce intelligence framework for integrating future skill demand and employee attrition prediction with risk-aware decision analytics
Prathap D L, Thimmaraju S N
Future Technology · 2026-07-31
This paper presents an explainable AI framework that unifies external labor-market skill demand and internal employee attrition prediction into a single workforce intelligence system. Using TF-IDF for skill clustering and comparing Logistic Regression, Random Forest, and XGBoost for attrition prediction, the study finds Logistic Regression performs best (ROC-AUC 0.7954; recall 0.7872 at a 0.40 threshold). SHAP analysis identifies frequent business travel, job level, lab technician role, and total years of work as the most influential attrition drivers. The resulting Workforce Risk Score combines normalized skill demand and attrition risk to provide actionable, summary-level decision support for workforce planning.
- Workforce
- Enterprise
Research
The "Fair Use" and "Fair Dealing" Dilemma in Large Language Model Pre-Training: A Comparative Analysis of US, EU, and UK Copyright Frameworks
Dr. Jyoti Garg Amaresh Patel
Economic Sciences. · 2026-07-31
This paper systematically compares how the United States, European Union, and United Kingdom copyright frameworks treat the use of copyrighted text in training large language models such as GPT-4, Claude, and Gemini. Drawing on recent landmark judicial decisions and legislative instruments including the EU AI Act and the UK's 2026 Copyright and AI Report, the authors find that the three jurisdictions differ fundamentally in design: the US relies on a post-hoc four-factor fair use balancing test, the EU employs a structured legislative opt-out framework, and the UK remains in unresolved policy flux. The paper concludes by proposing that an emerging international standard should combine the EU's structural clarity with US jurisprudential flexibility to create a regime that is both commercially viable and normatively sound. This matters for AI policy and enterprise deployment of LLMs, as legal uncertainty around training data directly affects how and where these systems can be built and commercialized.
- AI policy
- Enterprise
Research
Responsible Artificial Intelligence for Managing Vocational Certificate Education in Thailand
Chaimongkhol Pugsuwan
Zenodo (CERN European Organization for Nuclear Research) · 2026-07-31
This article develops a context-specific responsible AI management framework for Thailand's Vocational Certificate (PWC) level education, covering upper-secondary learners across diverse occupational fields. Drawing on an integrative review of Thai legal and policy documents, international standards, and peer-reviewed research, it identifies seven management domains where AI may add value and six cross-cutting risk clusters, including child rights, bias, and cybersecurity. The proposed framework comprises seven principles, four risk tiers, seven lifecycle gates, and a phased implementation roadmap, arguing that AI should augment rather than displace professional judgment in consequential decisions about learners. It offers actionable guidance for vocational education commissions, colleges, quality-assurance bodies, and employers, while setting a research agenda for equitable, child-centred AI use in Thai vocational education.
- Certifications
- Quality assurance
- AI policy
Research
COMPLIANCE THEATRE: RETHINKING EVALUATION AND ENFORCEMENT IN FRONTIER AI REGULATION
Matt Bartlett
The Cambridge Law Journal · 2026-07-31
This article argues that current AI regulation is built on faulty assumptions about our ability to evaluate general-purpose AI systems, a gap the author calls 'compliance theatre.' The technical literature does not support the evaluative capacity that nascent governance frameworks presuppose, and the rapid pace of frontier AI development has outpaced human experts' ability to reliably interpret AI behavior to existing legal standards. The author proposes a new paradigm called 'Sentinel Governance,' which emphasizes governance-oriented innovation and experimentation to supplement human oversight and prevent AI regulations from becoming mere checkbox exercises.
- AI policy
- Certifications
Research
Beyond Code Generation: AI Across the Product Development Lifecycle
Iuliia Mineeva
Zenodo (CERN European Organization for Nuclear Research) · 2026-07-31
This comparative case study examines AI adoption across the full product development lifecycle in a lean technology startup by comparing two similar projects that differed only in level of AI use. The AI-assisted project achieved a 35.8% reduction in overall labor effort, with the greatest gains in research, requirements preparation, documentation, and design, and additional efficiency improvements in software development and testing. The findings demonstrate that AI can support activities spanning market research and hypothesis validation through post-release improvement, while leaving key decisions to human experts.
- Enterprise
- Workforce
Research
AISPA: User-Centric System Prompt Auditing for Large Language Model Applications
Xiangning Lin, Shenzhe Zhu, Shu Yang et al.
arXiv (Cornell University) · 2026-07-30
AISPA is a user-centric auditing framework that systematically evaluates system prompts—hidden developer instructions that govern AI application behavior—across eight dimensions relevant to user interests. Applying this framework to 3,249 instructions from 88 commercial AI products, the authors find that while 98.9% of products include at least one protective instruction, only 24% cover all eight dimensions, and roughly 40% contain at least one instruction that works against user interests. System prompt design varies widely across organizations, with some averaging over 60 protective instructions per product and others fewer than 5. The findings underscore a significant transparency and accountability gap, calling for greater standardization and independent oversight of system prompts in commercial AI deployments.
- AI policy
- Quality assurance
Research
ORCA-bench: How Ready Are Language Model Agents for Oncall?
Albert Gong, Kyuseong Choi, Abhineet Agarwal et al.
arXiv · 2026-07-30
ORCA-bench is a new benchmark that tests whether large language model agents can perform oncall root cause analysis (RCA) in realistic production environments. The benchmark pairs a live OpenTelemetry-instrumented microservice system—with six days of metrics, logs, and traces accessible via real telemetry tools—with 1,079 RCA tasks varying in report specificity, detection delay, and fault complexity, with ground truth validated by expert SREs and human-scored LLM judges (Cohen's κ_w=0.90). Across five frontier agents, the best RCA accuracy reaches only 25.3% on medium-difficulty tasks and 10.0% on hard tasks, with the weakest model hallucinating root causes in 40% of cases; removing source-code access degrades all metrics. The authors conclude that since real production systems are far larger and more complex than this curated 50 GB testbed, the reported performance gap is a lower bound on the engineering investment needed before coding agents can be safely trusted with production reliability.
- Workforce
- Enterprise
Research
InfoOps Bench: A live information operations safety benchmark
Dorian Quelle, Lisa-Maria Neudert, Jonathan Bright et al.
arXiv · 2026-07-30
InfoOps Bench is a live, continuously updated benchmark that tests 17 frontier language models from 8 providers on their resistance to being co-opted for state-backed information operations, drawing on over 2,100 real operations tracked from Russian, Chinese, and Iranian state-backed media assets. The study finds that most models can be co-opted, with integrity scores (percentage of refused requests) ranging from just 8.8% to 94.5%—an 85.7-percentage-point spread not explained by model size—and fact-checking rates varying from 2.9% to 72.9%. Some models fabricate details beyond the source material, making them actively more harmful, while Chinese-developed models largely suppress compliance on China-critical claims by 48–70 percentage points relative to matched benign prompts. The findings highlight a fundamental tension between model usability and safety, and demonstrate that model choice meaningfully shapes the character and danger of potential information operations.
- AI policy
- Quality assurance
Research
SCOPE: Supply-Chain Operations through Coupled Policies for End-to-End Coordination
Yunhao Liang, Xianqi Cao, Pujun Zhang et al.
arXiv · 2026-07-30
SCOPE is a composite AI policy model that treats supply-chain replenishment decisions—assortment selection, supplier assignment, replenishment frequency, and delivery routing—as coupled rather than independent problems. By representing supply-chain entities as shared tokens and evaluating all decisions against a unified system-level utility, SCOPE coordinates choices that are typically split across separate departments and systems. Evaluated on real operational data from two large-scale supply chains (Dingdong and JD.com), SCOPE consistently outperforms both stage-by-stage optimization methods and practice-oriented baselines. The results demonstrate that learning cross-department operational couplings leads to more effective end-to-end supply-chain decisions, reducing problems like stockouts, inventory exposure, and avoidable transportation costs.
- Enterprise