News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
5608 items
Research
Manufactured Divisiveness: Decomposing the Hostile Content of Seven Social Media Influence Operations
Emilio Ferrara
arXiv · 2026-07-16
This paper challenges the common claim that state-backed social media influence operations are major sources of 'hate' content, arguing that widely used detectors conflate true hate speech with partisan attacks and geopolitical rhetoric—a measurement error that inflates hate rates roughly twofold. Analyzing 25 million tweets from seven government-attributed campaigns (8,275 accounts) in the Twitter Information Operations archive, the researchers use an LLM-based detector and an auditable classification rule to split hostile content into identity-based attacks (50.1%), partisan attacks (30.4%), and state/foreign-policy invective (19.5%), finding that only 18.7% meets a stricter definition of dehumanizing or inciting hate. The study further shows that six of seven campaigns fall into three distinct regimes—identity hate (Russian operations), geopolitical invective (Iranian operations), and partisan divisiveness (Venezuelan operations)—which a single aggregate 'hate' metric obscures. These findings have direct implications for how platforms, policymakers, and researchers measure and report on online influence operations and hate speech.
- AI policy
- Quality assurance
Research
Lower-Resource, Higher Scores: Language Bias in LLM Evaluators
Ej Zhou, Lucas Resck, Zheng Hui et al.
arXiv · 2026-07-16
This paper demonstrates that LLM-based evaluators — including trained reward models and LLM-as-a-Judge systems — exhibit a systematic language bias: lower-resource languages receive significantly higher scores than higher-resource ones, even when the instruction-response content is semantically identical. The bias is statistically significant and consistent across eight open-weight evaluators and frontier judges, yet remains invisible to standard pairwise accuracy metrics, with evaluators achieving above 90% pairwise accuracy while showing up to a 43% difference in acceptance rates across languages. This has direct safety implications: harmful content in lower-resource languages is more likely to pass safety filters. The authors find that model uncertainty is linked to the effect but cannot fully explain it, pointing to a structural, language-level misalignment in how these evaluators score multilingual content.
- Quality assurance
- AI policy
- Certifications
Research
BioTIER: A Refusal Benchmark for Targeted Biological Risk Mitigation
Eleanor M. Marshall, Pedro Medeiros, Peter Peneder et al.
arXiv · 2026-07-16
BioTIER is a benchmark designed to improve how large language models handle biological safety by distinguishing genuinely high-risk information from legitimate scientific content. It organizes 542 expert-curated prompts into three risk tiers — Catastrophe Avoidance, Biomedical DURC, and Related Biology — spanning a spectrum from extremely narrow high-risk topics to broad, benign biological knowledge. The benchmark aims to help developers implement more targeted refusal policies that block the small fraction of information posing catastrophic misuse risk while preserving access to beneficial scientific knowledge. This matters for AI policy and quality assurance because current models either over-refuse benign content or freely provide dangerous information, both of which represent failures in targeted mitigation.
- AI policy
- Quality assurance
- Certifications
Research
Beyond Generalist LLMs: Specialist Agentic Systems for Structured Code Workflow Execution
Harris Borman, Herman Wandabwa, Fusun Yu et al.
arXiv · 2026-07-16
This paper investigates whether specialist AI agents outperform general-purpose LLM-based agents for a specific software task: transforming Business Process Model and Notation (BPMN) diagrams into executable agentic workflows. The researchers introduce a specialist workflow system and benchmark it against generalist agents (Roo and Cline), finding that the specialist solution achieves 9–20 percentage points higher tool-use exactness, 2–4x better penalty-adjusted latency, 3x fewer tool-call errors, over 95% reduction in token generation cost, and eliminates repair iterations. The study also finds that generalist agents produce code inconsistently in both functionality and quality, raising concerns about their reliability in industrial settings. These results suggest that domain-specific agentic systems offer meaningful advantages for enterprise software automation tasks where reliability and maintainability are critical.
- Enterprise
- Quality assurance
- Workforce
Research
Do Generative AI Assistants Respect robots.txt? Tracing Web Access Beyond Visible Answers
Gabriel Lopez-Fonseca, David Rodriguez, Stefan Bechtold et al.
arXiv (Cornell University) · 2026-07-16
This paper presents a controlled empirical study of ten AI assistants with web-search capabilities, examining whether they comply with robots.txt restrictions that website owners use to control automated access. Using server-side logs and secret codes embedded in target pages, the researchers tested four access conditions across 200 trials, finding substantial variation: some assistants followed expected access rules, while others retrieved restricted content without checking robots.txt or used generic user-agents that obscured attribution. The study also reveals that retrieval behavior and answer correctness can diverge—assistants may access pages without surfacing the content, or fail to retrieve even permitted resources. These findings raise legal and governance concerns about content owner rights and call for updated, enforceable web governance standards in the era of search-augmented AI.
- AI policy
- Enterprise
Research
Tactile: Giving Computer-Using Agents Hands and Feet
Yong Liu, Zhenyi Zhong, Zhanpeng Shi
arXiv · 2026-07-16
Tactile is an open-source tool layer designed to give computer-using AI agents more reliable control over desktop applications. Instead of the brittle approach of predicting pixel coordinates from screenshots, Tactile converts UI evidence—including OS accessibility semantics, OCR-grounded text, and visual fallback regions—into structured, verifiable action targets. On macOSWorld-style tasks, adding Tactile improved Codex Success@100 from 41.1% to 50.0% overall and from 45.2% to 55.3% on accessibility-adapted tasks, with consistent gains across multiple agents including Codex, Claude Code, OpenCode, and Goose. The paper argues that reliable computer use requires not just stronger models but a reusable execution substrate that exposes software actions as semantic, verifiable, and auditable objects.
- Enterprise
- Quality assurance
Research
Cybersecurity Maturity and Risk Profiling of AI-Enabled Medical Sensors: A Cross-Manufacturer Comparative Analysis
Filip Tsvetanov
International Journal of Online and Biomedical Engineering (iJOE) · 2026-07-16
This paper proposes an integrated methodology for assessing cybersecurity maturity in AI-enabled medical sensors—including hospital patches, wearables, implantable cardiac devices, and continuous glucose monitors—using ISO 14971, NIST 800-30/53, and Analytic Hierarchy Process (AHP) weighting across ten criteria. The framework generates two composite indicators, a weighted security score and a risk profile score, enabling cross-manufacturer comparison. The analysis identifies critical vulnerabilities in communication security and AI modules of specific device categories, while finding greater cryptographic resilience in implantable systems. The findings are directly relevant to engineers, clinicians, and procurement organizations, and the authors call for unified AI-oriented security standards across the full medical sensor lifecycle.
- Certifications
- Quality assurance
Research
The role of artificial intelligence in detecting and preventing academic dishonesty in higher education: a systematic review
Walid Salamah
Frontiers in Education · 2026-07-16
This systematic review synthesizes 33 peer-reviewed studies (2012–2024) on how AI tools—including machine learning, NLP, stylometric analysis, and learning analytics—are being used to detect and prevent academic dishonesty in higher education. Four themes emerged: AI-based detection mechanisms, AI-based prevention strategies, ethical and institutional challenges (such as algorithmic bias, false positives, and data privacy), and critical gaps in evidence from developing and conflict-affected contexts. The review concludes that while AI shows significant potential for strengthening academic integrity, its effectiveness is constrained by bias, opacity, and variable institutional readiness, and that AI should complement rather than replace human judgment. The findings are directly relevant to quality-assurance and policy frameworks in higher education, particularly regarding governance and equitable deployment of AI integrity systems.
- Quality assurance
- AI policy
Research
Potencial de automatización laboral del modelo generativo de inteligencia artificial en las ocupaciones de México
María del Mar Oviedo Facundo, Luís Huesca Reynoso, David Castro Lugo
Revista de Economía Facultad de Economía Universidad Autónoma de Yucatán · 2026-07-16
This paper introduces the GPT Multimodal Automation Indicator (GMI/IGAM), a task-based metric designed to measure how much generative AI models like GPT-5 could automate occupations in Mexico. Applying the indicator to Mexico's 2019 SINCO occupational data using World Economic Forum methodology, the study finds that managerial, administrative, and sales divisions face low-to-medium automation potential (28–38%), while agricultural and services divisions face lower exposure (11–17%). Critically, some individual tasks show automation potential as high as 88%, meaning that even within low-risk occupations, specific tasks carry significant displacement risk. The findings suggest GPT-5 primarily complements rather than replaces occupations, but the concentration of highly automatable tasks within certain roles raises meaningful job displacement concerns for Mexico's workforce.
- Workforce
- AI policy
Research
Cheaper AI, More Informality? A Dual Labor Market Model for Developing Economies
Gabriel Montes‐Rojas, Fernando Toledo, Juan Manuel Rodríguez Repeti
arXiv (Cornell University) · 2026-07-16
This paper builds a small open economy DSGE model with a dual labor market to examine how falling AI prices affect formal versus informal employment in developing economies. The key finding is that outcomes hinge on whether imported AI capital substitutes for or complements formal workers: under substitution, cheaper AI erodes formal labor demand and pushes workers into informality, while under complementarity it expands formal employment, wages, and output. The model highlights that the same technological trend can either displace workers or drive formal-sector growth depending on the structural relationship between AI and human labor.
- Workforce
- AI policy
Research
Recent Advances in AI for Automated ICD Coding: A Systematic Literature Review
Abdul Rehman Khalid, Haider Ali, Kounen Fathima et al.
Journal of Medical Systems · 2026-07-16
This systematic review of 54 studies (drawn from 4,280 citations, 2019–2024) examines AI approaches to automating ICD code assignment from clinical text such as discharge summaries and electronic health records. The review traces a clear evolution from traditional machine learning to deep learning architectures—including convolutional, recurrent, transformer, and hybrid models—and finds that models perform better on frequent codes than on rare ones. Key persistent gaps include overreliance on single-language, single-institution datasets, poor rare-code prediction, limited model interpretability, and inconsistent evaluation protocols that impede cross-study comparison. The authors propose a 5P research agenda emphasizing Population Diversity, Performance Robustness, Prediction of Rare Codes, Provenance Transparency, and Practical Integration to guide deployment in real-world healthcare systems.
- Quality assurance
- Enterprise
Research
Can model attribution bridge AI's accountability gap in safety-critical domains?
Marc Juárez
Philosophical Transactions of the Royal Society A Mathematical Physical and Engineering Sciences · 2026-07-16
This paper examines 'model attribution'—a property intended to identify which specific ML model a user is interacting with when models are deployed as remote services—as a mechanism for closing AI accountability gaps in safety-critical domains. The authors review recent technical approaches to model attribution and find that limitations including computational costs, latency overheads, and insufficient guarantees make current solutions inadequate for safety-critical applications. They conclude that while model attribution is a meaningful step toward accountability, the field has not yet produced tools that meet the rigorous requirements of high-stakes settings. The work is relevant to ongoing efforts to establish robust governance and oversight mechanisms for AI in safety-critical systems.
- AI policy
- Quality assurance
Research
A STATISTICAL ANALYSIS OF THE SOCIAL IMPACT OF AI ON JOB DISPLACEMENT & REPLACEMENT IN SOUTH ASIA & AFRICA
Dr NR Jagannath
International Journal of Creative and Open Research in Engineering and Management · 2026-07-16
This paper statistically analyzes how AI-driven job displacement and replacement are unfolding across South Asian and African labour markets, drawing on data from the World Bank, IMF, ILO, WEF, and industry sources. It finds that formal, export-oriented sectors—IT-BPO, ready-made garments, customer support, and banking—face the highest automation exposure, with notable job losses already observed in Bangladesh and high susceptibility projected for India's IT-BPO workforce by 2030. Distributional impacts are pronounced: women in export manufacturing and routine services, youth, and lower-educated workers bear disproportionate displacement risk, while skills gaps and digital-infrastructure deficits limit access to emerging AI-adjacent roles. The paper concludes that net job creation is possible but conditional on reskilling investment, digital education, labour-market adjustment supports, and robust disaggregated monitoring systems.
- Workforce
- AI policy
Research
The Governance Reconstruction of the AI Era: From "Entering Society" to "Remaining in Society" — Establishing Occupational Transition and Retraining as a Core Public Service of Government
Leyi Zhang
Open MIND · 2026-07-16
This policy article argues that AI-driven occupational disruption—accelerated by generative AI and embodied robotics—is compressing job life cycles faster than prior waves of automation, undermining the historical assumption that technological unemployment is merely transitional. Drawing on evidence from the WEF Future of Jobs Report 2025, ILO 2025 assessments, and IBM reskilling projections, the author contends that occupational retraining must be elevated to a core state public service comparable to compulsory education. The proposed framework grounds government obligation in Rawlsian, Dworkinian, and capability-approach theories, identifies market failures justifying intervention, and recommends a funding mechanism financed by AI-driven fiscal dividends rather than automation taxes.
- Workforce
- AI policy
Research
Does the Responsible AI Guidance Adequately Protect Māori in Predictive Workplace Safety Regulation?
Nicola Knobel
New Zealand journal of health and safety practice. · 2026-07-16
This paper evaluates whether New Zealand's Responsible AI Guidance for the Public Service adequately protects Māori in predictive workplace safety regulation. Applying Critical Tiriti Analysis, responsive regulation theory, and post-colonial legal analysis, the paper finds the Guidance scores either Silent or Poor across all five Treaty of Waitangi criteria, lacking enforceable protections, Māori governance structures, and procedural fairness mechanisms. Drawing on case studies including the Woolworths warehouse strike and WorkSafe New Zealand's use of AI-enabled tools, the authors argue that reliance on soft-law principles is insufficient. The paper calls for reforms embedding co-governance, Māori Data Sovereignty, and legally mandated safeguards into public sector AI infrastructure.
- AI policy
- Workforce
Research
A Meta-synthesis of nursing students’ experiences with generative artificial intelligence-assisted learning
Siyi Du, Sha Wang, Fengxia Yan et al.
Frontiers in Medicine · 2026-07-16
This meta-synthesis of 18 qualitative studies examines nursing students' real-world experiences with generative AI-assisted learning, identifying 49 themes organized into four integrated findings: a dual experience of empowerment and challenges, internal conflict between technology and nursing humanism, differentiated user experiences amid practical constraints, and a general demand for supporting systems and educational reform. The findings indicate that nursing educators need to strengthen students' AI literacy, critical thinking, and ethical awareness, while curriculum designers should embed AI competencies into nursing programs without sacrificing humanistic care values. The study also calls on policymakers to establish clear governance frameworks and educational guidelines for responsible AI use in nursing education.
- Workforce
- AI policy
Research
Regulating Reality: Exploring Synthetic Media Through Multistakeholder AI Governance
Claire R. Leibowicz
Policy & Internet · 2026-07-16
This paper examines how stakeholders from civil society, industry, media, and policy sectors conceptualize and implement governance of AI-generated (synthetic) media. Drawing on 23 semi-structured interviews and real-world cases, the study finds that temporal perspectives and trust—both among stakeholders and between audiences and interventions—are central to effective governance, and that technical measures like AI labels have notable limitations. The findings are intended to inform both scholarly understanding of multistakeholder AI governance and the practical design of synthetic media policy.
- AI policy
Research
LLMs in Medical Education for Autism Caregivers: A Comparative Evaluation of Accuracy, Readability, Actionability, and Neurodiversity-Affirming Language
Shahid Akhtar Akhund, Asma Alsaleh, Naheed Haroon Kazi et al.
Healthcare · 2026-07-16
This study systematically compared three large language models—Google Gemini 1.5 Pro, ChatGPT (GPT-4o), and DeepSeek-V3—as health information tools for caregivers of children with autism spectrum disorder in Saudi Arabia. Across 24 clinically validated questions rated by expert raters, Gemini was most scientifically accurate, ChatGPT most readable, and DeepSeek most actionable, yet all three models fell below established AHRQ benchmarks for actionability and exceeded recommended readability grade levels for patient education. Accuracy gaps clustered around regional epidemiological, genetic risk, and financial queries, and no model demonstrated meaningfully neurodiversity-affirming language. The findings indicate that LLMs in their current form require substantial plain-language adaptation and cultural customization before they can reliably serve as caregiver education tools, and clinicians should actively guide families in interpreting LLM outputs.
- Workforce
- Quality assurance
Research
Conformity assessment procedures under the AI Act
Mattis Jacobs, Susanne Kuch, Dominic Deuber et al.
Technology and Regulation · 2026-07-16
This article critically analyzes the conformity assessment procedures mandated by the EU AI Act for high-risk AI systems. It argues that the Act's reliance on conformity assessment procedures borrowed from other EU harmonization legislation produces a risk classification that inadequately accounts for AI-specific risks and raises doubts about the suitability of those procedures for AI systems. The analysis also examines how procedure assignments affect the requirements imposed on conformity assessment bodies, identifying systemic regulatory blind spots in the AI Act's framework.
- Certifications
- AI policy
Research
Helping People Choose Careers in the Age of AI
Jennifer Steele, Isabella Cruz
arXiv (Cornell University) · 2026-07-16
This paper helps individuals make informed career decisions by analyzing how much different occupations are exposed to AI-driven task automation. The authors compare six existing projection models and develop a new empirical model using 2025 query data from Anthropic and OpenAI, finding significant variation across models but a general trend since 2020 where higher AI exposure correlates with higher salaries and job complexity. By averaging five models, they identify career tradeoffs across job categories, finding healthcare practice offers the best balance of higher pay and lower AI exposure. Jobs where AI acts as a complement rather than a substitute for human labor tend to be modestly better-paid, with implications depending on how fields adopt AI usage norms.
- Workforce
- AI policy
Research
The Most Exposed Sector Meets the Shock: AI Exposure and Firm-Level Labor Outcomes in Indian IT Services
Mihir Khanna
arXiv · 2026-07-16
This paper examines how AI exposure affected hiring and productivity at Indian IT-BPM firms around two technology shocks: the 2015-16 cloud/deep-learning wave and the 2022 LLM shock. Using a panel of 13 publicly listed IT-services firms and occupational AI-exposure scores, the author finds that after 2022, more-exposed firms significantly slowed net hiring (annual net hiring fell from 8.3% to 3.1% for high-exposure firms) while raising labor productivity, consistent with augmentation rather than outright displacement. The 2016 wave showed no comparable effect, and no significant hiring slowdown appeared for lower-AI-exposure engineering-R&D firms. The study provides the first firm-level evidence of LLM adoption effects in the global sector most concentrated in AI-exposed roles.
- Workforce
- Enterprise
Research
The Governance Reconstruction of the AI Era: From "Entering Society" to "Remaining in Society" — Establishing Occupational Transition and Retraining as a Core Public Service of Government
Leyi Zhang
Knowledge Commons (Lakehead University) · 2026-07-16
This policy paper argues that AI-driven occupational displacement—accelerated by generative AI and embodied robotics—is occurring at a scale and speed that invalidates the historical assumption that automation merely transitions workers rather than permanently displacing them. Drawing on empirical sources including the WEF Future of Jobs Report 2025, ILO assessments, and IBM reskilling projections, the author contends that governments should elevate occupational retraining from a marginal labor-market tool to a core public service comparable to compulsory education or public security. The proposed framework grounds this obligation in Rawlsian, Dworkinian, and Capability-approach ethics, identifies specific market failures justifying intervention, and recommends a funding mechanism financed by AI-driven fiscal dividends rather than automation taxes. The paper concludes that the state's historic role of helping citizens 'enter society' must be extended to helping them 'remain in society' through lifelong transition support.
- Workforce
- AI policy
Research
Context Contamination in LLM Analysis of Network Security Logs: Poison with Passive Prompt Injection and Mitigation Evaluation
Rabimba Karanjai, Ye Lu, Hemanth Hegadehalli Madhavarao et al.
arXiv (Cornell University) · 2026-07-16
This paper demonstrates that LLMs deployed in Security Operations Centers for log analysis are vulnerable to 'passive prompt injection,' where adversaries embed malicious instructions in network log fields that are later executed when analysts query the model. The authors introduce LogInject, a benchmark framework with nearly 13,000 log entries, and show attack success rates up to 88.2% across objectives like activity concealment, false positive generation, and output hijacking. A novel 'Context Stitching' technique fragments payloads across multiple log entries to evade filters while still achieving a 76.4% success rate. Layered defenses reduce attacks by 90.4% but leave an 8.4% residual vulnerability, underscoring the need for defense-in-depth and continued human oversight in security-critical LLM deployments.
- Quality assurance
- AI policy
Research
A study on the influencing factors of perceived artificial intelligence substitution risk
Xianghui Xing, Hongwei Dai, Yiwei Zhou et al.
Acta Psychologica · 2026-07-16
Using data from China's 2023 Social Survey, this study empirically identifies the key factors shaping workers' perceived risk of being replaced by AI. Results show that perceived unemployment risk, number of children, and unemployment insurance positively correlate with perceived AI substitution risk, while female gender, older age, higher job satisfaction, greater job skill level, and higher perceived socioeconomic status are negatively associated. Mediation analysis reveals that job satisfaction and perceived socioeconomic status reduce substitution fears partly by lowering perceived unemployment risk, and heterogeneity analysis shows these effects differ significantly across regions, urban-rural contexts, and work patterns. The findings offer empirical grounding for policies aimed at reducing AI-driven employment anxiety and improving workforce training systems.
- Workforce
- AI policy
Research
The Enterprise AI Governance Buyer's Guide
FERZ Inc., Edward Meyman
Zenodo (CERN European Organization for Nuclear Research) · 2026-07-16
This document presents Version 3.4 of the Enterprise AI Governance Buyer's Guide alongside a companion Procurement Fast Path, offering a vendor-neutral framework for evaluating AI governance claims in regulated enterprise environments. The framework formalizes a distinction between probabilistic and deterministic governance, centers on five canonical tests (Stop, Ownership, Replay, Escalation, and Provenance), and incorporates an Authorization Boundary Integrity Model to locate where a vendor's controls hold or fail. It maps governance requirements to major regulatory regimes including the EU AI Act, GDPR, HIPAA, DFARS, and NIST AI RMF, and is designed to help procurement teams, auditors, and risk officers obtain independently verifiable, tamper-evident evidence that AI decisions were authorized before execution. The guide matters because it provides structured, evidence-based due-diligence tooling for organizations that must demonstrate defensible AI governance in high-stakes sectors such as healthcare, financial services, defense, and government.
- Enterprise
- Quality assurance
- Certifications
- AI policy