News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
Generative Artificial Intelligence Productivity in Indonesian Micro, Small, and Medium Enterprises
Murniati, Sugeng Rianto, Ayman Ghazi Taher Nazzal et al.
Signifikan Jurnal Ilmu Ekonomi · 2026-08-07
This study surveyed 200 Indonesian micro, small, and medium enterprises (MSMEs) using PLS-SEM to examine how perceived usefulness, technological readiness, and organizational support affect generative AI adoption intensity and productivity. Perceived usefulness and technological readiness significantly increased GenAI adoption intensity, while perceived usefulness and organizational support directly boosted MSME productivity; however, GenAI adoption intensity and technological readiness alone did not show a direct productivity effect. The findings suggest that MSME development programs should prioritize AI literacy, workflow integration, and managerial support to realize measurable productivity gains from generative AI.
- Enterprise
- Workforce
Research
Poisoning Attacks in Federated Learning: An Accountability- Oriented Survey with Centralized Learning as a Baseline
Safiia Mohammed, Dima Alhadidi, Alioune Ngom
Journal of Cybersecurity and Privacy · 2026-08-07
This survey examines poisoning attacks in federated learning (FL)—including data poisoning, model poisoning, backdoor insertion, and Sybil behavior—using centralized learning as a baseline to explain how distributed architectures expand the attack surface. It evaluates countermeasures such as Byzantine-robust aggregation, anomaly detection, and verifiable aggregation protocols through an accountability-oriented lens focused on attribution, audit evidence, and forensic readiness. The key finding is that robustness alone is insufficient for trustworthy FL; defenses must also preserve evidence supporting independent verification, post-incident reconstruction, and governance review. The authors conclude that accountable FL systems should be designed as evidence-producing architectures, particularly for regulated and high-risk deployments.
- Quality assurance
- AI policy
Research
Artificial Intelligence in Higher Education: A Critical Synthesis of Student Learning, Assessment Transformation, and Academic Integrity
Ameen Talib
Asian Journal of Education and Social Studies · 2026-08-07
This critical narrative review synthesizes literature from 2019–2026 on how generative AI is reshaping student learning, assessment, and academic integrity in higher education. The evidence shows AI can expand access to feedback, explanation, and language support with some controlled studies reporting learning benefits, but effects are heterogeneous, often short-term, and highly sensitive to task design and human oversight. The review identifies assessment as the central institutional challenge, noting that automated plagiarism detection is unreliable and risks false accusation, while recommending an aligned model combining AI literacy, process-rich assessment, disclosure rules, and human judgment. The authors conclude that AI's educational legitimacy depends on institutions preserving epistemic agency, valid learning measurement, and equitable access.
- Quality assurance
- AI policy
Research
A mixed-methods analysis of the role of AI chatbots in research amongst university students in Nigeria: a case study of Delta State University, Abraka, 2020–2024
Uwomano Benjamin Okpevra, Esther Abasa
Research in Learning Technology · 2026-08-07
A mixed-methods study of 320 students and 27 faculty at Delta State University (Nigeria) finds that 78% of students use AI chatbots to brainstorm, summarize literature, and structure assignments, boosting productivity, yet over 60% admitted submitting AI-generated content without verification or attribution, raising serious academic integrity concerns. The study also highlights that Western-centric AI design limits cultural relevance and risks marginalizing indigenous Nigerian knowledge systems. The authors recommend institutional AI literacy programs, transparent usage policies, localized AI development, and alternative assessment strategies to ensure AI complements rather than replaces human reasoning in scholarly work.
- AI policy
- Quality assurance
Research
LLMs’ reshaping of people, processes, products, and society in software development: a qualitative exploration with early adopters
Benyamin T. Tabarsi, Heidi Reichert, Stephen D. Gilson et al.
Empirical Software Engineering · 2026-08-07
This qualitative study interviews sixteen early-adopter software professionals to examine how LLM-based tools reshape software development across people, processes, products, and society. Developers reported productivity gains from reduced mundane tasks, faster debugging, and streamlined search, but encountered a productivity-quality paradox where generated code was frequently discarded and effort shifted toward critical evaluation rather than writing. LLM adoption was highly phase-dependent—strong in implementation and debugging but weak in requirements gathering—and participants anticipated changes in hiring expectations, team practices, and computing education while stressing that human judgment and foundational skills remain essential. The findings offer actionable implications for developers, organizations, educators, and tool designers integrating LLMs into professional software practice.
- Workforce
- Enterprise
Research
AI failures in the eyes of the downstream developer: a first look at concerns, practices, and challenges
Haoyu Gao, Mansooreh Zahedi, Wenxin Jiang et al.
Empirical Software Engineering · 2026-08-07
This mixed-method study examines how downstream developers handle AI failure modes—such as data leakage and model bias—when building software that reuses pre-trained models. Drawing on interviews with 16 participants, a survey of 86 practitioners, and analysis of 874 AI incidents from the AI Incident Database, the study finds that while developers generally have strong awareness of potential failures, their actual practices during preparation and model selection are often inadequate. Key barriers include a lack of concrete guidelines and policies, poor documentation, and knowledge gaps, leading to highly variable mitigation practices across the development lifecycle. The authors offer recommendations for model contributors, developers, researchers, and policymakers to improve failure mitigation and reduce real-world harms.
- AI policy
- Quality assurance
Research
Vigilant Algorithms and the Contagion of Risk in Automated Opioid Risk Assessment
Amelie Lange, Benjamin Lipp
Sociology of Health & Illness · 2026-08-07
This ethnographic study examines NarxCare, a widely used algorithmic opioid risk-scoring system embedded in U.S. prescription monitoring programs. Through clinician interviews, interface walkthroughs, documentary analysis, and clinic observation, the authors find that NarxCare operates through multiple opacities—presenting scores as authoritative yet unexaminable—and collapses distinct concerns (overdose risk, liability, societal harm) into a single number. A key finding is a 'contagion of risk' logic whereby physicians themselves become risk subjects, their professional reputations tied to the scores of patients they treat. Rather than resolving opioid risk, the system reconfigures clinical authority and accountability, with clinicians responding by resisting, anticipating, or domesticating scores but unable to evade them entirely.
- Workforce
- AI policy
Research
Leveraging Artificial Intelligence and Natural Language Processing in Legal Epidemiology Studies: Opportunities and Challenges
Regen Weber-Fares, Fallon Julia Cochlin, Snigdha Peddireddy et al.
The Journal of Law Medicine & Ethics · 2026-08-07
This review paper examines how artificial intelligence and natural language processing (AI/NLP) can address a key bottleneck in legal epidemiology—the shortage of specialists able to conduct time-intensive analyses of how laws influence health outcomes. The authors argue that laws' standardized formats and terminology make them well-suited for AI/NLP, and they outline opportunities and challenges across tasks such as scoping legal documents, collecting primary legal data, developing coding schemes, and implementing quality controls. The paper identifies methodologic reporting gaps and proposes a research agenda to accelerate evidence-based policy work in this field.
- AI policy
- Quality assurance
Research
Lifelong learning in an AI-driven world: assistance, personalization and automation under scrutiny
Gerardo Alfredo Rodríguez, Andrés Chiappe, Fabiola Sáez-Delgado
Frontiers in Education · 2026-08-07
This scoping review of 189 articles examines whether AI's three headline promises—assistance, personalization, and automation—align with the equity-oriented and socially grounded goals of lifelong learning (LLL). Using systematic search procedures and statistical association tests, the authors find that AI-related literature shows weak alignment with LLL paradigms, regional contexts, and learner vulnerability profiles, even though personalization is the most prominently featured AI promise. The review concludes that current AI deployment in lifelong education tends to be technology-centred rather than connected to social justice, human development, or the needs of vulnerable populations, and calls for an analytical framework to better align AI initiatives with inclusive and transformative LLL goals.
- Workforce
- AI policy
Research
Artificial Intelligence in Journalism and Media Practice: A Systematic Review of Applications, Limitations, and Research Approaches
Jun Gui, Nasrullah Dharejo, Mumtaz Aini Alivi
Journalism and Media · 2026-08-07
This systematic review of 121 peer-reviewed articles (2020–2026) maps the state of AI and machine learning research in journalism and media across four domains: news production and automation (38%), audience perception and content analysis (24.8%), ethical and legal considerations (19.8%), and meta-research and implementation studies (17.4%). Key findings show that successful AI adoption requires workflow redesign and human–machine collaboration rather than full automation, that audience responses to AI-generated content vary by cultural context, and that copyright frameworks for AI-generated news remain legally contested across jurisdictions. The review develops an integrated theoretical framework identifying cultural context and disclosure practices as key moderators, and notes that legal-regulatory concerns pervade the entire literature. These findings matter for understanding how AI is reshaping professional journalism practice, governance, and the regulatory environment surrounding automated content.
- Workforce
- AI policy
Research
Parliamentary discourse and legislative approaches to artificial intelligence regulation: a comparative analysis of five jurisdictions
Elnur Beisenbayev, Bakhytzhan Bukharbayev, Zhengisbek Tolen et al.
Frontiers in Political Science · 2026-08-07
This comparative study examines how parliamentary debates on artificial intelligence translate into regulatory frameworks across five jurisdictions: the EU, US, Brazil, Kazakhstan, and the UK. Using critical discourse analysis and securitization theory, the authors identify distinct semantic cores shaping AI governance in each context—such as 'risk–fundamental rights' in the EU and 'dependency lock-in–delayed reflection' in the UK. The paper proposes an updated five-part typology of AI regulation and introduces concepts like 'sovereignty through partnership' and 'delayed reflection' as novel regulatory patterns. The findings matter for understanding how political discourse shapes divergent national approaches to AI policy and governance.
- AI policy
Research
Bias Analysis of L2 Speaking Assessment Systems Using Concept Activation Vectors
Arya Labroo, Mengjie Qian, Kate Knill
arXiv · 2026-08-06
This paper investigates fairness and bias in automated second-language (L2) speaking assessment systems by applying Concept Activation Vectors (CAVs) to two neural graders — a text-based BERT model and a multimodal Whisper-based model. The researchers probe whether irrelevant speaker attributes like first language (L1) or age are encoded in model representations and whether they influence predicted scores, using a gradient-based sensitivity metric. They also explore sparse autoencoders (SAEs) as a way to improve concept linearity, finding that SAEs make concepts more recoverable but reduce activation-space sensitivity, particularly in low-dimensional layers. The key takeaway is that concept recoverability and concept influence are distinct properties that both need to be examined when auditing bias in high-stakes speaking assessment systems.
- Quality assurance
- Certifications
Research
The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images
Zhiheng Wang, Bo Peng, Lai Wei et al.
arXiv · 2026-08-06
This paper investigates whether the visual evidence retrieved by multimodal LLMs using 'thinking-with-images' operations—such as crop-and-zoom—actually causes those models to produce better answers. Using a causal graph framework and three levels of intervention (policy, trajectory, and step), the authors introduce a metric called Visual Evidence Gain to isolate each observation's causal contribution. Across six models and five perception benchmarks, they find two systematic failure modes: 'Calling Without Looking' (retrieved observations have no causal effect) and 'Looking Without Planning' (observations are informative but calls are incoherently scheduled). The result is what they call the 'illusion of visual tool-use'—aggregate accuracy gains mask the fact that visual tool-use is causally ineffective across most rollouts, raising serious quality concerns about this paradigm's reliability.
- Quality assurance
Research
Improving the Realism of Synthetic Clinical Benchmarks Under Utility Constraints
Omid Bazgir, Md Nasir, Jacob Hoffman et al.
arXiv · 2026-08-06
This paper addresses a critical gap in enterprise AI evaluation: synthetic clinical benchmarks used to test healthcare AI agents can pass utility checks while remaining structurally unrealistic. The authors formulate 'utility-constrained realism improvement' — a framework for revising benchmarks to increase realism (measured via missingness structure, simplicity, structural plausibility, and population alignment) without falling below an operational utility floor. Applied to a care-gap benchmark built from Synthea-generated patients, they show the baseline benchmark is severely thin (79.44% sampled-pair missingness, 100% top-three token concentration), and that two deterministic revisions improve realism while maintaining utility, whereas naive densification preserves unrealistic templating. The findings argue that synthetic benchmark quality must be explicitly optimized, with utility treated as a constraint rather than as sufficient evidence of realism — a direct implication for how healthcare AI agents are evaluated before deployment.
- Enterprise
- Quality assurance
Research
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations)
Ro Encarnación, Tina Behzad, Emma Lurie et al.
arXiv · 2026-08-06
This paper audits common assumptions in large language model (LLM) benchmark evaluations by comparing two access modalities—ChatGPT's chat UI and OpenAI's API—with and without web search enabled, across 4,812 responses to 401 prompts from the BBQ and SafetyBench benchmarks. The study finds that enabling web search reduced accuracy by up to 8 percentage points, chat UI responses were less accurate than API responses when search was disabled, and repeated runs of the same prompt produced inconsistent responses in up to 21% of cases. The two modalities also grounded answers in different citations and showed inconsistent abstention behavior. These results demonstrate that reporting only simple accuracy metrics can obscure meaningful behavioral variation relevant to safety assessments, and the authors argue that evaluations should systematically account for modality, multi-run consistency, search conditions, and response-level behaviors to better reflect real-world AI deployment.
- Quality assurance
- AI policy
Research
Reducing belief in conspiracy theories as they unfold using large language models
Thomas H. Costello, Nathaniel Rabb, Michael Nicholas Stagnaro et al.
arXiv · 2026-08-06
This paper tests whether conversations with a large language model (LLM) can reduce belief in conspiracy theories as they emerge in real time. In two experiments conducted just after major crisis events (the July 2024 assassination attempt on Donald Trump and the September 2025 assassination of Charlie Kirk), U.S. adults with conspiratorial views engaged in multi-turn LLM conversations designed to lower those beliefs; compared to control groups, LLM-treated participants showed significantly reduced conspiracy belief. The study also found downstream effects, with treated participants showing lower belief in unrelated conspiracies one to two months later. The findings highlight the potential of scalable, cognitively-focused LLM interventions to counteract misinformation following high-profile societal events.
- AI policy
Research
Poli-Bias: Understanding and Measuring Large Language Model Biases in International Political Conflicts
Massi-Nissa Abboud, Aladin Djuhera, Elena Cabrio et al.
arXiv · 2026-08-06
This paper introduces Poli-Bias, a counterfactual evaluation framework designed to detect political bias in large language models (LLMs) by systematically swapping country identities in legally equivalent conflict scenarios. The framework decomposes response disparities into five interpretable dimensions—covering framing, argumentation, and legal reasoning—rather than reducing bias to a single metric. Testing across 13 contemporary LLMs, the authors find that country identities and user affiliations can systematically skew how equivalent actions are described, evaluated, and defended under international law. These findings matter for policy and quality-assurance efforts, as they reveal measurable unevenness and sycophancy in how AI systems handle sensitive geopolitical topics.
- Quality assurance
- AI policy
Research
Does Latent Context Help? A Controlled Evaluation of Inverse Reinforcement Learning in Arctic Shipping
Vaishnav Vaidheeswaran, Dilith Jayakody, Biruk Ambaw et al.
arXiv (Cornell University) · 2026-08-06
This paper evaluates inverse reinforcement learning (IRL) methods for modeling vessel behavior in Arctic shipping, using 3,186 AIS-derived voyages from 202 vessels across nine Arctic seasons. It finds that a nonlinear shared reward model improves held-out likelihood by 50.9% over a linear baseline, but adding vessel-specific latent context actually reduces performance by 16.5%, because apparent behavioral variation is largely explained by observable route and environmental features rather than hidden vessel-specific preferences. The study also shows that different evaluation metrics—predictive accuracy, route fidelity, and reward transfer—produce different model rankings, meaning no single metric suffices for assessing learned rewards. These findings support a more rigorous, transparent approach to AI-assisted navigation in safety-critical Arctic environments.
- Quality assurance
- AI policy
Research
FormBharo: Designing and Evaluating a Voice Agent for Conversational Form Filling in Rural India
Aman Dalmia, Sanskriti Midha, Jigar Doshi
arXiv · 2026-08-06
FormBharo is a voice agent designed to help low-literacy, Hindi-speaking rural mothers in India enroll in maternal and child health programs over a phone call, removing the burden from overstretched frontline health workers. The system pairs Large Language Models with deterministic rule-based validation and flow control, and is being piloted with NGO ARMMAN for antenatal and postnatal care enrollment. The authors release FormVoiceAgentBench, a benchmark with 3,760 multi-turn conversation tests across 960 simulated calls, finding that real-speech transcription errors can drop form completion rates by up to ~41 points and that component-level performance does not reliably predict end-to-end performance. Because no single model dominates on accuracy, cost, and latency simultaneously, the team uses Pareto-based weighted-sum scalarization to select a deployable configuration.
- Workforce
- Enterprise
Research
HERALD: Counterfactual Audits and Minimal Repairs for Proof-of-Retrieval Rewards
Zhuowen Liu, Bohan Cui, YinShang Guo et al.
arXiv · 2026-08-06
HERALD is an offline auditing framework that tests whether search-agent reward functions genuinely verify that cited evidence was actually retrieved, rather than just producing high scores. By applying counterfactual interventions on four model pools across three multi-hop QA benchmarks, the authors show that a baseline reward rejects obvious attacks like search deletion and fake IDs, but remains vulnerable to a 'citation-laundering' attack where a model cites corpus passages it never retrieved. Adding a targeted penalty for citing passages absent from retrieved evidence (R[L]) closes this vulnerability with zero empirical attack success rate, while improving citation precision by 2.02 points and support recall by 1.46 points without reducing natural citation behavior. The work matters because it provides a systematic method for identifying and minimally repairing exploitable gaps in AI reward design, directly informing how search-augmented AI systems can be made more trustworthy and harder to game.
- Quality assurance
Research
The em-dash em-beds in Congress: A population-level rise in em-dash frequency in U.S. congressional press releases at the dawn of the large-language-model era, 2021-2025
Przemysław Czuma
arXiv · 2026-08-06
This study examines whether large language models (LLMs) leave a measurable stylistic fingerprint in U.S. congressional press releases by tracking the frequency of unspaced em-dashes, a punctuation style common in LLM-generated text but unusual in AP-style press writing. Analyzing 146,239 releases from 480 House and Senate offices between 2021 and 2025, the researchers found that unspaced em-dash density more than doubled in 2025 compared to the 2021–2024 baseline, and the share of releases containing such an em-dash rose from roughly 13% to nearly 25%. The rise was consistent across parties and chambers, held within individual offices over time, and survived multiple falsification tests, suggesting broad diffusion of LLM-assisted writing as models matured. The authors caution that the em-dash is a population-level marker rather than a per-release authorship detector, and the design supports no causal claim.
- AI policy
Research
Seeing Is Not Deciding: Can Multimodal LLMs Act as Effective CEOs?
Yuyang Dai, Xueqing Peng, Yuxia Wang et al.
arXiv · 2026-08-06
This paper introduces C-SUITEBENCH, a multimodal benchmark that places nine frontier AI models in the role of a chief executive officer across five decision tasks in 50 business scenarios, comparing text-only versus multimodal conditions. The study finds that adding visual business evidence generally improves evidence-centric reasoning—especially in risk forecasting and board-facing justification—but uncovers a 'multimodal integration paradox' where visual inputs consistently degrade constrained resource allocation performance across all nine models. Ablation experiments trace this failure to signal crowding, where individual visual channels each help but their combination disrupts constraint satisfaction during decoding. The findings show that visual perception and constrained action are separable bottlenecks, warning that indiscriminate visual augmentation can harm high-stakes AI decision-making and motivating more selective grounding strategies.
- Enterprise
- Quality assurance
Research
Validity, Reliability, and Transparency in Artificial Intelligence Regulation
A. Mukundan, Debayan Gupta, Subhashis Banerjee
arXiv (Cornell University) · 2026-08-06
This article argues that existing data protection frameworks—including the EU AI Act—fail to address a fundamental risk from AI systems: unreliable or unjustified inferences produced even when data collection is lawful. The authors identify specific epistemological problems (construct validity, confounding, representativeness, distribution shift, and fairness trade-offs) and propose a structured regulatory framework in which validity of inference must be established as a precondition for proportionality assessment and deployment approval. Drawing on the Indian Supreme Court's Puttaswamy judgment on informational self-determination, the framework extends constitutional privacy protections from data collection to the legitimacy of data use. The authors call for AI-specific validity assessment, post-deployment monitoring, and structured proportionality assessments that weigh epistemic risk against potential benefit.
- AI policy
- Certifications
Research
Activity Frames: Deterministic Screen-Activity Compilation for Agent Memory and Replay
Nossa Iyamu
arXiv (Cornell University) · 2026-08-06
This paper presents Activity Frames, a deterministic, zero-model pipeline that compiles passively captured screen activity into compact, structured memory blocks for computer-use AI agents. On a single professional's corpus of 128,756 frames over 51 active days, the system achieves an 86x compression of a day's raw capture into a prompt-ready context block in 68 ms, enabling an agent to answer questions about the day at 98.4% accuracy compared to 66–80% for LLM-based summaries. The pipeline also measures two previously unquantified parameters of agent cost models—the Routine Overhead Ratio (60–343x) and delegable routine recurrence (roughly 7.7–9%)—and demonstrates that recognized routines can replay deterministically at zero model tokens. This matters for enterprise AI deployment by providing a concrete, auditable mechanism to reduce redundant frontier model inference costs and improve agent memory fidelity.
- Enterprise
- Workforce
Research
Subliminal Learning is Non-Semantic Distillation
Ethan Hadley, Eren Gultepe
arXiv · 2026-08-06
Subliminal Learning (SL) is a phenomenon where a language model (student) can inherit biases or behaviors from another model (teacher) through synthetic training data that appears unrelated or random, bypassing standard data auditing. This paper investigates the mechanisms behind SL, finding that adding Gaussian noise to model weights amplifies subliminal transfer by up to 1.9x (Gemma) and 1.3x (Llama), implicating non-semantic weight structures as a key driver. The authors also show that steering vectors applied to the teacher produce subliminal signals, and that student models inherit not just the semantic content of the bias but also the method used to instill it, while gradients of steered data show a linear correlation with teacher steering vectors, offering a potential path toward data auditing. These findings matter because synthetic data is increasingly central to AI training pipelines, and hidden signals that evade standard audits pose real risks to safe and predictable AI development.
- AI policy
- Quality assurance