News & Research
The latest AI research and news with real-world stakes. Each item is sourced, dated, summarized in plain English and tagged by impact area, and checked against its source before it appears.
Kind
Impact area
5849 items
- ResearcharXiv2026-05-16P
RTI-Bench: A Structured Dataset for Indian Right-to-Information Decision Analysis · Joy Bose
RTI-Bench is the first publicly released structured dataset for Indian Right to Information (RTI) administrative decisions, comprising over 1,500 Central Information Commission cases annotated with outcome labels, exemption citations, IRAC-style reasoning components, and procedural timelines. The dataset achieves 89% label coverage on one source corpus and 95.3% label precision on a manually reviewed sample of 50 cases. A zero-shot Mistral 7B baseline reaches 57.3% accuracy and 37.0% macro-F1 on outcome prediction, substantially above the majority-class baseline of 14.3% macro-F1. By making RTI decisions more accessible and machine-readable, the resource could help citizens and policymakers better understand administrative transparency decisions and assess the viability of appeals.
- ResearcharXiv2026-05-16QP
GPF-LiveNews: A Streaming Evaluation Protocol for Group-Conditioned Framing in Large Language Models · Mohd Ariful Haque, Fahad Rahman, Kishor Datta Gupta et al.
GPF-LiveNews introduces a streaming evaluation protocol designed to audit how large language models frame newly emerging news events differently depending on the identity group specified in the prompt. The protocol uses live BBC/Reuters news articles, 42 identity labels, and seven prompt families to generate and score response bundles with semantic-sensitivity and sentiment-disparity metrics. In a pilot spanning 12 monitoring runs and 23 models, Policy/Action prompts produced the strongest semantic variation across groups, while sentiment differences were relatively flat. The work matters for quality assurance and policy because it provides a repeatable, time-aware auditing tool for detecting group-conditioned framing shifts in deployed LLMs, complementing static bias benchmarks that cannot capture evolving model behavior.
- ResearcharXiv2026-05-16QP
Auditing Discriminatory Patterns in Mortgage Lending Through Association Rules and Fair Binning · Archit Rathod, Dhwani Chande, Het Nagda
This paper audits racial and gender disparities in U.S. mortgage lending using 103,481 cleaned applications from the HMDA 2023 dataset (Chicago metropolitan area). The authors build a three-stage pipeline combining fair binning, FP-Growth association rule mining, and K-Means clustering to test whether standard data preprocessing amplifies bias. Key findings include a 9.63% racial bias introduced by standard income binning, a 29.4% Price of Fairness when applying the epsilon-biased fair binning algorithm at epsilon=0.08, and a disparate impact audit flagging 10 out of 45 cluster-group pairs where Black applicants face significantly higher denial rates than White applicants even among financially similar groups. The work matters for policy and fairness oversight because it demonstrates that discriminatory patterns in lending can be embedded in preprocessing choices and detected through systematic auditing pipelines, even when racial bias does not appear as explicit high-support association rules.
- ResearcharXiv2026-05-16EQ
Genflow Ad Studio: A Compound AI Architecture for Brand-Aligned, Self-Correcting Video Generation · Debanshu Das, Lavi Nigam, Sunil Kumar Jang Bahadur et al.
Genflow Ad Studio introduces a compound AI system designed to address brand misalignment and temporal inconsistencies in generative video models for enterprise use. The architecture features a retrieval-based 'Brand DNA' extraction module that parameterizes video generation according to corporate identity guidelines, paired with an Adversarial Multi-Agent Quality Control loop in which evaluator agents iteratively critique generated frames and prompt refinements until a consensus is reached. According to the abstract, this multi-stage, self-correcting pipeline improved the yield of brand-compliant video generations from 42% to 89%, demonstrating a measurable improvement over single-pass monolithic approaches. The work is relevant to enterprises seeking scalable, controllable generative media production that reliably enforces rigid brand constraints.
- ResearcharXiv2026-05-16QP
State Contamination in Memory-Augmented LLM Agents · Yian Wang, Agam Goyal, Yuen Chen et al.
This paper investigates a failure mode called 'memory laundering' in LLM agents that rely on persistent memory: toxic or adversarial content can be compressed into memory summaries that pass standard toxicity detectors yet still influence future agent outputs in harmful ways. Using paired counterfactual multi-agent rollouts, the authors introduce the sub-threshold propagation gap (SPG) to measure hidden downstream behavioral differences caused by memory states that safety monitors classify as safe. The findings show that raw transcript reuse drives overt toxicity while compressed memory carries subtler, sub-threshold influence, and that sanitization must occur before unsafe content is compressed—cleaning only the final summary leaves laundered influence intact. This work argues that safety in memory-augmented agents must be treated as a state-control problem, with implications for how AI systems are monitored and governed in deployment.
- ResearchInternational Journal of Emerging Research in Science Engineering and Management2026-05-16WEQP
A Triple-Intelligence Framework for Sustainable AI-Driven Workforce Analytics: Integrating Artificial Intelligence, Human Judgment, and Organizational Governance · Praveen Kumar Guraja, Kamalamalini Nagasundaram, Manish Nalluri
This paper develops and validates a Triple-Intelligence Framework (TIF) that integrates AI, human judgment, and organizational governance to address risks in AI-driven workforce analytics. Based on a systematic literature review of explainable AI, algorithmic fairness, and people analytics governance from 2017–2025, the framework targets four high-risk decision domains: hiring/mobility, performance management, workforce planning, and remote/hybrid work analytics. The study finds that sustainable workforce analytics requires coordinated action across all three intelligence layers and offers a practical path aligned with Industry 5.0 principles. The work matters because it directly addresses documented harms including algorithmic opacity, automation bias, proxy-based discrimination, and employee surveillance.
- ResearchOsteoporosis International2026-05-16QCP
Consensus statement on the application of artificial intelligence in osteoporosis screening and management: perspectives from the Asia-Pacific region · Chun‐Feng Huang, Wen-Hui Fang, Kun-Hui Chen et al.
This consensus statement from Asia-Pacific multidisciplinary experts establishes 12 recommendations for the safe and equitable use of AI in osteoporosis screening and management, addressing a region where the condition is widely underdiagnosed due to limited access to DXA imaging. The panel defines appropriate AI applications such as imaging-based bone assessment and fracture risk prediction, while specifying minimum standards for model validation, transparency, data protection, and clinician training. The guidance concludes that properly validated AI can help identify high-risk patients who would otherwise go undiagnosed, but should complement rather than replace standard diagnostic methods and clinical judgment. The consensus also highlights the need for post-market surveillance, equity considerations, and alignment with local regulations across the Asia-Pacific region.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-05-16CP
Gold-Standard AGI: Outer AGI Superalignment · Aaron Turner
This paper proposes a theoretical framework for 'Gold-Standard AGI,' addressing the outer alignment problem for superintelligent AI systems by defining what it means for an AGI to pursue goals that maximally benefit all humanity without favoritism. The authors present an implementation-neutral solution to outer superalignment and introduce concepts of practical-maximal-alignment and practical-maximal-validation. The paper envisions these definitions forming the basis of an international certification standard, under which only formally-certified AGI systems could be lawfully deployed within relevant jurisdictions. It is written in an accessible, pedagogic style to reach non-technical audiences including AGI policymakers.
- ResearchDiscover Education2026-05-16WECP
Institutional transformation through artificial intelligence in higher education supporting Oman Vision 2040 · Malek Hamed Saif Alzakwani, Aziza Al Qamashoui, AlameluMangai Raman
This mixed-methods study examines how AI adoption in Oman's higher education institutions can align with the country's Vision 2040 national development goals, using Sociotechnical Systems Theory as its framework. Survey data from faculty across multiple specializations found that 86.2% believe AI has significant potential in higher education, 84.1% are aware of AI's contributions to research and teaching, and 53.5% support embedding AI as a core educational strategy. The paper argues that successful institutional transformation requires coordinated alignment of technological infrastructure, policy, faculty capacity, and ethical standards. The findings highlight AI's potential to prepare students for an AI-driven economy and to strengthen the national labor market.
- ResearchIEEE DataPort2026-05-16WQCP
"AAB AI Education Case Registry Dataset v1.0" · Winnie Han, Lei Xu
This dataset from the AI Assessment Board (AAB) provides a structured registry of documented AI education cases spanning classrooms, workforce training, community initiatives, teacher professional development, and robotics-enabled learning experiences. Each record captures implementation details such as country, organization type, learner age group, pedagogy, AI tool role, and observed outcomes, along with a provisional Evidence Maturity Index (EMI) classification. The dataset is designed to support comparative research, standards development, and transparent evidence preservation in AI education and AI literacy, without containing any personally identifiable or confidential data. It is intended to help researchers, educators, policymakers, and standards developers identify implementation patterns, documentation gaps, and emerging areas for further research.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-05-16WEP
Entry Barriers to the Labor Market in the Era of Generative Artificial Intelligence: A Critical Review with a Two-Dimensional Adjustment Framework (2022–2026) · González Tabarez Jose David
This critical review examines empirical evidence from 2022–2026 on how generative AI has reshaped hiring, wages, and junior job availability across technology, finance, consulting, and administration. Drawing on studies from Harvard, Stanford, Brookings, the WEF, and the ILO, the paper identifies a pattern of 'seniority-biased technological change,' where firms adopting generative AI cut entry-level hiring while retaining senior workers, compressing starting wages and narrowing access to first jobs. At the same time, generative AI complements the productivity of junior workers who do remain, creating a simultaneous substitution–complementarity paradox. The authors propose a two-dimensional framework based on institutional flexibility and AI adoption intensity that organizes findings into four adjustment regimes and flags the global South as the most critical gap for future research.
- ResearchZenodo (CERN European Organization for Nuclear Research)2026-05-16WEP
Entry Barriers to the Labor Market in the Era of Generative Artificial Intelligence: A Critical Review with a Two-Dimensional Adjustment Framework (2022–2026) · González Tabarez Jose David
This critical literature review synthesizes empirical evidence from 2022 to early 2026 on how generative AI is reshaping entry-level labor markets across technology, finance, consulting, and administration. Drawing on studies from Harvard, Stanford, IESE, Brookings, the WEF, and the ILO, the paper documents a 'seniority-biased technological change' pattern in which firms adopting generative AI cut junior hiring and compress starting wages while retaining incumbent workers. At the same time, evidence shows generative AI complements the productivity of entry-level workers who do remain, creating a simultaneous substitution-complementarity paradox. The authors introduce a two-dimensional framework organized by institutional flexibility and AI adoption intensity to explain variation in outcomes across countries and identify the global South as the most critical gap for future research.
- ResearcharXiv2026-05-15EP
CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows? · Haolin Chen, Deon Metelski, Leon Qi et al.
CHI-Bench (χ-Bench) introduces a benchmark for evaluating AI agents on realistic, end-to-end healthcare administrative workflows spanning provider prior authorization, payer utilization management, and care management. Each task requires an agent to navigate a high-fidelity simulator of 20 healthcare apps via 87 tools, follow a 1,290+ document managed-care operations handbook, play multiple roles with handoffs, and conduct multi-turn dialogs. Across 30 agent and model configurations, the best agent resolves only 28.0% of tasks, no agent exceeds 20% on strict pass^3, and performance collapses to 3.8% in single-session execution—demonstrating that current AI agents fall far short of automating complex, policy-dense healthcare operations. The authors suggest similar capability gaps are likely to appear in other policy-rich, multi-role enterprise domains.
- ResearcharXiv2026-05-15QP
To Trust or Not to Trust: Authors' Response to AI-based Reviews · César Leblanc, Lukas Picek
This paper reports findings from two pilot studies examining how authors at computer science venues perceive and respond to AI-generated peer review feedback. Most respondents (83.9%) found AI reviews useful, 80.4% said AI identified issues missed by human reviewers, and 82.1% incorporated at least some AI feedback into their camera-ready revisions. Despite this perceived value, authors trusted AI reviews less than human reviews and preferred AI to be used in a supervised or consent-based role, with 96.4% willing to use AI as a self-review tool before submission and 89.3% wanting advance notice when AI is used in formal review. The findings matter for research quality assurance and policy because they show both practical uptake and persistent concerns—including reports of inaccuracies and misleading comments—that need to be addressed before AI-assisted review is widely adopted.
- ResearcharXiv2026-05-15WQ
PromptDecipher: Supporting AI Tutor Authoring Through Editable Simulated Interactions · Miina Koyama, Ruiwei Xiao, John Stamper
PromptDecipher is a system designed to help educators author AI tutoring chatbots more effectively by restructuring the authoring workflow around direct correction-based interactions rather than abstract system prompt writing. The paper's formative study found that virtually no teachers systematically tested their bots before deploying them to students, highlighting a significant quality assurance gap. The system addresses this by letting teachers edit undesirable bot responses in a live chat preview, then automatically analyzing corrections, proposing targeted system prompt rewrites, and validating changes across pre-defined test scenarios — making QA a first-class activity. PromptDecipher is set to be deployed in an AI for Educators course enrolling hundreds of higher-education instructors, making it relevant to both workforce development and quality assurance in AI-assisted learning tools.
- ResearcharXiv2026-05-15QP
Voice "Cloning" is Style Transfer · Kaitlyn Zhou, Federico Bianchi, Martijn Bartelds et al.
This paper investigates whether widely-used voice cloning models faithfully reproduce an individual's voice and finds they do not. Instead, the models systematically apply style transfer, making cloned voices sound more authoritative, warm, customer-service-like, and human-like than their sources, as rated by human annotators. Annotators also report greater trust in cloned voices and greater willingness to disclose sensitive personal information to them. Additionally, voice cloning leads to homogenization of speaker characteristics, including reduced variance in accent, speaking rate, and audio embedding space, highlighting new limitations and risks of the technology.
- ResearcharXiv2026-05-15QP
Characterizing AI Fact-Checkers and Their Contributions on Community Notes · Yilin Gong, Siqi Wu
This study presents the first empirical analysis of AI fact-checkers operating on X's Community Notes platform, examining their volume, speed, coverage, and accuracy between September 2025 and May 2026. The researchers find that 20 AI writers produced 14.2% of all submitted notes, with their daily share rising to 44.8%, and that AI notes covered 74.4% of fact-checked posts not reviewed by any human. However, AI-generated notes are less likely to be rated as helpful than those written by human experts, though they outperform notes written by laypeople. The findings highlight both the scaling potential and quality limitations of AI-driven fact-checking, with implications for how human-AI collaborative content moderation systems should be designed and governed.
- ResearcharXiv2026-05-15EQ
Symphony for Speech-to-Text: Supporting Real-Time Medical Voice Interfaces · Arne Nix, Robert James, Lasse Borgholt et al.
Symphony for Speech-to-Text is a medical-grade speech recognition system designed for real-time and batch clinical transcription. It decomposes the transcription pipeline into specialized components for recognition, formatting, and contextual correction to accurately handle medical terminology, abbreviations, measurements, and clinical shorthand. Evaluations on public benchmark and medical speech datasets show Symphony substantially outperforms state-of-the-art systems in clinical settings while matching or exceeding them in general-domain settings. The authors also release a clinical benchmark dataset to support further validation and progress in medical speech recognition.
- ResearcharXiv2026-05-15EP
AI-Mediated Communication Can Steer Collective Opinion · Stratis Tsirtsis, Kai Rawal, Chris Russell et al.
This paper demonstrates that large language models (LLMs) used to mediate human-to-human communication—such as editing posts or explaining content—introduce directional political and social biases (e.g., nudging text toward gun control or against atheism). Through empirical audits of multiple LLM families and a mathematical model of opinion dynamics on real social network data, the authors show these biases can be amplified across networks and shift collective opinion at scale. An audit of X's 'Explain this post' feature finds evidence of pro-life bias in Grok's outputs on abortion content, traceable to specific design choices. The authors discuss implications for ongoing EU legislative efforts around AI regulation.
- ResearcharXiv2026-05-15WP
Prospective multi-pathogen disease forecasting using autonomous LLM-guided tree search · Sarah Martinson, Michael P. Brenner, Martyna Plomecka et al.
This paper presents an autonomous AI system that uses Large Language Model-guided tree search to generate, evaluate, and optimize disease forecasting models without manual expert curation. In a real-time prospective evaluation during the 2025-2026 US respiratory season, the system autonomously built models for influenza, COVID-19, and RSV, and its ensemble matched or outperformed the CDC hub ensembles out-of-sample. The system also handled data-scarce cold-start scenarios for RSV and incorporated design choices—such as log-scale metrics and an automated judge—to prevent reward hacking and maintain scientific fidelity. By removing the expert labor bottleneck in epidemiological modeling, this framework enables rapid, scalable deployment of disease forecasting across pathogens and geographies.
- ResearcharXiv2026-05-15QP
Marginal Alignment Does Not Guarantee Joint-Distribution Fidelity: An Official-Reference Audit of Nemotron-Personas-Korea with Cross-Locale Replication · Joonhyung Bae
This paper audits NVIDIA's Nemotron-Personas-Korea (NPK), a dataset of one million synthetic Korean personas, revealing that aligning a synthetic dataset with official demographic marginals does not guarantee that joint distributions across attributes like age, sex, occupation, and education are accurate. The authors introduce the Independence-Assumption Footprint (IAF), an audit tool that checks synthetic joint distributions against official or institutional references, and find that NPK passes marginal checks but fails on three joint distributions—including an over-flattened female representation in male-dominated occupations and an institutionally inconsistent age profile for military service. Testing across six additional NPK locales shows that diagnostic failures are locale-dependent rather than universal, complicating cross-locale comparisons. The findings argue that synthetic datasets used as stand-ins for real populations must accompany marginal alignment claims with explicit joint-distribution audits before downstream reuse.
- ResearcharXiv2026-05-15WE
Policy-Grounded Dynamic Facet Suggestions for Job Search · Dan Xu, Baofen Zheng, Qianqi Shen et al.
This paper presents a dynamic facet suggestion (DFS) system deployed at LinkedIn to help job seekers refine vague search queries. Because over 80% of LinkedIn job-related queries contain three or fewer keywords, the system surfaces personalized semantic attributes in real time based on the joint user-query context, using a policy-grounded, retrieval-augmented ranking framework that combines offline taxonomy curation, embedding-based candidate retrieval, and a distilled small language model for scoring. Offline evaluation shows high precision for generated suggestions, and online A/B tests demonstrate significant improvements in both suggestion engagement and job search outcomes, making the approach practically meaningful for connecting workers to relevant opportunities.
- ResearcharXiv2026-05-15QC
Fully Open Meditron: An Auditable Pipeline for Clinical LLMs · Xavier Theimer-Lienhard, Mushtaha El-Amin, Fay Elhassan et al.
Fully Open Meditron introduces the first fully open, end-to-end auditable pipeline for building large language model-based clinical decision support systems (CDSS). The pipeline includes a clinician-audited training corpus unifying eight public medical QA datasets, three clinician-vetted synthetic extensions (exam-style QA, guideline-grounded QA from 46,469 clinical practice guidelines, and clinical vignettes), and an evaluation protocol calibrated against 204 human raters. Applied to five fully open base models, the best variant (Apertus-70B-MeditronFO) improves +6.6 points over its base on aggregate medical benchmarks, while Gemma-3-27B-MeditronFO outperforms MedGemma on HealthBench (58% vs 55.9%), demonstrating that full transparency and auditability need not come at the cost of state-of-the-art clinical performance.
- ResearcharXiv2026-05-15WQ
Confirming Correct, Missing the Rest: LLM Tutoring Agents Struggle Where Feedback Matters Most · Tahreem Yasir, Wenbo Li, Sam Gilson et al.
This paper benchmarks seven LLM-based tutoring agents on their ability to give accurate feedback in propositional logic tasks, evaluating over 10,836 solution–feedback pairs with knowledge-graph-derived ground truth across three feedback conditions. The study finds that LLMs perform well when confirming correct (optimal) student steps but systematically fail where adaptive tutoring matters most: they over-reject valid but suboptimal reasoning and over-validate incorrect solutions. These diagnostic failures persisted across all tested models regardless of solution context, pointing to architectural rather than informational limitations, and accurate diagnosis did not reliably translate into pedagogically actionable feedback. The authors conclude that LLMs are better suited for hybrid architectures where knowledge-graph-grounded models handle diagnosis and LLMs support open-ended scaffolding and dialogue.
- ResearcharXiv2026-05-15QP
Formal Methods Meet LLMs: Auditing, Monitoring, and Intervention for Compliance of Advanced AI Systems · Parand A. Alamdari, Toryn Q. Klassen, Sheila A. McIlraith
This paper proposes a framework combining formal methods—specifically Linear Temporal Logic (LTL)—with machine learning to audit and monitor AI-enabled products and services for compliance with behavioral constraints such as safety rules, norms, and regulations. The techniques support both offline auditing (pre-deployment) and online runtime monitoring (post-deployment), and include predictive and intervening monitors that can preempt and mitigate violations by LLM-based agents while preserving task performance. Experimental results show that LTL-based auditing outperforms LLM baseline judges in detecting violations of temporally extended constraints, and that LLMs' own temporal reasoning degrades as event distance, constraint count, and proposition count increase. The work is directly relevant to AI governance, providing practical tools for developers, third-party evaluators, and regulators to enforce compliance throughout the AI development lifecycle.