News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
RTI-Bench: A Structured Dataset for Indian Right-to-Information Decision Analysis
Joy Bose
arXiv · 2026-05-16
RTI-Bench is the first publicly released structured dataset for Indian Right to Information (RTI) administrative decisions, comprising over 1,500 Central Information Commission cases annotated with outcome labels, exemption citations, IRAC-style reasoning components, and procedural timelines. The dataset achieves 89% label coverage on one source corpus and 95.3% label precision on a manually reviewed sample of 50 cases. A zero-shot Mistral 7B baseline reaches 57.3% accuracy and 37.0% macro-F1 on outcome prediction, substantially above the majority-class baseline of 14.3% macro-F1. By making RTI decisions more accessible and machine-readable, the resource could help citizens and policymakers better understand administrative transparency decisions and assess the viability of appeals.
- AI policy
Research
GPF-LiveNews: A Streaming Evaluation Protocol for Group-Conditioned Framing in Large Language Models
Mohd Ariful Haque, Fahad Rahman, Kishor Datta Gupta et al.
arXiv · 2026-05-16
GPF-LiveNews introduces a streaming evaluation protocol designed to audit how large language models frame newly emerging news events differently depending on the identity group specified in the prompt. The protocol uses live BBC/Reuters news articles, 42 identity labels, and seven prompt families to generate and score response bundles with semantic-sensitivity and sentiment-disparity metrics. In a pilot spanning 12 monitoring runs and 23 models, Policy/Action prompts produced the strongest semantic variation across groups, while sentiment differences were relatively flat. The work matters for quality assurance and policy because it provides a repeatable, time-aware auditing tool for detecting group-conditioned framing shifts in deployed LLMs, complementing static bias benchmarks that cannot capture evolving model behavior.
- Quality assurance
- AI policy
Research
Auditing Discriminatory Patterns in Mortgage Lending Through Association Rules and Fair Binning
Archit Rathod, Dhwani Chande, Het Nagda
arXiv · 2026-05-16
This paper audits racial and gender disparities in U.S. mortgage lending using 103,481 cleaned applications from the HMDA 2023 dataset (Chicago metropolitan area). The authors build a three-stage pipeline combining fair binning, FP-Growth association rule mining, and K-Means clustering to test whether standard data preprocessing amplifies bias. Key findings include a 9.63% racial bias introduced by standard income binning, a 29.4% Price of Fairness when applying the epsilon-biased fair binning algorithm at epsilon=0.08, and a disparate impact audit flagging 10 out of 45 cluster-group pairs where Black applicants face significantly higher denial rates than White applicants even among financially similar groups. The work matters for policy and fairness oversight because it demonstrates that discriminatory patterns in lending can be embedded in preprocessing choices and detected through systematic auditing pipelines, even when racial bias does not appear as explicit high-support association rules.
- AI policy
- Quality assurance
Research
Genflow Ad Studio: A Compound AI Architecture for Brand-Aligned, Self-Correcting Video Generation
Debanshu Das, Lavi Nigam, Sunil Kumar Jang Bahadur et al.
arXiv · 2026-05-16
Genflow Ad Studio introduces a compound AI system designed to address brand misalignment and temporal inconsistencies in generative video models for enterprise use. The architecture features a retrieval-based 'Brand DNA' extraction module that parameterizes video generation according to corporate identity guidelines, paired with an Adversarial Multi-Agent Quality Control loop in which evaluator agents iteratively critique generated frames and prompt refinements until a consensus is reached. According to the abstract, this multi-stage, self-correcting pipeline improved the yield of brand-compliant video generations from 42% to 89%, demonstrating a measurable improvement over single-pass monolithic approaches. The work is relevant to enterprises seeking scalable, controllable generative media production that reliably enforces rigid brand constraints.
- Enterprise
- Quality assurance
Research
State Contamination in Memory-Augmented LLM Agents
Yian Wang, Agam Goyal, Yuen Chen et al.
arXiv · 2026-05-16
This paper investigates a failure mode called 'memory laundering' in LLM agents that rely on persistent memory: toxic or adversarial content can be compressed into memory summaries that pass standard toxicity detectors yet still influence future agent outputs in harmful ways. Using paired counterfactual multi-agent rollouts, the authors introduce the sub-threshold propagation gap (SPG) to measure hidden downstream behavioral differences caused by memory states that safety monitors classify as safe. The findings show that raw transcript reuse drives overt toxicity while compressed memory carries subtler, sub-threshold influence, and that sanitization must occur before unsafe content is compressed—cleaning only the final summary leaves laundered influence intact. This work argues that safety in memory-augmented agents must be treated as a state-control problem, with implications for how AI systems are monitored and governed in deployment.
- Quality assurance
- AI policy
Research
A Triple-Intelligence Framework for Sustainable AI-Driven Workforce Analytics: Integrating Artificial Intelligence, Human Judgment, and Organizational Governance
Praveen Kumar Guraja, Kamalamalini Nagasundaram, Manish Nalluri
International Journal of Emerging Research in Science Engineering and Management · 2026-05-16
This paper develops and validates a Triple-Intelligence Framework (TIF) that integrates AI, human judgment, and organizational governance to address risks in AI-driven workforce analytics. Based on a systematic literature review of explainable AI, algorithmic fairness, and people analytics governance from 2017–2025, the framework targets four high-risk decision domains: hiring/mobility, performance management, workforce planning, and remote/hybrid work analytics. The study finds that sustainable workforce analytics requires coordinated action across all three intelligence layers and offers a practical path aligned with Industry 5.0 principles. The work matters because it directly addresses documented harms including algorithmic opacity, automation bias, proxy-based discrimination, and employee surveillance.
- Workforce
- Enterprise
- AI policy
- Quality assurance
Research
Consensus statement on the application of artificial intelligence in osteoporosis screening and management: perspectives from the Asia-Pacific region
Chun‐Feng Huang, Wen-Hui Fang, Kun-Hui Chen et al.
Osteoporosis International · 2026-05-16
This consensus statement from Asia-Pacific multidisciplinary experts establishes 12 recommendations for the safe and equitable use of AI in osteoporosis screening and management, addressing a region where the condition is widely underdiagnosed due to limited access to DXA imaging. The panel defines appropriate AI applications such as imaging-based bone assessment and fracture risk prediction, while specifying minimum standards for model validation, transparency, data protection, and clinician training. The guidance concludes that properly validated AI can help identify high-risk patients who would otherwise go undiagnosed, but should complement rather than replace standard diagnostic methods and clinical judgment. The consensus also highlights the need for post-market surveillance, equity considerations, and alignment with local regulations across the Asia-Pacific region.
- Quality assurance
- Certifications
- AI policy
Research
Gold-Standard AGI: Outer AGI Superalignment
Aaron Turner
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-16
This paper proposes a theoretical framework for 'Gold-Standard AGI,' addressing the outer alignment problem for superintelligent AI systems by defining what it means for an AGI to pursue goals that maximally benefit all humanity without favoritism. The authors present an implementation-neutral solution to outer superalignment and introduce concepts of practical-maximal-alignment and practical-maximal-validation. The paper envisions these definitions forming the basis of an international certification standard, under which only formally-certified AGI systems could be lawfully deployed within relevant jurisdictions. It is written in an accessible, pedagogic style to reach non-technical audiences including AGI policymakers.
- AI policy
- Certifications
Research
Institutional transformation through artificial intelligence in higher education supporting Oman Vision 2040
Malek Hamed Saif Alzakwani, Aziza Al Qamashoui, AlameluMangai Raman
Discover Education · 2026-05-16
This mixed-methods study examines how AI adoption in Oman's higher education institutions can align with the country's Vision 2040 national development goals, using Sociotechnical Systems Theory as its framework. Survey data from faculty across multiple specializations found that 86.2% believe AI has significant potential in higher education, 84.1% are aware of AI's contributions to research and teaching, and 53.5% support embedding AI as a core educational strategy. The paper argues that successful institutional transformation requires coordinated alignment of technological infrastructure, policy, faculty capacity, and ethical standards. The findings highlight AI's potential to prepare students for an AI-driven economy and to strengthen the national labor market.
- Workforce
- Enterprise
- AI policy
- Certifications
Research
"AAB AI Education Case Registry Dataset v1.0"
Winnie Han, Lei Xu
IEEE DataPort · 2026-05-16
This dataset from the AI Assessment Board (AAB) provides a structured registry of documented AI education cases spanning classrooms, workforce training, community initiatives, teacher professional development, and robotics-enabled learning experiences. Each record captures implementation details such as country, organization type, learner age group, pedagogy, AI tool role, and observed outcomes, along with a provisional Evidence Maturity Index (EMI) classification. The dataset is designed to support comparative research, standards development, and transparent evidence preservation in AI education and AI literacy, without containing any personally identifiable or confidential data. It is intended to help researchers, educators, policymakers, and standards developers identify implementation patterns, documentation gaps, and emerging areas for further research.
- Workforce
- Certifications
- AI policy
- Quality assurance
Research
Entry Barriers to the Labor Market in the Era of Generative Artificial Intelligence: A Critical Review with a Two-Dimensional Adjustment Framework (2022–2026)
González Tabarez Jose David
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-16
This critical review examines empirical evidence from 2022–2026 on how generative AI has reshaped hiring, wages, and junior job availability across technology, finance, consulting, and administration. Drawing on studies from Harvard, Stanford, Brookings, the WEF, and the ILO, the paper identifies a pattern of 'seniority-biased technological change,' where firms adopting generative AI cut entry-level hiring while retaining senior workers, compressing starting wages and narrowing access to first jobs. At the same time, generative AI complements the productivity of junior workers who do remain, creating a simultaneous substitution–complementarity paradox. The authors propose a two-dimensional framework based on institutional flexibility and AI adoption intensity that organizes findings into four adjustment regimes and flags the global South as the most critical gap for future research.
- Workforce
- AI policy
- Enterprise
Research
Entry Barriers to the Labor Market in the Era of Generative Artificial Intelligence: A Critical Review with a Two-Dimensional Adjustment Framework (2022–2026)
González Tabarez Jose David
Zenodo (CERN European Organization for Nuclear Research) · 2026-05-16
This critical literature review synthesizes empirical evidence from 2022 to early 2026 on how generative AI is reshaping entry-level labor markets across technology, finance, consulting, and administration. Drawing on studies from Harvard, Stanford, IESE, Brookings, the WEF, and the ILO, the paper documents a 'seniority-biased technological change' pattern in which firms adopting generative AI cut junior hiring and compress starting wages while retaining incumbent workers. At the same time, evidence shows generative AI complements the productivity of entry-level workers who do remain, creating a simultaneous substitution-complementarity paradox. The authors introduce a two-dimensional framework organized by institutional flexibility and AI adoption intensity to explain variation in outcomes across countries and identify the global South as the most critical gap for future research.
- Workforce
- AI policy
- Enterprise
Research
CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?
Haolin Chen, Deon Metelski, Leon Qi et al.
arXiv · 2026-05-15
CHI-Bench (χ-Bench) introduces a benchmark for evaluating AI agents on realistic, end-to-end healthcare administrative workflows spanning provider prior authorization, payer utilization management, and care management. Each task requires an agent to navigate a high-fidelity simulator of 20 healthcare apps via 87 tools, follow a 1,290+ document managed-care operations handbook, play multiple roles with handoffs, and conduct multi-turn dialogs. Across 30 agent and model configurations, the best agent resolves only 28.0% of tasks, no agent exceeds 20% on strict pass^3, and performance collapses to 3.8% in single-session execution—demonstrating that current AI agents fall far short of automating complex, policy-dense healthcare operations. The authors suggest similar capability gaps are likely to appear in other policy-rich, multi-role enterprise domains.
- Enterprise
- AI policy
Research
To Trust or Not to Trust: Authors' Response to AI-based Reviews
César Leblanc, Lukas Picek
arXiv · 2026-05-15
This paper reports findings from two pilot studies examining how authors at computer science venues perceive and respond to AI-generated peer review feedback. Most respondents (83.9%) found AI reviews useful, 80.4% said AI identified issues missed by human reviewers, and 82.1% incorporated at least some AI feedback into their camera-ready revisions. Despite this perceived value, authors trusted AI reviews less than human reviews and preferred AI to be used in a supervised or consent-based role, with 96.4% willing to use AI as a self-review tool before submission and 89.3% wanting advance notice when AI is used in formal review. The findings matter for research quality assurance and policy because they show both practical uptake and persistent concerns—including reports of inaccuracies and misleading comments—that need to be addressed before AI-assisted review is widely adopted.
- Quality assurance
- AI policy
Research
PromptDecipher: Supporting AI Tutor Authoring Through Editable Simulated Interactions
Miina Koyama, Ruiwei Xiao, John Stamper
arXiv · 2026-05-15
PromptDecipher is a system designed to help educators author AI tutoring chatbots more effectively by restructuring the authoring workflow around direct correction-based interactions rather than abstract system prompt writing. The paper's formative study found that virtually no teachers systematically tested their bots before deploying them to students, highlighting a significant quality assurance gap. The system addresses this by letting teachers edit undesirable bot responses in a live chat preview, then automatically analyzing corrections, proposing targeted system prompt rewrites, and validating changes across pre-defined test scenarios — making QA a first-class activity. PromptDecipher is set to be deployed in an AI for Educators course enrolling hundreds of higher-education instructors, making it relevant to both workforce development and quality assurance in AI-assisted learning tools.
- Workforce
- Quality assurance
Research
Voice "Cloning" is Style Transfer
Kaitlyn Zhou, Federico Bianchi, Martijn Bartelds et al.
arXiv · 2026-05-15
This paper investigates whether widely-used voice cloning models faithfully reproduce an individual's voice and finds they do not. Instead, the models systematically apply style transfer, making cloned voices sound more authoritative, warm, customer-service-like, and human-like than their sources, as rated by human annotators. Annotators also report greater trust in cloned voices and greater willingness to disclose sensitive personal information to them. Additionally, voice cloning leads to homogenization of speaker characteristics, including reduced variance in accent, speaking rate, and audio embedding space, highlighting new limitations and risks of the technology.
- AI policy
- Quality assurance
Research
Characterizing AI Fact-Checkers and Their Contributions on Community Notes
Yilin Gong, Siqi Wu
arXiv · 2026-05-15
This study presents the first empirical analysis of AI fact-checkers operating on X's Community Notes platform, examining their volume, speed, coverage, and accuracy between September 2025 and May 2026. The researchers find that 20 AI writers produced 14.2% of all submitted notes, with their daily share rising to 44.8%, and that AI notes covered 74.4% of fact-checked posts not reviewed by any human. However, AI-generated notes are less likely to be rated as helpful than those written by human experts, though they outperform notes written by laypeople. The findings highlight both the scaling potential and quality limitations of AI-driven fact-checking, with implications for how human-AI collaborative content moderation systems should be designed and governed.
- AI policy
- Quality assurance
Research
Symphony for Speech-to-Text: Supporting Real-Time Medical Voice Interfaces
Arne Nix, Robert James, Lasse Borgholt et al.
arXiv · 2026-05-15
Symphony for Speech-to-Text is a medical-grade speech recognition system designed for real-time and batch clinical transcription. It decomposes the transcription pipeline into specialized components for recognition, formatting, and contextual correction to accurately handle medical terminology, abbreviations, measurements, and clinical shorthand. Evaluations on public benchmark and medical speech datasets show Symphony substantially outperforms state-of-the-art systems in clinical settings while matching or exceeding them in general-domain settings. The authors also release a clinical benchmark dataset to support further validation and progress in medical speech recognition.
- Quality assurance
- Enterprise
Research
AI-Mediated Communication Can Steer Collective Opinion
Stratis Tsirtsis, Kai Rawal, Chris Russell et al.
arXiv · 2026-05-15
This paper demonstrates that large language models (LLMs) used to mediate human-to-human communication—such as editing posts or explaining content—introduce directional political and social biases (e.g., nudging text toward gun control or against atheism). Through empirical audits of multiple LLM families and a mathematical model of opinion dynamics on real social network data, the authors show these biases can be amplified across networks and shift collective opinion at scale. An audit of X's 'Explain this post' feature finds evidence of pro-life bias in Grok's outputs on abortion content, traceable to specific design choices. The authors discuss implications for ongoing EU legislative efforts around AI regulation.
- AI policy
- Enterprise
Research
Prospective multi-pathogen disease forecasting using autonomous LLM-guided tree search
Sarah Martinson, Michael P. Brenner, Martyna Plomecka et al.
arXiv · 2026-05-15
This paper presents an autonomous AI system that uses Large Language Model-guided tree search to generate, evaluate, and optimize disease forecasting models without manual expert curation. In a real-time prospective evaluation during the 2025-2026 US respiratory season, the system autonomously built models for influenza, COVID-19, and RSV, and its ensemble matched or outperformed the CDC hub ensembles out-of-sample. The system also handled data-scarce cold-start scenarios for RSV and incorporated design choices—such as log-scale metrics and an automated judge—to prevent reward hacking and maintain scientific fidelity. By removing the expert labor bottleneck in epidemiological modeling, this framework enables rapid, scalable deployment of disease forecasting across pathogens and geographies.
- AI policy
- Workforce
Research
Marginal Alignment Does Not Guarantee Joint-Distribution Fidelity: An Official-Reference Audit of Nemotron-Personas-Korea with Cross-Locale Replication
Joonhyung Bae
arXiv · 2026-05-15
This paper audits NVIDIA's Nemotron-Personas-Korea (NPK), a dataset of one million synthetic Korean personas, revealing that aligning a synthetic dataset with official demographic marginals does not guarantee that joint distributions across attributes like age, sex, occupation, and education are accurate. The authors introduce the Independence-Assumption Footprint (IAF), an audit tool that checks synthetic joint distributions against official or institutional references, and find that NPK passes marginal checks but fails on three joint distributions—including an over-flattened female representation in male-dominated occupations and an institutionally inconsistent age profile for military service. Testing across six additional NPK locales shows that diagnostic failures are locale-dependent rather than universal, complicating cross-locale comparisons. The findings argue that synthetic datasets used as stand-ins for real populations must accompany marginal alignment claims with explicit joint-distribution audits before downstream reuse.
- Quality assurance
- AI policy
Research
Policy-Grounded Dynamic Facet Suggestions for Job Search
Dan Xu, Baofen Zheng, Qianqi Shen et al.
arXiv · 2026-05-15
This paper presents a dynamic facet suggestion (DFS) system deployed at LinkedIn to help job seekers refine vague search queries. Because over 80% of LinkedIn job-related queries contain three or fewer keywords, the system surfaces personalized semantic attributes in real time based on the joint user-query context, using a policy-grounded, retrieval-augmented ranking framework that combines offline taxonomy curation, embedding-based candidate retrieval, and a distilled small language model for scoring. Offline evaluation shows high precision for generated suggestions, and online A/B tests demonstrate significant improvements in both suggestion engagement and job search outcomes, making the approach practically meaningful for connecting workers to relevant opportunities.
- Workforce
- Enterprise
Research
Fully Open Meditron: An Auditable Pipeline for Clinical LLMs
Xavier Theimer-Lienhard, Mushtaha El-Amin, Fay Elhassan et al.
arXiv · 2026-05-15
Fully Open Meditron introduces the first fully open, end-to-end auditable pipeline for building large language model-based clinical decision support systems (CDSS). The pipeline includes a clinician-audited training corpus unifying eight public medical QA datasets, three clinician-vetted synthetic extensions (exam-style QA, guideline-grounded QA from 46,469 clinical practice guidelines, and clinical vignettes), and an evaluation protocol calibrated against 204 human raters. Applied to five fully open base models, the best variant (Apertus-70B-MeditronFO) improves +6.6 points over its base on aggregate medical benchmarks, while Gemma-3-27B-MeditronFO outperforms MedGemma on HealthBench (58% vs 55.9%), demonstrating that full transparency and auditability need not come at the cost of state-of-the-art clinical performance.
- Quality assurance
- Certifications
Research
Confirming Correct, Missing the Rest: LLM Tutoring Agents Struggle Where Feedback Matters Most
Tahreem Yasir, Wenbo Li, Sam Gilson et al.
arXiv · 2026-05-15
This paper benchmarks seven LLM-based tutoring agents on their ability to give accurate feedback in propositional logic tasks, evaluating over 10,836 solution–feedback pairs with knowledge-graph-derived ground truth across three feedback conditions. The study finds that LLMs perform well when confirming correct (optimal) student steps but systematically fail where adaptive tutoring matters most: they over-reject valid but suboptimal reasoning and over-validate incorrect solutions. These diagnostic failures persisted across all tested models regardless of solution context, pointing to architectural rather than informational limitations, and accurate diagnosis did not reliably translate into pedagogically actionable feedback. The authors conclude that LLMs are better suited for hybrid architectures where knowledge-graph-grounded models handle diagnosis and LLMs support open-ended scaffolding and dialogue.
- Quality assurance
- Workforce
Research
Formal Methods Meet LLMs: Auditing, Monitoring, and Intervention for Compliance of Advanced AI Systems
Parand A. Alamdari, Toryn Q. Klassen, Sheila A. McIlraith
arXiv · 2026-05-15
This paper proposes a framework combining formal methods—specifically Linear Temporal Logic (LTL)—with machine learning to audit and monitor AI-enabled products and services for compliance with behavioral constraints such as safety rules, norms, and regulations. The techniques support both offline auditing (pre-deployment) and online runtime monitoring (post-deployment), and include predictive and intervening monitors that can preempt and mitigate violations by LLM-based agents while preserving task performance. Experimental results show that LTL-based auditing outperforms LLM baseline judges in detecting violations of temporally extended constraints, and that LLMs' own temporal reasoning degrades as event distance, constraint count, and proposition count increase. The work is directly relevant to AI governance, providing practical tools for developers, third-party evaluators, and regulators to enforce compliance throughout the AI development lifecycle.
- AI policy
- Quality assurance