News & Research
The latest AI research and news with real-world stakes. Each item is sourced, dated and summarized in plain English, tagged by impact area where one fits, and its summary is checked against the text it was written from.
8151 items
- ResearchAI & Society2026-07-02AI policy
Co-designing AI systems with value-sensitive citizen science · Sachit Mahajan, Dirk Helbing
This paper introduces Value-Sensitive Citizen Science (VSCS), a framework combining Value-Sensitive Design with citizen science to enable meaningful public participation in AI development. The framework uses a Participatory Value-Cognition Taxonomy and iterative scenario reasoning to translate community values into technical requirements, while embedding governance mechanisms for accountability throughout the AI lifecycle. The authors argue that this approach challenges top-down, monocultural AI design and addresses issues of epistemic justice, power asymmetries, and scalability, with practical implications for policymakers and practitioners building inclusive AI systems.
- ResearcharXiv (Cornell University)2026-07-02Quality assurance · AI policy · +1
The Eticas AI Risk Taxonomy: Open Infrastructure for Operationalizing AI Audits · Gemma Galdon Clavell, Pablo Accuosto, Usman Gohar
This paper presents the Eticas AI Risk Taxonomy v2.0.0, an open framework designed to bridge the gap between identifying AI risks and actually operationalizing them into structured audits with measurable, graded findings. The authors demonstrate the system end-to-end using PII leakage as a test case on GPT-4-0314, showing disclosure rates of 0%, 51%, and 84% under increasing adversarial conditioning, ultimately yielding a severity grade of E with a SYSTEMIC pattern. The taxonomy organizes 76 active subcategories across 10 categories and 20 sub-groups, with mappings to 18 external frameworks, and its upper layers are published under CC BY 4.0 as open semantic infrastructure. This work matters because it provides shared, standardized infrastructure for AI auditing that goes beyond risk catalogs to enable defensible, reproducible audit findings in high-stakes deployment contexts.
- ResearchVeredas do Direito Direito Ambiental e Desenvolvimento Sustentável2026-07-02AI policy · Privacy & Data Protection · +2
A INTELIGÊNCIA ARTIFICIAL GENERATIVA NA EDUCAÇÃO 5.0: LIMITES ÉTICO-JURÍDICOS, GOVERNANÇA ALGORÍTMICA E A RECONFIGURAÇÃO DA PERSONALIZAÇÃO DA APRENDIZAGEM · Luiz Fernando Ridolfi, Ana Alves Ramos, Cláudio Filipe Lima Rapôso et al.
This paper critically examines the ethical and legal limits of Generative Artificial Intelligence (GAI) in Education 5.0, a paradigm integrating technological innovation with human development and social responsibility. Through qualitative bibliographic and documentary research, the authors find that while GAI holds promise for personalized learning, inclusion, and pedagogical efficiency, its use raises serious concerns around algorithmic discrimination, decisional transparency, data protection, teacher autonomy, and fundamental rights. The study concludes that robust governance mechanisms—including accountability, human oversight, and ethical compliance—are necessary to ensure that AI-driven learning personalization aligns with democratic and constitutional principles. The findings have direct implications for AI policy and certification frameworks governing educational technology.
- ResearchPerspectives on Global Development and Technology2026-07-02Workforce · AI policy
Artificial Intelligence and Inequality: Policy Paths in a Polarized Future · Hassan Daliri
Using an agent-based simulation model grounded in theories of skill-biased technological change and labor market polarization, this study examines how AI and automation affect income inequality under different policy regimes. Results show that rapid automation without skills investment worsens inequality and displaces workers, while combining targeted skill subsidies with progressive taxation significantly reduces inequality and shifts employment toward AI-complementary sectors. The findings demonstrate that AI's inequality effects are policy-contingent rather than technologically predetermined, highlighting the importance of redistribution and adaptive human capital strategies for inclusive growth.
- ResearchAcademy Review2026-07-02Enterprise · AI policy · +1
IDENTIFICATION OF PROBLEMS ARISING FROM THE IMPLEMENTATION OF ARTIFICIAL INTELLIGENCE IN THE BUSINESS ENVIRONMENT · Oleh Havryliuk, Ihor Ponomarenko, Oleksandr Yakushev
This paper examines the key challenges businesses face when implementing artificial intelligence, with particular focus on generative AI in marketing, software engineering, and analytics. The authors identify major problem categories including data quality, integration with legacy systems, infrastructure limitations, workforce shortages of qualified AI specialists, cybersecurity threats, and data privacy concerns. The study also reviews the evolution of AI regulatory frameworks across the EU, USA, and China—including the GDPR and the EU AI Act—and proposes strategic development vectors such as federated learning, tamper-resistant models, and cooperation with regulators. The findings are relevant to enterprises seeking to balance AI adoption with risk mitigation and regulatory compliance.
- ResearcharXiv2026-07-01Quality assurance · Privacy & Data Protection
Janus: a Playground for User-Involved Agentic Permission Management · Natalie Grace Brigham, Eugene Bagdasarian, Tadayoshi Kohno et al.
Janus is a playground system for studying how users can be involved in managing the permissions that AI agents use when autonomously executing tool calls. The researchers implement six different permission management designs across a conceptual design space and evaluate them in three scenarios, finding that user input significantly strengthens privacy and security, AI augmentation can reduce cognitive load, and realistic factors like permission fatigue must be accounted for. No single design works best in all contexts, pointing to the need for context-sensitive approaches to permission management in agentic AI systems.
- ResearcharXiv2026-07-01Quality assurance
The Agentic Garden of Forking Paths · Jiacheng Miao, Jonathan K Pritchard, James Zou
This paper investigates how AI agents, when assigned different analytical personas, reproduce the ideological variation seen among human researchers analyzing the same dataset. Across four high-stakes domains, AI agents with different personas produced divergent and often opposing conclusions from identical data, replicating 72% of the human ideological gap observed in a study where 42 human teams analyzed the same immigration dataset. Critically, 86% of the AI-generated analyses passed independent AI review and 78% passed majority human expert review, suggesting the problem is not flawed analysis but selective exploration of a large space of defensible analytical choices. To address this, the authors introduce the 'm-value' (multiverse value) and 'Agentic Bootstrap,' a method that uses AI agents to sample plausible analysis paths and assess how extreme any single reported finding is relative to the full distribution of defensible analyses — offering a new criterion for evaluating scientific credibility.
- ResearcharXiv2026-07-01Workforce · Quality assurance
Grounded Optimization: A Layered Engineering Framework for Reducing LLM Hallucination in Automated Personal Document Rewriting · Shashank Indukuri, Adarsh Agrawal
This paper introduces Grounded Optimization, a five-layer engineering framework designed to reduce hallucinations when large language models rewrite resumes for applicant tracking systems. The framework combines temporal context validation, contamination detection, structural enforcement, prompt-level grounding, and an evaluator agent. Ablation experiments across three LLMs, four temperature settings, and six layer configurations on 25 synthetic resumes spanning 14 industries show that undefended baselines produce 2.48–5.36 detected hallucinations per resume, while the full framework reduces the overall detected hallucination rate to 0.04–0.24 and cuts temporal hallucinations by 50–95%. The work matters because it demonstrates that deterministic, layered defenses are necessary complements to prompt-level grounding, especially for weaker models or higher temperature settings, improving the reliability of AI-assisted document rewriting in high-stakes job application contexts.
- ResearcharXiv2026-07-01Quality assurance · Health
On the Utility and Factual Reliability of Pruned Mixture-of-Experts Models in the Biomedical Domain · Atsuki Yamaguchi, Szymon Palucha, Léo Bijar et al.
This paper examines how structured expert pruning of Mixture-of-Experts (MoE) language models affects both utility and factual reliability in the biomedical domain. Testing four MoE models across six pruning methods and multiple pruning ratios, the authors find that moderate pruning can preserve in-domain biomedical utility without an immediate drop in reliability, but extreme pruning raises hallucination risks. Performance degrades rapidly when pruned models are applied outside their target domain. The study concludes that evaluating pruned models on utility benchmarks alone is insufficient for safe deployment in high-stakes settings like biomedicine, where factual reliability must also be assessed.
- ResearcharXiv2026-07-01Enterprise · Quality assurance · +1
When Should Service Agents Reconsider? Difficulty-Routed Control in Customer-Service Operations · Qian Chen, Chengyuan Liu, Xin Yu
This paper addresses how autonomous AI customer-service agents—which now execute backend operations like refunds, cancellations, and order modifications—can balance speed for routine requests with safeguards for complex, error-prone ones. The authors propose a 'difficulty-routed' architecture that uses a lightweight router to send routine sessions down a fast baseline path while escalating operationally complex sessions to a more deliberative workflow featuring conflict-aware communication and write-triggered reconsideration. Evaluated on human-verified retail and airline tasks from the τ²-bench benchmark, the method improves reliability on conflicted service requests without broadly expanding interactions across all sessions. The findings suggest targeted control—concentrating deliberation before consequential backend writes rather than applying safeguards uniformly—is an effective design principle for enterprise AI service agents.
- ResearcharXiv2026-07-01Enterprise
Adoption and Impact of Command-Line AI Coding Agents: A Study of Microsoft's Early 2026 Rollout of Claude Code and GitHub Copilot CLI · Emerson Murphy-Hill, Jenna Butler, Alexandra Savelieva
This study examines Microsoft's early-2026 rollout of Claude Code and GitHub Copilot CLI across tens of thousands of engineers, finding that adoption spread primarily through social networks rather than demographic factors, and that retention was more closely tied to engineers' existing coding activity levels. Engineers who adopted these command-line AI coding agents merged roughly 24% more pull requests than they otherwise would have, a lift that persisted across the four-month observation window. The findings indicate that CLI coding agents produce measurable productivity gains rather than mere novelty effects, and that organizations should prioritize visible peer use as a central strategy when rolling out agentic tools at scale.
- ResearcharXiv2026-07-01Quality assurance
Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents? · Zhi Chen, Zhensu Sun, Yuling Shi et al.
This paper audits three repository-level coding-agent benchmarks—GSO, SWE-Perf, and SWE-fficiency—that measure performance optimization by comparing runtimes of agent-generated patches against reference patches. The authors find serious reliability problems: reference patches satisfy validity rules across machines for only a fraction of tasks (e.g., 11/140 for SWE-Perf), benchmark scoring rules strongly affect submission rankings with 9 of 28 pairwise comparisons disagreeing between the two benchmarks, and at least one public submission already matches or beats the reference patch on 85.3% of valid tasks. The findings suggest that leaderboard scores on these widely cited benchmarks can be misleading due to runtime instability, scoring rule sensitivity, and saturation, and the authors propose ways to identify tasks with more reliable signals and expose hidden performance gaps.
- ResearcharXiv2026-07-01Quality assurance
Distill to Detect: Exposing Stealth Biases in LLMs through Cartridge Distillation · Shayan Talaei, Abhinav Chinta, Devvrit Khatri et al.
This paper introduces Distill to Detect (D2D), a method for uncovering hidden preferential biases in large language models (LLMs) that have been covertly introduced somewhere in the model supply chain. D2D works by distilling the distributional difference between a suspected biased model and its unmodified base into a compact KV-cache prefix adapter (called a 'cartridge'), which concentrates and amplifies the bias signal until it becomes visible in generated text. The authors show D2D successfully surfaces stealth biases across multiple bias types and provide a theoretical explanation grounded in Fisher-weighted projection of the logit distribution shift. This matters for AI auditing and quality assurance because it offers a practical tool to detect hidden behavioral manipulations in deployed models that are otherwise invisible to text-based inspection or weight analysis.
- ResearcharXiv2026-07-01Enterprise · Quality assurance
TAG: A Lightweight Framework for Test-Driven Agentic Artifact Generation · Yaniv Melamed, Yoni Zukerman, Michal Shechter et al.
TAG is a lightweight framework for making LLM-generated structured artifacts—such as database queries, threat mappings, and entity schemas—reliable enough for production use. Its core principle is 'LLMs generate, we validate,' combining deterministic tests (for schema, syntax, and cross-reference checks) with LLM-based tests calibrated to replicate human expert judgment, so that when outputs fail, the LLM receives actionable error messages and refines its attempts. The framework is demonstrated on three artifact types in the cybersecurity domain—KQL query generation, MITRE ATT&CK mapping, and entity mapping—deployed in production at Microsoft Sentinel. By transforming manual human quality gates into scalable, reusable evaluation proxies, TAG offers a path to high-quality LLM outputs without sacrificing generation efficiency.
- ResearcharXiv2026-07-01Quality assurance
Skills Are Not Islands: Measuring Dependency and Risk in Agent Skill Supply Chains · Changguo Jia, Tianqi Zhao, Runzhi He et al.
This paper introduces Agent Skill Supply Chains (ASSCs), a framework for characterizing and analyzing the dependency graphs of skills used by Large Language Model (LLM) agents. The authors develop SkillDepAnalyzer, a tool that recovers skill metadata and dependency graphs from natural-language evidence, outperforming LLM-based baselines and package-centric Software Bill of Materials (SBOM) tools on the SKILL-DEP benchmark. Applying the tool to over 1.43 million skills, they find that skill metadata is 'activation-ready but governance-poor,' that dependency graphs contain concentrated reuse and hidden package inventory, and that security-relevant signals—including known malicious skills—are often hidden in dependencies rather than visible in the skill itself. The authors recommend typed dependency manifests, dependency-cluster management, risk-warning audit commands, and lockfile-like records to improve security and governance in LLM agent skill ecosystems.
- ResearcharXiv2026-07-01Quality assurance · Health · +1
Clinician-Level Agreement Without Clinical Caution: LLM Evaluator Limits in Medical AI Benchmarking · William Philipp, Finn Fassbender, Thorsten Langer et al.
This paper introduces MedQADE, a standardized open-response clinical benchmark for German comprising 3,800 items annotated by ten practising physicians and nine LLM evaluators, to test whether automated LLM-as-a-Judge systems replicate clinical calibration. The best-performing model (Gemini 3 Flash) reached statistical alignment close to the physician ceiling (κ = 0.694 vs. κ = 0.709), but automated evaluators showed near-absent clinical metacognition: unlike physicians, they assigned definitive scores in every case regardless of item difficulty rather than scaling abstention appropriately. The study also found systematic lineage-dependent biases where models preferentially scored architectural siblings, independent of language. The findings demonstrate that statistical alignment alone does not guarantee clinical caution, and that evaluator independence must be explicitly verified before LLM judges are trusted in medical AI benchmarking.
- ResearcharXiv2026-07-01Enterprise
Knowledge-Centric Information Systems · Mariano Garralda-Barrio
This paper argues that enterprise AI systems are driving a shift from traditional data engineering toward a new discipline called 'knowledge architecture,' in which organizational knowledge becomes active, executable infrastructure rather than a passive resource. The authors develop a conceptual model and taxonomy showing how classical data-engineering concepts must be redefined for knowledge artifacts: ETL becomes knowledge ingestion, change-data capture becomes knowledge change detection, lineage becomes provenance, and so on. Emerging formats such as LLM Wiki and the Open Knowledge Format (OKF) are cited as early evidence of this transition. The work matters for enterprises building AI systems that rely on agents, workflows, and models to retrieve, assemble, and act on organizational knowledge at scale.
- ResearcharXiv2026-07-01Quality assurance · AI policy · +1
From Runtime Records to Legal Findings: An Evidentiary-Adequacy Criterion for Agentic AI Oversight · Jeroen Janssen
This technical report proposes an 'evidentiary-adequacy criterion' for evaluating whether runtime logs and audit records generated by agentic AI systems are sufficient to support legally operative oversight findings. The criterion requires that a runtime record must carry both a typing that maps recorded events to legally relevant categories and the relational structure (e.g., provenance, authority, temporal validity) on which a legal determination depends — not merely that logs exist or are tamper-proof. The authors instantiate this criterion against EU AI Act oversight obligations, arguing that tamper-proof logs, generic process frameworks, and provenance structures alone are insufficient to establish the required findings. The work draws on concepts from requisite variety, the Good Regulator Theorem, and runtime verification theory to ground the argument formally.
- ResearcharXiv2026-07-01AI policy · Competition & Antitrust
Two AI Metrics Diverged: Will it Make All the Difference? · Alex Fogelson, Zachary A. Brown, Hans Gundlach et al.
This paper examines whether frontier AI models will grow increasingly unequal compared to smaller, cheaper models, or whether their capabilities will converge over time as compute scales. The authors show that the answer depends critically on how AI capabilities are measured: bounded performance metrics (those with a natural ceiling) tend to favor smaller 'meek' models, while unbounded metrics show frontier models maintaining or growing their lead indefinitely. They provide mathematical conditions for classifying metrics by this distinction and warn that many common metrics have bounded and unbounded counterparts that can suggest opposite conclusions. The policy implication is significant — if key capabilities like software engineering or rhetorical persuasiveness are unbounded in the terms that matter, frontier-level AI will remain concentrated among wealthy actors, while bounded capabilities will proliferate broadly.
- ResearcharXiv2026-07-01Quality assurance · Health · +1
Dynamic Bidirectional Pattern Memory: A Production-Scale Empirical Characterisation of Inference-Time Gating in Clinical NLP · Ali H. Lazem, William Teahan
This paper empirically evaluates a production-scale clinical NLP pipeline that pairs a large language model generator (Llama-3.3 70B) with a medical verifier (MMed-Llama-3.1 70B) across over 167,000 patient narratives, augmented with a lightweight inference-time memory that filters previously failed extractions. The study finds that learning filtering rules directly from verifier rejections failed silently at scale due to sparse signal, while a fixed clinical ontology filter successfully caught nearly 50,000 ontology-violating relations on a held-out set. Among five question-answering filter designs tested, only one succeeded — the version that checks whether a patient's extracted entities actually support the question asked, making it 1.84 times more likely to flag answers the verifier would reject. The core transferable finding is that a pre-generation gate is only selective when it tests the same evidence the verifier itself weighs, and the system is designed to flag rather than delete suspect extractions, preserving visibility for clinical review.
- ResearcharXiv2026-07-01AI policy · Algorithms & Automated Decisions
A field experiment of social influence and behavioral contagion with bots on Reddit · Hiroki Oda, Kinga Makovi, Taha Yasseri et al.
This field experiment on Reddit tested whether symbolic awards given by human or bot accounts—with different rationales (praising logic, emotional sensitivity, moral integrity, or a lottery)—could influence recipients' subsequent behavior and spread to other users. The study found that awards generally did not increase user activity or downstream impact, and bot-given lottery awards actually reduced them, though awards did encourage direct user-to-user communication. The results suggest online users are relatively resilient to simple behavioral manipulation by platform algorithms and artificial agents, but may remain vulnerable to more sophisticated human-simulating schemes. The authors conclude that transparent labeling of automated agents is essential for ethical platform governance.
- ResearcharXiv2026-07-01Quality assurance
Exploring the Semantic Gap in Agentic Data Systems: A Formative Study of Operationalization Failures in Analytical Workflows · Jalal Mahmud, Eser Kandogan
This paper investigates how AI agents powered by large language models fail when generating analytical workflows over databases. Across 236 analytical intents in finance, human resources, and public safety domains, the authors identify 153 recurring failures even when workflows were successfully generated and executed. These failures fall into five categories—comparative grounding, process reasoning, quantitative reasoning, role confusion, and policy grounding—revealing a semantic gap between what users intend analytically and what database schemas and data values actually represent. The findings suggest that future agentic data systems will need richer semantic representations to reliably translate analytical intent into accurate computation.
- ResearcharXiv2026-07-01Quality assurance
Phantom References: Hallucinated Citations That Survive Peer Review at Top-Tier Conferences · Mark Russinovich, Ram Shankar Siva Kumar, Ahmed Salem
This paper introduces RefChecker, an automated pipeline for detecting hallucinated citations—references to non-existent works or papers with substantially wrong author lists—in peer-reviewed AI and security conference proceedings (ICLR, ICML, NeurIPS, USENIX Security). The authors find that while hallucinated references are rare at the individual reference level (usually below 1%), proceedings are large enough that roughly one in twenty NeurIPS and USENIX Security papers in 2025 contains at least two likely hallucinated citations, with post-ChatGPT increases observed across several venues and hallucinations appearing even in award-winning papers. The study demonstrates that peer review alone does not reliably catch these integrity failures, but automated auditing is tractable at approximately $0.04 per paper, and the open-sourced tool enables routine pre-publication verification.
- ResearcharXiv2026-07-01Copyright & Creative Work
What's a Credit Worth? A Market Framework for Attribution-Aware Compensation in Generative Music · Luyang Zhang, Xirui Jiang, Junwei Deng et al.
This paper develops an economic framework for compensating music creators whose recordings are used to train generative AI models. Payment to each creator is based on a data-attribution score estimating their catalog's contribution to model outputs, with the informativeness (signal-to-noise ratio) of that score determining whether compensation takes the form of royalties or fixed-fee licensing. The authors show that more accurate attribution improves welfare for both creators and platforms, but under multi-platform competition a platform only captures those gains when its attribution signal is the most precise in the market. Empirical experiments with acoustic and symbolic music generation models confirm that noisy attribution pushes payments toward fixed-fee licensing and reduces welfare, motivating further research on better attribution methods.
- ResearcharXiv2026-07-01Quality assurance · Safety & Harms · +1
Safe Alone, Unsafe Together: Safeguarding Against Implicit Toxicity When Benign Images Combine · Jiaxian Lv, Shiyao Cui, Yingkang Wang et al.
This paper identifies and formalizes a new content-safety problem called multi-image implicit toxicity (MIIT), where individual images each appear benign but produce harmful semantics when viewed together—a pattern increasingly common in social media. The authors build MIIT-dataset, an automated image-only benchmark spanning seven risk categories, and train MiShield using progressively distilled reasoning supervision to detect these emergent harms with explicit entity-level analysis. Experiments show that MiShield-8B outperforms existing commercial moderation APIs and larger-scale models, demonstrating practical value for moderating multi-image content. This work directly advances the quality and reliability of AI-powered content moderation systems.