News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
Vision-language models for chest radiography do not always need the image
Mahshad Lotfinia, Sebastian Ziegelmayer, Lisa Adams et al.
arXiv · 2026-06-16
This paper challenges the assumption that high benchmark accuracy on chest radiograph tasks means medical vision-language models (VLMs) actually use the image. Using a causal audit involving image occlusion, irrelevant-region occlusion, and patient-image swapping, the authors show that a text-only model with no image access comes within 5.7 accuracy points of the best multimodal system, and a 119-billion-parameter multimodal model is statistically indistinguishable from a 7-billion text-only baseline. The audit categorizes nine systems: three that ignore the image entirely, one that is unstable, and five that use it selectively, with findings replicated across a second dataset, resolution, and prompt phrasing. The authors conclude that grounding audits—not accuracy—should gate clinical deployment, as reported confidence only flags ungrounded answers when a model genuinely uses the image.
- Quality assurance
- Certifications
Research
Mapping the Artificial Intelligence Divide in Africa: Infrastructure, Accessibility and Capacity
Abayomi O. Agbeyangi, Jose M. Lukose
arXiv · 2026-06-16
This paper empirically maps the 'AI divide' in Africa across three dimensions: physical infrastructure, accessibility, and human capacity. Key findings include only 38% internet penetration, less than 1% of global data centres located in Africa, high data costs relative to income, gender-based digital divides, and a lack of NLP models supporting African languages. Despite these barriers, the paper identifies positive grassroots trends such as local startups and university-led AI initiatives. Based on these findings, the authors offer concrete policy recommendations to foster a more equitable and comprehensive AI ecosystem across the continent.
- AI policy
- Workforce
Research
FacProcessTwin: An LLM-Based System for Process Twin Development
Yash Pulse, Yong-Bin Kang, Abhik Banerjee et al.
arXiv · 2026-06-16
FacProcessTwin is an LLM-based system that automates the development of process digital twins for manufacturing facilities by extracting process models from plant documentation and natural-language operator input, then binding those models to live operational data. In a real-world case study with an Australian food manufacturer covering 16 production process flows, the system achieved a mean F1 score of 95.2% for process model accuracy and reduced twin development time to roughly one-sixth of the manual effort. A human-in-the-loop governance layer ensures safety-critical data bindings remain correct: at ambiguous points where a baseline approach mis-binds 75% of the time, FacProcessTwin defers to the operator and achieves zero mis-bindings. The work demonstrates that LLMs can substantially lower the cost and time of deploying process twins while maintaining the accuracy and safety oversight required in manufacturing environments.
- Enterprise
- Quality assurance
Research
Understanding LLMs in Title-Abstract Screening: From Disagreements to Recommendations
Mika Mäntylä, Patricia Matsubara, Katia Romero Felizardo et al.
arXiv · 2026-06-16
This study investigates why large language models (LLMs) fail at title-abstract screening in systematic reviews (SRs) for software engineering, going beyond simple accuracy metrics to qualitatively analyze disagreements between LLMs and human researchers across six SRs and over 1,000 papers. Screening was performed independently by humans and LLMs in zero-shot mode, yielding Kappa agreement values ranging from 0.52 to 0.77, and qualitative analysis identified recurring failure causes including boundary ambiguity in key terms, keyword overemphasization, and incorrect topic inference. Based on these findings, the authors propose actionable recommendations such as validating semantic understanding before deployment, running multiple LLMs, and focusing validation on borderline cases. The work highlights that community-level normative guidelines for LLM use in systematic reviews are still needed, with implications for research quality assurance and evidence synthesis workflows.
- Quality assurance
Research
Scaling Enterprise Agent Routing: Degradation, Diagnosis, and Recovery
Kellen Gillespie, Robyn Perry
arXiv · 2026-06-16
This paper investigates how routing accuracy degrades in a deployed enterprise AI assistant as the number of specialized agents and tools grows. Testing three frontier language models on a catalog of 110 agents and 584 tools, the authors find that routing F1 on under-specified requests drops 16–23 percentage points as the catalog scales. They decompose this degradation into a retrieval gap and a confusion gap, and show that embedding-based shortlisting recovers 10–11 percentage points of F1 at full scale. A production annotation study with 1,435 human-labeled utterances confirms real-world recovery of 10–17 percentage points, validating the approach on live traffic.
- Enterprise
- Quality assurance
Research
LLM-as-Judge in Education: A Curriculum-Grounded Marking Pipeline
Xiwei Xu, Chen Wang, Jacky Jiang et al.
arXiv · 2026-06-16
This paper presents a curriculum-grounded LLM-as-Judge pipeline for automated marking of student responses in high-stakes exam preparation contexts. The pipeline grounds LLM outputs in authorised curriculum artefacts—such as syllabus verbs, performance band descriptors, glossary definitions, and marking-guideline principles—to generate question-specific rubrics and evaluate student answers. Preliminary evaluation shows the pipeline produces marking outcomes comparable to human tutors, with justifications more traceable to official curriculum standards. The system has been integrated into an online study platform, with early deployment data offering initial insights into operational usage and manual overrides.
- Quality assurance
- Enterprise
Research
Simulated Customers Never Walk Away: Decision Fidelity of LLM User Simulators Measured Against Real Purchase Outcomes
Liang Chen
arXiv · 2026-06-16
This paper investigates whether large language model (LLM) user simulators accurately replicate the decision-making behavior of real customers in high-stakes conversational settings. Using 2,790 production conversations between an LLM sales agent and real customers—including 793 with verified payment outcomes—the authors find a systematic 'disengagement deficit': simulators closely reproduce the behavior of actual buyers but significantly inflate non-buyers toward purchase-oriented engagement, halving expressed resistance (25.1% to 13.5%) and nearly doubling deliberation (21.9% to 40.1%). This bias persists across model families (e.g., DeepSeek: d=0.41, p=0.008) and is not resolved by simply instructing simulators that they may disengage. The findings matter because AI sales and persuasion agents trained or evaluated against such simulators will systematically overestimate funnel progress precisely among the customers most likely to walk away.
- Quality assurance
- Enterprise
Research
The Consistency Dilemma in LLMs: Generator-Evaluator Agreement and Vulnerability to Mistakes
Marina Mancoridis, Zoë Hitzig
arXiv · 2026-06-16
This paper introduces 'generator-evaluator self-consistency,' a measure of whether large language models apply concepts the same way when generating outputs as when evaluating those outputs. Testing 10 frontier models across 491 concepts, the authors find substantial variation in this self-consistency metric. Critically, in a clinical setting using physician-validated mistakes (Proniakin et al., 2025), models with higher self-consistency are paradoxically more vulnerable to mistakes—revealing a 'consistency dilemma' where being operationally consistent does not mean being safe to deploy. This finding has significant implications for agentic AI pipelines that rely on models self-evaluating their outputs without external verification.
- Quality assurance
- AI policy
Research
AIPatient Arena: EHR-grounded evaluation of large language models in end-to-end clinical consultation workflows
Jiahui Niu, Huizi Yu, Wenkong Wang et al.
arXiv · 2026-06-16
AIPatient Arena is an EHR-grounded evaluation framework that assesses large language models (LLMs) across eight dimensions of clinical competence in multi-turn physician-patient consultation workflows. Applied to a primary cohort of 437 patients and two validation cohorts, the framework found that LLMs performed well on interview questioning skills, ethical conduct, and clarity of explanation, but showed persistent weaknesses in handling ambiguous responses, information coverage, and diagnostic accuracy and reasoning. Process-based evaluation revealed recurring failures such as repetitive questioning, omission of past medical history, and inadequate handling of uncertainty. The paper argues that final-answer accuracy alone is insufficient for evaluating clinical readiness, and proposes this framework as a workflow-oriented pre-deployment evaluation tool for medical LLMs.
- Quality assurance
- Certifications
Research
PARSE: Provenance-Aware Retrieval Sanitization for Professional Domain LLM Agents
Aaditya Pai
arXiv · 2026-06-16
PARSE addresses a critical gap in AI security: existing prompt injection defenses tested on synthetic benchmarks fail to generalize to real enterprise documents such as SEC filings, Federal Register rules, and PubMed abstracts. The authors benchmark 122 tasks across five professional domains and show that paraphrasing—the strongest known synthetic-benchmark defense—produces no statistically significant attack reduction on real documents (p=0.500) while degrading utility from 91.8% to 82.8%. Their proposed system, PARSE, uses provenance-aware, fact-preserving sanitization to classify sentences by injection risk, extract structured facts, and verify preservation, achieving a 38% reduction in attack success rate (from 25.4% to 15.6%) at 86.9% utility with statistical significance (p=0.014). The findings carry a direct practical warning: AI security defenses for enterprise LLM agents must be evaluated on domain-matched real documents rather than synthetic proxies.
- Enterprise
- Quality assurance
Research
AI Transparency: Governance Compliance or Stakeholder Requirements?
Muneera Bano, Didar Zowghi
arXiv · 2026-06-16
This paper examines 92 AI transparency statements published by Australian Government agencies under a national AI governance mandate, finding that structural compliance with disclosure requirements does not equate to meaningful transparency for all stakeholders. The authors introduce the Risk–Control–Involvement–Need (RCIN) framework to classify stakeholders by their structural position and transparency needs, revealing that criteria serving high-control stakeholders are consistently met while criteria most critical for high-risk, low-control stakeholders are fewer and less substantively addressed. The authors call this the 'Transparency Illusion'—where compliant artefacts create an appearance of transparency without adequately serving those most exposed to AI-supported decisions. The study reframes transparency as a stakeholder-calibrated validation problem, with direct implications for how AI governance mandates are designed and assessed.
- AI policy
- Quality assurance
Research
Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
Xi Chu, Yupeng Hou
arXiv · 2026-06-16
This paper investigates how brands compete for recommendations within large language models (LLMs) using skincare products as a test case across GPT-4o-mini, Claude Sonnet, and Gemini 3 Flash. The authors find that well-known brands achieve a 'Conditional Monopoly' — receiving 100% of recommendations when products share identical specifications — but this dominance can be disrupted by a competitor gaining even a marginal rating advantage or by using authority-style marketing language, including fabricated clinical-evidence claims. When multiple brands simultaneously adopt the same generative engine optimization (GEO) strategies, individual payoffs collapse from +0.802 to +0.007 in the authors' payoff proxy, creating a social dilemma. The findings suggest GEO is not merely a security concern but an emerging marketing practice with significant implications for market competition and consumer information integrity.
- Enterprise
- AI policy
Research
Model Validation of Agentic AI Systems: A POMDP-Based Framework for Belief-State, Forecast, and Policy Validation
Matthew Francis Dixon
arXiv · 2026-06-16
This paper proposes a structured validation framework for agentic AI systems—autonomous agents that form beliefs, make forecasts, and take actions over time—grounded in Partially Observable Markov Decision Processes (POMDPs). It decomposes autonomous decision-making into distinct components (information, beliefs, forecasts, actions, and utility) so each can be validated independently, and develops a model-risk taxonomy covering state-space, filtering, forecast, policy, utility-specification, and parameter risks. A portfolio-management case study demonstrates the framework, with empirical results indicating that latent-state inference independently contributes to decision quality and that findings are robust across parameter values. The work provides a practical foundation for extending established model risk management concepts to agentic AI, with direct relevance to governance, monitoring, and validation of autonomous systems.
- Enterprise
- AI policy
- Quality assurance
Research
Artificial Intelligence, Skills, and Labor Mobility: Understanding the Transformation of Work
David Marguerit
Open Repository and Bibliography (University of Luxembourg) · 2026-06-16
This PhD dissertation examines how AI reshapes labor markets, skills, and education across the U.S. and Europe. It finds that augmentation AI (which enhances worker output) creates new work and raises wages primarily for high-skilled workers, while automation AI increases employment but depresses wages and harms low-skilled workers. In Europe, AI exposure shifts employer demand toward AI, Data, and Prediction skills while reducing demand for Social skills. AI-driven labor-market signals also propagate upstream to higher education, affecting student enrollment choices and program openings or closures depending on whether AI automates or augments the relevant field.
- Workforce
- AI policy
Research
Analytics for Quality Assurance for Item Pools (AQuAP): Monitoring and Maintaining Item Bank Health in AI-Driven Assessment Systems
Alina von Davier, Xiaowan Zhang, Yigal Attali et al.
arXiv (Cornell University) · 2026-06-16
This paper introduces AQuAP (Analytics for Quality Assurance for Item Pools), a dashboard environment designed to monitor item quality and item bank health in AI-driven, high-stakes educational assessments. AQuAP supports the Item Factory framework for automated and human-supported test development, translating psychometric concepts—such as Effective Bank Size, maximum exposure, and rarely-administered fraction—into operational quality assurance tools. The system is illustrated using the Duolingo English Test and aims to ensure item bank security, diversity, and efficiency in high-volume testing programs. This work matters because it shows how operational analytics can sustain the integrity and health of AI-generated item pools at scale.
- Quality assurance
- Certifications
Research
PriyankaPSurve/CEDAR42001: CEDAR-42001: From ISO/IEC 42001 Conformity to Architecture-Aware, Audit-Visible Assurance Posture for AI Cyber-Physical System
PriyankaPSurve
Zenodo (CERN European Organization for Nuclear Research) · 2026-06-16
CEDAR-42001 is a two-stage method that converts ISO/IEC 42001 audit evidence for AI-enabled cyber-physical systems into an architecture-aware assurance posture, going beyond simple conformity assessment. The method adds architectural attribution, maturity profiling, risk-proportionate targets, and action recommendations to each audit row, revealing that while 89.9% of audit rows were conforming, only 34.3% reached a baseline High-Assurance category. Applied retrospectively to the 2023 Cruise robotaxi incident, the method mapped documented concerns across governance, perception, decision-making, and human oversight to layer-specific actions. This matters because it helps organizations identify where audit evidence warrants deeper technical assurance, organizational improvement, or remediation beyond what conformity certification alone reveals.
- Certifications
- Quality assurance
- AI policy
Research
Agentic AI and Circular Procurement Performance: An Empirical Study
Surajit Bag, Susmi Routray, Andrea Chiarini et al.
Business Strategy and the Environment · 2026-06-16
This empirical study examines how adopting agentic AI in industrial purchasing affects circular procurement performance, drawing on resource-orchestration theory and data from a developing nation analyzed via covariance-based SEM and Process analytical methods. The findings show firms benefit most from agentic AI when three conditions are met: structured resource scanning and evaluation processes, integrated cross-functional procurement systems, and established supplier collaboration routines. The study is the first to theorize the relationship between agentic AI adoption and circular procurement performance, offering practical guidance for purchasing and supply managers on formulating policies and standard operating procedures.
- Enterprise
- AI policy
Research
Operationalizing NIST AI RMF 1.0 for Federal Training and Academic AI Deployers
Ruchir Bakshi
Zenodo (CERN European Organization for Nuclear Research) · 2026-06-16
This paper operationalizes the NIST AI Risk Management Framework (AI RMF 1.0) specifically for federal instructional-design, training, and academic units that deploy AI tools such as adaptive learning systems, AI tutoring, and AI-text detection—organizations not explicitly addressed by the existing framework. The authors develop a use-case AI RMF Profile that applies a deployer lens to all 72 Playbook subcategories, retaining 71 as in-scope and providing applicability analysis alongside blank current-state and target-state fields for adopters to complete. A reproducible build script ensures the Profile stays aligned with the machine-readable NIST Playbook source, eliminating transcription drift. The work matters because it lowers the operationalization burden for a broad class of public-sector AI deployers seeking structured, voluntary risk management guidance.
- AI policy
- Certifications
- Enterprise
Research
The Emergence of Governability Assurance for Autonomous Systems: Evidence from AI Assurance, Safety Cases, Continuous Assurance, Standards, and Autonomous Systems Research (2026)
Andreas Blumer
Zenodo (CERN European Organization for Nuclear Research) · 2026-06-16
This paper investigates whether 'Governability Assurance' has emerged as a recognized, distinct category within assurance frameworks for autonomous systems. Reviewing evidence across AI assurance, safety cases, continuous assurance, dynamic certification, standards, and regulatory initiatives, the authors find that while confidence in safety, security, compliance, and trustworthiness is relatively well-developed, assurance specifically targeting observability, controllability, intervention capability, accountability, recoverability, and legitimate authority remains comparatively underdeveloped. The paper identifies this gap—termed the 'Governability Assurance Gap'—and argues that although conceptual foundations exist across adjacent assurance disciplines, Governability Assurance is not yet widely recognized as a distinct field. The authors conclude that Governability Assurance may represent an emerging discipline warranting dedicated attention as autonomous systems become more prevalent.
- Certifications
- AI policy
- Quality assurance
Research
Operationalizing NIST AI RMF 1.0 for Federal Training and Academic AI Deployers
Ruchir Bakshi
OSF Preprints (OSF Preprints) · 2026-06-16
This paper develops a sector-specific AI Risk Management Framework (RMF) Profile tailored for federal training, instructional-design, and academic units that deploy AI tools such as tutoring systems, adaptive learning platforms, and AI-text detection—organizations the NIST AI RMF 1.0 does not explicitly address. The authors apply a deployer lens to all 72 subcategories of the NIST AI RMF Playbook, retaining 71 as in-scope and providing applicability analyses paired with blank current- and target-state fields for adopters to complete. A reproducible build script ensures the profile stays faithful to the official machine-readable Playbook, reducing transcription drift. The work offers a voluntary, reusable template to help these organizations operationalize federal AI risk management guidance without constituting a certification or compliance mandate.
- AI policy
- Certifications
- Workforce
- Enterprise
Research
Managing Election Misinformation: Comparing Artificial Intelligence Policies Across Indian News Organizations
Tanisha Mathur, P Anil Kumar
Zenodo (CERN European Organization for Nuclear Research) · 2026-06-16
This paper compares AI-related misinformation policies across three Indian media institutions—The Quint, the Press Information Bureau, and the Press Council of India—finding a fragmented regulatory landscape with no consistent standard for handling generative AI and deepfakes during elections. The Quint prohibits journalists from using generative tools, the government mandates synthetic media tagging and removal, and the Press Council relies solely on journalistic ethics with no technical rules. The study concludes that existing self-regulation is too slow for the pace of modern elections and proposes a minimal three-point interim policy focused on human verification and disclosure. The findings are relevant to how media policy frameworks can be strengthened to address AI-driven political misinformation.
- AI policy
- Quality assurance
Research
Managing Election Misinformation: Comparing Artificial Intelligence Policies Across Indian News Organizations
Tanisha Mathur, P Anil Kumar
Zenodo (CERN European Organization for Nuclear Research) · 2026-06-16
This paper examines how Indian news organizations and regulatory bodies address AI-generated election misinformation by comparing three policy frameworks: The Quint's internal digital newsroom rules, the government's Press Information Bureau regulations, and the Press Council of India's ethics code. The study finds a fragmented regulatory landscape—The Quint bans generative AI content creation, the government mandates synthetic media tagging and removal, while the Press Council relies solely on journalistic ethics with no technical rules. The authors conclude that existing self-regulation is too slow for fast-moving elections and propose a minimal three-point interim policy focused on human verification and disclosure. The research highlights a significant policy gap, noting that legacy newspapers largely do not publish technology policies, limiting comprehensive analysis.
- AI policy
- Quality assurance
Research
UG AIMS: A Scalable AI Governance and Certification Framework for Africa
Erich Barlow
Zenodo (CERN European Organization for Nuclear Research) · 2026-06-16
This white paper introduces the Uganda AI Management and Assurance Scheme (UG-AIMS), a scalable AI governance and certification framework designed to translate international AI governance principles—anchored in ISO/IEC 42001—into auditable, locally relevant controls for Uganda and the broader African market. The framework addresses governance risks including algorithmic bias, explainability, data protection, cybersecurity, and accountability across sectors such as healthcare, agriculture, and financial services. UG-AIMS is proposed as a regional reference model that pairs a common standards-based baseline with jurisdiction-specific legal overlays, aligned with the African Union's continental AI strategy. The paper recommends piloting the framework through a multi-stakeholder working group, sector pilots, a certification handbook, and capacity building for auditors, regulators, and vendors.
- Certifications
- AI policy
- Quality assurance
- Enterprise
Research
BRIDGING DETERMINISTIC CERTIFICATION AND PROBABILISTIC AI: A HYBRID ASSURANCE FRAMEWORK FOR SAFETY-CRITICAL AVIONICS
Shyamala Bai Kotin
Zenodo (CERN European Organization for Nuclear Research) · 2026-06-16
This paper proposes the Hybrid Deterministic-Probabilistic Assurance Framework (HDPAF), a five-layer architecture designed to bridge the gap between traditional aviation certification standards (DO-178C, ARP4754A) and the probabilistic nature of AI and machine learning systems used in safety-critical avionics. The framework addresses key challenges including dataset governance, constrained AI model development, robustness testing, runtime monitoring with deterministic fallback, and alignment with FAA and EASA regulatory roadmaps. Rather than replacing existing standards, HDPAF extends them to provide a traceable, lifecycle-aware pathway for certifying AI components in high-criticality aviation environments. This work matters because it offers a structured method to close the regulatory gap that currently limits the safe deployment of AI in avionics.
- Certifications
- Quality assurance
- AI policy
Research
AI Adoption Across a Multinational Workforce: Sociotechnical Conditions for GenAI Acceptance in Human Resources
Dalia Ali, Maria José Rodríguez Velázquez, Manoel Horta Ribeiro et al.
arXiv (Cornell University) · 2026-06-16
This paper examines GenAI adoption in a multinational tech company's HR department during a live transition from a legacy search system to a GenAI-supported one, using search logs, surveys, and interviews. Findings show adoption varied based on employees' roles, spoken language, and tenure, and that trust in GenAI answers was built through source-checking, comparing systems, and seeking colleague input. The research demonstrates that factors like situational fit, search literacy, content quality, and employee training shape whether workers benefit from or are left behind by AI tools. The authors argue organizations must design AI systems with context-sensitive inclusivity and treat organizational knowledge infrastructure as part of AI infrastructure to ensure accountability and usability in high-stakes settings like HR.
- Workforce
- Enterprise
- AI policy