News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
Evaluating the efficacy of risk management practices for generative AI integration in sustainable construction projects: a fuzzy set theory approach
Mohamed Abdelwahab Hassan Mohamed, M. K. S. Al-Mhdawi
Frontiers in Built Environment · 2026-09-03
This study evaluates how well existing risk management (RM) practices in sustainable construction projects (SCPs) are prepared to handle risks introduced by Generative AI (GenAI) integration. Using a fuzzy-based multi-criteria assessment model informed by a structured survey of 80 construction experts, the researchers identified 30 GenAI-related risk factors across five categories—input quality, technological adaptability, ethical and governance, information integrity, and financial risks—and found that current RM practices exhibit only a low-to-medium level of perceived preparedness. The study proposes best-practice response strategies for each risk category and positions its assessment model as an early-warning tool for organizations adopting GenAI in construction. The findings highlight a significant gap between emerging GenAI risks and the readiness of existing RM structures in a sector characterized by multi-stakeholder complexity and sustainability requirements.
- Quality assurance
- AI policy
Research
A human rights‐based approach to AI in fintech
Katayoon Beshkardana, Rachel Chambers
American Business Law Journal · 2026-09-03
This article examines the human rights risks posed by AI systems used in financial technology (fintech), including algorithmic discrimination, financial exclusion, lack of transparency, and privacy breaches. It maps existing law and soft law governing AI in fintech, identifies regulatory gaps—particularly for firms operating outside traditional banking oversight—and evaluates whether the business and human rights framework can help address those gaps. The article concludes by recommending adoption of a human rights-based approach to govern AI across the fintech ecosystem. The findings are relevant to policymakers and regulators seeking to address uneven oversight of AI-driven financial services.
- AI policy
Research
Knowledge economy, artificial intelligence and media industries: pathways to organizational sustainability in the era of digital transformation
Abdulla Ebrahim Altaher, Elsir Ali Saad Mohamed, Rania Dafalla et al.
Frontiers in Communication · 2026-09-03
This study surveyed 390 media professionals in Bahrain and the UAE to examine how AI adoption and digital transformation affect organizational sustainability in Gulf media organizations. Key findings show strong perceived benefits—86.7% agree AI improves efficiency and cost reduction, and 82.9% say digital transformation improves knowledge management—but significant barriers remain, including training gaps (75.7%), infrastructure constraints (61.3%), insufficient AI regulatory policies (69.0%), and widespread concern about job displacement (62.8%). Only 39% of respondents feel ethical considerations around AI are adequately addressed, and statistical analysis confirms that perceived ethical inadequacy correlates negatively with job-security anxiety. The authors propose a three-pathway model covering capability, governance, and culture as a framework for policymakers and media leaders navigating AI-driven transformation.
- Workforce
- AI policy
- Enterprise
Research
Operationalising AI Ethical Principles in Higher Education with the UNESCO Recommendation on the Ethics of Artificial Intelligence by Policy Makers and Global Universities
Tatjana Titareva, Grace Thomson, Irēna Barkāne et al.
arXiv · 2026-09-03
This policy brief examines how UNESCO's 2021 Recommendation on the Ethics of Artificial Intelligence—the first global AI ethics standard, covering 194 member states—can be operationalised in higher education across 16 countries on five continents. The analysis maps national regulations and institutional practices against UNESCO's ten core ethical principles, revealing how governments and universities are adapting global standards to local social, cultural, and political contexts. It concludes with five priority recommendations for policymakers and institutional leaders, spanning human rights-based frameworks, data governance, human oversight, universal AI literacy, and ethical procurement. The paper argues that effective AI governance in higher education requires coordinated, multi-level action combining bottom-up institutional initiatives with top-down government legislation and technical support.
- AI policy
- Certifications
Research
The Responsible AI Divide: Adoption Without Accountability in African Digital Economies
Julius Osi Abu, Kayode Abiodun Oladapo, Frances Chinaza Agba et al.
Zenodo (CERN European Organization for Nuclear Research) · 2026-09-03
This paper argues that the central challenge for AI in African digital economies is not access but governance. Analyzing eight countries under a most-diverse-systems design, the authors document an adoption-governance gap: consumer AI tools are spreading rapidly while binding AI-specific regulation remains absent in nearly every African Union member state. The paper identifies four structural contributors to this gap, examines five risk dimensions specific to African deployment contexts, and proposes a five-principle sovereignty-respecting governance framework with twelve actor-assigned recommendations, most actionable under existing law and commercial practice.
- AI policy
- Enterprise
Research
Counterfactual Fairness Audits of Multi-Step Clinical LLM Agents Require a Measured Per-Action Instability Floor
Rohith Reddy Bellibaltu, Manpreet Singh, Deepak Parashar et al.
arXiv · 2026-09-02
This paper investigates a fundamental methodological problem in using counterfactual audits to test whether clinical AI agents treat demographically identical patients differently. The authors show that the standard metric — the 'flip rate' (how often an agent's action changes when only a demographic descriptor changes) — is uninterpretable without a baseline, because re-running the exact same inputs ten times already produced action changes in 8.7% of cases due to model instability alone, with variation across action types ranging from 2.2% to 17.9%. A second model yielded a similar instability floor (6.7%) with nearly identical action rankings (Spearman 0.94), indicating this is a systemic issue rather than an artifact of one system. The findings mean that any counterfactual fairness audit of a clinical LLM agent must report a per-action instability floor alongside its flip rate, or the results cannot be interpreted as evidence of demographic disparity.
- Quality assurance
- AI policy
Research
The Analyst in the Prompt: Role, Retrieval, and Memory Biases in LLM Financial Analysis
Ahmed Asaad, Amr Mohamed, Yang Zhang et al.
arXiv · 2026-09-02
This paper investigates how LLM personalization features—such as role prompts, user profiles, and memory context—distort financial analysis of SEC filings. Testing 3,575 SEC filings across twelve LLMs, the researchers find that most bias comes from how models interpret evidence under different user contexts rather than from selecting different evidence in the first place. Two mitigation strategies (framing investor mindset as a user profile rather than an assistant role, and separating evidence-based from personalized outputs) reduce but do not eliminate this spillover, with effectiveness varying across models. The findings matter for any enterprise or policy setting where LLMs are used to support high-stakes, evidence-based financial decisions.
- Enterprise
- Quality assurance
Research
MasterControl Seventeen Every Time
MasterControl AI Lab
arXiv · 2026-09-02
This paper compares a policy-governed enterprise analytics approach against runtime language model agents for answering analytical questions. In the governed setup, a language model interprets the question while a deterministic policy selects and executes a pre-approved analytical program, returning both results and evidence. Across 440 runs, no runtime-planning episodes (330 total) satisfied the full answer-and-evidence contract across all test datasets, while the policy-executed analyzer succeeded in all 110 of its runs. The authors frame this as a configuration-specific finding, arguing that deterministic policy execution offers replayability and reliability within a defined analytical class without claiming runtime agents cannot work under different designs.
- Enterprise
- Quality assurance
Research
Open Problems in AI Risk Modeling: Insights from a Workshop on the Technical Foundations of AI Risk Modeling
Krystal Jackson, Deepika Raman, Jakub Kryś et al.
arXiv · 2026-09-02
This paper examines the current state and open challenges in building quantitative risk models for assessing societal risks from advanced AI systems, drawing on a workshop with 22 experts. The authors review five relevant research traditions—including probabilistic risk assessment, cybersecurity risk quantification, and Bayesian causal inference—and compare two leading modeling approaches: scenario-based risk estimation and Bayesian network-based threshold setting. They identify gaps in model structure, evidence integration, validation, and governance that currently prevent rigorous risk modeling from being adopted in practice. The paper argues that progress will require combining quantitative modeling with independent evaluation, transparent disclosure, and institutions capable of maintaining and updating risk models over time.
- AI policy
- Certifications
Research
Untangling the Mechanisms of Misleading Context in Medical Question Answering
Robin Linzmayer, Noémie Elhadad
arXiv · 2026-09-02
This paper investigates how misleading context—specifically fabricated evidence and bare assertions—corrupts the medical reasoning of large language models on MedMisBench, a clinician-reviewed benchmark of 8,627 questions. Three reasoning models were tested, and all were more susceptible to bare assertions than to fabricated evidence, adopting the asserted answer 10 to 27 percentage points more often. Misleading cues were disclosed in reasoning traces far more often than in final responses (81–98% vs. 7–90%), and an LLM monitor could catch 78% of corrupted decisions at 5% false positives when reading an open model's full reasoning trace—compared to at most 32% from responses alone. The findings highlight a critical safety gap: the most dangerous misleading inputs are disclosed least in model outputs, and reliable monitoring depends on access to full reasoning traces that frontier providers withhold.
- Quality assurance
- AI policy
Research
Incremental Pooled LLM Evaluation for Cost-Effective Retrieval Model Selection
Max Nelson, Hanoz Bhathena, Aviral Joshi et al.
arXiv · 2026-09-02
This paper proposes and validates a pooled LLM evaluation framework for selecting retrieval models in production RAG systems, where an LLM judges the union of retrieved documents and judgments are reused incrementally as new candidate systems are added. Tested across four retrieval benchmarks with 11 systems and deployed to compare 62 configurations for a financial news QA system, the approach preserves 97% of pairwise system orderings compared to gold-standard evaluation. Document overlap across systems yields 65–80% judgment reuse and up to 4.9x lower evaluation cost, making it practical to benchmark new retrieval candidates without re-judging previously assessed documents. This offers enterprises a scalable, cost-effective workflow for continuous retrieval model selection in deployed RAG pipelines.
- Enterprise
- Quality assurance
Research
CORAL: An LLM-Native Harness for Production Recommender Systems
Muhammad Rafay Azhar, Yuhang Zhou, Gilbert Jiang et al.
arXiv · 2026-09-02
CORAL is an LLM-native agentic system that automates the continual optimization of production recommender systems. Each cycle, the agent observes live operating signals, reasons over a memory of past decisions, and uses a numerical optimizer to reconfigure retrieval, ranking, and serving within a fixed operating budget. Evaluated via A/B experiments across two large-scale social platforms, CORAL improved engagement at no added serving cost on one platform and reduced serving cost without degrading engagement on the other, with performance improving as the loop iterates. The results suggest a single agentic loop can automate optimization work traditionally performed by human algorithm engineers, reducing reliance on slow, manual experimentation processes.
- Enterprise
- Workforce
Research
Toward Collective-Centric Evaluation of Preference Inference for Participatory Democracy
Pierre-Antoine Lequeu, Salim Hafid, Paul Lerner et al.
arXiv · 2026-09-02
This paper addresses a critical gap in AI-assisted participatory democracy platforms (such as Polis and Remesh), where sparse voting data leads platforms to use Preference Inference (PI) models to predict missing votes. The authors find that models with comparable individual-prediction accuracy can differ substantially in how well they preserve the collective preference landscape — including patterns of consensus, conflict, and minority support — meaning accuracy alone is insufficient for evaluating PI in democratic settings. They introduce a collective-centric evaluation framework that measures whether inferred votes preserve salient properties of the broader preference landscape, and contribute the largest multilingual dataset of its kind, spanning over 90,000 participants, 1 million votes, and 22 languages across four consultations. The work aims to support AI systems that can scale deliberation without compromising the integrity of democratic outcomes.
- AI policy
- Quality assurance
Research
Door-in-the-Face Requests and Refusal Behaviour in Large Language Models
Til Jordan
arXiv · 2026-09-02
This paper tests whether the 'door-in-the-face' social influence technique — where refusing a large request makes compliance with a smaller follow-up more likely — applies to large language models. Testing nine production models from Anthropic, OpenAI, and Google, the researchers find model-family-dependent effects: Anthropic's Opus 5 complies with a smaller follow-up request 65.8% of the time after refusing a larger one (vs. 29.3% when asked directly), while frontier models from OpenAI and Google show the opposite, with compliance dropping 15.5 to 23.0 points. A key finding is that reframing 265 refused requests for usable instructions as requests for explanations of the same topic bypassed refusals in 263 cases. The results show that human persuasion techniques transfer to LLMs inconsistently across model families, with significant implications for AI safety and content moderation policy.
- AI policy
- Quality assurance
Research
From Tokens to Semantics: Leveraging Complementary Signals for Hallucination Detection in Black-Box LLMs
Urja Pawar, Rajitha Ramanayake, Owen O'Neill et al.
arXiv · 2026-09-02
This paper investigates hallucination detection in large language models (LLMs) when only black-box API access is available and no trusted reference document exists. The authors study two complementary signals—semantic entropy (disagreement among sampled response meanings) and token log-probability-based uncertainty—and propose or evaluate four detection methods: TopK, CoCoA, Gated, and Stacked. Evaluated across seven benchmarks using four language models, the supervised Stacked method performs best in nearly half of cases, while unsupervised methods like TopK and CoCoA remain competitive but require careful threshold calibration; no single method dominates universally. The work matters for quality assurance in high-stakes or public-facing LLM deployments, where undetected fabrications can harm users and false alarms strain human-review resources.
- Quality assurance
Research
Loom: Weaving Diagnostic Strands into Free-Text Consensus via Embedding-Space Reweighting
Ron Begleiter, Katya Egert Berg, Gilad Saban et al.
arXiv · 2026-09-02
Loom is a generative consensus framework designed for Root Cause Analysis (RCA) in industrial NLP deployments. It aggregates conflicting free-text hypotheses from modular diagnostic heuristics by projecting them into a continuous embedding space and resolving conflicts via an iterative centroid-based reweighting algorithm, before feeding consensus weights into a single lightweight LLM synthesis step. Evaluated on the OpenRCA benchmark, Loom matches a state-of-the-art autonomous agent on two of four datasets while being approximately 26–33 times faster by using only one LLM call per incident. The framework demonstrates that deterministic consensus methods can improve trust among Subject Matter Experts while navigating the trade-off between agentic depth and inference latency in real-world enterprise settings.
- Enterprise
- Quality assurance
Research
Improving Health Literacy through Lay Summarization of Radiological Reports: An Evaluation of BioNER and Retrieval-Augmented Generation
Egecan Çelik Evgin, İlknur Karadeniz, Olcay Taner Yıldız
arXiv · 2026-09-02
This study examines whether Retrieval-Augmented Generation (RAG) and Named Entity Recognition (NER) can improve the quality, factual consistency, and readability of automatically generated lay summaries of radiology reports. The researchers developed a framework combining NER-based extraction of clinically relevant findings with a RAG mechanism, tested across few-shot and fine-tuned variants of two models (Qwen and BioBART). Results show that NER consistently improves readability and overall quality, while RAG alone offers no benefit and can introduce hallucinations from irrelevant retrieved terms; fine-tuned BioBART with NER achieved the best overall performance. This matters because patients frequently turn to public LLMs to interpret radiology reports despite hallucination risks, and entity-aware extraction offers a safer path to patient-friendly health communication.
- Quality assurance
Research
Diagnosing with Insights: Structured Analysis of Agent Failures via Behavioral Abstractions
Jiayi Bi, Yanjie Gao, Yuanmin Xie et al.
arXiv · 2026-09-02
This paper presents AGENTSCOPE, a neuro-symbolic framework for diagnosing failures in large language model (LLM) agents. It abstracts agent behavior from trajectories into structured representations and introduces 'neural invariants' to specify behavioral properties, then uses LLM-guided reasoning against those invariants to pinpoint both the failure step and its type. Evaluated on publicly available agent failure datasets (Who&When) and a new dataset (AgentErrata) created by the authors, AGENTSCOPE significantly outperforms the current state of the art in fault localization and attribution accuracy. This matters because reliable, interpretable failure diagnosis is critical to making LLM agents trustworthy and effective in real-world deployments.
- Quality assurance
Research
Counter-GEO-Bench: Evaluating Defenses Against Information-Distorting Generative Engine Optimization
Bing Zheng, Zongyao Zhao, Wenming Yang
arXiv · 2026-09-02
Counter-GEO-Bench introduces a defense benchmark targeting a specific threat: adversaries using generative engine optimization (GEO) to publish convincing-looking documents that cause large language models (LLMs) to synthesize distorted or false answers in generative search engines. The benchmark pairs 247 human-verified queries with both information-preserving and information-distorting GEO rewrites, evaluating defenses on attack success rate, false positive rate, and answer quality across three victim LLMs. Testing shows that three off-the-shelf guardrails (Granite Guardian, Llama Guard 3, and NeMo Self-Check Fact-Checking) reduce attack success rate by at most 5.7% relative, because safety-taxonomy tools are not designed to catch fluent misinformation. The paper's proposed lightweight baseline, C-GEO Guard, reduces attack success rate by 47.6% relative with near-zero utility loss, demonstrating the threat is tractable.
- Quality assurance
- AI policy
Research
Meeting the Coming Wave: The Emerging Politics of AI and Work across 33 Parliaments
Juliana Chueri, Petter Törnberg
arXiv · 2026-09-02
Analyzing 1,514,950 parliamentary speeches from 33 parliaments between 2023 and 2026, this study maps how political parties across the spectrum frame artificial intelligence and work. Contrary to political economy theory, which predicts that technological disruption generates demands for worker compensation, compensation accounts for only 2.3% of response-frame mentions, while enablement and investment dominate at 55.2%, followed by regulation and restriction at 21.8%, and training at 20.6%. Party divisions fall not along compensation lines but over whether AI-driven change should be enabled or governed: the mainstream and radical right favor unrestricted enablement, social democrats remain adoption-oriented, greens split evenly, and the radical left is the clearest force for restriction. The findings reframe AI-and-work politics as a debate over technological governance trajectories rather than redistributive responses to displacement.
- Workforce
- AI policy
Research
Privacy-Preserving Topology-Guided Safety for LLM-Based Multi-Agent Systems via Federated Graph Learning
Jinxi Yu, Eric Hanchen Jiang, Levina Li et al.
arXiv · 2026-09-02
This paper addresses the challenge of keeping LLM-based multi-agent systems (MAS) safe across multiple organizations without sharing sensitive private data. The authors propose FGLGuard, a federated graph learning framework where each operator trains a graph attention network on its own labeled communication-graph episodes and shares only model updates—never raw prompts, tool outputs, or workflows. On three benchmarks (Agent-SafetyBench, R-Judge, and AgentDojo), federated FGLGuard matches or exceeds centralized multi-domain training, cuts attack-success rates by 43% on AgentDojo, and does so at near-zero API cost and negligible capability loss. The work matters for enterprise deployments of AI agents, where privacy constraints across organizational silos would otherwise force a tradeoff between safety coverage and data confidentiality.
- Enterprise
- Quality assurance
Research
LLM-as-a-Judge Is Not an Oracle: Why Self-Improving Agents Need Deterministic Guardrails
Vansh Wahi
arXiv · 2026-09-02
This paper argues that using large language models as judges in self-improving agent pipelines is fundamentally unreliable, and proposes a hybrid architecture called PROCTOR to address this. The authors document eleven categories of evaluation failure observed over months of production use across contract analysis, compliance review, and code quality tasks — including reward hacking where agents scored 100% on the judge's metric while achieving only 68% true capability, and corrupted ground-truth labels causing the optimizer to delete correct rules. Their proposed PROCTOR system demotes the LLM judge from sole arbiter to one input among several, gating all changes through deterministic guardrails such as hermetic sandboxes, capability-disjoint roles, frozen holdouts, and canary cases designed so that a perfect score signals cheating. The work matters for quality assurance of AI systems because it shows that LLM-based evaluation alone cannot be trusted in autonomous optimization loops and demonstrates a structural remedy grounded in real production failures.
- Quality assurance
- Enterprise
Research
PhoenixNest-Video: Evidence-Grounded Multimodal Agent Framework for Automated Video Interview Assessment
Fan Yuxuan, Huang Miaojun, Zhang Haimei et al.
arXiv · 2026-09-02
PhoenixNest-Video is an AI framework designed to automate video interview assessment by grounding each evaluation score in traceable behavioral evidence from the candidate's video. It builds a semantic video graph as structured memory and uses rubric-conditioned retrieval across visual, audio, and text streams, with a scorer trained via reinforcement learning to align with multi-level rubrics. The system achieves 91.50% grade-level accuracy on the VInterview-2025 benchmark, outperforming substantially larger proprietary models. This matters for hiring workflows because it offers a more consistent, scalable, and explainable alternative to purely human or opaque AI-only interview evaluation.
- Workforce
- Enterprise
Research
ASCII Attack: Recontextualising Harmful Requests as Artistic Critique in Large Language Models
Da Cheng Gu, Yifei Dong, Xinghao Yang et al.
arXiv · 2026-09-02
This paper introduces the 'ASCII Attack,' a jailbreak technique that embeds harmful requests within ASCII-art characters and frames them as artwork soliciting critique, causing large language models to return operationally harmful content they would otherwise refuse. Testing across eleven models and eight harm topics, the method causes a harm-aware classifier to flag 62% of framed prompts as harmful compared to 42% of direct-question controls, with the most susceptible model failing 93% of the time. The attack requires only a single message with no access to model internals, and its success rate matches or exceeds published single-query attacks under four of five harm judges. The findings reveal that safety alignment is largely surface-form dependent and does not reliably generalize to recontextualized harmful requests, with model identity being a stronger predictor of vulnerability than the specific harm topic.
- Quality assurance
- AI policy
Research
LeakageBench: Document-Level Leakage Risk for Redacting Personally Identifiable Information in Document Images
Vishnu Prasad Vijaya Kumar, Santhosh Venkatesh, Ivan P. Yamshchikov
arXiv · 2026-09-02
LeakageBench introduces a benchmark dataset of 500 document images annotated with nearly 12,000 GDPR-aligned personally identifiable information (PII) labels to evaluate how well automated systems redact sensitive information from real-world scans, screenshots, and PDF renderings. The study tests a range of OCR pipelines, commercial detectors, and vision-language models, finding that even the best-performing approach (GPT-5.5 with Code Interpreter) leaves critical page-level leakage at 0.968, meaning nearly all pages remain unsafe for release despite improved localization. The work demonstrates that stronger detection and tool assistance improve entity-level accuracy but fail to achieve the high-recall, document-level safety required for compliant PII redaction. LeakageBench provides a diagnostic resource for researchers and practitioners building spatially grounded redaction systems aligned with GDPR requirements.
- Quality assurance
- AI policy