News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
Whose readiness counts? Disagreement within and between sectors in perceived AI and robotics preparedness
Peng Wang
arXiv · 2026-08-24
This paper investigates how much information is lost when AI and robotics readiness is summarized as a single score for a sector or organization. Using a card-based survey of 982 respondents producing 15,200 readiness evaluations across 17 AI and robotics challenges, the study finds that 60% of variation occurs at the response level, and sector-level differences account for only about 2% of total variation—meaning aggregate scores conceal substantial disagreement both within and between sectors. Manufacturing shows the highest mean readiness overall, yet sub-applications within it are judged very differently, and technical respondents (CS, AI/ML) consistently rate readiness higher than non-technical respondents. The authors recommend that readiness reporting retain application-level disagreement, disclose whose judgements form the average, and address ethics, cybersecurity, and capability needs rather than collapsing them into a single score.
- AI policy
- Enterprise
Research
EviSafe: Evidence-Grounded Safety Evaluation for Vision-Language Models
Xuetong Li, Gaofeng Liu
arXiv · 2026-08-24
EviSafe introduces a new benchmark framework for evaluating the safety of vision-language models (VLMs) that goes beyond simply checking whether a model refuses or complies with a request. The framework assesses whether models are safe for the right reasons by jointly examining natural responses, explicit grounding in textual and visual evidence, and behavioral sensitivity to counterfactual changes in safety-critical inputs. Testing eleven VLMs on EviSafeBench—a controlled benchmark with 1,181 image-text scenarios and 2,452 counterfactual variants across eight safety domains—reveals large performance gaps: natural severity accuracy ranges from only 27.6% to 52.8% and diagnostic consistency from 6.1% to 29.3%, indicating that current VLMs are not reliably safe for the right multimodal reasons. These findings motivate moving evaluation of AI safety beyond simple refusal counts toward evidence-grounded assessment.
- Quality assurance
- AI policy
Research
FIDES: A Concordance Protocol for LLM-Generated Trading Strategies
Arther Tian, Alex Ding, Simon Wu et al.
arXiv · 2026-08-24
FIDES is a measurement protocol designed to check whether the three outputs of an LLM-generated trading strategy—its natural-language rationale, its executable code, and its actual backtest track record—are internally consistent. The protocol scores three 'concordance gaps' (say-to-do, do-to-real, and say-to-result) by executing model-generated code in a sandboxed, out-of-sample backtest across 40 strategies on 8 US ETFs from 2023–2024. Key findings show that concordance does not predict profitability (only 2 of 40 strategies beat buy-and-hold), that LLM self-assessment is poorly calibrated (32 of 40 strategies claimed to beat buy-and-hold but only one did), and that changing the judge model flips concordance scores on more than half of items. The work matters for quality assurance and enterprise adoption of LLMs in finance, as it exposes systematic gaps between what models claim, what their code implements, and what actually occurs in testing.
- Quality assurance
- Enterprise
Research
Expectations and Practices around AI Disclosure in CS Research
Arati Mohapatra, Danish Pruthi
arXiv · 2026-08-24
This paper examines the state of AI disclosure policies and practices in computer science research. The authors find that disclosure policies at top CS venues are highly under-specified, and through a survey of 109 CS researchers, they show that disclosures are deemed most necessary for tasks involving research design and low human involvement. An analysis of 13,867 disclosure statements from EMNLP 2025 and ICLR 2026 reveals a major gap between researcher expectations and actual disclosure practices—for example, writing assistance is frequently disclosed despite being viewed as less necessary. The authors offer recommendations for authors and policymakers to better align disclosure policies with the research community's expectations.
- AI policy
- Quality assurance
Research
AI emotional support is better only when chosen, but shifts preferences even when it is not
Yaoxi Shi, Cathy Mengying Fang, Guy LabanPattie Maes et al.
arXiv · 2026-08-24
This paper investigates what happens when people seeking emotional support receive a source (human or AI) that matches or contradicts their preference. Across three experiments with nearly 2,000 participants, AI emotional support was rated as superior only when participants had actively chosen it; however, exposure to AI support increased willingness to choose AI again regardless of whether the pairing was congruent with their original preference. A 28-day longitudinal study conducted with OpenAI (N=981) further found that daily AI conversations progressively shifted preferences toward AI and away from humans, but only when conversations became personal. The findings suggest that emotional support choices are path-dependent, meaning AI interactions can cumulatively redirect people away from human connection even without initial preference for AI.
- Workforce
- AI policy
Research
Beyond Surface Cues: Disentangling Sociocultural Signals in Multilingual LLMs
Yuanjun Feng, Tanzhou Liu, Stefan Feuerriegel et al.
arXiv · 2026-08-24
This paper investigates whether multilingual large language models (LLMs) genuinely reflect sociocultural understanding or merely rely on surface-level cues such as identity labels, names, and source-language wording. Using a human-validated, multi-agent audit of 89,253 outputs from 12 LLMs across English, French, and Chinese—spanning 18 occupations and three task conditions—the authors disentangle three distinct signals: social bias reproduction, differential group representation, and cross-cultural patterning. They find that removing direct identity cues sharply reduces identity-label prediction in English and Chinese but has a much smaller effect in French, and that the ability to identify source language drops substantially after translation and name masking. The work matters for quality assurance and policy because it shows that standard multilingual audits may mistake surface shortcuts for genuine cultural grounding, potentially producing misleading conclusions about bias and cross-cultural variation in deployed AI systems.
- Quality assurance
- AI policy
Research
Large language models simulate intersectional synthetic identities with a budget of one to two dimensions
Virgile Rennard, Christos Xypolopoulos
arXiv · 2026-08-24
This paper stress-tests the use of large language models (LLMs) as synthetic survey respondents by comparing simulated opinions to 21 million real response distributions drawn from 15 waves of Pew Research's American Trends Panel. The authors find that while real respondents' opinions grow more distinctive as multiple identities intersect, LLMs effectively collapse multi-dimensional personas to a single dimension—one feature explains a two-feature persona's responses better than the additive combination in 75–82% of subgroups, and adding a third identity dimension contributes almost nothing. Critically, the models tend to systematically discard race and religion, which are the strongest actual drivers of opinion. This means LLMs cannot reliably represent rare or complex intersectional populations, undermining their use for survey research and policy analysis targeting diverse subgroups.
- AI policy
Research
Proxy reliance in large language model decisions is uncalibrated to predictive evidence
Zengqing Wu, Chuan Xiao
arXiv · 2026-08-24
This paper investigates how large language models (LLMs) use demographic proxy attributes in decisions—such as clinical triage and lending—by comparing actual model reliance on proxies against a computable reference for how much reliance the evidence actually warrants. Testing four LLMs on a clinical-ranking task with known ground truth, the researchers find that proxy reliance is poorly calibrated: models rely on uninformative proxies when they shouldn't, under-rely on informative proxies, and that social-field labeling suppresses reliance in ways that are fragile and easily reversed by in-context examples. Critically, standard accuracy-based audits fail to detect any of these miscalibrations, suggesting current evaluation methods are insufficient for distinguishing discrimination from sound inference in high-stakes AI deployments.
- AI policy
- Quality assurance
Research
FLARE: A Systematic, Uncertainty-Aware Framework for Evidence-Based Adoption of Artificial Intelligence in Healthcare
Jacob Idoko, Siddhartha Paudel, Mariana Bento et al.
arXiv (Cornell University) · 2026-08-24
FLARE is a proposed framework that combines fuzzy logic, time-driven activity-based costing, and return-on-investment analysis to evaluate whether adopting AI in healthcare is economically worthwhile, not just technically accurate. It was demonstrated through a case study of AI-assisted large vessel occlusion detection in a CT stroke pathway, identifying a break-even threshold of approximately 3,992 patients per year and positive first-year ROI at typical annual volumes of about 5,000 patients. The framework makes explicit the roles of patient volume, verification time, infrastructure choices, and workflow design in determining economic benefit, providing decision support for clinicians, administrators, and policymakers assessing AI adoption viability.
- AI policy
- Enterprise
Research
Same Agent, Different Answers: A Repeat-Aware Audit of Corpus-Induced Answer Churn in Retrieval-Augmented QA
Jingjie Ning, Xueqi Li
arXiv · 2026-08-24
This paper identifies and measures 'accuracy-blind answer churn' in retrieval-augmented question-answering (RAG) systems — the phenomenon where expanding a document index causes a system to return different answers to the same questions, even when all other settings remain fixed. The authors introduce a 'Snapshot Compatibility Audit' methodology that isolates true index-induced answer changes from ordinary generation variability by comparing cross-snapshot disagreement against same-snapshot repeat disagreement. Across preregistered studies on Natural Questions and TriviaQA benchmarks, they find meaningful excess churn (up to 10.25 percentage points semantically) even when aggregate accuracy metrics change by only small amounts, meaning gains and losses can mask substantial answer instability. The findings argue that RAG system releases should audit answer-level compatibility — not just aggregate utility — when corpus updates are made.
- Quality assurance
- Enterprise
Research
Beyond the Harness: End-to-End Optimization of Context Artifacts for Enterprise Text-to-SQL
Kate Gwimm, Carson Eisenach
arXiv · 2026-08-24
This paper examines how to improve large language model performance on enterprise Text-to-SQL tasks by optimizing the knowledge-base context fed to the model, rather than the model itself. Using a query-DAG decomposition approach applied to production SQL logs, the authors find that retrieved knowledge-base context—specifically distilled 'SQL reference cards' built from historical query profiles—yields larger accuracy gains (roughly 12–25% AST similarity improvement) than tuning the retrieval harness (roughly 3–12%) on a benchmark of 5,176 production queries from a major online retailer. On the public BEAVER benchmark, the best optimized variant combining reference cards and raw historical SQL scores 9.00% versus 6.33% for the comparable baseline on a held-out subset of 300 queries. The findings suggest that how enterprise context artifacts are constructed and retrieved is a critical bottleneck for deploying LLMs in real-world database environments.
- Enterprise
Research
SDoH-Aware Narrative Anchoring Bias in Medical LLMs for Trustworthy Clinical Decision Support
Ahnaf Atef Choudhury, Ramkrishna Saha
arXiv · 2026-08-24
This paper investigates whether medical large language models change their answers to clinical questions based on how a patient's case is narrated, a phenomenon the authors call 'SDoH-aware narrative anchoring bias.' Using a counterfactual dataset (NarrativeShield SDoH MedQA) where the same clinical case is presented in different persona-based narratives but with a fixed correct answer, the study evaluates three Qwen2.5 models (1.5B, 3B, and 7B parameters) across 300 cases and 8,100 total responses. The best-performing model (Qwen2.5 7B) achieves 56.33% accuracy and 40.33% correct consistency, yet still exhibits a narrative sensitivity error of at least 31.67%, meaning it gives inconsistent answers to medically equivalent cases depending on patient voice. The findings argue that trustworthy clinical decision support systems must be assessed not only on average correctness but also on response stability across equivalent patient narratives.
- Quality assurance
- AI policy
Research
DelistBench: Evaluating Search-Enabled LLMs for Auditable Corporate-Event Database Completion
Xuan Yao, Li Shuping, Dai Yang et al.
arXiv · 2026-08-24
This paper introduces DelistBench, a 1,200-record benchmark for evaluating search-enabled large language models on the task of reconstructing corporate delisting event records from public sources—a process the authors call Search-to-Record. Across five models tested in closed-book and web-enabled conditions, enabling web access raised announcement-date accuracy within seven days by 34–48 percentage points and event-status accuracy by roughly 2.8–21.7 points, with the best system achieving 81.5% overall joint accuracy. Economy web systems reached 75.9–78.3% joint accuracy at just 4.5–6.6% of the API cost of the most expensive system. The findings offer concrete guidance for financial institutions seeking to audit vendor corporate-event databases, including recommendations on triage calibration, recall preservation, and routing ambiguous cases to human review.
- Enterprise
- Quality assurance
Research
Does Rank Still Matter? Position Bias When AI Agents Shop on Our Behalf
Davood Wadi, Yu Ma
arXiv · 2026-08-24
This paper investigates whether search-result ranking still influences purchasing decisions when AI agents shop on behalf of consumers, rather than humans browsing sequentially. Using a randomized experiment across 5,000 AI agent sessions with 100 hotel listings and four large language models, the researchers find that position still weakly predicts which listings are inspected, but in a non-monotonic pattern where middle-positioned listings are least likely to be examined. Unlike humans, AI agents search more deeply and always complete a purchase, and while position bias affects inspection for some models, all models converge on the same top-performing listing. The key finding is that for agentic search, the content and attributes shown on a results page matter more than where a listing is placed, with significant implications for how businesses optimize their digital presence.
- Enterprise
Research
AI Safety Guard: Design, Prototype Implementation, and Validation Roadmap for a Privacy-Preserving Multi-Sensory Edge-AI Driver Drowsiness System
Kar Ho TSO
Gamification Fatigue and Continuance Intention in Technology‑Enhanced Learning Platforms A Post‑Adoption Extension of TAM · 2026-08-24
This paper presents AI Safety Guard, an edge-AI prototype for detecting driver drowsiness using non-contact facial-landmark analysis and multi-sensory alerts (auditory and optional olfactory), designed to run locally on Raspberry Pi-class hardware without retaining video. The study adopts a design-science and safety-by-design methodology, outlining a five-phase validation programme covering bench metrology, public-dataset evaluation, simulator experiments, closed-track trials, and regulatory readiness. The paper explicitly distinguishes artefact feasibility from safety efficacy and frames the prototype as a supplementary warning device requiring independent validation before road deployment. This matters for quality assurance and certification because it provides an evidence-bounded blueprint for translating a student-developed prototype into a testable, privacy-preserving driver-monitoring system aligned with functional-safety and human-machine-interface requirements.
- Quality assurance
- Certifications
Research
AI ‐Assisted Recognition of Actual Ureteral and Pancreatic Injuries in Laparoscopic Colorectal Surgery: A Randomized Video Review Study
Shunjin Ryu, Teppei Kamada, Kai Neki et al.
Annals of Gastroenterological Surgery · 2026-08-24
This randomized video-review study involving 150 surgeons across two institutions found that AI-enhanced surgical videos (using the EUREKA system) significantly improved surgeons' recognition of actual ureteral injuries during laparoscopic colorectal surgery compared to unenhanced videos, while not increasing false-positive adverse-event judgments. Pancreatic injury recognition also improved among surgeons without advanced laparoscopic certification. The authors conclude this provides proof-of-concept evidence that AI-generated anatomical overlays could support timely identification of intraoperative adverse events, though prospective studies are needed to confirm whether this translates to better patient outcomes.
- Quality assurance
- Certifications
Research
University Job Stress Management System: An Explainable Artificial Intelligence Approach for Nigerian Universities
Olutomisin M. Orogbemi, Temi E. Ologunorisa
International Journal of Scientific Research and Modern Technology. · 2026-08-24
This study developed and validated a Job Stress Management System (JSMS) for Nigerian university staff using Explainable AI techniques (SHAP, LIME) on ensemble machine learning models (XGBoost, Random Forest). A mixed-methods study across 427 staff at 15 universities identified major stressors including excessive workload (68.2%) and poor remuneration (62.7%), and a 12-week pilot with 120 participants showed a mean 32.7% reduction in self-reported stress, with 87.3% of users attributing their trust in the system to its explainability. The findings demonstrate that XAI-powered, culturally tailored tools can deliver scalable, personalized stress management in resource-constrained institutional settings, with direct implications for university HR administrators and policymakers.
- Workforce
- AI policy
Research
DEVELOPMENT OF AN AI-ENABLED PERFORMANCE MONITORING AND EVALUATION FRAMEWORK FOR ACADEMIC QUALITY ASSURANCE IN NIGERIAN HIGHER EDUCATION INSTITUTIONS
Luqman Muhammed Audu, SULEMAN MARIAM ITSEMEH
International Journal of Engineering Innovation and Technology Research · 2026-08-24
This study developed and validated an AI-enabled framework for academic quality assurance in Nigerian higher education institutions, addressing the shortcomings of manual, periodic, compliance-driven monitoring systems. Using Design Science Research methodology, the framework integrates seven components—including AI analytics, predictive analytics, natural language processing, and decision support dashboards—grounded in international standards and Nigerian regulatory requirements. Expert evaluation yielded a grand mean acceptance score of 4.73 out of 5.00 and a content validity index of 0.94, indicating strong expert endorsement. The authors conclude the framework can shift quality assurance from reactive compliance to proactive, data-driven institutional management.
- Quality assurance
Research
EGAMA-RC: Risk-Calibrated Evidence-Gated Adaptive Malware Analysis for Robust and Interpretable Memory-Forensic Triage
Isaac Kofi Nti
arXiv (Cornell University) · 2026-08-24
EGAMA-RC is a risk-calibrated, evidence-gated framework for memory-forensic malware triage that goes beyond classification accuracy to address uncertainty, novelty, robustness, interpretability, and review cost. The system uses SHAP-guided feature refinement, a pool of models, and adversarial/open-family testing to route low-risk samples automatically while escalating uncertain or novel cases for analyst review. Across three malware datasets, the hybrid gate accepts 93.12% of samples with 99.86% accuracy and a 0.136% false-accept rate, while XGBoost provides fast-path inference at sub-millisecond latency. The work demonstrates that dependable malware analysis requires risk-calibrated routing and novelty awareness, not accuracy metrics alone.
- Quality assurance
Research
Frontiers in FinTech: Multimodal Foundation Models for Financial Reporting and Decision Science
Yulu Huang, Niannian Yu, Yaxin Yang et al.
arXiv (Cornell University) · 2026-08-24
FinVision is a multimodal large language model system designed to process heterogeneous financial data—including PDFs, Excel files, chart images, and scanned documents—for financial reporting and decision support. The system combines vision-language models with domain-specific financial reasoning, featuring a cross-modal consistency validator, two-stage training on valuation methodologies (DCF, P/E, P/B, P/S), and a natural-language decision pipeline incorporating portfolio theory and real-time risk monitoring. Validated on 200 listed companies, FinVision achieved a 19% reduction in valuation error, and a user study with 48 professionals showed a 51% reduction in task completion time. The paper discusses implications for audit automation, financial reporting quality, and broader access to expert-level financial analysis.
- Enterprise
- Quality assurance
Research
THE RELATIONSHIP BETWEEN ARTIFICIAL INTELLIGENCE AND PUBLIC SERVICE QUALITY: AN ANALYTICAL STUDY OF THE PUBLIC SECTOR IN THE KINGDOM OF SAUDI ARABIA
Amani Bani Alkahtani
Advances and Applications in Statistics · 2026-08-24
This study surveyed 327 employees across five top-ranked Saudi Arabian government agencies to examine how AI use relates to public service quality. Findings show that most employees have basic-to-intermediate AI knowledge, with over half using AI tools daily or frequently, and a positive correlation between AI proficiency and usage frequency. Employees broadly perceived AI as improving service quality by speeding up tasks, reducing errors, and enhancing decision-making. The study recommends expanded AI training, digital infrastructure investment, and institutional innovation culture to maximize public sector AI benefits.
- Quality assurance
- Workforce
Research
An Empirical Analysis of Motivation and Challenges of Industry 4.0 Adoption in Malaysian ICT SMEs
Nurulizwa Rashid, Yasmin Dania Khairul Hisham, Samer Ali Al-Shami et al.
International Journal of Computer Information Systems and Industrial Management Applications · 2026-08-24
This study surveys 50 managers from Malaysian ICT SMEs that have adopted at least one Industry 4.0 technology, finding that the primary motivations for adoption are faster time-to-market, cost savings, customer requirements, and competitive pressure. Medium-sized firms show more advanced adoption (e.g., big data analytics and AI), while micro enterprises focus on foundational technologies like cloud computing and cybersecurity. Key barriers include low strategic awareness, insufficient staff knowledge, and gaps in ongoing employee training. The authors conclude that strategic alignment and human capital development are critical to enabling successful digital transformation in this sector.
- Enterprise
- Workforce
Research
Driver State Monitoring Systems: A Comprehensive Review Bridging Academic Research and Industrial Deployment
Xiuwei Zhang, Zhongxiang Feng, Zeyang Cheng et al.
Human-Centric Intelligent Systems · 2026-08-24
This comprehensive review of 171 publications on driver state monitoring (DSM) systems identifies a persistent gap between academic research and real-world industrial deployment. While laboratory-based approaches achieve high accuracy (over 96% for physiological signals, 95% for behavioral recognition), four fundamental constraints limit practical implementation: reactive rather than predictive paradigms, modular architectures that limit human-machine synergy, universal models that ignore individual variability, and the mismatch between controlled research conditions and real-world constraints such as computational limits and privacy regulations. Commercial systems favor robustness and regulatory compliance—driven by standards like EU GSR 2024/2026—over algorithmic sophistication. The authors propose four strategic directions including proactive probabilistic risk prediction, human-in-the-loop adaptive architectures, personalized digital twin monitoring, and industry-aligned methodologies to close this gap.
- Quality assurance
- Certifications
- AI policy
Research
Aviation 5.0: catalyst for workforce sustainability
Gülnaz Bülbül, Ewa Niechwiej-Szwedo, Suzanne Kearns et al.
Aviation · 2026-08-24
This paper introduces Aviation 5.0, an aviation-specific adaptation of the Industry 5.0 paradigm built on human-centricity, sustainability, and resilience, and presents a five-layer conceptual framework linking these principles to enabling technologies and workforce outcomes. It distinguishes Aviation 5.0 from its predecessor by a value-driven orientation—asking not what technology can do but what it should do and for whom—within the boundary conditions of safety primacy, international regulation, and certified human capital. The framework is applied to workforce sustainability, addressing threats such as demographic shifts, technological disruption, and declining career attractiveness, and concludes with policy, practice, and research recommendations to operationalize the framework globally.
- Workforce
- AI policy
Research
From Language Models to Agentic AI: A Survey of Autonomous, Action-Enabled, and Collaborative LLM Agents
Sparsh Bajoria -, Sumit Ranjan, Adhitya M et al.
Cognitive Computation · 2026-08-24
This survey paper maps the rapidly evolving landscape of LLM-based agentic AI systems, presenting a unified taxonomy that characterizes agents along four dimensions: autonomy, tool use, collaboration, and safety-governance. The authors identify key problems in the current literature—fragmented terminology, inconsistent architectures, and weak evaluation standards—and propose a modular reference architecture to guide system design and comparison. The paper places particular emphasis on enterprise and safety-critical deployment considerations, including access control, human-in-the-loop oversight, and policy enforcement, making it directly relevant to organizations seeking to deploy these systems reliably.
- Enterprise
- AI policy
- Quality assurance