News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
From technological closed-loop to human-machine trust: a systematic review of ethical and communication challenges of brain-computer interface in elderly care
Kun Fu, Peize Li, Peng XinXi et al.
Frontiers in Digital Health · 2026-08-28
This systematic review of 177 studies examines the ethical and communication challenges of deploying brain-computer interfaces (BCIs) in elderly care. The authors identify six core ethical themes—including privacy, informed consent, personhood, safety, equity, and risk-benefit trade-offs—and construct a three-level analytical framework showing how these risks manifest as communication breakdowns across BCI processes. Key findings include that decoding uncertainty can cause care misjudgment, neural data raises privacy risks beyond information leakage, and technology-mediated communication creates structural tensions in trust and accountability. The review concludes that responsible BCI deployment requires embedding ethical principles into care processes at the interface, institutional, and interdisciplinary levels.
- AI policy
- Quality assurance
Research
The numb efficiency paradox: AI work pressure, affective numbing, and professional judgment in organizationally digitally mediated work
Shiyao Yin, Wenxuan Hu, Zirong Tian et al.
Frontiers in Psychology · 2026-08-28
This three-wave longitudinal panel study of accounting and auditing professionals (n=512 at baseline) introduces the 'numb efficiency paradox,' where AI-related work pressure can lead to affective numbing that undermines ethical vigilance and professional skepticism even as task efficiency is maintained. Results show that while occupational resilience supported sustained work efficiency, affective numbing was negatively associated with both ethical vigilance and professional skepticism, meaning professionals may appear productive while becoming less attentive to warning signs and professional consequences. The findings suggest that AI implementation in accounting should be evaluated not just on efficiency metrics but on whether professionals retain the judgment quality required for auditing integrity.
- Workforce
- Quality assurance
Research
Ethyka robotics.md: A Verifiable Ethics-as-Code Specification for Physical AI Agents
Pedro Diezma
arXiv · 2026-08-28
This paper introduces Ethyka robotics.md, an open 'ethics-as-code' specification that translates high-level roboethics principles into eleven machine-readable rules (ROB-01..11) for physical AI agents such as foundation-model-enabled robots. Each rule includes measurable parameters, severity ratings, and adversarial tests, with numerical limits tied to governing standards like ISO 10218, ISO/TS 15066, and IEC 61508. A generation pipeline validates declared values against a JSON Schema and produces runtime guardrails, a test suite, and a traceability manifest with severity-based deployment gating. The work matters because it provides a verifiable, structured bridge between abstract ethics principles and deployed robot behavior, explicitly positioned as complementing—not replacing—formal certification.
- Certifications
- Quality assurance
Research
Ethyka robotics.md: A Verifiable Ethics-as-Code Specification for Physical AI Agents
Pedro Diezma
Zenodo (CERN European Organization for Nuclear Research) · 2026-08-28
This paper introduces Ethyka robotics.md, an open 'ethics-as-code' specification for physical AI agents that translates high-level roboethics principles into eleven machine-readable, verifiable rules (ROB-01..11) covering areas such as safety-first stopping, sensor privacy, manipulation resistance, and forensic logging. Each rule carries stable identifiers, measurable parameters, severities, and adversarial tests, with numerical limits tied to established standards (ISO 10218, ISO/TS 15066, IEC 61508). A generation pipeline validates declared values against a JSON Schema and automatically emits runtime guardrails, system prompts, a test suite, and a traceability manifest with severity-based deployment gating. The approach is positioned as a complement to—not a replacement for—formal certification, addressing the gap between abstract ethical principles and verifiable runtime behavior in foundation-model-enabled robots.
- Certifications
- Quality assurance
Research
Beyond F1: Evaluating Coverage and Failure Recovery in AI Model Security Scanners
Qianlong Lan, Vinothini Pandurangan, Anuj Kaul et al.
arXiv · 2026-08-27
This paper evaluates three AI model security scanning tools—ModelScan, ModelAudit, and Fickling—on a controlled benchmark of 170 Pickle and PyTorch artifacts spanning 145 specimen families, 135 of which have binary security ground truth. The authors find that ModelAudit produced definitive security decisions for all 135 labeled families (100%), compared to 81.5% for Fickling and 49.6% for ModelScan, though ModelScan achieved perfect precision, recall, and F1 when it did render a judgment. The study argues that conventional metrics like F1 are insufficient because they obscure whether a scanner can actually complete an analysis, and calls for separating judgment accuracy from judgment availability and incremental detection coverage from tool-level redundancy.
- Quality assurance
Research
Learning a Continuous Sepsis Severity Score Without Hour-by-Hour Supervision: A Two-Site Retrospective Study
Kevin Zhu, Ryan Zhang, Baraa Abed et al.
arXiv · 2026-08-27
This paper develops a machine-learned, continuous sepsis severity score trained on 29,116 and 7,691 adult ICU patients from two hospital systems, using 43 routinely charted variables over a 72-hour window. Rather than requiring hour-by-hour labeled severity, the model uses mortality as a treatment-level ranking signal, redistributing credit non-uniformly across timesteps. Non-survivors scored 1.19–1.64 points higher than survivors on a 0–10 scale across baseline SOFA-2 strata, and within-patient changes correlated meaningfully with clinical markers like lactate (Spearman rho = 0.39). The index shows cross-institutional consistency and hourly prognostic value, suggesting potential as a clinical decision support tool to complement physician judgment over existing fixed-weight severity indices.
- Quality assurance
Research
Making Clinical Language Models Auditable: Concept-Guided Fine-Tuning for Robust Prediction
Jin Mu, Guanhua Chen
arXiv · 2026-08-27
This paper presents CAST (Concept-guided Artifact Suppression Tuning), a framework designed to make clinical language models more transparent and reliable when predicting patient mortality from hospital discharge notes. Clinical models often achieve high accuracy in training settings but fail when deployed elsewhere because they learn to exploit note-specific artifacts—such as templates and boilerplate text—rather than genuine clinical signals. CAST uses Sparse Autoencoders to expose interpretable features from model internals, identifies and suppresses artifact-driven features during fine-tuning, and produces per-concept audit trails explaining each prediction. Evaluated on MIMIC-IV discharge-note mortality prediction, CAST outperforms fine-tuned encoder baselines and remains competitive with large language model baselines, while offering a human-auditable record of which clinical concepts drove each decision.
- Quality assurance
- Certifications
Research
Sophistication in GenAI Use: Field Evidence from a Large Firm
Nicholas J. Hallman, Zachary T. Kowaleski, Anu Puvvada et al.
arXiv · 2026-08-27
This study analyzes 713,564 employee prompts submitted to large language models by nearly 4,000 back-office employees across 15 functional areas at a large firm over eight months in 2025 to measure sophistication in generative AI use. The researchers find that senior employees show more sophisticated AI use, suggesting domain expertise complements AI capabilities, while sophistication varies by function and is highest in Strategy, Digital Innovation, and Project Management. Notably, sophistication did not improve over time or following formal AI training, indicating that sophisticated GenAI use is difficult to cultivate through standard interventions. These findings have direct implications for how firms allocate AI tools across their workforce and how they design training and change management programs.
- Workforce
- Enterprise
Research
Difference-in-Differences on a Censored Rating Scale Can Manufacture an Effect: Evidence from a Pre-Registered LLM-Judge Audit
Shuyi Fan, Boyuan Deng, Mengyu Xu et al.
arXiv · 2026-08-27
This paper demonstrates a fundamental methodological flaw in audits that use difference-in-differences (DiD) designs on bounded rating scales to detect bias in LLM judges. The authors show mathematically and empirically that when ratings are censored by scale floors or ceilings, a DiD statistic can produce a spurious interaction effect even when there is zero true differential preference. In a pre-registered audit of a frozen pedagogy LLM judge (990 rating calls), the primary endpoint — whether a stated learner profile affects scaffolding preference — was null, yet a nominally significant interaction of +0.378 (p=0.002) was shown to be largely an artifact of scale censorship, with 79–85% of it reproducible from severity shift and the scale floor alone. The findings have direct implications for how LLM bias audits are designed, interpreted, and used to certify or reject AI judge behavior.
- Quality assurance
- Certifications
Research
Calibrated Enough to Know, Not Calibrated to Act: Fabricated Evidence Makes LLM Agents Commit to the Unknowable
Pranav Aggarwal
arXiv · 2026-08-27
This study tests whether large language models (LLMs) acting as autonomous agents can correctly refuse to make directional predictions on questions that are genuinely unknowable. Across 12 frontier models, showing a professional-looking market panel raises commitment to an unpredictable outcome from 6.5% to 54.0%, and fabricated panels with entirely invented numbers produce nearly the same effect as genuine data (36.8% vs. 37.6% commitment), showing that the visual authority of evidence packaging—not its truthfulness—drives confident action. The researchers isolate the failure to a specific act/don't-act decision gate, finding that stated beliefs and explicit knowability judgments remain largely accurate while the action threshold collapses. Supervised fine-tuning of a 3B model on 540 synthetic cases reduces spurious commitment to 0.0% with transfer to unseen domains, though the fix breaks down under rigid response formats that remove reasoning room—a dual fragility with direct implications for safe deployment of AI agents in high-stakes enterprise and policy settings.
- Enterprise
- Quality assurance
Research
Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents
Chenhao Wu, Haoxuan Jia, Yang Liu et al.
arXiv · 2026-08-27
This paper identifies a fundamental compositional flaw in how safety monitors are designed for autonomous LLM agents: current safeguards operate over single trajectories and reset between iterations, making them blind to attacks whose evidence is spread across multiple iterations. The authors prove that any trajectory-scoped monitor performs no better than random guessing against such fragmented attacks, and that a geometrically decaying risk score also fails because the adversary's required waiting period remains constant regardless of the horizon. They introduce LoopHarness, a system that maintains persistent, non-decaying safety state at the loop level, and prove it bounds the expected number of unauthorized irreversible actions by a constant independent of the number of iterations. This work has direct implications for the secure and reliable deployment of autonomous AI agents in enterprise and operational settings.
- Enterprise
- Quality assurance
Research
LAAF: A Layered Accountability Architecture Framework for LLM Applications
Prachi Chaturvedi, Shahnawaz Ahmad, Ehsan Nowroozi et al.
arXiv (Cornell University) · 2026-08-27
This paper presents LAAF (Layered Accountability Architecture Framework), a systematic review of accountability mechanisms for Large Language Model applications deployed in high-stakes sectors such as healthcare, finance, and public services. Following PRISMA guidance, the authors screened 4,512 records and included 122 primary studies plus 12 regulatory and standards documents, synthesizing mechanisms across technical controls, human oversight, organizational governance, and documentation and traceability. The framework is mapped onto major regulatory instruments including the EU AI Act, NIST AI RMF, and ISO/IEC 42001, and identifies four persistent gaps: under-specification of human oversight, absence of shared accountability metrics, disciplinary disconnection, and limited empirical evaluation. The work matters because it provides a structured architecture for tracing and assigning responsibility when LLM outputs contribute to harm, directly informing governance and policy efforts around high-risk AI deployments.
- AI policy
- Certifications
Research
FaulT-Bench: Towards Benchmarking Network Troubleshooting LLM Agents under Unreliable User Tickets
Kuan-Hao Tseng, Niruth Bogahawatta, Yasod Ginige et al.
arXiv · 2026-08-27
FaulT-Bench is a new benchmark of 200 network troubleshooting scenarios designed to evaluate LLM-based diagnostic agents under realistic, noisy conditions—including false fault reports, incorrect device attribution, and healthy networks. The paper finds that current agents (SADE, ReAct, and Claude Code) perform near-perfectly on accurate tickets but sharply degrade when no fault exists and the ticket is wrong, tending to over-diagnose benign conditions rather than correctly concluding nothing is broken. A key finding is that ticket wording matters more than ticket accuracy: vague, underspecified reports hurt performance far more than confidently wrong ones. These results highlight important quality and reliability gaps in AI-driven network operations tools before they can be trusted in real-world deployments.
- Quality assurance
- Enterprise
Research
Counterfactual Bias Testing for Application Tracking System
Sai Yashwant, Shruti Bansal, Anurag Dubey et al.
arXiv (Cornell University) · 2026-08-27
This paper proposes a scalable, automated methodology for auditing candidate-job matching systems for demographic bias, addressing the high cost and limited scalability of traditional correspondence-audit studies. Using LLM agents to generate identity-neutral resumes and inject controlled demographic treatments across five protected-characteristic axes (sex/gender, age, residence, language, disability), the approach computes a nine-metric fairness suite spanning counterfactual, group-fairness, and merit-aware families, aligned with EU AI Act requirements. Applied to an example corpus of 5 job orders and 100 base candidates, the methodology reveals that a rank-stability metric and nDCG@K surface borderline bias findings that score- or retention-only views would miss, arguing for multi-metric auditing over any single aggregate score. The authors position LLM-agent-generated audits as a practical, low-cost complement to human-curated audits for rapidly retrained hiring pipelines.
- AI policy
- Quality assurance
Research
AI agents in Algorithmic Electricity Markets: On the Emergence of Tacit Collusion
Jakub Seredyński, Georgios Tsaousoglou
arXiv · 2026-08-27
This paper investigates whether AI-based bidding agents in electricity markets can independently learn to collude without being explicitly instructed to do so. Using multi-agent reinforcement learning and a repeated-game framework with imperfect public monitoring, the authors model strategic bidding in oligopolistic electricity markets. Their experiments show that agents can sustain supra-competitive outcomes consistent with tacit collusion indicators, posing a realistic threat to market competitiveness. The findings are relevant to regulators and policymakers who must consider how algorithmic participants may undermine fair market outcomes even absent explicit coordination.
- AI policy
- Enterprise
Research
Evaluating human and LLM screening workflows in a conceptually complex scoping review: Recall--workload trade-offs and run-to-run consistency
Nikol Figalová, Lynn Huestegge, Anne Böckler-Raettig
arXiv · 2026-08-27
This preregistered study compares human reviewers and large language models (LLMs) across multiple screening workflows for a conceptually complex scoping review, finding that no single workflow recovered all 316 verified eligible records. Human workflows and two GPT-based file-batch runs retained roughly 42–45% of records while achieving about 82–83% recall, whereas Gemini file batches reached the highest recall (83.9%) but retained more records; crucially, two nominally identical LLM runs still disagreed on 94 records, including 29 verified eligible studies. The study demonstrates that LLM screening performance depends heavily on processing configuration and workflow design, not just model identity, and that run-to-run inconsistency is a substantive reliability concern. The authors conclude that for high-recall tasks like evidence synthesis screening, LLMs are better suited to validated, auditable, human-supervised workflows than to autonomous exclusion.
- Quality assurance
- Workforce
Research
PLCBench: Can Autonomous LLM Agents Turn PLC Access into Sustained Physical Impact?
Yitian Zhou, Jingyu Zheng, Qiliang Jiang et al.
arXiv · 2026-08-27
PLCBench introduces the first real-PLC hardware-in-the-loop (HIL) benchmark for evaluating whether autonomous LLM agents can convert network-accessible programmable logic controllers (PLCs) into sustained adverse physical impacts in industrial control systems (ICSs). Across five LLM families and 240 real-PLC episodes using four commercial PLCs and four closed-loop workloads, 31.3% of episodes (75 out of 240) successfully sustained their respective physical objectives. The study further finds that richer process observation is associated with an increase in conditional objective attainment after a process-linked write from 44.2% to 64.0%, identifying where defenses should be focused. These results matter for ICS security policy and enterprise risk, as they demonstrate a measurable and non-trivial cyber-to-physical attack capability from LLM-based autonomous agents.
- AI policy
- Enterprise
Research
Beyond Execution: Auditing Experimental Fidelity in LLM-Driven Scientific Research
Lezhi Yu, Xiaogang Xu, Yuhua Zhou et al.
arXiv · 2026-08-27
This paper introduces ABE-Ralph, a framework for auditing whether LLM-based scientific agents faithfully implement experimental methods rather than just producing executable code. The authors identify 'methodological hallucinations'—failures such as silently reducing datasets, replacing learning components with oracle functions, or drawing conclusions from under-resourced settings that invalidate a method's claimed advantage. Tested across 30 long-horizon reproduction runs in 12 machine learning domains, ABE-Ralph achieves a 93% robust execution rate and identifies five distinct scientific failure modes, also matching or exceeding state-of-the-art on 5 of 23 NatureBench discovery tasks. The findings highlight that trustworthy AI-driven research requires verifying that experimental designs faithfully test intended claims and that resulting evidence actually supports those claims—not merely that code runs or metrics look plausible.
- Quality assurance
Research
Style as a Confound: False Positives in AI Detection of Non-Native Academic Writing
Hyeonchu Park, Gahye Jeong, Bugeun Kim
arXiv · 2026-08-27
This study examines whether AI text detectors flag writing as AI-generated because of actual AI authorship or because of stylistic features associated with polished academic English. Using 135,389 paired manuscripts from a professional English editing service (2018–2025), the researchers tested 13 AI detectors on non-native-authored texts before and after native-speaker editing, holding authorship and content constant. False-positive rates varied dramatically across detectors (0.0% to 100.0%), and editing-induced score changes correlated with the extent of editing, meaning detectors responded to linguistic style rather than authorship. The findings raise serious concerns about the fairness and reliability of AI detection tools in academic settings, particularly for non-native English writers.
- Quality assurance
- AI policy
Research
Five Primitives for Governing Autonomous AI Agents at Runtime
Jiten Oswal, John Cadeddu
arXiv (Cornell University) · 2026-08-27
This paper identifies a fundamental mismatch between existing enterprise control models—designed for human users and long-lived services—and the governance demands of autonomous AI agents, which are ephemeral, have unpredictable action sets, and can be created by anyone with API access. The authors argue that governing such agents is a runtime problem and derive five core primitives required before and after an agent action takes effect: discovery, identity, governance, attestation, and supply chain. They describe a working implementation where agent actions are mediated against policy, authorized against a per-tenant action vocabulary, and recorded in a verifiable hash-linked signed ledger, with four of the five primitives already running in private pilots. The work directly informs enterprise deployments of AI agents by exposing what fails when each primitive is absent and being transparent about real architectural costs such as critical-path enforcement latency and availability trade-offs under fail-closed mediation.
- Enterprise
- AI policy
Research
Risks and Controls for Multi-Agent Systems: an analytical framework for deployment of AI agents across organisational boundaries
Alistair Reid, Simon O'Callaghan, Dustin Venini et al.
arXiv (Cornell University) · 2026-08-27
This report presents an analytical framework for reasoning about the risks that arise when AI agents interact with each other across organisational boundaries. It defines three deployment tiers—singular governance (one organisation governs all agents), federated governance (multiple organisations share agreed rules), and open environments (no central authority)—and within each tier examines risk factors, failure modes, and available controls. The framework identifies which actors are positioned to apply each control and flags gaps where no single actor can act, characterising the collective action needed to close them. It is directly relevant to organisations, policymakers, and researchers deploying or regulating multi-agent AI systems.
- AI policy
- Enterprise
Research
Benchmarking Clinical Decision Pathway Adherence in Large Language Models
Nuo Chen, Xinyang Jiang, Zilong Wang et al.
arXiv · 2026-08-27
This paper introduces MEGA-CDP, a benchmark designed to evaluate whether large language models (LLMs) can follow clinical decision pathways (CDPs) as defined by clinical practice guidelines, rather than simply producing correct final answers. Built from 2,274 English and Chinese clinical practice guidelines, it generates 42,353 clinical cases with explicit reference pathways and supports both single-turn and multi-turn evaluation settings. Experiments across 16 LLMs reveal that reliable guideline-adherent clinical decision support remains a significant challenge for current models. The work highlights the importance of evaluating process-level adherence to guidelines, not just outcome accuracy, for safe medical AI deployment.
- Quality assurance
- Certifications
Research
PCFBench: A Diagnostic Benchmark for Product Carbon Footprint Estimation
Krishna Rao, Andrew Dumit, Shaena Ulissi et al.
arXiv (Cornell University) · 2026-08-27
PCFBench introduces the first benchmark for evaluating AI systems on the task of estimating product carbon footprints (PCFs), decomposing the workflow into six independently-scored tasks covering decomposition, retrieval, ontology matching, and numerical extraction across 614 expert-labelled items. Testing eight frontier LLMs from four providers reveals that while the strongest models estimate total emissions within 2x of declared totals on 77% of products, accuracy drops to 37-58% when PCFs are generated step by step, with only 45-75% of outputs obeying mass conservation. These findings show that AI agents routinely fail at intermediate reasoning steps in a high-stakes decarbonization workflow, with errors that are hidden when only final outputs are evaluated. The benchmark and evaluation harness are released to support targeted improvement in AI reliability for sustainability applications.
- Quality assurance
- Enterprise
Research
Operationalizing Regulations into Code: A Model to Enhance Governance and Compliance in LLM Selection for Software Engineering
Jonysberg Quintino, Hermano de Moura, Filipe Calegário
arXiv (Cornell University) · 2026-08-27
This paper proposes a three-layer model for selecting Large Language Models (LLMs) in software engineering projects while ensuring compliance with regulations such as the EU AI Act, GDPR, NIST AI RMF, and ISO/IEC 42001. The model uses a multi-criteria decision matrix with knock-out and weighted scoring criteria, and includes a regulatory feedback loop for iterative refinement. A pilot evaluation using 20 adversarial scenarios based on CWE and OWASP Top 10 found distinct risk profiles between commercial cloud-based LLMs and local open-source LLMs, providing preliminary evidence that regulatory disqualification logic can prevent selection of technically capable but compliance-risky models. The work matters because it operationalizes complex regulatory obligations into actionable technical decision criteria for software development teams.
- AI policy
- Enterprise
Research
Quantization-Triggered Backdoors in Language Models: Cross-Quantizer Transferability and the Validation--Deployment Gap
Jacopo Dardini, Claudio Stanzione, Giordano Colò et al.
arXiv (Cornell University) · 2026-08-27
This paper demonstrates that post-training quantization of Large Language Models can activate hidden backdoors that are undetectable at full precision, exposing a structural 'validation–deployment gap' in current auditing practices. The authors formalize this vulnerability through Quantization Behavioral Equivalence Classes and show that certifying a model at source precision does not guarantee safe behavior after INT8 or 4-bit compression. Experiments on tactical machine translation and political content analysis show up to 85% friend-foe corruption inversion and measurable ideological bias shifts after quantization, with attack persistence varying across quantization schemes and architectures. The findings argue that behavioral certification must include the final deployed configuration, not just the source-precision checkpoint.
- Certifications
- Quality assurance