News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
5221 items
Research
What Makes a Great Co-Worker in an AI-Native Workplace?
Rudrajit Choudhuri, Max Meijer, Sam Yu-Te Lee et al.
arXiv (Cornell University) · 2026-09-12
This study investigates what knowledge workers value in human and AI co-workers within AI-native workplaces, drawing on 22 interviews and a survey of 1,534 employees at a multinational technology company. The researchers developed the BACI framework—75 co-worker qualities spanning Benevolence, Ability, Cooperativeness, and Integrity—and identified 11 co-worker archetypes, revealing disagreements over whether AI should exhibit warmth, take initiative, or own outcomes. The study also produces a taxonomy of AI work etiquette covering expectations around preparing, sharing, and taking responsibility for AI-supported work. The findings carry direct implications for worker-centric AI design and how organizations structure human-AI collaboration.
- Workforce
- Enterprise
Research
Digitalization pathways for food loss and waste prevention in agri-food supply chains: Current evidence, challenges and future directions
Esteban Pérez-García, Hani A. Alfheeaid, Esther Sanjuán Velázquez et al.
Trends in Food Science & Technology · 2026-09-12
This critical review examines how Industry 4.0 and Agriculture 4.0 digital technologies—including IoT, AI, big data analytics, blockchain, and consumer-facing platforms—contribute to food loss and waste (FLW) prevention across agri-food supply chains. The authors find that IoT-enabled monitoring and smart logistics provide the strongest direct evidence for reducing spoilage, while AI applications in inventory management are moderately supported, but evidence for blockchain and dynamic pricing remains limited or context-dependent. Persistent barriers include fragmented data infrastructure, interoperability gaps, unequal digital capabilities, and uncertain economic returns. The review concludes that digital tools can meaningfully reduce FLW only when combined with organizational, behavioral, and governance strategies, and calls for standardized impact metrics and longitudinal system-level evaluations.
- Enterprise
- Quality assurance
Research
CognitiveAI_Assurance_Framework (CAAF V1.3)
Furaha Marwa
Zenodo (CERN European Organization for Nuclear Research) · 2026-09-12
The CognitiveAI Assurance Framework (CAAF V1.3) is an evidence-based methodology for evaluating the trustworthiness, safety, and operational readiness of high-risk and autonomous AI systems across eight domains including performance, explainability, fairness, privacy, security, and human oversight. It combines control implementation scores with evidence confidence to produce auditable assurance scores, and uses Critical Assurance Gates to prevent serious deficiencies from being masked by strong aggregate results. The framework is designed to convert responsible AI principles into measurable, decision-ready requirements for governance, deployment, and ongoing risk management in regulated environments. Its tiered risk classification and continuous monitoring features make it applicable to certification and policy contexts worldwide.
- Certifications
- AI policy
- Quality assurance
Research
Sustainable Career Readiness in the GenAI Era: Student Perceptions of Automation, Entry-Level Employment, and Pedagogical Support
Vasso Stylianou, Despo Ktoridou, Andreas Savva et al.
Sustainability · 2026-09-12
A survey of 153 undergraduates finds that students are broadly aware that generative AI will displace routine entry-level tasks—such as data entry, basic research, and simple coding—yet report weaker confidence in how well their academic programs are preparing them for an AI-augmented labor market. Students prioritize human-centered skills like critical thinking, creativity, and communication alongside AI literacy and hands-on digital tool experience. The study identifies an 'awareness-preparedness gap' and calls on higher education institutions to reform curriculum design, assessment, experiential learning, and ethical AI literacy to close it. The findings carry direct implications for how universities structure career readiness programming in the generative AI era.
- Workforce
Research
Key Considerations of Artificial Intelligence in Cognitive Behavioural Therapy
Marcus Chad
arXiv · 2026-09-12
This paper examines the integration of conversational AI into cognitive behavioural therapy (CBT) and finds a fundamental tension between AI's scalability benefits and the core requirements of effective therapy, namely genuine therapeutic alliance and ethical accountability. The analysis shows that while AI features like perceived anonymity can boost initial patient engagement and disclosure, they also prevent authentic rapport-building, and sycophantic AI behaviors may undermine therapeutic progress. The authors identify serious clinical safety risks from unregulated AI deployment, including documented ethical violations and an emerging phenomenon called 'AI-psychosis.' Based on this evidence, the paper argues for rigorous regulatory oversight and a hybrid clinician-AI model where AI handles data-intensive tasks while human therapists maintain the therapeutic relationship.
- AI policy
- Quality assurance
Research
Generative AI, AI literacy and the policy–practice gap in Australian higher education
Werner Botha, Liu Fei Tan, Binoy Appukuttan
Educational Studies · 2026-09-12
This study combines a survey of 462 students with academic-integrity case records from 2022–2026 at an Australian university to examine how generative AI is used and regulated in higher education. Students primarily used GenAI as a comprehension and planning scaffold, but use divided sharply by student status, with international students reporting heavier use across all thirteen tasks surveyed—especially language-dependent ones. Integrity records showed international students and commencing students (at twice the rate of continuing students) were over-represented in cases, many of which closed without a misconduct finding. The study concludes that current Australian policy treats AI literacy as integrity-supporting behaviour rather than as a developmental, transdisciplinary capability, revealing a significant policy–practice gap.
- AI policy
- Certifications
Research
From pilots to plots: A critical review of artificial intelligence for smallholder agriculture in South Asia
Amar Singh, Vinod Kumar Shukla, Chatter Singh
Outlook on Agriculture · 2026-09-12
This critical review examines AI applications targeting smallholder farmers (farms under two hectares) in South Asia across advisory, diagnostic, precision, credit, and market uses, drawing on 31 peer-reviewed sources. The evidence base is heavily skewed toward technical performance benchmarks, where near-perfect accuracies are reported on curated datasets, while only five experimental studies evaluate actual farmer outcomes—two of them in South Asia—with effects ranging from null to modest. Crucially, none of the outcome studies evaluates a machine-learned system, so the added value of AI over simpler digital tools remains unmeasured. The review attributes this gap to structural barriers including connectivity and gender divides, linguistic diversity, weak public extension integration, and misaligned evaluation incentives, and calls for outcome-based evaluation standards and better integration of AI into existing advisory institutions.
- Workforce
- AI policy
Research
Explainable Machine Learning for Predicting Indonesian Vocational School Accreditation: Geographic Validation, Probability Calibration, and Subgroup Auditing
Muhamad Riyan Maulana, Putu Sudira, Priyanto Priyanto et al.
Journal of Computing Theories and Applications · 2026-09-12
This study builds a machine-learning framework to predict the accreditation class (C, B, or A) of Indonesian vocational schools using administrative data on 14,134 schools. The authors use province-disjoint validation, probability calibration, and subgroup auditing to test whether the CatBoost model generalizes geographically; on a locked holdout of 1,536 schools from 8 unseen provinces, the calibrated model achieves a macro ROC-AUC of 0.757 and expected calibration error of 0.049. Key predictors include school size, teacher resources, and program diversity, though performance varies across school groups and provinces. The authors emphasize the framework is suited for calibrated preliminary screening only, not for replacing professional accreditation assessors.
- Certifications
- Quality assurance
Research
Epistemic dependence in AI-mediated learning within higher education: a framework for student judgement and responsibility
Yiran Du, Yijia Yuan
Assessment & Evaluation in Higher Education · 2026-09-12
This conceptual paper introduces the Epistemic Dependence in AI-Mediated Learning (ED-AIL) framework to analyze how generative AI can undermine students' own knowledge-related judgements in higher education. The framework identifies a mechanism by which AI outputs become 'epistemically directive,' displacing student verification, interpretation, and justification across five learning domains. Drawing on social epistemology, self-regulated learning, and assessment theory, the authors distinguish ED-AIL from related constructs like AI literacy and automation bias. The paper offers diagnostic implications for pedagogy, assessment, policy, and learning assurance centered on preserving students' epistemic responsibility.
- Quality assurance
- AI policy
Research
Artificial Intelligence and Its Impact on Building an Effective Cybersecurity Protection Framework An Analytical Legal Study in Light of Saudi Legislation, International Standards, and Governance Frameworks
Yussri Abdalla
International Journal For Multidisciplinary Research · 2026-09-12
This legal-analytical study examines whether existing Saudi and international regulatory frameworks adequately govern AI use in cybersecurity, comparing Saudi legislation against standards such as ISO/IEC 27001:2022, NIST CSF 2.0, NIST AI RMF 1.0, and the EU AI Act. The findings show that AI substantially improves threat detection, incident response, and cyber risk management, but that effectiveness depends on an integrated governance framework covering legal compliance, institutional accountability, and algorithmic transparency—not technology alone. The study concludes that Saudi Arabia has a broadly aligned regulatory environment but needs a more specialized legal framework addressing AI liability, human oversight, and risk assessment to support its Vision 2030 digital transformation goals.
- AI policy
- Certifications
Research
Testing the Kill Switch: A Conformance-Based Approach to Agentic AI Containment Assurance
Naveen Sundaresan
Zenodo (CERN European Organization for Nuclear Research) · 2026-09-12
This paper proposes a conformance-based audit framework for verifying that 'kill switch' or stop mechanisms in agentic AI systems actually work under real conditions. It defines a Target of Evaluation and five testable control families—trigger recognition, authority, cessation, latency, and failure resilience—along with evidence requirements that go beyond procedural self-attestation. The framework is designed to fill a gap left by existing regulations and standards (EU AI Act, NIST AI RMF, ISO/IEC 42001, and others), which require human oversight capabilities but do not specify how independent auditors can verify them. The authors also propose a vendor-neutral assurance harness with pluggable adapters to generate the evidence such auditable criteria would require.
- Certifications
- AI policy
Research
EAIMS: Enterprise AI Maturity Standard
Elias Naserkhaki
Zenodo (CERN European Organization for Nuclear Research) · 2026-09-12
EAIMS 1.1.0 is a release of the Enterprise AI Maturity Standard that extends the 1.0.x baseline with nine new normative requirements covering adversarial and agentic governance for material AI systems, including controls for blast radius assessment, agent credential boundaries, runtime abuse evidence, and synthetic identity control. The update also extends the G3 Autonomous/Agentic gate family with five new gates and strengthens managed-service-provider dependency requirements. The release underwent automated specification, backward-compatibility, and dependency validation, though the authors explicitly note it does not claim independent third-party assessor validation, accreditation, or regulatory approval. This matters because it provides a structured, versioned framework enterprises can use to assess and govern AI systems against adversarial and agentic risk.
- Enterprise
- Certifications
Research
Convergence of Digital Transformation in Accounting and Auditing Research and Training: Preliminary Survey and Solutions
V Nguyen
American Journal of Economics and Business Innovation · 2026-09-12
This paper surveys the integration of AI and digital tools into accounting and auditing education in Vietnam, using a Technology Acceptance Model (TAM)-based questionnaire analyzed with SPSS 20.0. Despite participants recognizing AI's potential to improve teaching quality and student employability, actual adoption rates were very low—only 8% for AI and 28% for data analytics—with key barriers being lack of training (80%), insufficient technical support (72%), and inadequate infrastructure (64%). The authors propose solutions including curriculum redesign, faculty development, and the establishment of legal and policy frameworks to accelerate digital transformation in the sector. The findings underscore the need to close digital skill gaps to prepare graduates for labor market demands and support broader economic development.
- Workforce
- AI policy
Research
The New Disease and the Machine Room: Keynes's 1930 Essay, Artificial Intelligence, and the Data-Center Buildout
Doug Doucette
Zenodo (CERN European Organization for Nuclear Research) · 2026-09-12
Doucette rereads Keynes's 1930 essay on technological unemployment as a framework for analyzing AI's early labor-market effects and the U.S. data-center buildout of the mid-2020s. The paper finds that AI-related labor disruption is concentrated in junior cognitive work and entry-level hiring rather than aggregate unemployment, while data centers generate many construction jobs but few permanent ones, creating local conflicts over water, power, and public consent. Drawing on Keynes's distinction between short-run maladjustment and long-run abundance, the paper proposes three policy tests: keeping entry-level career paths open, assigning infrastructure costs to large-load operators, and tracking whether productivity gains reach typical workers' pay. The conclusion argues that failure to meet these tests would be institutional rather than technical.
- Workforce
- AI policy
Research
Testing the Kill Switch: A Conformance-Based Approach to Agentic AI Containment Assurance
Naveen Sundaresan
Zenodo (CERN European Organization for Nuclear Research) · 2026-09-12
This paper proposes a conformance-based audit framework for verifying that 'kill switch' or stop mechanisms in agentic AI systems actually function as required under real operating conditions. Existing regulatory frameworks—including the EU AI Act, NIST AI RMF, ISO/IEC 42001, and Singapore's MAS and IMDA guidelines—mandate human oversight and intervention capability but provide no method for independent auditors to verify these controls work. The authors define a Target of Evaluation, five testable control families (trigger recognition, authority, cessation, latency, and failure resilience), machine-readable control representations, and evidence requirements that distinguish measured system behavior from procedural self-attestation. The approach aims to convert descriptive containment guidance into auditable, independently verifiable criteria using a vendor-neutral assurance harness with pluggable adapters.
- Certifications
- AI policy
- Quality assurance
Research
Reconceptualizing Age Assurance as a Sociotechnical Problem: Connecting Evidence, Evaluation, Claims, and Decisions
Renkai Ma, Prakriti Dumaru, Thomas Synaepa-Addison et al.
arXiv (Cornell University) · 2026-09-11
This paper argues that age assurance—systems designed to verify or estimate whether a user is a child—should be understood as a sociotechnical process rather than a purely technical problem. Reviewing 85 publications from 2020 through early 2026, the authors find that the field uses shared terminology inconsistently, that rights and privacy receive more scholarly attention than accuracy, error, and fairness, and that institutional actors are rarely held accountable when systems fail or users lack remedies. To address these gaps, the authors introduce the Age-Assurance Process Framework, which maps how evidence is evaluated, translated into age-related claims, and used in access or eligibility decisions.
- AI policy
- Quality assurance
Research
Same Patient, Different Order: Action-Level Reliability of Clinical LLM Agents Under Repeated Runs
Rohith Reddy Bellibatlu, Manpreet Singh, Zhoutian Han et al.
arXiv · 2026-09-11
This paper exposes a critical reliability gap in clinical AI agent benchmarks: the same inputs fed to a large-language-model clinical agent can produce materially different medical orders—different tests, medications, or referrals—across repeated runs, even when the benchmark reports the same pass/fail verdict each time. The authors introduce a 'same-input rerun' methodology with six reliability metrics and apply it to 1,000 runs across 50 tasks using two small open-weight models at varying temperatures, finding that under the 8B model at temperature 0.7 all 43 ordering groups produced a different set of orders across five identical runs, and in 22 of those cases the benchmark reported the same failing verdict despite materially different behavior. The study also finds that orders sometimes reach different endpoints, with one rejected by the record server while the agent was told it succeeded. The authors call for repeated-run evaluation, action-level stability reporting, and execution-faithful environment feedback as necessary safeguards before clinical LLM agents can be responsibly assessed or deployed.
- Quality assurance
- Certifications
Research
Harmfulness Propagation Dynamics: Layer-wise Trajectories of Adversarial Intent in Large Language Models
Noor Islam S. Mohammad, Uluğ Bayazıt
arXiv · 2026-09-11
This paper identifies 'Harmfulness Propagation Dynamics' (HPD), a phenomenon where harmful prompts cause a model's internal representation of harm to rise monotonically across transformer layers, while benign prompts stay flat or oscillatory. Building on this, the authors introduce HERALD, a lightweight input moderator that extracts a seven-dimensional feature vector from these cross-layer trajectories and classifies it with a tiny 288-parameter MLP, achieving an average F1 of 89.3 on OLMo2-7B across eight benchmarks and outperforming prior guard models on adversarial jailbreak detection (98.4 vs. 96.9 F1). The method requires only 262 KB of storage and adds negligible compute overhead, while also producing per-instance audit trails that reveal when and how harmfulness emerges within the model. This matters for AI safety and quality assurance by providing an efficient, interpretable, and reproducible mechanism for detecting harmful or adversarial inputs before they are acted upon.
- Quality assurance
News
Lawyer fined $5K over AI-hallucinated witnesses in a murder case
theverge.com · 2026-09-11
The Verge reports that New Mexico's Supreme Court has sanctioned attorney Stephen Aarons for submitting an AI-generated appellate brief containing fabricated witnesses and false police testimony in a murder case. The court fined him $5,000 and held him in contempt for failing to verify the factual claims and legal authority in the brief, which included invented witness testimony and false details about a shooter's appearance. A justice questioned how Aarons could have been unaware of the risks of using AI to draft legal documents.
- AI policy
- Quality assurance
Research
Governing at Machine Speed: An Adaptive Intelligence Architecture for Real-Time AI Policy Enforcement
Sandeep Bokkasam, B. Durgalakshmi
arXiv (Cornell University) · 2026-09-11
This paper diagnoses what it calls the 'attestation deficit'—a structural gap in which organizations adopt AI widely but cannot produce auditable, tamper-evident evidence that their governance policies are actually enforced within regulatory timelines. Drawing on the Stanford 2026 AI Index Report (362 documented incidents), the IBM/Ponemon 2026 Cost of a Data Breach study (USD 4.99M average breach cost, 92% of breached organizations lacking access controls), and the EY/AIUC-1 Consortium survey (38% end-to-end monitoring, 17% agent-to-agent coverage), the authors argue the failure is organizational and architectural rather than technical. To address this, they propose AGIL (Adaptive Governance Intelligence Layer), a conceptual five-layer architecture covering shadow AI detection, behavioral risk classification, inline policy enforcement at sub-100ms latency, continuous tamper-evident audit trail generation, and ML-driven cross-jurisdictional policy evolution. The authors explicitly note AGIL is a theoretical framework and that empirical validation through controlled deployment remains future work.
- AI policy
- Enterprise
News
ChatGPT-using lawyer punished for citing fake testimony from made-up witnesses
arstechnica.com · 2026-09-11
Ars Technica reports that the New Mexico Supreme Court held attorney Stephen Aarons in direct contempt for filing an AI-generated appellate brief containing fabricated witness testimony and false legal citations in a murder appeal. Aarons admitted he did not verify the brief's factual or legal content before signing and filing it, and never disclosed these failures to his client. The court referred him to a disciplinary board, finding he showed no remorse and insufficient concern for his client, who is serving a life sentence for murder.
- Quality assurance
- AI policy
Research
Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures
Harsh Raj, David Lee, Anas Mahmoud et al.
arXiv · 2026-09-11
This paper addresses the challenge of automated root-cause attribution (RCA) for failures in long-horizon AI agent tasks, where execution logs are too large for human review. The authors identify that existing one-shot LLM-based RCA methods perform poorly on lengthy traces because relevant evidence is sparse and distributed, causing the model to settle on early diagnoses. They propose 'Continual Search,' an iterative framework that repeatedly nudges the LLM judge to keep examining unresolved evidence across successive turns, and introduce MegaRCA-Mix, a new benchmark of 50 human-annotated long-horizon failure trials. Across multiple benchmarks and model families, Continual Search consistently improves attribution accuracy — raising GPT-5.5's F1 score by over 40% on MegaRCA-Mix — and shows that effective search strategy can allow lower-tier models to outperform higher-tier counterparts.
- Quality assurance
- Enterprise
Research
Hindsight Bias in Clinical Temporal Reasoning: How Future Data Exposure Affects Large Language Model Judgment
Misaki Matsuura, Sayantan Kumar, Ojas Kadam et al.
arXiv · 2026-09-11
This paper introduces a paired benchmark to measure 'hindsight bias' in large language models (LLMs) used for clinical reasoning—specifically, whether models give better-seeming answers when they can see future patient outcomes that real clinicians wouldn't have had access to at the time of decision-making. Using 171 PubMed Central case reports (covering sepsis and GLP-1/diabetes cases), the authors test four LLMs (GPT, Gemma, GLM, and Opus variants) under prospective versus full-timeline conditions, finding that exposure to complete timelines consistently shifts model responses toward outcome-consistent 'hindsight traps.' Critically, temporal masking—hiding future data—reduces this bias without hurting accuracy, suggesting that standard retrospective evaluations may systematically overstate model reliability for real-world clinical use.
- Quality assurance
- Certifications
Research
Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models
Utkarsh Soni, Syed Shariyar Murtaza, Yifan Nie et al.
arXiv · 2026-09-11
This paper introduces Tasks over Application Manuals (TAM), a benchmark designed to test whether large language models can follow long, rule-dense procedural documents across two real-world domains: ICD-10-CM clinical coding and U.S. federal sentencing guidelines. Each task requires navigating tens of thousands of interdependent rules across lengthy manuals to produce an exact answer. Even the best-performing approaches—including retrieval-augmented generation, ReAct-style prompting, and an agent harness on GPT-5—achieve only 1% exact-match accuracy on clinical coding and 15.5% on sentencing tasks, revealing that current LLMs are far less capable of reliable procedural reasoning than existing short-horizon benchmarks suggest. These findings have direct implications for high-stakes domains like medical coding and legal sentencing, where rule-following accuracy is essential.
- Quality assurance
- Certifications
Research
Scaling Clinical Judgment to Evaluate Medical AI
Thomas A. Buckley, Zahir Kanjee, Peter G. Brodeur et al.
arXiv · 2026-09-11
This paper introduces PrecepTron, a fine-tuned large language model designed to replicate physician-level evaluation of AI-generated clinical responses at scale. Trained via low-rank adaptation (LoRA) on a 32-billion-parameter model using a small set of physician-labeled examples, PrecepTron is validated against GRAND-ROUNDS, a new benchmark comprising 9,217 physician scores from 11 physicians across seven studies. The authors show that standard 'LLM-as-a-judge' approaches frequently disagree with physicians and with each other, while PrecepTron achieves physician-level consistency and successfully reproduces headline findings from five influential studies published in JAMA, Science, and Nature Medicine without requiring new human grading. This work matters for quality assurance in medical AI by enabling reproducible, large-scale assessment of clinical reasoning in LLMs that was previously infeasible due to the high cost and limited scale of human physician evaluation.
- Quality assurance
- Certifications