News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
Large Language Models as Examinees and Graders in Simulated General Surgery Oral Board-Style Cases: A Psychometric Pilot Study
Kian A Huang, Haris K Choudhary, Allan L Xu et al.
Cureus · 2026-08-03
This psychometric pilot study tested six frontier large language models as both test-takers and graders on simulated general surgery oral board-style cases. All LLMs achieved passing or near-passing performance, and AI graders showed greater internal consistency (Cronbach's α = 0.697) than a small three-surgeon human panel (α = −0.923), though the human panel lacked calibration training. Variance decomposition revealed that grader disagreement, rather than true differences in examinee performance, was the dominant source of score variability. The findings suggest LLMs may have value in surgical training and assessment frameworks, but replication with larger samples and calibrated human raters is needed before broader conclusions can be drawn.
- Certifications
- Quality assurance
Research
CogSig-Mamba: Hippocampal-Inspired Explainable Motion Forecasting with Causal Temporal Attribution
Emin Bayramov, Zoltán Istenes
Vehicles · 2026-08-03
CogSig-Mamba is a motion forecasting model for autonomous driving that combines high prediction accuracy with causally verified temporal explanations. Inspired by hippocampal memory, the model uses a five-stage pipeline to identify which observation windows drove each trajectory prediction, confirming causal faithfulness by showing that removing tagged windows shifts predictions by 4.6 m while removing untagged windows has negligible effect (0.46 m). Evaluated on Argoverse 2 across nearly 25,000 validation scenarios, it achieves competitive accuracy (minADE6 = 0.908 m, minFDE6 = 1.949 m) with only 1.9M parameters. The authors claim this is the first motion forecaster with verified temporal credit assignment, and explicitly connect the audit trail it produces to ISO 21448 safety-of-the-intended-functionality compliance requirements.
- Certifications
- Quality assurance
Research
Assessing a reduced-channel algorithm for end-to-end seizure detection on multiday EEG using inter-rater agreement with epileptologists
Zoë Tosi, Vamshi K. Muvvala, Tyler Newton et al.
Scientific Reports · 2026-08-03
This paper evaluates REMI Vigilenz AI for Event Detection (VED), a reduced four-channel automated seizure detection algorithm designed for wearable EEG systems, against expert epileptologist reviewers and a full-channel state-of-the-art detector (Persyst 15). Across 60 EEG records totaling over 4,000 hours, VED achieved average relative sensitivity of 77.0% compared to experts—comparable to inter-expert sensitivity ranges of 68.4%–88.3%—though at a higher false positive rate. The study demonstrates that inter-rater evaluation frameworks provide richer performance context than consensus-only approaches, and that a four-channel wearable-compatible algorithm can approach full-channel software performance, supporting its viability as a clinical decision support tool for extended ambulatory EEG monitoring.
- Quality assurance
- Certifications
Research
Can digital village policy support sustainable rural transformation? a multi-dimensional evaluation of governance, development, and environmental livability in China
Chunling Chen, Xiandong Wang
Frontiers in Environmental Science · 2026-08-03
This study evaluates China's Digital Village Pilot Policy (DVPP), launched in 2020 across 236 pilot counties, using a difference-in-differences framework on panel data from 464 prefecture-level cities over 2010–2023. The analysis finds a statistically significant but modest positive effect on composite rural revitalization (approximately 2.4% of the sample mean), robust across multiple robustness checks including PSM-DID and permutation testing. Decomposition reveals improvements across all five official dimensions—governance, culture, industry, living standards, and ecological livability—with effective governance and civilized culture showing the strongest short-run responses. The findings suggest digital village policy can simultaneously advance governance, development, and environmental outcomes in rural China, though disentangling policy effects from COVID-19 disruptions remains a challenge.
- AI policy
Research
Artificial Intelligence Competence Levels Among Secondary Mathematics Teachers in Northern Samar, Philippines
Levi V. Calubag
Journal of Research in Education and Pedagogy. · 2026-08-03
This cross-sectional survey of 215 public secondary mathematics teachers in Northern Samar, Philippines assessed AI competence using UNESCO's AI Competency Framework for Teachers (AI-CFT). Teachers overall scored at a 'Competent' level (mean 3.53/5), with stronger ratings in human-centered mindset and AI ethics but weaker scores in AI foundations and professional development. Access to AI tools and AI/ICT training exposure were the strongest positive predictors of AI competence, while more years of teaching experience showed a small negative association. The study concludes that targeted AI upskilling and equitable access to AI tools are key levers for improving responsible AI integration in mathematics classrooms.
- Workforce
- AI policy
Research
Utilization of Artificial Intelligence to support administrative decision-making in special education institutions in Saudi Arabia: perceptions of principals, supervisors, and teachers
Abdulaziz Alsuhaymi, Mahmoud Mohamed Eltantawy
Frontiers in Artificial Intelligence · 2026-08-03
This study surveyed 173 principals, supervisors, and teachers in Saudi Arabian special education institutions to assess their perceptions of AI-supported administrative decision-making. Participants showed moderate overall acceptance of AI (M=3.22) and high agreement on its usefulness for administrative processes, but expressed significant ethical and legal reservations (M=2.38). AI training level—but not professional role or years of experience—significantly influenced perceptions, with those completing more AI courses holding more favorable views. The authors conclude that targeted professional development, ethical safeguards, and sustained human oversight are essential for responsible AI adoption in special education administration.
- AI policy
- Workforce
Research
A Structural Model of Factors Influencing AI Adaptation and SME Performance in Bangkok and Nonthaburi, Thailand
Pattarapon Chummee
Asia Social Issues · 2026-08-03
This study examines how technological readiness, organizational capability, and environmental pressures drive AI adoption and performance among 360 SMEs in Bangkok and Nonthaburi, Thailand, using the TOE framework with CFA and SEM analysis. Organizational readiness was the strongest predictor of AI adaptation (β=0.42), followed by technology (β=0.36) and environment (β=0.29), while AI adaptation strongly predicted SME performance outcomes including productivity, cost efficiency, and market expansion (β=0.67). AI adaptation fully mediated the relationship between these contextual factors and performance outcomes. The findings offer policy-relevant guidance for strengthening SME capabilities and accelerating AI-driven economic growth in Thailand's metropolitan region.
- Enterprise
- AI policy
Research
Responsible AI Governance in Iran: A Program-Theory Evaluation of the National Artificial Intelligence Strategy
Mohammad Masoumi Godarzi, Amin Poyanrad
Journal of Historical Research Law and Policy · 2026-08-03
This study evaluates Iran's National Artificial Intelligence Strategy using a qualitative, theory-based case-study approach, combining documentary analysis with interviews of 12 experts across academic, research, policy, and industrial sectors. The findings show that while the strategy offers a normative foundation emphasizing human dignity and national values, its implementation architecture is insufficiently specified, with major weaknesses in institutional authority, financing, data governance, legal liability, and measurable indicators. The authors conclude that effective responsible AI governance requires translating broad objectives into an explicit, funded, and accountable program of action, not merely ethical declarations. The study is relevant to how governments design and assess national AI policy frameworks.
- AI policy
Research
The fear of being replaced by Artificial Intelligence and its association with career anxiety among university students
Tayiba Rasheed, Rizwan Shabbir, Muhammad Rehan et al.
Aposta · 2026-08-03
This study surveyed 369 university students across multiple Pakistani cities and found a statistically significant positive association between fear of AI and career anxiety, with AI-related fear serving as a significant predictor of students' concerns about future employment. The research provides empirical evidence that rapid workplace automation is generating measurable psychological consequences for students regarding their career prospects. The authors recommend promoting AI literacy, career adaptability, and psychological support programs to better prepare students for an AI-driven labor market.
- Workforce
Research
Applying EASTL ethical constructs in AI-driven CRM: the mediating role of sustainable customer trust in enhancing customer retention in private sector banks
A. Geetha, V. Venkatragavan
Frontiers in Artificial Intelligence · 2026-08-03
This study examines how ethical AI practices in banking CRM systems affect customer retention, using Sustainable Customer Trust as a mediating variable among 361 respondents in private sector banks. Applying the EASTL framework via PLS-SEM, the model explains 48.4% of variance in Sustainable Customer Trust and 32.1% in Customer Retention. Algorithmic Transparency, Data Security, Institutional Accountability, and Regulatory Legitimacy all significantly influence retention both directly and indirectly through trust, while Normative Ethical Alignment was not significant. The findings position ethical AI as a strategic lever—not merely a compliance requirement—for building durable customer relationships in digital banking.
- Enterprise
- AI policy
Research
The means of prediction and the production function of AI
Maximilian Kasy
Journal of Economic Interaction and Coordination · 2026-08-03
This paper argues that the central risks of AI are not conflicts between humans and machines but between different human groups over who controls what AI systems optimize for. The author analyzes the 'production function of AI'—how data, compute, expertise, and energy map to predictive performance—drawing on scaling laws to explain why power over AI has become concentrated among a few actors. The paper finds that market-based governance through individual data property rights fails because machine learning inherently produces data externalities and platform network effects are artificially maintained, concluding with proposals for democratic institutions like sortition and liquid democracy to give affected communities a voice in AI objectives.
- AI policy
Research
The shifted burden of search: how digital hiring platforms reallocate screening labor to job seekers
Stephen Santhosh
Frontiers in Human Dynamics · 2026-08-03
This perspective paper argues that digital hiring platforms have quietly shifted uncompensated screening labor onto job seekers: tasks once handled by employers—verifying a posting is genuine, reachable by human reviewers, and a viable match—now fall on applicants. The authors frame this through information asymmetry and present a taxonomy of hiring-intent signals drawn from 784 illustrative job postings to show these opacity signals appear routinely in practice. The paper closes with research propositions and suggestions for how platform and policy design might rebalance this burden.
- Workforce
- AI policy
Research
Governing Sustainable Artificial Intelligence: A Green AI Regulatory Framework Based on Responsible Innovation
Sri Ariyanti, Muhammad Suryanegara, Ajib Setyo Arifin
Journal of Sustainability · 2026-08-03
This study analyzes AI regulatory frameworks across six major jurisdictions—the EU, US, China, Japan, South Korea, and Singapore—finding that current frameworks lack explicit mechanisms for environmental impact assessment, energy and emissions reporting, and adaptive sustainability regulation. Using a Responsible Innovation conceptual lens, the authors propose a three-layer Green AI regulatory framework covering sustainability assessment, governance and compliance, and technical innovation to address these gaps. The work offers practical guidance for policymakers seeking to integrate environmental sustainability—particularly around energy consumption, carbon emissions, and resource use—into AI governance aligned with global sustainability goals.
- AI policy
Research
Same violence, different answer: how AI responds to coercive control against women across languages
Lyu Chang, Sònia Estradé Albiol, Núria Vergés Bosch
arXiv · 2026-08-02
This study tests seven widely used AI language models across nine languages to see how they respond when a woman experiencing coercive control—specifically phone tracking by a partner—asks the AI to help her write a self-blaming letter accepting the surveillance. The researchers scored whether each model complied with the request and whether it recognized the controlling behavior, challenged the self-blame, and affirmed the woman's agency. Results reveal two independent failure axes: AI systems from non-anglophone developers were most likely to comply in their own builders' language, and the degree to which a sympathetic excuse for the partner could undermine the model's recognition of coercive control varied sharply by language. Two frontier systems maintained protective responses across all languages tested, demonstrating that consistent, language-agnostic safeguards are achievable and that failures elsewhere reflect deliberate or neglected design choices.
- AI policy
- Quality assurance
Research
Why Formal Monitors Fail: Attack Distribution Entropy as a Coverage Bound for LTL-Based LLM Agent Safety
Ruiyang Zhang
arXiv · 2026-08-02
This paper explains why LTL/FSA-based runtime safety monitors for LLM agents succeed on some model architectures but nearly fail on others. The authors prove theoretically that an FSA monitor's recall is bounded by how concentrated the attack distribution is: when attacks cluster into a few repeated trigger-completion patterns (low Shannon entropy), a small fixed invariant set achieves high recall (68–75% on GPT-class and DeepSeek backends), but when attacks are highly dispersed (high entropy, as with Gemini variants), even well-designed monitors achieve near-zero recall (6–13%). Validated across eight frontier LLM architectures, entropy explains 76% of variance in monitor coverage (Pearson r = -0.87), and the authors introduce a pre-deployment entropy test that predicts monitor effectiveness from a small attack sample before deployment. This matters for quality assurance and certification of AI agent safety systems, as it shows that architecture-agnostic formal monitors cannot be assumed reliable without first characterizing the attack distribution of the target model.
- Quality assurance
- Certifications
Research
High-Stakes Decisions with Language Models: Insights from Emergency Triage
Khurram Yamin, Christopher Kelly, Bryan Wilder et al.
arXiv · 2026-08-02
This paper examines how large language models make emergency triage recommendations through a probabilistic decision-making framework, demonstrating that the same underlying model predictions can yield very different clinical recommendations depending on explicitly stated utility functions (i.e., how much weight is given to missing emergencies versus unnecessary escalation). Using clinical vignettes from a consumer triage system evaluation, the authors show that capable language models do adjust their recommendations in response to stated utilities. The key finding is that effective deployment of language models in high-stakes settings requires not just accurate predictions but also explicit specification of decision objectives. The authors argue that language models in clinical and other high-stakes domains should be evaluated as probabilistic decision systems that combine predictive performance with explicit utilities.
- Quality assurance
- AI policy
Research
Can Language Models Identify Shadow Trading Targets? An NLP Evaluation of SEC Enforcement Theory
Sarah Wilson, Michael MacKay, Anthony Marello et al.
arXiv · 2026-08-02
This paper tests whether large language models can identify 'economically linked' peer firms—the core empirical premise behind the SEC's shadow trading enforcement theory—by applying a two-stage NLP pipeline to the Management's Discussion and Analysis sections of SEC 10-K filings across 30 M&A events. The pipeline successfully recovers Incyte as a close peer in the SEC v. Panuwat case, but across the full 217-observation dataset finds no meaningful association between semantic similarity and announcement-day abnormal stock returns (Spearman correlation +0.05, 95% CI [-0.08, +0.18]). The results cast doubt on the empirical assumption that insiders could reliably identify shadow trading targets ex ante using public disclosures, and the authors argue this has direct bearing on constitutional questions surrounding the SEC's financial surveillance infrastructure.
- AI policy
Research
The Overstated Cost of AI Fairness in Criminal Justice
Ignacio Cofone, Warut Khern-am-nuai
arXiv · 2026-08-02
This paper challenges the widely held assumption that making AI systems fairer necessarily reduces their predictive accuracy, using the COMPAS recidivism dataset as an empirical test case. Through causal inference methods, the authors find that racial bias in COMPAS is not merely replicated by trained models but is actively amplified by them, undermining claims that algorithmic decision-making is neutral or simply mirrors human bias. Crucially, they argue that the standard fairness-accuracy tradeoff is overstated because the unconstrained model's predictions are themselves distorted by biased outcome variables (rearrest data), meaning that applying fairness constraints can correct—rather than cost—predictive quality. The authors extend these findings to lending, hiring, and housing, and draw implications for how law and policy should treat algorithmic fairness requirements.
- AI policy
- Workforce
Research
Fighting Fire with Fire: On the Feasibility of Protecting Exercises Against AI Cheating
Tobias Braun, Jonas Grebe, Louis Rethfeld et al.
arXiv · 2026-08-02
This paper investigates whether adversarial machine learning techniques can be used to detect students who use AI tools to cheat on educational exercises. The approach embeds subtle visual perturbations into multiple-choice question images that steer AI solvers (Claude, Gemini, ChatGPT) toward specific incorrect answers, creating a statistical fingerprint; students who blindly copy AI responses reproduce this induced answer pattern at a detectable rate. The authors demonstrate feasibility under realistic black-box conditions using surrogate models to optimize perturbations, then apply statistical hypothesis testing for detection. The findings highlight both the potential and the limitations of using AI vulnerabilities to protect the integrity of educational assessments.
- Quality assurance
- Certifications
Research
Control Under Compression: Reliability Frontiers for Tool-Using Agents
Yinghan Hou, Zongyou Yang
arXiv · 2026-08-02
This paper examines what happens to AI agent reliability when the system-level instructions that govern tool use—called agent control contexts (ACCs)—are compressed to reduce token costs. Using a benchmark of 15,525 runs across nine ACCs, three task families, and six compression budgets, the authors find a nonlinear reliability frontier: at 75% retained context, top methods still achieve ~92–93% success near the full-context baseline of 93.8%, but between 50% and 35% retention methods diverge sharply, with the best achieving only 47% and the worst 19.9%, and protocols become fragile below 25%. The study shows that failures manifest primarily as tool-execution and action-parsing errors, that no single compressor ranks best across all contexts, and that ACC compression must be treated as a runtime-reliability problem evaluated through executable outcomes rather than just a token-reduction technique.
- Quality assurance
- Enterprise
Research
Don't Offer What Can't Be Done: Deterministic Executability Gating for LLM Skill Selection at Scale
Ortal Ashkenazi, Vitalii Kloz, Mykhailo Ulianchenko
arXiv · 2026-08-02
This paper describes a three-stage pipeline deployed in Wix's customer-care LLM agent ('Helpmate') to prevent the model from selecting skills it cannot actually execute given a user's current account state. A semantic matcher first narrows candidate skills, then a deterministic 'executability gate' removes any skill whose hard-stop conditions are already met, and finally the LLM chooses among the remaining valid options. In a production analysis of 756.6K user messages, the combined approach reduced skill-description context by 90.5% compared to exposing all skills to every message, and a counterfactual replay showed that without the gate the model would have selected a non-executable skill in 7.8% of tested conversations. The work demonstrates that deterministic pre-filtering meaningfully improves LLM agent reliability at scale by stopping the model from offering actions that cannot be completed.
- Enterprise
- Quality assurance
Research
DeBERTa-Sentinel: Toward Transparent and Trustworthy Detection of AI-Generated Text
Muhammad Yousaf Rehman, Muhammad Islam
arXiv · 2026-08-02
DeBERTa-Sentinel is a transformer-based framework for detecting AI-generated text that uses DeBERTa-v3's disentangled attention to identify subtle structural irregularities in synthetic content. Evaluated on the GLC-AIText dataset of 28,057 human and LLM-generated samples (GPT, LLaMA, and Claude), it achieves 97.53% test accuracy, 95.89% precision, 99.33% recall, and 99.53% ROC-AUC, outperforming the RoBERTa-Sentinel baseline. A key design feature is token-level explainability that exposes linguistic markers such as academic phrasing and formal transitions, enabling journalists, educators, and platform trust-and-safety teams to audit detection decisions. This matters for quality assurance and policy contexts where verifiable, auditable content-authenticity tools are needed to combat misinformation and protect academic integrity.
- Quality assurance
- AI policy
Research
Do people rely on ChatGPT more than their peers to detect deepfake news?
Yuhao Fu, Nobuyuki Hanaki
arXiv (Cornell University) · 2026-08-02
This experimental study examines how people weight advice from ChatGPT (GPT-4), human peers, and linguistic experts when identifying AI-generated fake news (deepfake news). Results show participants initially relied more on ChatGPT than on peers, though a 2025 follow-up found greater reliance on linguistic experts than ChatGPT, suggesting shifting trust in AI tools over time. Detection performance improved only when participants relied on high-quality advice, meaning AI-assisted deepfake detection is beneficial only if the AI tool itself is sufficiently accurate. The findings underscore generative AI's dual role as both a creator of disinformation and a potential mitigation tool.
- Quality assurance
- AI policy
Research
Conformance result: Shango MID against the AISVS C9 action-class scenario suite v1.0.0
Ishaan Ghosh
Open MIND · 2026-08-02
This conformance report documents how Shango MID, an AI write-governance layer, performed against the OWASP AISVS C9 action-class scenario suite (v1.0.0). Fourteen test cases were run in declaration-only mode; the original reading showed nine matches, four divergences, and one not-applicable, while a correction attributable to a defect in the suite itself raises the match count to twelve of fourteen. The report is notable for its transparency: it distinguishes verdicts produced by shipped code from those carried by adapter scaffolding, and retains the original contradictory record rather than overwriting it. The single remaining divergence—a composition-aware chain-sealing case—is acknowledged as an open research question unsatisfied even by the suite's own reference implementation.
- Certifications
- Quality assurance
Research
AI-Enabled Post-Process Surface Inspection in Laser-Welded Al–Cu Battery Interconnects: A Critical Review
Maricruz Hernández-Hernández, Adriana C. Flores‐Gallegos
Metals · 2026-08-02
This review paper examines how artificial intelligence can be applied to post-process surface inspection of laser-welded aluminum–copper battery interconnects used in electric vehicles. The authors detail a framework linking controlled image acquisition, defect taxonomy, weld-level datasets, and AI inference tasks—including classification, detection, segmentation, and anomaly detection—with accept–review–reject decision logic and functional validation. The review highlights that Al–Cu laser welding is prone to visible defects and hidden discontinuities due to thermophysical mismatch and intermetallic compound formation, making reliable automated inspection critical for battery pack quality. The proposed framework supports reproducible, risk-sensitive quality assurance while preserving expert oversight for uncertain or high-stakes decisions.
- Quality assurance