News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
The Severance Problem: LLMs are Unaware of the Person Beyond the Prompt
Dor Litvak, Liu Leqi
arXiv · 2026-07-15
This paper identifies the 'Severance Problem': AI personal assistants lack an explicit representation of what they don't know about the user beyond the immediate prompt, which the authors argue underlies failure modes like sycophancy, overconfidence, and hallucination. They propose the 'Severance Schema,' a structured way to encode the model's ignorance about the user across dimensions such as physicality, temporality, and interiority. Tested across five model families, models using the schema show reduced sycophancy, harmful advice, and hallucination, and are more likely to ask clarifying questions instead of confidently extrapolating from incomplete information. This matters for AI assistants deployed in personal and consequential decision-support contexts, where overconfident or sycophantic outputs can cause real harm.
- Workforce
- Enterprise
- Quality assurance
- AI policy
Research
Implicit Reasoning Steering via Concept Chaining
Xiao Ye, Sanika Chavan, Yuxi Huang et al.
arXiv · 2026-07-15
This paper investigates a vulnerability in large language model reasoning called 'Concept Chaining,' where short natural-language paragraphs linking question entities to a target answer through intermediate concepts can systematically bias a model toward designated answers. By continuing pretraining on these connection paragraphs, the researchers show that a model's answer preferences on multiple-choice questions can be covertly redirected without explicit instructions or direct answer cues. The key finding is that reasoning brittleness in LLMs—evidenced by inconsistent answers across repeated sampling—is not just an evaluation artifact but a practical exploit channel through which indirect, ordinary-looking text can amplify latent biases. This has significant implications for AI policy and quality assurance, as it demonstrates that models can be manipulated through subtle data poisoning that is harder to detect than direct paraphrases.
- Quality assurance
- AI policy
Research
Privacy Leakage in Federated Learning in Radiology Reports: A Comparative Evaluation of Tokenizer-Driven Privacy Risks
Santhosh Parampottupadam, Andres Martinez, Dimitrios Bounias et al.
arXiv · 2026-07-15
This study investigates how much sensitive patient information can be reconstructed from shared model gradients in federated learning (FL) systems trained on radiology reports. Using a GPT-2-style transformer trained across six simulated FL clients on nearly 370,000 clinical text documents, researchers tested three tokenizers (GPT-2, RadBERT, LLaMA-2) under a malicious server scenario and found exact sentence reconstruction rates of 31–44% across conditions, with RadBERT recovering the most clinical terminology. Critically, no tokenizer prevented leakage, and the authors conclude that tokenizer choice is a privacy-relevant decision — not just a utility one — with larger batch sizes only partially mitigating risk. The findings suggest that additional safeguards such as differential privacy and secure aggregation are likely necessary to meet HIPAA and GDPR requirements for FL in radiology NLP.
- AI policy
- Quality assurance
- Enterprise
Research
Instant NuRec: Feed-Forward 3D Gaussian Reconstruction for Driving Scene Simulation
NVIDIA, :, Jiahui Huang et al.
arXiv · 2026-07-15
Instant NuRec is a feed-forward neural reconstruction model that converts short multi-view driving footage into a fully simulatable 3D Gaussian Splatting world in a single forward pass, taking roughly 1.5 seconds per 10–20-second multi-camera scene. The model produces layered static and dynamic scene representations, a sky cubemap, and per-camera corrections, and achieves a PSNR 2.01 dB above the strongest evaluated baseline on the Waymo Open Dataset. By dramatically accelerating scene reconstruction without per-scene tuning, it lowers the cost and time required for closed-loop autonomous driving simulation, supporting safer and more scalable policy evaluation.
- Enterprise
- Quality assurance
Research
How Artificial Intelligence LLM Engines Shape the Global Conflict Information Environment
Jason Miklian
arXiv (Cornell University) · 2026-07-15
This paper examines how large language model answer engines handle questions about armed conflicts, testing five leading systems across 28 conflicts and scoring 5,460 answers against documented evidence. The authors find that conflicts with thinner information records prompt more hallucination, misattribution, and miscounting, and that these thin records are structurally vulnerable to Generative Engine Optimization (GEO), a form of information warfare where partisan actors manipulate the sources AI engines draw from. An analysis of 1,048 websites used by these engines finds that state-partisan digital capture of AI sourcing is already underway and growing. The paper argues this has significant policy implications, calling for renewed investment in deep local monitoring and translation-based research that AI tools cannot replicate.
- AI policy
Research
Accountability perspectives in the context of AI in intelligence and security investigations
Jorge Constantino, Floris van Krimpen, Sarah van Gerwen et al.
Frontiers in Political Science · 2026-07-15
This paper investigates how intelligence and security professionals in the Netherlands, Belgium, Norway, and the UK conceptualize accountability when using AI in digital investigations. Drawing on semi-structured interviews, the study finds that accountability challenges are seen as broad socio-legal-technical problems involving AI and data governance, human oversight, and organizational resources—not as issues stemming from AI alone. A key tension identified is balancing national security imperatives against fundamental rights protections in democratic societies. The findings offer empirical grounding for debates about how accountability is actually perceived and operationalized in AI-supported security contexts.
- AI policy
Research
Building AI self-efficacy in EFL teachers: A mixed-methods study of professional development using an artificial intelligence-based lesson planning tool
Ferdi Çelik
Uluslararası Türk Eğitim Bilimleri Dergisi · 2026-07-15
This mixed-methods study examined whether a two-week professional development program using Diffit, an AI-driven lesson planning tool, improved AI self-efficacy among 35 in-service EFL teachers. Quantitative results from the Wilcoxon signed-rank test showed a statistically significant increase in AI self-efficacy from pretest to posttest, while qualitative analysis identified seven themes explaining this change, including efficiency gains, quality control concerns, and ethics and data issues. The findings offer empirical support for AI self-efficacy as a meaningful variable in educational technology research and provide practical guidance for designing professional development that helps language educators integrate AI tools sustainably.
- Workforce
Research
AI-Augmented Human Resource Management? Insights from German companies
Yannick Kalff, Katharina Simbeck
arXiv (Cornell University) · 2026-07-15
This study investigates how German companies are integrating AI—including generative AI and predictive analytics—into Human Resource Management, drawing on interviews, group discussions, and a survey of 410 respondents. The findings show that while AI tools improve HR analytics capabilities and hold strategic potential for talent development, their adoption is primarily driven by efficiency and rationalization goals rather than people-centered transformation. Implementation is shaped by factors such as digital infrastructure, co-determination frameworks, and concerns around data governance and algorithmic transparency. The research highlights the ambiguous role of AI in HR, where the promise of augmented predictive capabilities often serves cost-cutting and process-streamlining ends.
- Workforce
- Enterprise
Research
“Already Outdated and Still Under Review”: Mapping the Landscape of Research on Adolescent Development and AI Chatbot Use
Anne J. Maheux, Caitlin Mbuakoto, Madeline Valentino et al.
arXiv · 2026-07-15
This mixed-methods study surveyed 141 developmental researchers and professionals and interviewed 15 to map priorities for studying adolescent AI chatbot use. Experts identified top research priorities including AI literacy, AI as a relational actor, misinformation, and overreliance, while warning against vague 'screen time' metrics. Key barriers include AI's rapid pace outstripping academic timelines, limited industry transparency, and funding gaps. The study recommends three immediate policy guardrails: enforceable youth-centered regulation, developmentally informed product design, and multidimensional AI literacy education.
- AI policy
Research
Evaluation of AI Tools in Terms of Sustainable Workforce Productivity and Their Impact on Organizational Outcomes Using an Intuitionistic Fuzzy Approach
Adis Puška, Jurica Bosna, Darko Božanić
Journal of Operations Intelligence · 2026-07-15
This study develops a multi-criteria decision-making framework to evaluate six AI tools across eight sustainability-oriented criteria, aiming to help organizations improve productivity without compromising employee well-being. Using an intuitionistic fuzzy approach combined with SWARA and MABAC methods, the researchers found that improving individual productivity and enhancing employee engagement are the most important criteria, while AI tools focused on learning, development, and decision-support rank highest for sustainable organizational efficiency. The framework offers practical guidance for managers balancing technological adoption with long-term human resource development. Results were validated through comparative and sensitivity analyses.
- Workforce
- Enterprise
Research
Extreme Cognitive Assistance and Open Futures
Nathaniel Sharadin
Journal of Ethics and Social Philosophy · 2026-07-15
This philosophical paper argues that providing children with 'extreme cognitive assistance'—AI help that is ubiquitous, domain-general, and substitutes for their own cognitive effort—violates a moral obligation to preserve children's open futures. The author contends that when children routinely offload cognitive work to such systems, core regulatory capacities like inhibitory control, working memory, and planning fail to develop normally. The paper warns that major technology companies are explicitly investing hundreds of billions of dollars toward building exactly these kinds of systems, and that even existing, less-than-maximal AI assistance may already be causing developmental harm. It concludes skeptically that compensating sources of cognitive exercise are unlikely to offset the damage.
- AI policy
- Workforce
Research
AI-augmented judgement in teacher performance assessment: Evidence from a human-AI moderation workflow
Zara Ersozlu, Susan Ledger, Mark Babic et al.
Contemporary Educational Technology · 2026-07-15
This study designed and evaluated a human-AI workflow for marking teaching performance assessment (TPA) portfolios, where an AI agent assists assessors by providing structured, evidence-based feedback. Testing compared human-only marking against AI-supported marking across dimensions including time, accuracy, cognitive load, feedback quality, usability, bias, and fairness. Results showed the AI-supported workflow reduced marking time and perceived workload while improving consistency and fairness in the moderation process. Human assessors retained final decision-making authority, with the AI positioned as a 'third eye' that supports rather than replaces human judgement.
- Quality assurance
- Certifications
Research
Governance and Security-by-Design: Embedding Safety and Alignment into Agentic AI Systems
Himanshu Joshi, Shivani Shukla
arXiv · 2026-07-15
This paper proposes a governance and security-by-design framework that embeds safety and alignment mechanisms directly into agentic AI architectures rather than relying on external oversight. Across 800 experiments, the authors identify three critical failure modes: AI-generated code introduces memory safety vulnerabilities in 42.7% of efficiency-focused cases and cryptographic flaws in 21.1% of security-focused cases, security issues worsen by 37.6% after five feedback iterations, and domain expertise degrades by 47% as irrelevant context accumulates. Their multi-agent governance architecture, validated through industry partnerships, achieves a 40% reduction in post-deployment safety incidents. The findings argue that autonomous AI systems require governance to be built into their architecture to remain safe and aligned as complexity scales.
- AI policy
- Quality assurance
Research
Caught in Between Protection and Autonomy: A Scoping Review of Youth’s Right to Participation in Artificial Intelligence
Lauriane Lalande, Sarah Bouhouita-Guermech, Hazar Haidar
Youth · 2026-07-15
This scoping review examines how scientific literature conceptualizes young people's right to participate in AI governance and design. Analyzing 15 sources from 2010–2024, the study finds that while frameworks like the OECD AI Principles and UNESCO's AI Ethics Recommendation exist, they rarely incorporate youth perspectives despite children's widespread exposure to AI in educational and social contexts. The literature frames youth participation as both a children's rights requirement under UN General Comment No. 25 and a condition for legitimate digital governance, yet a major gap remains between youth exposure to AI systems and their actual influence over decisions. The authors argue that integrating youth participation into AI governance and education could lead to more inclusive and developmentally informed technological ecosystems.
- AI policy
Research
Artificial Intelligence and the Future of the Labour Economy:A Multi-Criteria Expert Evaluation of Institutional Models of Adaptation
Jurica Bosna, Adis Puška, Darko Božanić
Journal of Soft Computing and Decision Analytics · 2026-07-15
This study uses a hybrid fuzzy-rough multi-criteria decision-making framework (combining SiWeC and WASPAS methods) to evaluate six institutional models for adapting labour markets to AI and automation. Expert judgments across eight criteria—most notably innovation/productivity incentives and institutional feasibility—were applied in a Croatian case study, finding that mass retraining and participatory AI capital models are the most suitable responses. The research supports evidence-based policymaking for workforce adaptation in the AI era while also contributing a methodological advance by explicitly handling uncertainty and imprecision in expert evaluations.
- Workforce
- AI policy
Research
The Missing Link to Safe AI Homologation
Mark Locherer, Michael Kordovan, Bernd Buxbaum et al.
arXiv · 2026-07-15
This paper examines the challenges of certifying AI-based safety-related systems under the EU AI Act and Machinery Regulation, arguing that traditional safety assurance approaches are insufficient. The authors introduce the concept of an 'efficient assurance argument' as a structured means to demonstrate regulatory compliance throughout a system's lifetime, covering risk management, design, verification, and validation. They conclude that this assurance argument acts as the critical missing link in AI homologation, serving both engineering teams and notified bodies responsible for type approval.
- Certifications
- AI policy
Research
Addressing benchmarking gaps in large language models for health and medicine with dynamic red-teaming
Jiazhen Pan, Bailiang Jian, Paul Hager et al.
Nature Health · 2026-07-15
This paper introduces DAS (Dynamic, Automatic and Systematic), a red-teaming audit framework that continuously stress-tests large language models across four safety-critical dimensions: robustness, privacy, bias, and hallucination. Applied to 15 state-of-the-art LLMs, the framework revealed a severe 'benchmarking gap': despite median MedQA accuracy above 80%, 94% of correct answers failed under dynamic robustness testing, privacy leaks were elicited in 86% of scenarios, cognitive bias altered recommendations in 81% of fairness tests, and hallucination rates exceeded 74%. The findings suggest that high static benchmark scores may reflect superficial memorization rather than genuine safety, and that continuous adversarial auditing is necessary before LLMs are deployed in consumer health or clinical settings.
- Quality assurance
- Certifications
Research
ЭКОНОМИЧЕСКАЯ ОЦЕНКА ВНЕДРЕНИЯ ИНСТРУМЕНТОВ ИСКУССТВЕННОГО ИНТЕЛЛЕКТА В БИЗНЕС-ПРОЦЕССЫ МАЛОГО ПРЕДПРИЯТИЯ
Егоров Д. Б.
CyberLeninK (CyberLeninka) · 2026-07-15
This paper develops a methodological framework for economically evaluating AI tool adoption in small business processes. It argues that simply equating saved labor time with monetary savings overstates AI project returns, since freed-up capacity does not automatically translate into cost reductions or additional output. The proposed model integrates total cost of ownership with effects from labor time savings, marginal revenue growth, error reduction, and loss prevention, alongside correction coefficients for actual utilization, quality, and result attribution. A pilot algorithm and model example in a small business services firm demonstrate that investment attractiveness depends heavily on how freed time is actually used and whether commercial effects are sustained.
- Enterprise
Research
Rethinking Penetration Testing for AI-Enabled Systems: From Resource Compromise to Behavioral Objective Violation
Mohammad Allahbakhsh, Mohammad Hassan Bahari, Moslem Attar-Raouf
arXiv (Cornell University) · 2026-07-15
This paper argues that traditional penetration testing—focused on infrastructure and resource compromise—is insufficient for AI-enabled systems, where adversaries can alter system behavior through prompt injection, data poisoning, sensor manipulation, and other influence pathways without ever breaching underlying infrastructure. The authors reframe penetration testing as objective-driven behavioral evaluation, defining success as the feasible induction of AI-governed behavior that violates operational objectives under an explicit threat model. They propose a structured testing workflow and illustrate the approach with an AI-enabled security operations center assistant. The framework provides practical guidance for evaluating adversarial risk in deployed AI systems.
- Quality assurance
- Certifications
Research
CAVA: Canonical Action Verification and Attestation for Runtime Governance of Agentic AI Systems
Zexun Wang
arXiv (Cornell University) · 2026-07-15
CAVA introduces a runtime-semantics layer that converts heterogeneous agentic AI activity—spanning coding hooks, browser automation, API gateways, and workflow engines—into canonical runtime action objects, enabling consistent governance across incompatible runtime records. The system formalizes how an approved action is identified, how approval is bound to execution evidence, and how an independent verifier can later reproduce the same action identity. A 384-variant benchmark covering semantic equivalence, wrapper bypass, tamper detection, and runtime portability validates the reference implementation. The work addresses a foundational gap in deployer-side AI governance: ensuring that what was approved is verifiably what was executed.
- AI policy
- Enterprise
Research
From promise to practice: artificial intelligence in mental health care in the MENA region
Sara El Hajj, Ahmad Nsouli, Mohamad Wehbe et al.
Frontiers in Psychiatry · 2026-07-15
This narrative review examines the use of AI-driven conversational tools—including large language models and psychotherapy chatbots—for mental health care in the Middle East and North Africa (MENA) region, where treatment gaps reach 80–95% due to provider shortages, stigma, and cultural barriers. The review finds that while AI tools offer high accessibility and user engagement for low-intensity support, their effectiveness is constrained by linguistic mismatches such as Arabic diglossia and poor alignment with locally grounded expressions of distress. A paradox emerges in which stigma and privacy concerns drive users toward anonymous AI tools yet simultaneously limit trust in their clinical reliability, reinforcing preference for hybrid human-oversight models. The authors conclude that current systems are insufficiently adapted to the MENA context and call for culturally grounded, dialect-sensitive, and clinically supervised approaches.
- AI policy
- Workforce
Research
Human-like conversational agents as social partners: a scoping review of socioaffective mechanisms, well-being outcomes, risks and governance in the post-Turing era
Qian Li, Han Geng, Xin Hu et al.
Frontiers in Artificial Intelligence · 2026-07-15
This scoping review synthesizes 58 sources on how large-language-model-based companion agents produce human-like social behavior and what psychosocial consequences follow. Therapeutic chatbot studies showed the strongest evidence for short-term symptom reduction, while evidence for sustained loneliness reduction in open-domain companions remained preliminary. The review identifies risks including dependency, displacement of human relationships, sycophancy, manipulation, and harms to vulnerable users, and proposes a 'relational safety stack' as an evaluation and governance agenda for managing psychosocial impact as these systems scale.
- AI policy
- Quality assurance
Research
Human Capital Readiness for Cloud–AI–Data Center Ecosystems: A Bibliometric Review
Rinaldi Noor, Agus Rahayu, Ratih Hurriyati et al.
Human Resources Management and Services · 2026-07-15
This bibliometric review examines 783 Scopus-indexed journal articles (2010–2025) to map the intellectual landscape of human capital readiness for cloud, AI, and data center ecosystems. The study finds that publication activity accelerated after 2019 and again after 2023, reflecting growing scholarly attention to digital workforce readiness. Analysis reveals a persistent divide between technology-oriented research (AI, cloud computing, Industry 4.0) and human-centered research (digital skills, competencies, workforce readiness), and identifies gaps in Asia-Pacific scholarship despite the region's expanding digital infrastructure role. The authors argue for repositioning human capital readiness as a strategic HR capability integrating digital talent development, HR analytics, and workforce planning.
- Workforce
- Enterprise
Research
Algorithmic Management and Employee Outcomes: Dual Effects of Control and Capability on Autonomy and Career Security
Stefan Milojević, Miroslav Knežević, Aleksandra Vujko
Administrative Sciences · 2026-07-15
This study of 441 hotel employees in Bavaria, Germany finds that algorithmic management has two distinct effects on workers: algorithmic control reduces perceived autonomy, while AI-supported skill development increases both autonomy and career security. Using structural equation modeling and mediation analysis, the research shows that the control-oriented and developmental dimensions of AI have meaningfully different consequences for employee work experiences. The findings highlight the importance of distinguishing between these two dimensions when designing or studying AI-enabled management systems in labor-intensive service organizations.
- Workforce
Research
Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)
Olivier Martinez
arXiv (Cornell University) · 2026-07-15
This critical survey reviews 45 studies on Generative Engine Optimization (GEO)—the practice of making content more likely to be cited or used by AI-powered generative search engines. The authors find that while some content-level tactics (such as topical relevance and position in context) reliably affect citation within a fixed retrieval setting, no reviewed technique demonstrates a stable, long-term causal effect on organic discoverability or user behavior across platforms. The survey identifies key methodological weaknesses across the field, including inconsistent terminology, low reproducibility, run-to-run variability, and fidelity gaps, and proposes a formal multistage model and evidence standards to guide future work. The findings matter for enterprises and content creators seeking to maintain visibility in AI-mediated information environments, as the evidence base for most GEO tactics remains narrow and context-dependent.
- Enterprise
- Quality assurance