News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
AI-Enabled Energy Management Systems Designed for Audit Repeatability and Regulatory Verification
Vishnu Vardhan Reddy Kavuluri
Journal of Wireless Mobile Networks Ubiquitous Computing and Dependable Applications · 2026-06-30
This paper presents an architectural framework for AI-based energy management systems designed to satisfy regulatory audit requirements. Key mechanisms include data pipeline versioning, model ephemerality, and deterministic execution environments, implemented using DVC and ONNX in a discrete-event simulation. The approach achieved 100% decision replay identity across multiple assessment rounds and maintained full historical data integrity through immutable audit trails. The findings suggest that prioritizing deterministic inference over real-time learning enables AI energy systems to meet regulatory compliance standards.
- Quality assurance
- Certifications
- AI policy
Research
What the AI Race Has Given Us and What It Requires Next
Xufeng Zhang
Open MIND · 2026-06-30
This article examines how the competitive dynamics driving AI development produce both public benefits—such as rapid capability gains, wider experimentation, and diffusion of technical knowledge—and systemic risks including opacity, market concentration, misuse, ecological burdens, and inherited model failures. The authors argue that the core problem is not competition itself but competition without adequate institutional steering, and they develop a normative framework for converting competition-driven outputs into durable public goods. Proposed mechanisms include shared evaluation infrastructure, staged openness, lifecycle governance, public participation, distributive accountability, and international coordination, all embedded in incentive-changing tools such as regulation, procurement, auditing, compute oversight, and reciprocal assurance. The paper is broadly relevant to AI governance, enterprise deployment risks, workforce and societal impacts, and certification and audit practices.
- AI policy
- Enterprise
- Quality assurance
- Certifications
- Workforce
Research
Penetration of artificial intelligence into the software development life cycle: An empirical labor market analysis
O. V. Stoyanova, Ivan S. Okuskov
Business Informatics · 2026-06-30
This empirical study analyzes 182,447 IT job postings to measure how AI competency requirements have penetrated different phases of the Software Development Life Cycle (SDLC). Using weighted clustering aligned to industry standards (SWEBOK, BABOK, ISO/IEC/IEEE 12207), the researchers find that AI competency penetration varies across eight functional clusters, ranging from 0.34% in support roles to 1.70% in design roles, with early SDLC phases like requirements analysis lagging behind design, development, and architecture clusters. The study confirms that AI skills are increasingly demanded across all SDLC phases, with management roles slightly above the market average penetration rate. The findings have direct implications for workforce training, educational curriculum updates, and revision of professional standards in an era of generative AI proliferation.
- Workforce
- Enterprise
- Certifications
- AI policy
Research
Artificial Intelligence, Firm Heterogeneity, and Labor Market Adjustment: Evidence from Service and Industrial Tech Sectors in Italy
Ebrima K. Ceesay, Mamadou Salieu Jallow, Cosimo Magazzino et al.
Muslim Business and Economic Review · 2026-06-30
This study examines how AI patent stock affects employment, wages, productivity, and labor cost shares across Italian firms from 2005 to 2024, using fixed-effects regressions on firm-level panel data. Results reveal a dual pattern: pooled models show positive effects of AI-related innovation on labor market outcomes, but firm-level analysis finds divergent impacts—UniCredit (banking) sees reduced employment and labor cost shares with rising productivity, while Zerynth (industrial tech) experiences consistently negative effects on employment, wages, and labor cost shares. The findings point to strong task-level substitution and heterogeneous firm-level outcomes, underscoring the need for targeted reskilling and labor-augmenting innovation policies during Italy's digital transition.
- Workforce
- Enterprise
- AI policy
Research
Not All Micro, Small, and Medium Enterprises Are Equal: A UTAUT-Based Multi-Group Analysis of AI Marketing Tool Adoption in the Philippines
Nelson Guillen
British Journal of Business Sciences · 2026-06-30
This study examines whether AI marketing tool adoption drivers differ across micro, small, and medium enterprises (MSMEs) in the Philippines, finding that firm size and industry sector meaningfully moderate how UTAUT factors influence adoption. Surveying 387 MSME owner-managers in Metro Manila, the research shows that performance expectancy matters more for medium enterprises while social influence is more influential for micro enterprises, and that service-sector firms are more sensitive to ease-of-use than product-based firms. The findings challenge the common practice of treating MSMEs as a uniform group and argue for segmented, size- and industry-specific approaches to supporting AI adoption in small business contexts. This has direct implications for enterprise policy and support programs targeting AI uptake among Philippine MSMEs.
- Enterprise
- AI policy
- Workforce
Research
Prompting GPT-5 on Scrum Certification Questions: An Empirical Accuracy Study
Mirko Perkusich, Danyllo Albuquerque, João Paiva et al.
arXiv · 2026-06-29
This paper empirically tests how different prompting strategies affect GPT-5's accuracy on 993 Scrum certification-style questions aligned to the Professional Scrum Master (PSM) standard. All three techniques — zero-shot, chain-of-thought, and citation-based — exceeded 85% accuracy, with the citation-based approach performing best at 89.1%. Errors clustered around misalignment with the Scrum Guide, out-of-scope content, and outdated interpretations, with multi-select and interpretive questions proving least reliable. The findings suggest that prompt engineering can modestly but consistently improve LLM reliability for Agile learning and certification preparation.
- Certifications
Research
How Human Feedback Shapes AI-generated Community Notes
Soham De, Isaac Slaughter, Jiawei Guo et al.
arXiv · 2026-06-29
This paper analyzes X's Collaborative Notes system, in which an LLM drafts fact-checking notes that human contributors then iteratively refine, studying a corpus of 19,146 collaborative notes and 211,850 instances of human feedback. The authors find that factual corrections and additional context from humans are most likely to be incorporated into AI drafts, while subjective policy judgments rarely are, and that human feedback — especially challenges to a note's main claim from more active contributors — measurably improves helpfulness. However, collaborative notes reach 'helpful' status and are displayed on the platform at lower rates than human-only or AI-only notes, with limited human participation as the key bottleneck. Rather than replacing other note types, collaborative notes tend to fill a complementary role by targeting posts that neither humans nor AI alone have addressed.
- Quality assurance
- AI policy
Research
Training Therapeutic Judges and Multi-Agent Systems for Human-Aligned Mental Health Support
Mizanur Rahman, Abeer Badawi, Elahe Rahimi et al.
arXiv · 2026-06-29
This paper presents TheraAlign, a two-stage AI framework for improving the therapeutic quality of large language model responses in mental health support. Stage I introduces TheraJudge, an open-source evaluator trained on human-annotated data that achieves high agreement with clinician ratings (ICC = 0.87–0.95) across seven psychological dimensions including Safety, Relevance, and Empathy. Stage II introduces TheraAgent, a multi-agent system with Critic, Coach, and Therapist roles that uses TheraJudge's evaluations to refine responses, yielding a +0.43 improvement in human-rated therapeutic quality on a 5-point scale and a 94% recovery rate for low-quality outputs. The findings suggest that acting on human-aligned evaluation signals—rather than simply scaling up generation—is key to safer and more effective mental health AI.
- Quality assurance
- AI policy
Research
SPINE: Bridging the Cyber-Physical Gap with Agentic AI
Minkyu Ham, Dongho Kim, Chan Lee et al.
arXiv · 2026-06-29
SPINE is an agentic AI framework designed to close the 'deployment gap' in embodied robotics — the difficult, expert-dependent process of calibrating and debugging physical robot platforms. The system uses two orchestrated multi-agent workflows (a profile builder and a debugger) to systematically bring bimanual robots to operational status with minimal robotics expertise. Tested across two distinct bimanual robot platforms (DOBOT X-Trainer and AgileX PiPER), SPINE improved operationalization success from 75% to 100% and reduced mean time-to-teleoperation from 16 min 45 s to 13 min 47 s compared to unstructured use of Claude Code, while matching or exceeding expert-level bug resolution. These results suggest agentic AI can meaningfully reduce the need for specialized human expertise in deploying physical robot systems at scale.
- Workforce
- Enterprise
Research
Less Deliberate in Teams: Student LLM Use Across Individual and Collaborative Work
Sehrish Basir Nizamani, Zannah Ziew, Saad Nizamani et al.
arXiv · 2026-06-29
This semester-long study of 96 undergraduate students in computing courses tracked how LLM usage changed across individual homework and team project milestones. LLM usage dropped by 42.7 percentage points when students transitioned from individual to team work, and students also wrote fewer and simpler prompts, used fewer intentional prompting strategies, and checked AI-generated output less carefully — including a 19.4 percentage-point drop in running tests on AI-generated code during team assignments. A within-student analysis found that 18.9% of consistent individual LLM users stopped using them entirely in teams, while only 3.2% moved the other direction. The findings suggest that collaborative settings are associated with reduced deliberate and quality-conscious AI engagement, pointing to a gap in how computing courses currently support students at the transition to team work.
- Workforce
- Quality assurance
Research
Using AI Agents to Automate Black-Box Audits of Personalization Algorithms at Scale
Alessandro Morosini, Sarah H. Cen, Andrew Ilyas et al.
arXiv · 2026-06-29
This paper introduces a framework that uses generative AI agents as synthetic users to conduct scalable, black-box audits of personalization algorithms on online platforms. Each agent is assigned a fixed persona grounded in demographic and political survey data, allowing auditors to experimentally vary platform-visible signals (such as age, gender, or location) independently of behavior—enabling counterfactual causal analysis that traditional audit methods cannot achieve. In a case study deploying 1,120 agents on X after the 2024 U.S. election, the researchers found that X's algorithmic feed amplifies toxic, polarizing, political, and right-leaning content relative to the chronological feed, with amplification varying by user ideology, while demographic signals affect content delivery differently across subgroups. The work establishes generative AI agents as a practical new tool for algorithmic auditing at scale.
- AI policy
- Quality assurance
Research
Security--Fidelity Tradeoffs: The Hidden Cost of Prompt Injection Defense
Mitchell Hermon, Rahul Gupta, Weitong Ruan et al.
arXiv · 2026-06-29
This paper reveals a fundamental tradeoff in defending large language models against indirect prompt injection attacks: defenses that resist injected instructions tend to do so by suppressing untrusted text, which degrades performance on tasks—like translation or document editing—that require faithfully processing that text. The authors introduce SecFid, a benchmark of 1,168 examples across 48 configurations designed to distinguish between a model executing an injection, processing it as data, or ignoring it, making 'fidelity' a measurable quantity alongside security. Results show no model or defense achieves both goals simultaneously: the highest-fidelity model reaches 96.5% fidelity but only 47.8% security, while the most secure defenses achieve 99.3% security but only 71.0%–73.9% fidelity. The paper argues that reporting security alone hides the real cost of these defenses, and that the correct balance depends on deployment-specific costs of a hijack versus a dropped span.
- Quality assurance
- Enterprise
Research
A Single Rewrite Suffices: Empirical Lessons from Production Skill Description Optimization
Yangqiaoyu Zhou, Mohammad Alqudah, Kwei-Herng Lai et al.
arXiv · 2026-06-29
This paper tackles 'skill collision' in enterprise AI agents, where overlapping natural-language skill descriptions cause routing errors as agents scale to many specialized skills. The authors deploy an automated description optimization pipeline on a production group chat agent with 9 skills and 372 regression cases, achieving 79.2% F1—matching manually tuned descriptions at 79.4% F1—while cutting per-skill engineering effort from 120 minutes to 3.8 minutes (a 32x speedup). Ablation studies on both the production system and ToolBench reveal that a single LLM rewrite using available false-positive and false-negative cases captures most of the improvement, with other design choices each contributing less than 0.5% F1 difference. The authors also identify a diagnostic signal (a large train-validation F1 gap) to flag cases requiring architectural rather than text-level fixes.
- Enterprise
- Workforce
Research
Understanding and Evaluating Claw-like Agent Security Through a Computer-Systems Lens
Peizhi Niu, Wenjie Qu, Shangding Gu et al.
arXiv · 2026-06-29
This paper examines security vulnerabilities in always-on, system-level AI agents (termed 'Claw-like' agents, e.g., OpenClaw) that have persistent access to credentials, files, tools, and external services. The authors develop SafeClawArena, a benchmark of 406 adversarial tasks spanning four attack surfaces—Skill Supply-Chain Integrity, Persistent State Exploitation, Cross-Boundary Data Flow, and Indirect Prompt Injection—evaluated across three agent platforms and five frontier LLMs. Key findings show attack success rates as high as 70%, with malicious Plugins succeeding in 100% of cases regardless of the underlying LLM, while the most hardened platform (SeClaw) reduces GPT-5.4's attack success rate from 70% to 22%, partly through utility-security tradeoffs. These results reveal that current defenses are inadequate for agentic systems operating with OS-level privileges and point to the urgent need for stronger security mechanisms analogous to decades of classical cybersecurity research.
- Quality assurance
- Certifications
Research
Words Speak Louder Than Code: Investigating Cognitive Heuristics in LLM-Based Code Vulnerability Detection
Asif Shahriar, Hongyu Cai, Hadjer Benkraouda et al.
arXiv · 2026-06-29
This paper presents the first systematic study of how cognitive heuristics—halo, framing, and anchoring effects—bias LLM-based code vulnerability detection. By holding code fixed and varying only surrounding context, the researchers show that all eight evaluated LLMs are susceptible, with average susceptibility rates of 33.2% for framing, 23.5% for anchoring, and 18.4% for halo effects across three programming languages. Vulnerabilities requiring semantic reasoning are more susceptible than those detectable by pattern matching, and models frequently flip verdicts without correctly identifying the actual vulnerability. A proof-of-concept black-box cognitive attack demonstrates that up to 97% of previously detected vulnerabilities can be suppressed, highlighting a serious and exploitable reliability gap in AI-driven security tooling.
- Quality assurance
- Enterprise
Research
AI Premium
Nicola Borri, Yukun Liu, Aleh Tsyvinski
arXiv · 2026-06-29
Using 380 trillion tokens of AI consumption data from over 400 large language models on the OpenRouter platform (roughly 2% of global monthly AI token consumption), this paper constructs an 'AI Factor' from growth in tokens, dollars, and users, and estimates firm-level 'AI Betas' from stock return comovement to characterize an 'AI Premium.' The study finds that firms whose returns covary more with AI activity earn higher subsequent returns—a value-weighted long-short strategy yields 64.1 basis points per week—with the premium concentrated in intensive, frontier-oriented AI use (closed-source models, paying users, long prompts) and extending into consumer-facing and capital-heavy sectors, but absent in emerging markets including China. Crucially for workforce impacts, occupations one standard deviation higher in interaction-and-communication content have 0.36-standard-deviation higher market-implied AI exposure, while analytical, scientific, and operations-control skills show negative exposure, suggesting AI is reshaping labor market valuations unevenly across job types.
- Workforce
- Enterprise
Research
Entity Binding Failures in Tool-Augmented Agents
Rahul Suresh Babu, Shashank Indukuri
arXiv · 2026-06-29
This paper identifies and studies 'entity binding failures' in tool-augmented AI agents — cases where an agent selects the correct tool but acts on the wrong real-world entity (e.g., emailing the wrong person or updating the wrong account). Across a controlled evaluation of 60 tasks, five model backends, and six tool-use methods, all methods achieved 0% wrong-tool error, yet action-oriented baselines still produced wrong-entity actions in 24–26% of runs. Entity-aware mechanisms (such as confidence-gated binding and clarification under ambiguity) eliminated wrong-entity actions but reduced direct task completion by deferring when uncertain. The findings demonstrate that safe, reliable tool use in enterprise workflows requires not just correct tool selection but also accurate binding of natural-language references to the correct real-world entities.
- Enterprise
- Quality assurance
Research
SIMAX: A Scalable and Interpretable Framework for Multi-Fidelity and Annotated Clinician-Patient Dialogue Simulation
Zhuhan Bao, Rui Yang, Bohao Yang et al.
arXiv · 2026-06-29
SIMAX is a framework that automatically generates large-scale, annotated simulated clinician-patient dialogues to support the development and evaluation of AI-driven clinical communication coding systems. The framework produces controlled dialogues across specialties, personas, and accent conditions, with reference behavioral labels drawn from two codebooks covering overall communication quality and specific countable behaviors. Evaluated on 3,388 simulated dialogues, SIMAX demonstrated reasonable speech naturalness, high transcription fidelity, and acceptable clinical realism in human review, and it was able to expose insufficient sensitivity in a downstream communication coding system. By providing reproducible, labeled training and validation data, SIMAX addresses the cost and scalability bottlenecks that currently limit quality assessment of AI ambient scribes and clinical communication tools.
- Quality assurance
Research
Can LLMs Rank? A Tale of Triads and Triage
Gaurab Pokharel, Shafkat Farabi, Patrick J. Fowler et al.
arXiv · 2026-06-29
This paper examines whether large language models (LLMs) can reliably rank people for high-stakes decisions such as homelessness service allocation and emergency department triage, where scarce resources must be prioritized. The authors propose using two complementary consistency measures—the coefficient of consistency (ζ), which counts circular triads in pairwise comparison tournaments, and inter-run ranking distance metrics like Kendall's τ—to evaluate LLM reliability before committing to a ranking. Testing three leading LLMs across these two prioritization tasks, they find that the models show considerably different performance profiles across the two consistency dimensions, meaning neither measure alone is sufficient. The paper offers practical guidelines for practitioners on how to assess LLM consistency before deploying such systems in consequential ranking contexts.
- AI policy
- Quality assurance
Research
Whose Side Is Your Agent On? Multi-Party Principal Loyalty in LLM Agents
Bojie Li, Noah Shi
arXiv · 2026-06-29
This paper examines a largely overlooked problem in LLM agents that operate in multi-party settings—where an agent must stay loyal to the principal who hired or briefed it while also interacting with a counterparty whose interests may conflict, such as in vendor negotiations or employee mediation. The authors introduce PrincipalBench, a 75-item multi-turn benchmark, and evaluate 13 frontier models, finding a sharp divide: some models selectively refuse adversarial probes while following legitimate principal requests (≤20% harm), while others over-refuse broadly (53.6–75.3% harm), a distinction invisible to standard single-turn safety evaluations. Two mitigation mechanisms are proposed and tested: a prompt-time loyalty scaffold (seven prioritized rules derived from 50+ failure cases) and a per-token-KL distillation recipe for transferring behavior from a large teacher model to smaller open-weight student models. A key structural finding is that both mechanisms only trade off along a leak/over-refusal axis rather than jointly improving both, suggesting a fundamental tension that neither approach resolves.
- Quality assurance
- Enterprise
Research
EMPATH: A Multilingual Auditor-Judge Benchmark for Safety Evaluation of Emotional-Support Chatbots
Camilo Chacón Sartori
arXiv · 2026-06-29
EMPATH is a multilingual, multi-turn benchmark designed to evaluate the safety of emotional-support chatbots, targeting gaps that fixed-prompt, single-language benchmarks miss. It uses an auditor model to simulate help-seeking users across 140 seed instructions and 34 personas, while a judge model scores full transcripts on 19 metrics spanning crisis handling, therapeutic quality, conversational integrity, emotional safety, and cultural adaptation—currently implemented in Mexican Spanish and US English. The study finds that aggregate scores across three frontier models differ by less than 0.74 points, but per-metric profiles diverge by up to six points, and run-to-run reliability varies dramatically—one model swings 2 to 10 points on a crisis metric across identical runs—establishing reproducibility as a distinct, per-model safety property. These findings matter for quality-assurance and certification efforts around AI systems deployed in sensitive mental-health contexts, as they reveal that standard rubrics inflate scores and that reliability must be measured independently of average performance.
- Quality assurance
- Certifications
Research
CaresAI at CT-DEB26: Detecting Dosing Errors In Clinical Trials Using Domain-Specific Transformer Embeddings and Classification Models
Leon Hamnett, Favour Igwezeke, Joseph Itopa Abubakar et al.
arXiv · 2026-06-29
This study applies domain-specific biomedical transformer models—including BioBERT, ClinicalBERT, PubMedBERT, and MedCPT—to automatically detect dosing errors in clinical trial protocols. By combining these text embeddings with categorical features and feeding them into classical machine learning and neural network classifiers, the system achieved ROC-AUC scores between 0.821 and 0.853, with BioBERT outperforming alternatives under a logistic regression baseline at 0.794. The findings show that domain alignment of the language model matters more than combining multiple embeddings, and that automated dosing error detection can support safety monitoring and regulatory decision-making in clinical trials.
- Quality assurance
- AI policy
Research
EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures
Buğra Alperen Uluırmak, Rifat Kurban
arXiv · 2026-06-29
EvalSafetyGap introduces a conceptual framework and hybrid survey to address a core measurement problem in large language model (LLM) evaluation: benchmark scores and safety metrics can improve on paper while the underlying properties they are meant to capture remain unverified. The paper synthesizes evidence across eight streams—including benchmark validity, reward hacking, jailbreak robustness, and governance—spanning work from 2018 to 2026, and organizes findings around 'Goodhart's Law' dynamics using two new constructs: an Instability Decomposition and an Alignment Trilemma. A structured ten-model audit finds that the association between capability and adversarial robustness is statistically indeterminate (Pearson r = +0.232, p = 0.520), and that the apparent safety gap between open and closed models is driven mainly by governance and disclosure practices rather than behavioral robustness. The framework offers a shared vocabulary and evidence map to support more transparent, auditable AI evaluation and alignment practices.
- Quality assurance
- AI policy
Research
Not-quite-human tastes: the stylized omnivorousness of LLM survey surrogates
Xiangyu Ma, Mengmi Zhang, Shannon Ang et al.
arXiv · 2026-06-29
This study tests whether large language models (LLMs) from OpenAI, Anthropic, and DeepSeek can serve as stand-ins for human survey respondents in the domain of cultural taste. The researchers generated over 277,000 synthetic survey responses mimicking participants from the Survey of Public Participation in the Arts and found that LLM surrogates systematically overestimate cultural engagement, fail to reproduce the complex relational structure of real tastes, and distort known associations between taste and social categories like age, class, gender, and race. The findings raise serious concerns about the validity of 'synthetic' survey panels already being marketed by market research companies, as well as the risk of LLM-generated responses contaminating conventional survey data.
- AI policy
- Enterprise
Research
SpreadsheetBench 2: Evaluating Agents on End-to-End Business Spreadsheet Workflows
Jian Zhu, Yuzheng Zhang, Zeyao Ma et al.
arXiv · 2026-06-29
SpreadsheetBench 2 is a benchmark that evaluates AI agents on realistic, end-to-end business spreadsheet workflows — covering task generation, debugging, and visualization — built from authentic financial reports and corporate filings and validated by domain experts. The 321 tasks average 11.8 worksheets and require roughly 594 cell modifications each, far exceeding the scope of prior benchmarks that test only isolated operations. Evaluating eight frontier large language models, the best system achieves only 34.89% overall task accuracy, with debugging accuracy as low as 12%, revealing that current AI remains unreliable for real-world spreadsheet automation. Failure analysis identifies insufficient spreadsheet inspection and incorrect target-cell selection as the dominant bottlenecks, pointing to concrete directions for improvement.
- Enterprise
- Quality assurance