News & Research
The latest AI research and news with real-world stakes — each item sourced, dated, summarized in plain English, and tagged by impact area. Every item is checked against its source before it appears.
Research
AWARE-FX: An Auditable Knowledge-Guided AI System for Measuring Corporate Foreign-Exchange Hedging Disclosure
Qi Wang
arXiv · 2026-07-30
AWARE-FX is an auditable AI/NLP system that extracts and scores corporate foreign-exchange hedging disclosures from annual reports. Built on a professional-source lexicon, negation and accounting-status logic, domain-specific financial encoders, and an audit ledger, it processes 543,527 snippets across 24,909 Hong Kong firm-years from 2008–2025. FinBERT achieves higher mean F1 in seven of eight encoder comparisons (temporal F1 ranging from 0.702 to 0.872), and abstaining on the 20% least-confident observations raises F1 by 0.050–0.077; a general-purpose LLM (Qwen3-8B) performs unevenly across label types, showing domain constraints remain necessary. The system's strict FX score is negatively associated with FX exposure in baseline and stress periods, providing external construct validation, while its modular architecture keeps retrieval, classification, uncertainty handling, and aggregation separately auditable.
- Quality assurance
- Enterprise
Research
Is Solving Better Than Evaluating GenAI Solutions?
Ethan Dickey, Marios Mertzanidis, Alexandros Psomas
arXiv (Cornell University) · 2026-07-30
This randomized A/B crossover study (N=220) in a junior-level algorithms course compared having students evaluate flawed GenAI-generated solutions against traditional problem solving across six assignments. The study found no statistically significant differences in midterm scores, final exam scores, or overall course grades between the two approaches, though students earned higher homework scores when evaluating GenAI solutions. The advantage from evaluation tasks did not transfer to summative assessments, and most students reported no change in study habits—though those who did adapt their strategies found the GenAI-evaluation assignments more helpful. The authors conclude that GenAI-evaluation activities can be introduced without broad performance losses, but meaningful learning gains likely require deliberate scaffolding beyond simple error diagnosis.
- Quality assurance
Research
Evaluating Agentic Bioinformatics through Function, Evidence, and Validation
Phuc Pham, Truong-Son Hy
arXiv (Cornell University) · 2026-07-30
This paper introduces the Function–Evidence–Validation (FEV) framework for evaluating AI agents that plan and execute biological data analyses. Rather than judging agents solely on final outputs or benchmark scores, FEV examines the full inspectable workflow trajectory for demonstrated operations, traceable evidence, and use-case-specific validation. Applying FEV to 109 agentic systems and 28 benchmark resources across genomics, proteomics, drug discovery, and other bioinformatics domains, the authors find that planning and tool execution have advanced faster than replayability, provenance tracking, and prospective empirical testing. The work argues that agentic bioinformatics must be assessed on workflow correctness rather than final-answer correctness to achieve scientific accountability and auditability.
- Quality assurance
Research
Using Large Language Models for Idea Generation in Innovation
Lennart Meincke, Karan Girotra, Gideon Nave et al.
arXiv (Cornell University) · 2026-07-30
This study compares new product ideas generated by GPT-4 (using zero-shot and few-shot prompting) against ideas produced by university students in a product design course. AI-generated ideas outperformed human ideas on average purchase intent, and were seven times more likely to rank in the top 10% of all ideas — a figure the authors call a conservative estimate given AI's higher productivity. However, AI ideas were rated as less novel and showed greater pairwise similarity, especially with few-shot prompting, indicating a less diverse solution space. The findings suggest LLMs offer a substantial advantage for high-quality idea generation in new product development, with some trade-offs around diversity and novelty.
- Enterprise
Research
WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization
Fanzhe Wei, Li Liu
arXiv (Cornell University) · 2026-07-30
WitCert introduces a provably sound runtime monitor for KV-cache quantization in large language model serving, providing per-layer, per-head, per-step upper bounds on the quality degradation caused by cache compression. The system operates at two tiers—a deterministic worst-case bound valid for any quantizer and a tighter probabilistic certificate for a specific INT8 scheme with core theorems machine-checked in Lean 4. Integrated into the SGLang serving framework, meter-driven gating empirically restores output quality floors; for example, raw-cast FP8 recovers from 22.8 to 79.7 on hard RULER tasks, and the certified INT8 cache achieves 1.88× more KV tokens at the same memory. This matters because it shifts KV quantization validation from offline benchmark averages to live, per-request risk observability and automated repair.
- Quality assurance
- Certifications
Research
Cloned Voices, Real Consequences: Evaluating Bias in Political Deepfake Detection for Electoral Integrity in Brazil
Lucas Rafael Stefanel Gris, Daniel Casanova, Frederico Santos de Oliveira et al.
arXiv (Cornell University) · 2026-07-30
This paper introduces ParlaSpoof-BR, an audio deepfake dataset built from Brazilian Chamber of Deputies recordings and augmented with synthetic speech from text-to-speech and voice conversion models. The authors benchmark state-of-the-art deepfake detectors on this dataset, finding that current systems struggle to generalize consistently to Brazilian Portuguese political speech, with methodological factors like synthesis model choice dominating over demographic disparities. The work highlights the risks AI-generated disinformation poses to electoral integrity and provides a domain-specific benchmark for developing more robust detection tools in an underrepresented political context.
- AI policy
- Quality assurance
Research
To Add Is Machine, To Delete Is Human: Measuring and Mitigating Deletion Avoidance in LLM Code Editing
Amir M. Ebrahimi, Mohammed Mehedi Hasan, Aaditya Bhatia et al.
arXiv (Cornell University) · 2026-07-30
This paper identifies 'deletion avoidance' in large language models—a systematic tendency to retain code that should be removed during editing—and quantifies it across leading models on SWE-bench Verified. Across five top models, deletion recall against developer patches reaches at most 71.7%, and models locate the right file over 92% of the time but cut the exact required line in under 52% of cases. The authors introduce a new benchmark (CanItDelete) of 200 real-commit tasks requiring only deletions, and show that retrofitting existing tests to check for removal drops frontier model pass rates from 63.2% to 41.9%. Their findings suggest deletion avoidance is an undertrained behavior rather than a fundamental limitation, with post-training on deletion tasks improving both deletion and broader code-editing performance.
- Quality assurance
- Enterprise
Research
From Process to Evidence: How Computing Can Ground Appropriate Reliance on Legal AI
James Williams
arXiv (Cornell University) · 2026-07-30
This paper examines how courts—specifically the New York court system—are responding to over 1,500 cases involving AI hallucinations in legal filings, finding that official guidance repeatedly calls for evidence (such as error rates and do-not-use lists) that does not yet exist. Instead, courts substitute procedural requirements like training mandates and checklists, placing the greatest burden on those least equipped, such as legal aid programs and self-represented litigants. The authors map legal duties onto human-computer interaction concepts of 'appropriate reliance' and argue the computing community must develop task taxonomies, shared error metrics, benchmarks, and test harnesses to supply the empirical grounding the justice system currently lacks. Without such evidence, oversight of legal AI risks widening rather than closing the justice gap.
- AI policy
- Quality assurance
Research
Effect of evaluation prompt strategies on LLM-as-a-judge reliability in critical care
Jia-Yu Yan, Wing-Sum Chan, Ching‐Tang Chiu et al.
Anaesthesiology Intensive Therapy · 2026-07-30
This study tested whether the way evaluation prompts are structured affects how reliably large language models can score AI-generated critical care reports compared to human clinicians. Using 90 ICU clinical reports evaluated under three prompt strategies with GPT-4o and o3-mini, the researchers found that a bottom-up incremental prompting approach produced the closest alignment with clinician ratings (ICC = 0.94, mean deviation 0.1), while a top-down decremental strategy showed significantly lower agreement (ICC = 0.82, mean deviation 4.7). The findings demonstrate that prompt design meaningfully influences both scoring patterns and human-AI concordance, highlighting the need for standardized prompt architectures before LLM-as-a-judge systems can be reliably used in clinical AI evaluation.
- Quality assurance
Research
Three futures for the diagnostic radiologist: A structured disagreement about what AI actually changes
Jan Beger, Amine Korchi, Christoph A. Agten
European Journal of Radiology Artificial Intelligence · 2026-07-30
This paper presents three independently authored 2035 job descriptions for diagnostic radiologists, written by two radiologists and a health IT professional to capture optimistic, trade-off, and stratification perspectives on how imaging AI will reshape the profession. All three scenarios agree that routine workloads will be AI-managed and radiologists will bear accountability for AI output, but they diverge sharply on headcount, career security, and whether the profession expands, concentrates, or stratifies into differentiated tiers. The paper concludes that AI will not eliminate diagnostic radiologists, but that workforce outcomes depend on health system decisions not yet made. It is directly relevant to workforce planning in radiology, highlighting that both optimism and economic caution are simultaneously defensible.
- Workforce
- AI policy
Research
FORENSIC ACCOUNTING TECHNOLOGIES AND OCCUPATIONAL FRAUD MITIGATION IN THE NIGERIAN MARITIME SECTOR
Malik Sidique Muhammad, Ajape Mohammed Kayode
Journal of Business Management and Accounting · 2026-07-30
This study finds that forensic accounting technologies—including AI-driven analytics, blockchain traceability, cybersecurity resilience, regulatory enforcement, and organizational technology readiness—each have statistically significant positive effects on occupational fraud detection in Nigeria's maritime sector. AI-driven forensic analytics was the strongest predictor (β = 0.391, p < 0.001), and together the five factors explained 63% of the variance in fraud detection effectiveness across 100 surveyed maritime organizations. The findings support integrated adoption of these technologies to strengthen transparency, accountability, and governance in the sector.
- Enterprise
- Quality assurance
Research
Occupational Convergence or Divergence? Mapping Labor Market Structural Shifts Driven by AI Penetration
Rafiazka Hilman, Julia Koltai
arXiv (Cornell University) · 2026-07-30
Using large-scale job vacancy data from ten countries and a combination of NLP, large language models, and bipartite network analysis, this paper finds that AI-related skill demand is heavily concentrated in a narrow STEM technical core—roughly three-quarters to four-fifths of AI vacancies—centered on skills like Python, SQL, and machine learning. While AI-exposed occupations converge around these competencies, that convergence does not spread across the broader labor market, leaving most occupations largely untouched. The result is occupational stratification rather than democratization: AI reinforces existing advantages for technically skilled workers and raises entry barriers, particularly for those without prior technological backgrounds. These findings have direct implications for workforce planning, training policy, and understanding how AI shapes career pathways globally.
- Workforce
- AI policy
Research
Legal Concerns Regarding the Use of Artificial Intelligence in Tax Proceedings
Tess Veldhoven
Open MIND · 2026-07-30
This legal study examines the constitutional and ethical risks of deploying AI systems in tax proceedings, focusing on how the 'black box' phenomenon undermines fair hearing rights, effective remedy rights, and the obligation to state reasons. Drawing on the Dutch tax authority scandal and US tax authority practices, the author shows how algorithmic discrimination rooted in historical data conflicts with GDPR principles and creates unclear liability for tax authorities. The study situates these concerns within the EU AI Act's classification of public-authority profiling as high-risk and proposes safeguards including explainable AI models, mandatory impact assessments, external auditing, and the human-in-the-loop principle. The findings matter for policymakers and regulators navigating how to govern AI-driven administrative decision-making lawfully and fairly.
- AI policy
- Certifications
Research
Designing for all? Accessibility of native android interfaces from large language models
Daniel Mesquita Feijó Rabelo, Júlia Holanda Muniz, Kiev Gama et al.
Universal Access in the Information Society · 2026-07-30
This paper evaluates how large language models (LLMs) like ChatGPT and GitHub Copilot generate native Android UI code with respect to established accessibility guidelines. Across four empirical studies covering seven mobile UI types, the researchers identified 702 accessibility-related issues, finding that Jetpack Compose produced more accessible interfaces than other layout approaches, and English-language prompts led to fewer errors. Counterintuitively, prompts that explicitly requested accessibility often introduced more problems, suggesting current LLMs struggle to correctly interpret and apply accessibility directives. The findings highlight the need for better prompt engineering and more robust LLM code generation to ensure AI-assisted mobile development produces genuinely accessible software.
- Quality assurance
- Enterprise
Research
Digital Labor and Social Protection in the Platform Economy
Laura Kolar Vasudeva
Digital social sciences. · 2026-07-30
This paper examines how digital labor platforms—spanning ride-hailing, delivery, microtask, content moderation, and freelance work—create a structural mismatch with employment-based welfare systems originally designed for stable, full-time jobs. The authors trace the history of the employment welfare state, analyze algorithmic management practices, and survey legal classification battles across the US, UK, EU, and Global South, including gendered and cross-border dimensions of the resulting welfare deficit. Emerging policy responses such as portable benefits schemes, universal basic income proposals, algorithmic transparency mandates, and the EU Platform Work Directive are assessed. The paper argues that addressing platform workers' welfare deficit ultimately requires decoupling social protection from the traditional employer-employee relationship and reorganizing it around the worker as an individual.
- Workforce
- AI policy
Research
Prompting for Pragmatics: Improving the Cultural Sensitivity of LLM Translations for Business Emails
Helene Tenzer, Oumnia Abidi, Stefan Feuerriegel
Management International Review · 2026-07-30
This study examines whether large language models can produce culturally appropriate English-to-Japanese translations of workplace emails, comparing three prompting strategies: naive translation prompts, audience-targeted prompts specifying the recipient's cultural background and role, and instructional prompts providing explicit guidance on Japanese communication norms. Using both linguistic analysis and native speaker evaluations, the researchers find that naive prompts yield limited cultural adaptation, while both audience-targeted and instructional prompts significantly outperform the baseline. Notably, a threshold effect emerges: instructional prompts produce stronger textual adaptation but do not generate significantly higher appropriateness or compliance ratings than the simpler audience-targeted approach. The practical implication is that lightweight audience-targeted prompting is sufficient to meaningfully improve cultural sensitivity in LLM-mediated workplace translation.
- Workforce
- Enterprise
Research
Governance für KI-Agenten: Ein risikobasierte Perspektive für Transparenz und Nachvollziehbarkeit
Bennet Santelmann
HMD Praxis der Wirtschaftsinformatik · 2026-07-30
This paper presents a structured literature review of 33 peer-reviewed studies (2022–2026) on governance requirements for agentic AI systems—LLM-based agents that autonomously plan tasks, use tools, and intervene in business processes. It synthesizes how traceability, auditability, and accountability form an interconnected evidence-and-responsibility chain, and identifies recurring mismatches: insufficient audit-robustness of evidence, inadequate stakeholder-specific explainability, and unclear attribution in multi-agent settings. Building on these findings and the risk-based requirements of the EU AI Act, the authors develop a risk-informed governance perspective linking traceability infrastructure, governance controls, and accountability models to support the design of auditable agentic AI in enterprise contexts.
- Enterprise
- AI policy
Research
A Unified Framework for Human–AI Collaboration in Security Operations Centers with Trusted Autonomy
Ahmad Mohsin, Helge Janicke, Ahmed Ibrahim et al.
ACM Transactions on Internet Technology · 2026-07-30
This paper proposes a tiered autonomy framework for Human-AI collaboration in Security Operations Centers (SOCs), defining five levels of AI autonomy—from manual to fully autonomous—each mapped to human oversight roles and task-specific trust thresholds. The framework is validated through a simulated cyber range featuring an LLM-based AI assistant, demonstrating reduced alert fatigue and improved incident response coordination. The work addresses limitations of existing SOC automation approaches, which tend to treat autonomy as binary and lack formal structures for managing trust and human-in-the-loop decision-making. The findings are relevant to enterprise cybersecurity operations seeking to augment rather than replace human analysts with adaptive, explainable AI.
- Enterprise
- Workforce
Research
Artificial Intelligence in Educational Administration: A Systematic Evidence Review and Exploratory Meta-Analysis of Empirical Studies, 2020–2025
Matyoqubovich Sobirov
International Education Trend Issues · 2026-07-30
This systematic review and exploratory meta-analysis synthesizes empirical research (2020–2025) on AI use in educational administration, covering tasks such as automating administrative work, supporting decision-making, and optimizing resources. Across 11 studies representing over 2,400 units, only two were statistically compatible for meta-analysis, yielding a pooled Hedges' g between 0.65 and 1.00 depending on estimation method, but with substantial heterogeneity (I²=90.2%). Narrative findings suggest potential gains in processing time, reporting accuracy, and resource allocation, though most studies were cross-sectional, single-site, or perception-based. The authors conclude that AI can enhance educational administration under the right conditions but that the current evidence base is insufficient for strong causal or universal claims, calling for controlled multisite designs and rigorous governance frameworks.
- Enterprise
- AI policy
Research
AI-powered talent chain management with multi-agent systems for industry and innovation growth
Rong‐Fu Wang, Xiufen Zeng, Fuchao Li et al.
Scientific Reports · 2026-07-30
This paper proposes an AI-powered multi-agent framework for predicting an individual's next career transition using sequences of ESCO occupation codes and job-title text from career histories. The system integrates structured occupational taxonomy embeddings, temporal career dynamics encoding, and uncertainty-aware refinement to handle noisy or ambiguous records. Evaluated on two benchmark datasets (KARRIEREWEGE and DECORTE), the model outperforms the strongest baseline by 2.7–2.9 percentage points in NDCG@10 while using fewer parameters, demonstrating improved accuracy and deployment feasibility. These results are relevant to workforce planning and talent mobility analysis at both organizational and policy levels.
- Workforce
- Enterprise
Research
When one teaches, two learn: the bidirectional learning process between supervisors and PhD students
Pierre Boutros, Michele Pezzoni, Sotaro Shibayama et al.
Economics of Innovation and New Technology · 2026-07-30
This paper examines knowledge transfer between PhD supervisors and students in French STEM doctoral programs (2010–2018), focusing specifically on AI knowledge. Using data on 40,852 PhD graduates, it finds that having an AI-knowledgeable supervisor makes a student 10 percentage points more likely to write an AI thesis, while supervisors without prior AI knowledge who are exposed to AI-focused students become 15 percentage points more likely to publish AI-related work. The findings confirm that learning flows in both directions, challenging the conventional assumption that knowledge moves only from supervisor to student.
- Workforce
Research
Unpacking pre-service teachers’ career self-efficacy: impact of artificial intelligence knowledge, anxiety, and digital literacy
Chinedu Hilary Joseph, Mensah Prince Osiesi, Nomanesi Madikizela-Madiya et al.
SN Social Sciences · 2026-07-30
This study of 470 pre-service teachers in Nigeria finds that AI knowledge and digital literacy positively predict career self-efficacy, while AI anxiety has a significant negative effect. Digital literacy also mediates the relationship between AI anxiety and career self-efficacy, meaning higher digital literacy buffers the psychological harm of AI anxiety on career confidence. The authors conclude that teacher education programmes should embed structured AI education and digital literacy training in their curricula to strengthen future teachers' professional confidence.
- Workforce
Research
Embedding Children’s Rights by Design
Elemegious Mugamba
Quaderns IEE · 2026-07-30
This article examines how children's rights can be embedded into AI systems used in youth justice contexts—such as risk assessment, diversion, probation, and child protection—across Europe. Drawing on international and EU law frameworks including the UN Convention on the Rights of the Child, GDPR, the Law Enforcement Directive, and the EU AI Act, it develops a three-condition normative model requiring child-centered purpose justification, rights-embedded system design, and continuous institutional oversight. Using Ireland as a case study, the model assigns legally grounded obligations to legislators, authorities, technology providers, and courts throughout the AI lifecycle, including auditing mechanisms and enforceable consequences for non-compliance. The work matters because existing regulatory frameworks are fragmented and adult-centered, leaving children's distinctive legal status inadequately protected in high-stakes AI-driven decisions.
- AI policy
- Certifications
Research
The impact of advanced robotics on employees' performance in large-scale operations: evidence from the National Centre for Artificial Intelligence and Robotics (NCAIR), Abuja, Nigeria
Queen Ladi Patrick, I. T. Ndulue
British journal of interdisciplinary research. · 2026-07-30
This quantitative study of 157 employees at the National Centre for Artificial Intelligence and Robotics (NCAIR) in Abuja, Nigeria finds that advanced robotics deployment has strong positive effects across four workforce dimensions: robotic process automation boosts employee productivity (β=.741), human-robot collaboration improves job satisfaction (β=.683), robotic deployment aids skill development (β=.659), and robotics integration enhances operational efficiency (β=.712), with all effects statistically significant at p<.001. Grounded in the Technology-Organization-Environment framework and analyzed via regression and ANOVA, the study recommends ongoing training programs, structured human-robot collaboration plans, change management, and real-time performance monitoring tools. The findings matter because they provide empirical evidence—from a leading African AI and robotics institution—that robotics adoption can improve rather than simply displace employee performance in large-scale operations.
- Workforce
- Enterprise
Research
Artificial intelligence's use in external auditing: Evidence from systematic literature review
Jimoh Adams Lukman, Audrey H Legodi
Edelweiss Applied Science and Technology · 2026-07-30
This systematic literature review synthesizes evidence from 130 peer-reviewed articles on AI applications in external auditing, identifying three dominant research streams: machine learning and fraud detection, audit analytics and big data, and AI adoption and governance. The study finds that AI techniques such as machine learning, neural networks, and natural language processing enhance fraud detection, risk assessment, and audit quality, while raising concerns about algorithmic bias, transparency, and professional skepticism. The authors develop an integrated framework linking AI applications to audit quality and provide a research agenda for auditors, regulators, and organizations seeking to implement AI responsibly in audit processes.
- Quality assurance
- AI policy