arXiv:2607.23813cs.CLcs.AI2026-07

构建500小时金融财报语音评测基准,支持细粒度评估。

Earnings25: A Comprehensive 500-Hour Speech Benchmark for Finance

  • 基于真实财报录音构建双测试集,覆盖全时长与分段样本。
  • 提供带角色标签的对齐转录,支持按说话人/行业评估。
  • 适合金融领域ASR研究者,可比对模型在真实场景表现。

我们提出Earnings25,一个面向金融领域英语财报电话会议的自动语音识别(ASR)评测基准,可在真实条件下进行评估。该基准包含两个互补测试集:(i) testset-full,498小时的S&P 500公司2025年第四季度完整财报通话;(ii) testset-segmented,从2025年美国财报通话中采样得到的290个片段,总计46小时,实现行业平衡。基准提供对齐转录和结构化元数据,包括说话人角色、行业标签和通话结构,支持超越整体词错误率(WER)的说话人与行业感知评估。报告了Whisper与Parakeet-TDT在标准化评分下的可复现基线。

原文摘要 · Abstract (English)

We introduce Earnings25, a finance-domain benchmark for evaluating automatic speech recognition (ASR) on English-language earnings calls under realistic conditions. Earnings25 comprises two complementary test sets: (i) testset-full, 498 hours of full English-language S&P 500 earnings calls from Q4 2025, and (ii) testset-segmented, a 46-hour industry-balanced set of 290 segments sampled from English-language U.S. earnings calls in 2025. The benchmark provides aligned transcripts and structured metadata, including speaker roles, industry labels, and call structure, enabling speaker- and industry-aware evaluation beyond aggregate word error rate (WER). We report reproducible baselines for Whisper and Parakeet-TDT using standardized scoring.

语音识别金融AI评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。