新测试集让语音识别模型学会按用户偏好输出格式。
Preference-ASR: A Preference-Aware Test Set for Benchmarking ASR in the Era of Speech LLMs
- 用大模型辅助构建带风格指令的测试数据,模拟真实使用场景。
- 不同模型在各类偏好下排名变化,暴露传统评估忽略的差异。
- 适合研究语音大模型输出可控性或评测系统个性化能力者。
主流语音识别测试集在数字、冗余词、实体和大小写方面标准不一,而通用归一化器会抹除用户在意的格式差异。现有基准无法衡量模型是否遵循用户对输出风格的偏好。我们提出Preference-ASR,一个评估语音识别系统在四种类别(归一化、实体、冗余词、大小写)中遵循自然语言偏好指令能力的测试集。该数据集基于七个开源语料库,通过两阶段大模型辅助流程并经人工验证构建。采用偏好感知归一化器,仅跳过与当前指令匹配的处理步骤进行评估。对四个模型的基准测试显示,不同偏好类型下模型排名显著变化,揭示了传统评估掩盖的质量差异。数据集已公开发布。
原文摘要 · Abstract (English)
Popular ASR test sets adopt inconsistent conventions for numbers, disfluencies, entities, and casing, while standard normalizers erase the format distinctions users care about. Current benchmarks therefore cannot measure whether a model follows user preferences for output style. We introduce PreferenceASR, a test set evaluating ASR systems on their ability to follow natural-language preference instructions across four categories: normalization, entities, disfluencies, and case. Built from seven open-source corpora via a two-stage LLM-assisted pipeline with human verification, it is evaluated with a preference-aware normalizer that selectively skips steps matching the active instruction. Benchmarking four models shows rankings shift across preference types, exposing quality differences traditional evaluation obscures. We publicly release the dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。