arXiv:2605.00022cs.CLcs.AI2026-05ACL

用50个精选音频样例高效评估大音视频模型,兼顾性能与用户偏好。

Putting HUMANS first: Efficient LAM Evaluation with Human Preference Alignment

论文配图:Putting HUMANS first: Efficient LAM Evaluation with Human Preference Alignment
图 1 · 摘自论文原文
  • 仅用50个样本(0.3%数据)即可可靠评估大音视频模型。
  • 精选样本训练的回归模型与用户偏好相关性达0.98,优于全量数据。
  • 适合追求高效、贴近真实体验的语音模型评测者使用。

大音视频模型(LAMs)快速涌现,亟需高效模型评估方法,但全面基准测试成本高昂。本文研究最小样本集是否可有效评估LAMs,同时降低开销与数据冗余。在40项任务、18个模型上对比10种子集选择方法,发现仅50个样本(占总量0.3%)即可实现与完整基准得分超过0.93的皮尔逊相关性。为验证评估结果是否反映实际用户满意度,我们收集了776条真实语音助手对话中的人类偏好评分,发现完整基准与子集得分与人类偏好相关性仅为0.85。进一步在精选子集上训练回归模型,相关性提升至0.98,优于随机子集及全量数据训练模型。结果表明,在回归建模中优质子集胜过海量数据。我们开源该加权子集,构建HUMANS基准,作为兼顾模型性能与用户偏好的高效代理评估工具。

原文摘要 · Abstract (English)

The rapid proliferation of large audio models (LAMs) demands efficient approaches for model comparison, yet comprehensive benchmarks are costly. To fill this gap, we investigate whether minimal subsets can reliably evaluate LAMs while reducing costs and data redundancy. Analyzing 10 subset selection methods with 18 audio models across 40 tasks covering major LAM evaluation dimensions, we show that subsets of just 50 examples (0.3% of data) can achieve over 0.93 Pearson correlation with full benchmark scores. To understand how well these scores align with what practitioners ultimately care about, user satisfaction, we collect 776 human preference ratings from realistic voice assistant conversations, finding that both subsets and full benchmark achieve only 0.85 correlation with human. To better predict preferences, we trained regression models on these selected subsets, achieving 0.98 correlation -- outperforming regression models trained on both random subsets and the full benchmark. This demonstrates that in regression modeling, well-curated subsets outpredict the full benchmark, showing quality over quantity. We open-source these regression-weighted subsets as the HUMANS benchmark, an efficient proxy for LAM evaluation that captures both benchmark performance and user preferences.

音频模型高效评估用户偏好子集选择

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。