arXiv:2502.15919cs.CLcs.SD2025-02被引 5

对比静态评测与真实交互,发现现有评估无法准确反映用户偏好。

Mind the Gap! Static and Interactive Evaluations of Large Audio Models

  • 通过7500次真实用户交互收集数据,分析语音模型实际使用表现。
  • 静态评测与用户偏好相关性低,最大相关系数仅0.33。
  • 建议开发更贴近用户需求的交互式评估方法,尤其针对问答和年龄识别任务。

随着AI聊天机器人普及,语音交互为语义与社交信号传递提供了高效途径,推动了大音频模型(LAMs)的发展。然而,要使模型发展契合用户目标,需明确用户需求与偏好以建立可靠评估指标。本研究提出一种交互式评估方法,收集了来自484名参与者的7500次LAM交互数据。通过主题建模识别音频界面的主要应用场景,结合用户偏好排序与定性反馈,判断哪些模型最符合用户需求。进一步分析发现,单一静态基准测试与交互表现的相关性极弱(τ≤0.33),即使整合多个粗粒度特征,预测能力也有限(R²=0.30)。在二十个语音问答与年龄预测数据集中,仅有两个表现出显著正相关。结果表明,当前评估体系亟需改进,以更好匹配用户真实偏好。

原文摘要 · Abstract (English)

As AI chatbots become ubiquitous, voice interaction presents a compelling way to enable rapid, high-bandwidth communication for both semantic and social signals. This has driven research into Large Audio Models (LAMs) to power voice-native experiences. However, aligning LAM development with user goals requires a clear understanding of user needs and preferences to establish reliable progress metrics. This study addresses these challenges by introducing an interactive approach to evaluate LAMs and collecting 7,500 LAM interactions from 484 participants. Through topic modeling of user queries, we identify primary use cases for audio interfaces. We then analyze user preference rankings and qualitative feedback to determine which models best align with user needs. Finally, we evaluate how static benchmarks predict interactive performance - our analysis reveals no individual benchmark strongly correlates with interactive results ($τ\leq 0.33$ for all benchmarks). While combining multiple coarse-grained features yields modest predictive power ($R^2$=$0.30$), only two out of twenty datasets on spoken question answering and age prediction show significantly positive correlations. This suggests a clear need to develop LAM evaluations that better correlate with user preferences.

大音频模型用户评估交互评测语音接口

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。