研究发现语音和文本评价偏好差异大,需分别设计评估方法。
Same Words, Different Judgments: How Preferences Vary Across Modalities
- 对比100个相同语义内容的文本与语音评价,控制变量
- 每类模态需约9人评分才能达良好一致性(ICC≈0.80)
- 语音评价更贴近用户、偏差小,跨模态一致性接近随机
基于偏好的强化学习是对齐人工智能系统与人类偏好的主流框架。然而,当前评估协议主要针对文本设计,尚未验证在语音场景中的适用性。本文首次开展基于ICC的跨模态控制研究,比较100个相同语义内容在文本与语音模态下的人类与合成偏好标注。结果表明,单类模态内达到良好一致性(ICC(2,k) ≈ 0.80)需约9名评判者;但两类模态在偏好表达上存在显著差异:语音评判者决策阈值更窄、长度偏差更低、更关注用户导向标准,跨模态一致性接近随机水平。我们还证明,合成评分可有效预测人评一致性,可用于刺激选择的早期信号及人类标注的代理。这些发现表明,音频偏好数据的评估协议应采用模态特异性设计,而非直接沿用文本方案。
原文摘要 · Abstract (English)
Preference-based reinforcement learning (PbRL) is the dominant framework for aligning AI systems to human preferences. However, evaluation protocols for such data were designed for text and have not been validated for speech. We present the first ICC-based, controlled cross-modal study of human and synthetic preference annotations, comparing text and audio evaluations of identical semantic content across 100 prompts. We show that achieving $\textit{good}$ agreement within either modality (ICC(2,$k$) $\approx$ .80) requires $\sim$9 raters. At the same time, modalities show marked differences in how people report preferences: audio raters exhibit narrower decision thresholds, reduced length bias, and more user-oriented evaluation criteria, with near-chance cross-modality agreement. We demonstrate that synthetic ratings can be used to effectively predict inter-rater agreement, thus serving as an early signal for stimulus selection and proxy for human annotations. Together, these findings argue that evaluation protocols for audio preference data require modality-specific design rather than direct adaptation from text.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。