评测情绪描述偏好时,模型只需看字数和生成器就能猜对人类选择,根本没看视频。
Style over Substance: A Shortcut Audit of Emotion-Description Preference Evaluation

- 仅凭描述长度和生成器身份,就能达到与专业模型相当的准确率
- 超过66%的人类偏好结果可由生成器风格先验解释,且模型仍受此风格影响
- 当前评测易被文字长度和生成器特征干扰,需改进配对与评估设计
情绪描述偏好正成为多模态情绪理解的标准评估指标,如EmoPrefer基准。该评测假设判断偏好需基于视频的跨模态理解。我们通过内容无关探测进行系统性捷径审计发现:仅使用描述长度和生成器身份的逻辑回归模型(无需处理文本、视频或音频),在EmoPrefer-V2上表现达65.8 WAF,接近7B参数文本与视听判别器微调后的66.8 WAF。生成器身份可从描述中以99.5%准确率恢复;所有候选对均来自不同生成器;66%的人类偏好与生成器胜率先验一致。当人类标注与先验冲突时,训练判别器仍遵循风格先验,在63%至80%的样本中做出相同选择。在长度匹配子集上,不同媒体配置无显著提升;而一种解耦风格捷径的诊断方法显示,内容头性能接近随机。这表明当前评分可在不验证描述与视频一致性的情况下达成。建议未来采用源平衡配对、严格长度控制、反刻板切片报告及多标注者共识。
原文摘要 · Abstract (English)
Preference over model-generated emotion descriptions is emerging as a standard evaluation metric for multimodal emotion understanding, exemplified by the MER2026 MER-Prefer track on EmoPrefer. Such benchmarks assume that predicting the preferred description requires grounded cross-modal understanding of the video. We conduct a systematic shortcut audit of EmoPrefer using content-blind probes. A simple logistic regression using only description length and generator identity, without processing the text, video, or audio, performs comparably to LoRA-finetuned 7B text and audio-visual judges (65.8 versus 66.8 WAF on EmoPrefer-V2). Generator identity is recoverable from description text with 99.5 percent accuracy, every candidate pair contrasts two distinct generators, and the human preference labels agree with a fold-exclusive per-generator win-rate prior on 66 percent of the evaluated pairs. When the human label conflicts with this prior, trained judges still follow the style prior on 63 to 80 percent of the pairs. On a length-matched subset that neutralizes verbosity bias, the tested media configurations yield no statistically significant improvement, while an ODIN-inspired diagnostic that decouples the style shortcut leaves its content head near chance. These results do not imply that human preferences are inherently stylistic or that the descriptions contain no emotional information. Instead, they show that the current scores can be reached without verifying either description against the video. We recommend source-balanced pairing, strict length control, counter-stereotypical sliced reporting, and multi-annotator consensus for future cross-generator evaluations. Code is available at https://github.com/jiabingyang01/EmoPrefer-Audit.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。