现有方法难捕捉主观写作偏好,新数据集揭示推理链模型更有效
Beyond Correctness: Evaluating Subjective Writing Preferences Across Cultures
- 用生成式奖励模型输出推理过程,准确率提升至81.8%
- 传统序列模型仅52.7%准确率,零样本模型为53.9%
- 跨文体差异大,小模型与大模型表现无显著差距
当前偏好学习方法在标准基准上表现优异,但当客观质量信号被移除时性能大幅下降。我们构建了WritingPreferenceBench数据集,包含1,800对人工标注的偏好样本(1,200英文,600中文),覆盖8种创意写作体裁,且响应在客观正确性、事实准确性及长度上保持一致。在此基准上,基于序列的奖励模型准确率仅为52.7%,零样本语言模型判别器达53.9%。相比之下,生成显式推理链的生成式奖励模型准确率达81.8%。我们观察到模型在不同体裁间存在显著方差:单个模型准确率范围从18.2%到81.8%,平均标准差达10.1%。该方差不随模型规模变化,27B参数模型未显著优于8B版本。结果表明,当前RLHF方法主要学习识别客观错误,而非捕捉主观质量偏好(如创意、文风魅力和情感共鸣),成功偏好建模可能需要中间推理表示而非直接分类。
原文摘要 · Abstract (English)
Current preference learning methods achieve high accuracy on standard benchmarks but exhibit significant performance degradation when objective quality signals are removed. We introduce WritingPreferenceBench, a dataset of 1,800 human-annotated preference pairs (1,200 English, 600 Chinese) across 8 creative writing genres, where responses are matched for objective correctness, factual accuracy, and length. On this benchmark, sequence-based reward models--the standard architecture for RLHF--achieve only 52.7% mean accuracy, while zero-shot language model judges perform at 53.9%. In contrast, generative reward models that produce explicit reasoning chains achieve 81.8% accuracy. We observe high within-model variance across genres: individual models range from 18.2% to 81.8% accuracy across different writing categories, with standard deviations averaging 10.1%. This variance persists regardless of model scale, with 27B parameter models showing no consistent improvement over 8B variants. Our results suggest that current RLHF methods primarily learn to detect objective errors rather than capture subjective quality preferences (e.g., creativity, stylistic flair, and emotional resonance), and that successful preference modeling may require intermediate reasoning representations rather than direct classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。