用细粒度属性评估大模型推理理由,更精准揭示好坏原因。
Rethinking Human Preference Evaluation of LLM Rationales
- 从已有研究提取关键推理理由属性,构建评价体系。
- 通过SHAP分析发现解释性、逻辑连贯性等属性主导人类偏好。
- 用属性评分替代二元判断,让模型比较更细致可信。
大语言模型常生成自然语言推理理由,有助于复杂任务表现和用户可解释性,但其评估仍具挑战。现有方法依赖人类或模型进行二元偏好判断,但结果模糊且缺乏洞察。本文重新思考这一评估范式,提出三个问题:(1) 好的推理理由具备哪些属性?(2) 人类偏好能否由这些属性解释?(3) 属性评估能否克服二元比较的局限?我们基于文献提炼关键属性,通过自动指标、模型判断与人工标注进行评估。利用SHAP分析标准人类偏好数据集MT Bench与Chatbot Arena,识别出最能解释偏好的属性。进一步采用属性特异性ELO评分重新评估模型生成理由,获得更精细的模型对比与洞见。结果表明,细粒度属性评估能更准确刻画理由质量,推动未来向更可解释、可靠的评估实践发展。
原文摘要 · Abstract (English)
Large language models (LLMs) often generate natural language rationales -- free-form explanations that help improve performance on complex reasoning tasks and enhance interpretability for human users. However, evaluating these rationales remains challenging. While recent work has relied on binary preference judgments from humans or LLM judges, such evaluations are often opaque and coarse-grained, offering limited insight into what makes one rationale better than another. In this work, we rethink preference evaluation for LLM-generated rationales by asking: (1) What attributes define good rationales? (2) Can human preferences be explained by these attributes? (3) Can attribute-based evaluation overcome the limitations of binary comparisons? We identify a set of key rationale attributes from prior literature and assess them using automatic metrics, LLM judgments, and human annotations. We then analyze two standard human preference datasets MT Bench and Chatbot Arena using SHAP to identify which attributes best explain human preference outcomes. Finally, we re-evaluate model-generated rationales using attribute-specific ELO scores, revealing more nuanced model comparisons and insights. Our findings suggest that fine-grained attribute evaluations can better characterize rationale quality and guide future research toward more interpretable and reliable evaluation practices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。