为评估大模型解释质量,提出新评分标准并构建2.6万条带标注数据集。
Rubrik's Cube: Testing a New Rubric for Evaluating Explanations on the CUBE dataset
- 设计教育领域启发的评分框架Rubrik,评估解释的清晰度与简洁性。
- 发现大模型解释质量差主要因冗长而非逻辑或用词问题。
- 适合关注AI可解释性、评测方法研究的研究者使用。
大语言模型(LLMs)在解释生成任务中广泛应用,但其生成的解释常不可靠,用户难以辨别优劣。为此,我们提出Rubrik's CUBE——一种教育启发式评分标准,并构建包含26,000条解释的数据集,这些解释由人类及六种开源与闭源大模型撰写,并经该评分标准进行质量标注。数据集涵盖两类推理任务与两类语言任务,具备足够多样性,可用于有效验证所提评分体系。实验表明,解释质量受任务类型和感知难度影响,低质量主要源于表达冗长,而非连贯性或词汇选择问题。完整数据集、评分标准与代码已公开于https://github.com/RubriksCube/rubriks_cube。
原文摘要 · Abstract (English)
The performance and usability of Large-Language Models (LLMs) are driving their use in explanation generation tasks. However, despite their widespread adoption, LLM explanations have been found to be unreliable, making it difficult for users to distinguish good from bad explanations. To address this issue, we present Rubrik's CUBE, an education-inspired rubric and a dataset of 26k explanations, written and later quality-annotated using the rubric by both humans and six open- and closed-source LLMs. The CUBE dataset focuses on two reasoning and two language tasks, providing the necessary diversity for us to effectively test our proposed rubric. Using Rubrik, we find that explanations are influenced by both task and perceived difficulty. Low quality stems primarily from a lack of conciseness in LLM-generated explanations, rather than cohesion and word choice. The full dataset, rubric, and code are available at https://github.com/RubriksCube/rubriks_cube.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。