用关键词引导评分,让AI评价故事更准更省资源。
CoKe: Customizable Fine-Grained Story Evaluation via Chain-of-Keyword Rationalization
- 先生成关键词再写解释,引导模型精准预测故事评分
- 在StoryER数据集上相关性超人类,2倍优于GPT-4
- 小模型实现高精度,参数量远低于大模型
使用语言模型评估创意文本(如人工撰写的故事)始终具有挑战性,主要源于多标注者评分的主观性。为模拟人类思维过程,思维链(CoT)生成自由文本解释以指导模型预测,自一致性(SC)则对多个生成解释进行预测平均。本研究发现,现有自一致性方法因生成‘流畅外观’解释与实际提升评分预测之间存在目标不匹配,导致效果不佳。为此,我们提出链式关键词(CoKe)方法,在生成自由文本解释前,先生成一系列关键词来引导评分预测。随后,生成多样化的关键词序列,并聚合对应得分。基于小型微调评估模型的CoKe在StoryER数据集上不仅达到人类水平性能,且与人工标注者的相关性显著超过GPT-4,提升2倍,同时所需参数量大幅减少。
原文摘要 · Abstract (English)
Evaluating creative text such as human-written stories using language models has always been a challenging task -- owing to the subjectivity of multi-annotator ratings. To mimic the thinking process of humans, chain of thought (CoT) generates free-text explanations that help guide a model's predictions and Self-Consistency (SC) marginalizes predictions over multiple generated explanations. In this study, we discover that the widely-used self-consistency reasoning methods cause suboptimal results due to an objective mismatch between generating 'fluent-looking' explanations vs. actually leading to a good rating prediction for an aspect of a story. To overcome this challenge, we propose $\textbf{C}$hain-$\textbf{o}$f-$\textbf{Ke}$ywords (CoKe), that generates a sequence of keywords $\textit{before}$ generating a free-text rationale, that guide the rating prediction of our evaluation language model. Then, we generate a diverse set of such keywords, and aggregate the scores corresponding to these generations. On the StoryER dataset, CoKe based on our small fine-tuned evaluation models not only reach human-level performance and significantly outperform GPT-4 with a 2x boost in correlation with human annotators, but also requires drastically less number of parameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。