提出多指标评估框架,更好比较文本生成的解码策略。
Towards Better Open-Ended Text Generation: A Multicriteria Evaluation Framework
- 用偏序关系构建基准测试,实现解码方法的相对排序。
- 设计新综合指标,平衡连贯性、多样性与困惑度等矛盾指标。
- 适合关注文本生成质量评估的研究者和实践者使用。
由于大型语言模型的兴起,开放式文本生成已成为自然语言处理中的重要任务。然而,由于常用指标(如连贯性、多样性、困惑度)之间存在权衡,评估模型质量和解码策略仍具挑战性。本文针对开放式文本生成的多指标评估问题,提出了新的相对与绝对排名方法。具体而言,我们采用基于偏序关系的基准测试方法,并引入一种新综合指标,以平衡现有自动评估指标,提供更全面的文本生成质量评估。实验表明,所提方法能稳健比较不同解码策略,为开放式文本生成任务中的模型选择提供有效指导。论文还展望了文本生成评估方法的未来方向,并公开了代码、数据集与模型。
原文摘要 · Abstract (English)
Open-ended text generation has become a prominent task in natural language processing due to the rise of powerful (large) language models. However, evaluating the quality of these models and the employed decoding strategies remains challenging due to trade-offs among widely used metrics such as coherence, diversity, and perplexity. This paper addresses the specific problem of multicriteria evaluation for open-ended text generation, proposing novel methods for both relative and absolute rankings of decoding methods. Specifically, we employ benchmarking approaches based on partial orderings and present a new summary metric to balance existing automatic indicators, providing a more holistic evaluation of text generation quality. Our experiments demonstrate that the proposed approaches offer a robust way to compare decoding strategies and serve as valuable tools to guide model selection for open-ended text generation tasks. We suggest future directions for improving evaluation methodologies in text generation and make our code, datasets, and models publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。