通过多维度评估发现,微调比模型大小对对话能力提升更关键。
The Oracle Has Spoken: A Multi-Aspect Evaluation of Dialogue in Pythia
- 设计多种基于语言学理论的对话评估指标,覆盖不同细粒度维度。
- 微调后多数指标得分快速饱和,模型规模影响较小。
- 相同评估模型导致多个指标趋势相似,引发测量可靠性质疑。
对话是大型语言模型(LLMs)的标志性能力之一。尽管广泛应用,但很少有研究能区分后训练过程中对话行为背后的特定成分。本文采用一套基于模型的综合评估指标,每个指标针对对话的一个细粒度方面,受语言学理论启发。我们评估了预训练的Pythia模型在不同模型规模下,以及在对话数据集上进行监督微调后的表现变化。结果显示,原始模型规模对大多数指标影响较弱,而微调则迅速使所有但最小模型的得分趋于饱和。出乎意料的是,许多指标表现出非常相似的趋势,特别是当它们依赖于同一评估模型时,这引发了其测量特定维度可靠性的疑问。为此,我们进一步分析了得分分布、指标相关性及生成回复中的词频,以解释观察结果。
原文摘要 · Abstract (English)
Dialogue is one of the landmark abilities of large language models (LLMs). Despite its ubiquity, few studies actually distinguish specific ingredients underpinning dialogue behavior emerging during post-training. We employ a comprehensive suite of model-based metrics, each targeting a distinct fine-grained aspect of dialogue, motivated by linguistic theory. We evaluate how the performance of pre-trained Pythia models changes with respect to each of those dimensions, depending on model size and as a result of supervised fine-tuning on conversational datasets. We observe only a mild impact of raw model size on most metrics, whereas fine-tuning quickly saturates the scores for all but the smallest models tested. Somewhat contrary to our expectations, many metrics show very similar trends, especially if they are all rooted in the same evaluator model, which raises the question of their reliability in measuring a specific dimension. To that end, we conduct additional analyses of score distributions, metric correlations, and term frequencies in generated responses to help explain our observations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。