arXiv:2409.15979cs.CL2024-09被引 3

微调大模型提升无参考文本评估效率与精度。

Finetuning LLMs for Comparative Assessment Tasks

  • 用软概率微调模型,让输出更贴近真实对比分布。
  • 在少量对比对上仍保持高精度,超越现有方法。
  • 适合需要高效文本生成评估的科研与工业场景。

自然语言生成的自动化评估是一项挑战性任务。指令微调的大规模语言模型(LLMs)在无参考评估中展现出潜力,尤其是在比较评估方面。然而,成对比较的二次计算复杂性限制了其可扩展性。为解决此问题,已有研究通过零样本LLM概率应用比较策略来实现高效的比较评估。本文提出一种针对比较评估的LLM微调框架,旨在使模型输出对齐目标比较概率分布。通过在软概率上训练,该方法在保持高效比较子集性能的同时,提升了当前最优表现。

原文摘要 · Abstract (English)

Automated assessment in natural language generation is a challenging task. Instruction-tuned large language models (LLMs) have shown promise in reference-free evaluation, particularly through comparative assessment. However, the quadratic computational complexity of pairwise comparisons limits its scalability. To address this, efficient comparative assessment has been explored by applying comparative strategies on zero-shot LLM probabilities. We propose a framework for finetuning LLMs for comparative assessment to align the model's output with the target distribution of comparative probabilities. By training on soft probabilities, our approach improves state-of-the-art performance while maintaining high performance with an efficient subset of comparisons.

大模型评估对比学习微调生成质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。