arXiv:2506.23979cs.CL2025-06被引 1

用分类体系自动生成多语言偏好数据,提升大模型指令遵循能力。

TaP: A Taxonomy-Guided Framework for Automated and Scalable Preference Data Generation

  • 基于结构化分类体系,自动构建多样化偏好数据。
  • 在小规模数据上训练的模型性能超越大规模现有数据训练的模型。
  • 适合需要多语言、高质量偏好数据的研究者使用。

监督微调和偏好微调大型语言模型需要高质量数据集以提升其指令遵循能力和与人类偏好的一致性。然而,构建此类数据集成本高昂,且大多数公开数据集为英文。为此,我们提出一种分类体系引导的偏好数据生成框架(TaP),实现跨语言、自动化、可扩展的数据集构建。TaP利用结构化分类体系对数据组成进行细粒度控制,确保多样性与广泛覆盖。我们使用TaP生成的数据集对多个大型语言模型进行监督与偏好微调。实验表明,基于TaP数据训练的模型性能优于现有开源数据集训练的模型。值得注意的是,使用TaP生成的数据训练的模型性能超过使用规模达180倍更大的开源数据集训练的模型。

原文摘要 · Abstract (English)

Conducting supervised and preference fine-tuning of large language models (LLMs) requires high-quality datasets to improve their ability to follow instructions and align with human preferences and values. However, constructing such datasets is resource-intensive, and most publicly available datasets are in English. To address these challenges, we propose the \underline{\textbf{Ta}}xonomy-Guided \underline{\textbf{P}}reference Data Generation (TaP) framework for automated, scalable preference dataset construction across languages. TaP uses a structured taxonomy to provide fine-grained control over dataset composition, ensuring diversity and broad coverage. We use TaP-generated datasets to perform supervised and preference fine-tuning on multiple LLMs. Experimental results demonstrate that LLMs trained on TaP-generated datasets outperform those trained on existing open-source datasets. Remarkably, LLMs trained on TaP-generated datasets outperform models trained on an open-source dataset that is 180$\times$ larger.

偏好学习数据生成多语言大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。