arXiv:2502.11541cs.CLcs.AI2025-02ACL被引 12

无需强模型,通过多粒度对比训练提升复杂指令遵循能力

MuSC: Improving Complex Instruction Following with Multi-granularity Self-Contrastive Training

  • 分粗细粒度构建约束感知偏好数据,实现无依赖自对齐
  • 在多个开源模型上显著超越现有自对齐方法,复杂指令任务提升明显
  • 适合希望低成本提升模型指令遵循能力的研究者与开发者

复杂指令遵循需处理多重约束,对大语言模型至关重要。现有方法依赖更强模型(如GPT-4)构建数据,限制应用范围。本文提出多粒度自对比训练(MuSC)框架,不依赖强模型即可提升复杂指令对齐能力。在粗粒度层面,基于指令分解与重组构建约束感知偏好数据;在细粒度层面,采用动态标记级监督进行令牌级偏好优化。实验在开源模型上验证,结果表明该方法在复杂与通用指令遵循基准上均取得显著提升,优于此前自对齐方法。

原文摘要 · Abstract (English)

Complex instruction-following with elaborate constraints is imperative for Large Language Models (LLMs). While existing methods have constructed data for complex instruction alignment, they all rely on a more advanced model, especially GPT-4, limiting their application. In this paper, we propose a Multi-granularity Self-Contrastive Training (MuSC) framework, to improve the complex instruction alignment without relying on a stronger model. Our method is conducted on both coarse and fine granularity. On coarse-granularity, we construct constraint-aware preference data based on instruction decomposition and recombination. On fine-granularity, we perform token-aware preference optimization with dynamic token-level supervision. Our method is evaluated on open-sourced models, and experiment results show our method achieves significant improvement on both complex and general instruction-following benchmarks, surpassing previous self-alignment methods.

指令遵循自对齐多粒度开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。