arXiv:2509.02075cs.CLcs.AI2025-09被引 1

指令微调让大模型更精准控制生成文本长度。

How Instruction-Tuning Imparts Length Control: A Cross-Lingual Mechanistic Analysis

  • 通过分析模型内部组件贡献,发现微调后深层注意力头更专注长度控制。
  • 英语中后期注意力头贡献显著增强,意大利语则由最后一层MLP补偿调节。
  • 揭示了跨语言下模型对长度约束的适应性机制,适合研究模型可解释性者参考。

遵循明确的长度约束(如精确字数生成)仍是大语言模型面临的重要挑战。本研究对比基础模型与指令微调模型在英语和意大利语中的长度控制生成表现,采用基于直接逻辑归因的累积加权归因(Cumulative Weighted Attribution)分析性能及内部组件贡献。结果表明,指令微调显著提升长度控制能力,主要通过深层模型层组件的专门化实现。具体而言,在英语中,指令微调模型后期层的注意力头呈现越来越强的正向贡献;在意大利语中,注意力贡献较弱,但最后一层MLP的正向作用更明显,暗示补偿机制存在。研究显示,指令微调重构了后期层以适应任务要求,组件级策略可能随语言上下文动态调整。

原文摘要 · Abstract (English)

Adhering to explicit length constraints, such as generating text with a precise word count, remains a significant challenge for Large Language Models (LLMs). This study aims at investigating the differences between foundation models and their instruction-tuned counterparts, on length-controlled text generation in English and Italian. We analyze both performance and internal component contributions using Cumulative Weighted Attribution, a metric derived from Direct Logit Attribution. Our findings reveal that instruction-tuning substantially improves length control, primarily by specializing components in deeper model layers. Specifically, attention heads in later layers of IT models show increasingly positive contributions, particularly in English. In Italian, while attention contributions are more attenuated, final-layer MLPs exhibit a stronger positive role, suggesting a compensatory mechanism. These results indicate that instruction-tuning reconfigures later layers for task adherence, with component-level strategies potentially adapting to linguistic context.

指令微调长度控制可解释性多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。