不需标注数据,用自然语言精准修改神经元,提升模型在复杂叙事中听指令的能力。
Sparse Activation Editing for Reliable Instruction Following in Narratives
- 通过自然语言定位并编辑关键神经元,实现无需训练的指令遵循优化
- 在1212个真实叙事案例上测试,显著提升指令遵循准确率
- 适用于复杂故事场景,适合需要高可靠性指令执行的研究与应用
复杂叙事语境常挑战语言模型的指令遵循能力,现有基准也无法充分反映此类困难。为此,我们提出Concise-SAE——一种无需训练的框架,仅通过自然语言指令即可识别并编辑与指令相关的神经元,无需标注数据。为全面评估该方法,我们构建了FreeInstruct,一个包含1,212个样本的多样且真实的基准,突出叙事丰富场景下的指令遵循挑战。尽管最初针对复杂叙事设计,Concise-SAE在多种任务上均表现最优,且未牺牲生成质量。
原文摘要 · Abstract (English)
Complex narrative contexts often challenge language models' ability to follow instructions, and existing benchmarks fail to capture these difficulties. To address this, we propose Concise-SAE, a training-free framework that improves instruction following by identifying and editing instruction-relevant neurons using only natural language instructions, without requiring labelled data. To thoroughly evaluate our method, we introduce FreeInstruct, a diverse and realistic benchmark of 1,212 examples that highlights the challenges of instruction following in narrative-rich settings. While initially motivated by complex narratives, Concise-SAE demonstrates state-of-the-art instruction adherence across varied tasks without compromising generation quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。