通过精准定位模型电路,实现低资源语言迁移中的稳定适应。
Beyond Transfer Accuracy: Faithful Circuits for Controlled Low-Resource Adaptation
- 用标签平衡激活均值和任务相关性评分,无须反事实条件发现模型电路。
- 在低资源跨语言情感分析中,性能接近全量微调且遗忘更少。
- 适合需要保留旧知识、追求稳定迁移的轻量级模型适配场景。
现有电路发现方法依赖带清晰反事实的模板任务,难以应用于多样化自然文本。本文将上下文分解法(CD-T)扩展至非结构化场景,通过标签平衡的激活均值与任务方向相关性评分,实现无需反事实的电路发现。利用这些电路设计电路目标监督微调(CT-SFT),仅更新任务相关头和层归一化参数。在NusaX跨语言情感分类任务上,CT-SFT在低资源条件下表现优异,虽非电路稀疏更新或全量微调有时能达相同准确率,但只有CT-SFT能显著降低灾难性遗忘,保持源语言及关联任务性能。在XNLI上的扩展验证了该方法在更广泛任务与模型族中的有效性,表明电路靶向适配是比全局微调更安全、因果基础更强的替代方案。
原文摘要 · Abstract (English)
Existing circuit discovery methods rely on templated tasks with clean counterfactuals, limiting their use on diverse natural text. We adapt Contextual Decomposition for Transformers (CD-T) for unstructured settings via label-balanced activation means and task-directional relevance scoring, enabling counterfactual-free circuit discovery. We leverage these circuits for Circuit-Targeted Supervised Fine-Tuning (CT-SFT), restricting parameter updates to task-relevant heads and LayerNorm. Experiments on NusaX cross-lingual sentiment transfer show that CT-SFT is highly competitive for low-resource adaptation. While non-circuit sparse updates and full fine-tuning sometimes match target accuracy through capacity recruitment, CT-SFT uniquely minimizes catastrophic forgetting, preserving source-language and related-task performance. Extensions to XNLI confirm these findings hold across broader tasks and model families, demonstrating that circuit-targeted adaptation provides a safer, causally grounded alternative to global fine-tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。