让大模型学会何时该纠正、何时该放弃,避免错误修正反而出错。
To Steer or Not to Steer? Mechanistic Error Reduction with Abstention for Language Models
- 自动优化纠正方向与力度,不依赖人工调参
- 在多个数据集上实现安全有效纠错,性能不降反升
- 可叠加使用,适合想提升现有纠错方法的开发者
我们提出机制性错误减少与弃权框架(MERA),一种通过选择性、自适应干预来缓解语言模型误差的系统方法。与依赖固定、手动调节干预强度的现有方法不同,MERA通过优化干预方向,并校准干预时机与程度,从而在无法自信修正时选择弃权,理论上可提升性能或安全回避。在多种数据集和语言模型家族上的实验表明,MERA实现了安全、有效的非退化错误修正,优于现有基线。此外,MERA可叠加应用于已有纠错技术之上,进一步提升效果,展现出通用且高效的机制激活调控潜力。
原文摘要 · Abstract (English)
We introduce Mechanistic Error Reduction with Abstention (MERA), a principled framework for steering language models (LMs) to mitigate errors through selective, adaptive interventions. Unlike existing methods that rely on fixed, manually tuned steering strengths, often resulting in under or oversteering, MERA addresses these limitations by (i) optimising the intervention direction, and (ii) calibrating when, and how much to steer, thereby provably improving performance or abstaining when no confident correction is possible. Experiments across diverse datasets, and LM families demonstrate safe, effective, non-degrading error correction, and that MERA outperforms existing baselines. Moreover, MERA can be applied on top of existing steering techniques to further enhance their performance, establishing it as a general-purpose, and efficient approach to mechanistic activation steering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。