提出自一致性优化,解决大模型回答不一致、盲目迎合等问题
Position: It's Time to Optimize LLMs for Self-Consistency
- 将多种改进方法统一为一致性优化框架
- 通过跨输入推理检测并修复模型错误行为
- 适合关注模型可靠性与逻辑一致性的研究者
尽管语言模型的预训练和后训练流程日益复杂,仍存在诸多关键缺陷:模型过度依赖用户提问方式(盲目迎合)、逻辑泛化不完整,以及自信但错误的回答。我们认为,这些失败源于贯穿整个流程的根本假设——可独立评估单个输出对的表现。许多模型问题只有在分析跨输入响应关系时才能被发现。本文提出以自一致性为核心框架来理解这些问题。我们观察到,针对对抗鲁棒性、事实一致性等不同目标的多种技术,均可视为一种通用‘一致性优化’程序的特例,并可用标准优化工具处理。随后,我们展望了通过一致性优化可能实现的新模型特性,并讨论了构建普遍一致的大模型的意义,包括其带来的能力提升与引发的质疑。
原文摘要 · Abstract (English)
Despite ever-increasing sophistication in language model (LM) pre- and post-training pipelines, many important failures persist: models overcondition on user framing ("sycophancy"), exhibit incomplete logical generalization, and produce confident but incorrect responses. We argue that these failures arise from a modeling assumption permeating all aspects of the pipeline: that behavior can be specified and evaluated independently on single-output pairs. Many model failures are difficult, if not impossible, to detect without reasoning about relationships between a model's responses across inputs. In this position paper, we propose self-consistency as a framework for understanding these failures. We first observe that a wide variety of techniques designed to improve specific aspects of LM behavior-targeting properties as diverse as adversarial robustness and factual coherence-can be understood as special cases of a common "consistency optimization" procedure and addressed with a standard set of optimization tools. We next outline a set of new model properties that could be achieved by optimizing for consistency, and conclude with a discussion of what it would mean to develop generally consistent LMs, including the capabilities they would enable and the objections they raise.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。