通过筛选合理选项,让大模型答题更稳定可靠。
Inclusion-of-Thoughts: Mitigating Preference Instability via Purifying the Decision Space

- 只保留合理选项重构题目,减少干扰项影响。
- 在多个基准上显著提升推理表现,计算开销极低。
- 适合关注模型决策透明度与稳定性研究者。
多项选择题(MCQ)广泛用于评估大语言模型(LLMs)。然而,模型仍易受合理干扰项影响,导致在正确与错误答案间频繁波动,表现出偏好不稳定性。本文提出包含思维(Inclusion-of-Thoughts, IoT)方法,一种渐进式自过滤策略,旨在缓解这种认知负担,使模型更聚焦于合理选项。该方法仅使用合理选项重构题目,为比较判断提供可控环境,从而评估模型在扰动下的内部推理稳定性。通过显式记录过滤过程,IoT还增强了模型决策的可解释性。大量实证评估表明,IoT在多个算术、常识推理及教育基准上显著提升链式思维(chain-of-thought)性能,且计算开销极小。
原文摘要 · Abstract (English)
Multiple-choice questions (MCQs) are widely used to evaluate large language models (LLMs). However, LLMs remain vulnerable to the presence of plausible distractors. This often diverts attention toward irrelevant choices, resulting in unstable oscillation between correct and incorrect answers. In this paper, we propose Inclusion-of-Thoughts (IoT), a progressive self-filtering strategy that is designed to mitigate this cognitive load (i.e., instability of model preferences under the presence of distractors) and enable the model to focus more effectively on plausible answers. Our method operates to reconstruct the MCQ using only plausible option choices, providing a controlled setting for examining comparative judgements and therefore the stability of the model's internal reasoning under perturbation. By explicitly documenting this filtering process, IoT also enhances the transparency and interpretability of the model's decision-making. Extensive empirical evaluation demonstrates that IoT substantially boosts chain-of-thought performance across a range of arithmetic, commonsense reasoning, and educational benchmarks with minimal computational overhead.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。