arXiv:2507.10124cs.AI2025-07

用‘你可能错了’引导大模型自省,发现隐藏偏见

Could you be wrong: Debiasing LLMs using a metacognitive prompt for improving human decision making

  • 在模型回答后加入'你可能错了'提示,激发自我反思
  • 能暴露初始回答中未察觉的偏见与矛盾证据
  • 适合希望提升模型可信度的研究者和使用者

识别大模型中的偏见仍是持续挑战。由于模型不断演进,当前正确的内容未来可能失效,因此需要能超越现有模型的通用去偏策略。借鉴人类决策心理学中的元认知干预方法,本文提出一种基于‘你可能错了’的提示机制。该提示在大模型生成回应后触发,使其主动补充解释、错误、偏见、反例及替代方案,这些内容在原始回答中均未显现。实验基于近期关于大模型偏见的论文问题,涵盖隐性歧视与元认知缺陷。结果表明,该提示可有效揭示模型与用户对提示理解不一致的问题,并纠正看似合理但不完整的推理。研究证明,人类心理机制为提示工程提供了新路径,可借力长期验证有效的认知改进策略。

原文摘要 · Abstract (English)

Identifying bias in LLMs is ongoing. Because they are still in development, what is true today may be false tomorrow. We therefore need general strategies for debiasing that will outlive current models. Strategies developed for debiasing human decision making offer one promising approach as they incorporate an LLM-style prompt intervention designed to bring latent knowledge into awareness during decision making. LLMs trained on vast amounts of information contain information about potential biases, counter-arguments, and contradictory evidence, but that information may only be brought to bear if prompted. Metacognitive prompts developed in the human decision making literature are designed to achieve this, and as I demonstrate here, they show promise with LLMs. The prompt I focus on here is "could you be wrong?" Following an LLM response, this prompt leads LLMs to produce additional information, including why they answered as they did, errors, biases, contradictory evidence, and alternatives, none of which were apparent in their initial response. Indeed, this metaknowledge often reveals that how LLMs and users interpret prompts are not aligned. Here I demonstrate this prompt using a set of questions taken from recent articles about LLM biases, including implicit discriminatory biases and failures of metacognition. "Could you be wrong" prompts the LLM to identify its own biases and produce cogent metacognitive reflection. I also present another example involving convincing but incomplete information, which is readily corrected by the metacognitive prompt. In sum, this work argues that human psychology offers a new avenue for prompt engineering, leveraging a long history of effective prompt-based improvements to human decision making.

大模型偏见元认知提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。