arXiv:2506.16064cs.CL2025-06被引 2

让大模型自我批评并优化回答,提升诚实与有用性。

Self-Critique-Guided Curiosity Refinement: Enhancing Honesty and Helpfulness in Large Language Models via In-Context Learning

  • 通过自省与修正两步提示,实现无需训练的响应优化。
  • 在HONESET数据集上,优质回答率提升,差回答减少,H²得分最高增4.3%。
  • 适合追求高可信度输出的AI应用开发者使用。

大语言模型在自然语言任务中表现出色,但持续生成诚实且有用的输出仍是挑战。本文从两个方向应对:一是对十款主流大模型(含OpenAI、Meta、Google的闭源与开源模型)进行综合基准评估;二是提出一种新颖的提示策略——自批判引导的好奇心精炼提示。该策略通过引入轻量级的上下文内自省与修正步骤,使模型无需额外训练即可自我优化。在HONESET数据集上,以GPT-4o为裁判,采用H²(诚实性与有用性)框架评估,所有模型均表现提升,差回答减少,高质回答增加,相对好奇心驱动提示的H²得分提升1.4%至4.3%。结果表明,结构化自精炼是一种可扩展、免训练的提升模型可信度的有效策略。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated robust capabilities across various natural language tasks. However, producing outputs that are consistently honest and helpful remains an open challenge. To overcome this challenge, this paper tackles the problem through two complementary directions. It conducts a comprehensive benchmark evaluation of ten widely used large language models, including both proprietary and open-weight models from OpenAI, Meta, and Google. In parallel, it proposes a novel prompting strategy, self-critique-guided curiosity refinement prompting. The key idea behind this strategy is enabling models to self-critique and refine their responses without additional training. The proposed method extends the curiosity-driven prompting strategy by incorporating two lightweight in-context steps including self-critique step and refinement step. The experiment results on the HONESET dataset evaluated using the framework $\mathrm{H}^2$ (honesty and helpfulness), which was executed with GPT-4o as a judge of honesty and helpfulness, show consistent improvements across all models. The approach reduces the number of poor-quality responses, increases high-quality responses, and achieves relative gains in $\mathrm{H}^2$ scores ranging from 1.4% to 4.3% compared to curiosity-driven prompting across evaluated models. These results highlight the effectiveness of structured self-refinement as a scalable and training-free strategy to improve the trustworthiness of LLMs outputs.

大模型优化提示工程可信生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。