用不确定性检测自动识别需澄清的提问,降低人工标注成本。
CLAIM: Leading Open-domain Active Clarification of Large Language Models with Uncertainty Measurement

- 通过多模型答案分歧熵衡量查询不确定性,自动生成标注数据。
- 结合监督与强化学习,让模型学会在开放域主动提出有效追问。
- 无需人工偏好标注,适合真实场景中大模型的智能交互优化。
在开放域人机交互中,大语言模型常面临模糊或不完整的用户提问,直接回答易导致泛化错误或信息量低。现有方法依赖人工标注或偏好对齐来决定何时澄清及澄清哪个方面,成本高且泛化差。为此,我们提出CLAIM框架,通过多个模型答案分歧的熵来量化查询不确定性,构建高质量合成数据,训练统一的澄清决策模型。该框架采用熵驱动的合成数据生成流程,结合语义聚类与推理判断,实现澄清需求的可靠自动标注。训练时将澄清过程建模为结构化决策生成问题,融合监督微调(SFT)与组相对策略优化(GRPO)。实验表明,CLAIM可在无手动标注数据下学习稳定且通用的澄清策略,为真实开放域交互提供低成本、强鲁棒的主动理解方案。
原文摘要 · Abstract (English)
In open-domain human-computer interaction scenarios, large language models (LLMs) frequently encounter user queries that are ambiguous or incomplete. In such cases, directly producing an answer often leads to overgeneralized, erroneous, or low-information responses. In contrast, asking clarifying questions can substantially improve interaction quality. However, existing approaches still rely heavily on manually annotated data or preference alignment to address two fundamental challenges: when clarification is necessary, and which aspect of the query should be clarified. This reliance incurs high annotation costs and limits generalization. To address these challenges, we propose CLAIM, an uncertainty-driven framework for active clarification learning in open-domain settings. CLAIM eliminates the need for explicit human preference annotations by quantifying query uncertainty through the entropy induced by answer disagreements across multiple models. This uncertainty signal is then used to construct high-quality synthetic data, enabling the training of a unified clarification decision model through a combination of supervised learning and reinforcement learning. Specifically, we propose an entropy-driven synthetic data generation pipeline that integrates entropy-based uncertainty estimation with semantic clustering and reasoning-based judgments, enabling reliable automatic annotation of clarification requirements. To train CLAIM, we formulate the clarification process as a structured decision generation problem and adopt a training paradigm that combines supervised fine-tuning (SFT) with group-relative policy optimization (GRPO). Experimental results demonstrate that CLAIM can learn stable and generalizable clarification strategies without relying on manually labeled data, offering a low-cost and robust solution for proactive understanding in real-world open-domain interactions with LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。