用细粒度属性训练大模型问出好问题,临床诊断错误降56.6%
ALFA: Aligning LLMs to Ask Good Questions A Case Study in Clinical Reasoning
- 将好问题拆解为清晰、相关等可量化属性,逐项优化
- 在医疗场景中使诊断错误率降低56.6%,提问胜率超64%
- 适合需要精准提问的专家领域,如医疗、法律
大语言模型在不确定时往往无法提出有效问题,影响其在需主动获取信息决策领域的可靠性。我们提出ALFA框架,通过(i)将“好问题”分解为理论支撑的细粒度属性(如清晰性、相关性),(ii)可控生成属性特定的问题变体,(iii)基于偏好优化对齐模型,使其显式学习在这些属性上提出更优问题。以临床推理为例,构建了包含17,000条真实医患交互及80,000对属性特定偏好标注的MediQ-AskDocs数据集,并设计新型专家标注的互动医疗问答任务评估提问能力。经ALFA对齐的模型在MediQ-AskDocs上诊断错误率比当前最优指令微调模型降低56.6%,提问胜率达64.4%,且具备强泛化能力。结果表明,以结构化细粒度属性引导提问,是提升大模型在专业领域表现的可扩展路径。
原文摘要 · Abstract (English)
Large language models (LLMs) often fail to ask effective questions under uncertainty, making them unreliable in domains where proactive information-gathering is essential for decision-making. We present ALignment via Fine-grained Attributes, (ALFA) a framework that improves LLM question-asking by (i) decomposing the notion of a "good" question into a set of theory-grounded attributes (e.g., clarity, relevance), (ii) controllably synthesizing attribute-specific question variations, and (iii) aligning models via preference-based optimization to explicitly learn to ask better questions along these fine-grained attributes. Focusing on clinical reasoning as a case study, we introduce the MediQ-AskDocs dataset, composed of 17k real-world clinical interactions augmented with 80k attribute-specific preference pairs of follow-up questions, as well as a novel expert-annotated interactive healthcare QA task to evaluate question-asking abilities. Models aligned with ALFA reduce diagnostic errors by 56.6% on MediQ-AskDocs compared to SoTA instruction-tuned LLMs, with a question-level win-rate of 64.4% and strong generalizability. Our findings suggest that explicitly guiding question-asking with structured, fine-grained attributes offers a scalable path to improve LLMs, especially in expert application domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。