用行为标记训练临床大模型,让其在主动与被动间灵活切换。
BehaviorSFT: Behavioral Token Conditioning for Clinical Agents Across the Proactivity Spectrum
- 引入行为标记,动态控制模型在主动与被动间的决策
- 在行为基准上达97.3%宏F1,主动任务提升至96.5%
- 临床医生评测显示更自然、恰到好处的主动干预
大型语言模型作为临床助手需精细调整行为。尽管擅长响应式任务(如诊断推理),但在主动行为(如未被提示时发现关键缺失信息或风险)方面表现不佳。我们构建了BehaviorBench,一个覆盖从被动问答到主动干预(如澄清模糊信息、标记遗漏关键数据)的综合性评估数据集。实验揭示大模型在主动行为上表现不一。为此,我们提出BehaviorSFT,通过行为标记显式引导模型在行为谱系中动态选择行为。该方法显著提升性能,在BehaviorBench上达到最高97.3%的总体宏F1,并改善主动任务得分(如Qwen2.5-7B-Ins从95.0%提升至96.5%)。关键的是,盲评临床医生确认,经BehaviorSFT训练的代理展现出更真实的临床行为,平衡了有益主动性(如及时、相关建议)与必要克制(如避免过度干预),优于标准微调或显式指令模型。
原文摘要 · Abstract (English)
Large Language Models (LLMs) as clinical agents require careful behavioral adaptation. While adept at reactive tasks (e.g., diagnosis reasoning), LLMs often struggle with proactive engagement, like unprompted identification of critical missing information or risks. We introduce BehaviorBench, a comprehensive dataset to evaluate agent behaviors across a clinical assistance spectrum, ranging from reactive query responses to proactive interventions (e.g., clarifying ambiguities, flagging overlooked critical data). Our BehaviorBench experiments reveal LLMs' inconsistent proactivity. To address this, we propose BehaviorSFT, a novel training strategy using behavioral tokens to explicitly condition LLMs for dynamic behavioral selection along this spectrum. BehaviorSFT boosts performance, achieving up to 97.3% overall Macro F1 on BehaviorBench and improving proactive task scores (e.g., from 95.0% to 96.5% for Qwen2.5-7B-Ins). Crucially, blind clinician evaluations confirmed BehaviorSFT-trained agents exhibit more realistic clinical behavior, striking a superior balance between helpful proactivity (e.g., timely, relevant suggestions) and necessary restraint (e.g., avoiding over-intervention) versus standard fine-tuning or explicit instructed agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。