用人类反馈迭代优化语言模型的诊疗策略,提升罕见病诊断准确率。
Policy Iteration with Human Feedback: Bringing Post-Training RL to In-context Learning
- 基于语言模型构建可版本化策略,结合专家评审实现持续改进。
- 在罕见病诊断中,召回率最高提升32.7个百分点,跨多个大模型有效。
- 适合医疗AI、具身智能等需高可靠性决策的场景使用。
生成式预训练建立了可复用的任务表征;后续工作表明固定模型可通过指令和示范实现上下文学习。政策迭代与人类反馈(PIHF)在此基础上,采用广义策略迭代的循环评估与优化结构。PIHF以预训练语言模型为执行基础,将持续修正机制转移至版本化的自然语言策略与工具集。语言模型评鉴器与临床专家共同分析推理与工具使用轨迹,定位重复错误并生成候选修正方案;专家可重新解读证据,并掌握准入与回滚权限,通过Recall@1和Recall@5验证修正后结果。在累计消融实验与超罕见病基准测试中,基于PIHF的策略在一家专有执行器及三个开源权重执行器(参数量3亿至490亿)上均提升了性能,其中GPT-5.4 Recall@1提升32.7个百分点,Qwen3.6-35B提升31.1个百分点,差异为1.7个百分点。这些结果证明,预训练语言模型可作为固定权重执行基底,支持专家引导下的罕见病诊断策略开发。
原文摘要 · Abstract (English)
Generative pretraining established reusable task representations; later work on language-based task conditioning and in-context learning showed that a fixed model could adapt its behavior from instructions and demonstrations. Policy Iteration with Human Feedback (PIHF) builds on this development and the recurrent evaluate-and-improve structure of generalized policy iteration. PIHF uses a pretrained language model as its execution substrate and moves persistent revision to a versioned natural-language policy and tool set. A language-model critic and clinical expert review complete-panel reasoning and tool-use trajectories to localize recurrent failures and form candidate revisions; the expert may reinterpret the evidence and retains authority over admission and rollback, while Recall@1 and Recall@5 validate outcomes after candidate execution. Across cumulative ablations and ultra-rare-disease benchmarks, a PIHF-derived policy improved Recall@1 in one proprietary executor and three open-weight executors spanning 3 to 49 billion active parameters. Gains were 32.7 percentage points for GPT-5.4 and 31.1 points for Qwen3.6-35B, a difference of 1.7 points. These results support the feasibility of using pretrained language models as fixed-weight execution substrates for expert-guided policy development in rare-disease diagnosis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。