arXiv:2606.10279cs.AIcs.CL2026-06

用合成理由数据微调反而降低阿尔茨海默病预测准确率

Supervised Fine-tuning with Synthetic Rationale Data Hurts Real-World Disease Prediction

  • 用合成临床理由数据做监督微调,但效果比仅用标签差
  • 504种配置实验显示性能普遍下降,跨模型和数据量均成立
  • 理由本身准确,但训练时引入了与判别目标冲突的叙事逻辑

以合成临床理由数据进行监督微调,常被认为能提升语言模型在临床预测任务中的表现,因其不仅教会模型预测结果,还解释其依据。我们在纵向健康记录上测试了五年内阿尔茨海默病及相关痴呆症(ADRD)的预测,通过大规模控制实验(504种配置),发现基于理由的SFT始终显著降低预测性能,相比仅使用标签的微调。该性能下降在不同模型家族和数据规模下持续存在,且无法通过使用推理导向的基础模型缓解。关键的是,这一失败并非源于理由质量差:人类专家验证表明生成理由在医学上准确、忠实于患者特定证据;少样本实验也显示,相同理由作为推理时演示可提升性能,但作为训练目标却有害。我们识别出根本原因在于叙述合理性与判别式优化之间的结构冲突。希望本研究为理解何时及如何使用理由监督提供更精准指导,推动高风险临床预测中语言模型的负责任发展。

原文摘要 · Abstract (English)

Supervised fine-tuning with synthetic rationale data is widely assumed to improve language model performance on clinical prediction tasks by teaching models not just what to predict but why. We test this assumption on five-year Alzheimer's disease and related dementias (ADRD) prediction from longitudinal health histories. Across a large-scale controlled experiment of 504 configurations, we find that rationale-based SFT consistently and substantially hurts prediction performance relative to label-only fine-tuning. The degradation persists across model families and data scales, and is not resolved by using a reasoning-oriented base model. Crucially, the failure is not explained by poor rationale quality: human expert annotation confirms that the generated rationales are medically accurate and faithfully grounded in patient-specific evidence, and few-shot experiments show that the same rationales improve performance when used as inference-time demonstrations rather than training targets. We identify the root cause as a structural conflict between narrative plausibility and discriminative optimization. We hope our work paves the path toward a more precise understanding of when and how rationale-based supervision helps and when it does not, guiding the responsible development of language models for high-stakes clinical prediction.

临床预测语言模型微调策略阿尔茨海默病

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。