LLM triage errors come from output format,而非医学知识不足
Internal Representation, Not Clinical Knowledge: Where Apparent LLM Triage Failures Originate

- 用稀疏自编码器分析发现,医学特征在两种输出格式下表现一致
- 多选题错误主要源于选项顺序与格式设计,非知识缺失
- 模型常选相邻等级,说明是输出机制问题而非临床理解失败
患者语音的临床分诊基准显示,消费者级大模型在受限多选输出下存在高误漏诊率,但在自由文本输出中表现不同。我们探究输出格式是否改变模型的临床表征,或仅影响从同一表征到答案的映射。基于Gemma 3 4B/12B IT和Qwen3-8B模型,使用稀疏自编码器(SAE)特征分析发现:在相同临床叙事下,医学特征在两种格式中均激活,但在多选决策标记处全部沉默。三种独立方法(自然语言自编码器重构、决策标记对数归因、关键特征表征)一致表明:驱动决策的是结构与格式特征,而非医学特征。行为层面,多选惩罚在结构化与自然语言输入中反转;选项顺序打乱排除位置偏差;错误以‘差一位’为主(选择相邻分诊等级),而非知识错误。因此,失败根源在于输出格式,而非临床表征。
原文摘要 · Abstract (English)
Patient-voiced clinical-triage benchmarks report high under-triage rates for consumer LLMs for constrained multiple-choice output, yet the same cases score differently with free-text. We ask whether output format changes the model's \emph{clinical representation} or only the mapping from a preserved representation to an answer. Using sparse-autoencoder (SAE) features in Gemma 3 4B/12B IT and Qwen3-8B, we find the same medical features fire on the shared clinical narrative under both formats but go {silent} at the multiple-choice decision token in all the cases at every model. Three independent methods (natural-language autoencoder verbalization, decision-token logit attribution, and top-feature characterization) agree that scaffold and format features, but not medical features, drive the decision logits. Behaviorally, the multiple-choice penalty inverts under both structured and natural-language input, option-order shuffle rules out positional bias, and the gap is dominated by off-by-one decision (the model picks an adjacent acuity letter to the gold answer) rather than knowledge failure. Thus, the failure originates in the output format and not in the clinical representation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。