arXiv:2608.04772cs.CLcs.AI2026-08

用眼科指南当监督信号,零标注训练出高效眼病电话分诊模型。

Guideline-as-Oracle: Zero-Annotation Training of an Ophthalmic Telephone Triage Agent

论文配图:Guideline-as-Oracle: Zero-Annotation Training of an Ophthalmic Telephone Triage Agent
图 1 · 摘自论文原文
  • 将眼科指南转为70条规则表,作为唯一监督信号生成3000组对话。
  • 模型在201例测试中与标准答案一致率从61.7%提升至74.1%,突发病例召回率从9.5%升至69.0%。
  • 无需前沿大模型,适合医疗对话系统低资源训练场景。

多轮医疗对话系统的规模化监督困难,因专家对话标注成本高且临床数据受隐私限制。本文提出Guideline-as-Oracle(GAO),将美国眼科学会指南转化为70行可操作规则表,作为3000组训练对话的唯一实例级监督信号,仅保留人工标注用于评估。为将规则转化为对话,我们归纳八种构建策略,包括引用行层级分配、单事实边界对、仅元数据修复和标签修复,并界定其证据状态:标注机制、空值、混淆或仅整体评估。在90亿参数主干模型上微调后得到GAO-Triage,使与201例操作参考的一致性从61.7%提升至74.1%(精确McNemar p=0.0046),突发病例召回率从9.5%增至69.0%;该提升在第二组随机种子和患者模拟器下仍持续。七种通用系统均未在两个指标上全面超越GAO-Triage,且推理时无需前沿模型。打乱标签-对话对应关系导致模型退化为固定流程预测器,表明信号源于指南驱动的分配而非对话表面形式。标签修复与训练后期安全性能下降的消失同步发生。

原文摘要 · Abstract (English)

Scaling supervision for multi-turn medical agents is difficult because expert dialogue annotation is costly and clinical conversations are privacy-restricted. We introduce Guideline-as-Oracle (GAO), which compiles American Academy of Ophthalmology guidance into a 70-row operational rule table and uses it as the sole source of instance-level supervision for 3,000 training dialogues, reserving human labeling for evaluation. Because converting rules into dialogues is itself a design problem, we catalog eight construction strategies, including cited-row tier assignment, one-fact boundary pairs, metadata-only repair, and label repair, and characterize the evidential status of each: labeling mechanism, null, confounded, or evaluated only as a package. Fine-tuning a 9B backbone on this corpus yields GAO-Triage, improving agreement with a 201-case operational reference from 61.7% to 74.1% (exact McNemar p=0.0046) and emergent-case recall from 9.5% to 69.0%; the gains persist across a second seed and patient simulator. None of the seven general-purpose systems we test dominates GAO-Triage on both metrics, and GAO-Triage requires no frontier model at inference time. Permuting label-dialogue assignments collapses the model to a constant-routine predictor, indicating that the signal lies in guideline-derived assignment rather than dialogue surface form. Label repair coincides with the disappearance of a late-training safety degradation.

医疗对话零标注规则引导眼病分诊

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。