通过局部编辑推理轨迹,让大模型按规程做医疗决策
Answer Engineering: Local Trajectory Editing for Protocol-Constrained Decision Making in Large Language Models
- 在生成过程中局部修改推理路径,不重训模型
- 临床判断准确率从42.0%提升至80.7%,合规率超80%
- 适合需严格遵循流程的医疗、法律等高风险场景
大语言模型在关键领域可能生成看似自信但违反规程的答案。本文提出Answer Engineering,一种无需重训练、不修改权重的确定性运行时干预层,通过局部规则引导,在自回归生成过程中对可见推理轨迹进行精准修正。在突发性耳聋(SSNHL)临床基准测试中,未引导生成的合规率仅为54.5%,而经局部编辑后提升至83.5%;在传导性听力损失对照条件下,接受率从1.6%升至58.9%,整体平衡准确率由42.0%增至80.7%。结果表明,通过可审计的运行时控制可显著提升规程遵守度,但也揭示了规则覆盖范围、触发可靠性及诊断优先生成模式带来的局限性。
原文摘要 · Abstract (English)
Large language models can produce confident but protocol-invalid answers in domains where procedural compliance is critical. This paper presents Answer Engineering, a deterministic runtime and authoring layer that applies localized rule-guided interventions to the visible reasoning trajectory during standard autoregressive generation, without retraining, modifying model weights, or performing global search. The method is evaluated on a controlled clinical benchmark for sudden sensorineural hearing loss (SSNHL), where correct management depends on protocol-consistent interpretation of symptom timing, Weber/Rinne tuning-fork findings, and otoscopic findings. In the benchmark, step-by-step reasoning shifted rather than eliminated errors: compliant outcomes for SSNHL decreased from 54.5% under unguided generation to 25.1%, while acceptance on the conductive contrast condition increased from 1.6% to 58.9%. Local trajectory editing increased SSNHL compliance to 83.5% and conductive-case adherence to 77.9%, raising balanced accuracy from 42.0% under reasoning-only generation to 80.7%. The results support a systems-level view in which protocol adherence can be improved through auditable runtime control of reasoning trajectories, while also identifying limitations caused by rule coverage, trigger reliability, and persistent diagnosis-first generation dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。