让大模型决策可调,临床分级更安全可靠
STEER: Inference-Time Risk Control via Constrained Quality-Diversity Search
- 通过约束质量多样性搜索生成多样化语言人格
- 推理时仅需调节一个参数,即可控制决策保守程度
- 适合对安全性和可调性要求高的医疗等场景
大型语言模型在追求平均正确率时易出现模式坍缩,导致在存在多种合理回答的任务中行为单一。这一问题在临床分诊等序数决策场景中尤为严重,标准对齐方法会削弱根据上下文调整特异性与敏感性的能力。我们提出STEER(基于进化集成精炼的可调微调),一种无需训练的框架,重新引入可调控的决策控制。STEER通过离线的、受约束的质量-多样性搜索构建一组自然语言人格,兼顾行为覆盖度并满足最低安全、推理和稳定性阈值。推理时,用户指定风险百分位,系统通过单一可解释参数选择对应人格,实现决策保守性的单调调节。在两个临床分诊基准上,STEER相比温度采样和静态人格集成展现出更广的行为覆盖;相比代表性后训练方法,在明确紧急病例上保持更高准确率,同时对模糊决策提供相当的控制力。结果表明,STEER是一种能保留领域能力的安全可控决策范式。
原文摘要 · Abstract (English)
Large Language Models (LLMs) trained for average correctness often exhibit mode collapse, producing narrow decision behaviors on tasks where multiple responses may be reasonable. This limitation is particularly problematic in ordinal decision settings such as clinical triage, where standard alignment removes the ability to trade off specificity and sensitivity (the ROC operating point) based on contextual constraints. We propose STEER (Steerable Tuning via Evolutionary Ensemble Refinement), a training-free framework that reintroduces this tunable control. STEER constructs a population of natural-language personas through an offline, constrained quality-diversity search that promotes behavioral coverage while enforcing minimum safety, reasoning, and stability thresholds. At inference time, STEER exposes a single, interpretable control parameter that maps a user-specified risk percentile to a selected persona, yielding a monotonic adjustment of decision conservativeness. On two clinical triage benchmarks, STEER achieves broader behavioral coverage compared to temperature-based sampling and static persona ensembles. Compared to a representative post-training method, STEER maintains substantially higher accuracy on unambiguous urgent cases while providing comparable control over ambiguous decisions. These results demonstrate STEER as a safety-preserving paradigm for risk control, capable of steering behavior without compromising domain competence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。