用模型自信心过滤生成推理链,低成本构建生物推理数据集
Towards Label-Free Biological Reasoning Synthetic Dataset Creation via Uncertainty Filtering
- 用模型置信度替代真实标签筛选推理路径
- 过滤后数据使模型准确率提升,接近真实标签训练效果
- 适合湿实验数据稀缺的生物医学领域研究者
合成思维链广泛用于训练大型推理模型(LRMs),通过提供逐步监督提升泛化能力。然而多数方法需依赖真实标签来生成或筛选这些链,而在生物学等湿实验数据稀缺的领域,这一过程成本高昂。本文提出无需标签的替代方案:基于不确定性的过滤机制,利用模型自身置信度(如自一致性、预测困惑度)作为外部标签的替代信号,采样多个推理路径并保留低不确定性子集。在生物扰动预测任务中,该方法显著提升模型准确率;使用过滤后数据进行有监督微调,性能优于未过滤的合成数据,缩小与真实标签训练的差距,并超越多个强基线模型。消融实验表明,按类别过滤可校正类别间不确定性差异,混合不确定性指标能生成更高质量数据集。结果表明,模型内部置信度是高效推理数据构建的强大信号,适用于标注成本高的领域。
原文摘要 · Abstract (English)
Synthetic chain-of-thought (CoT) traces are widely used to train large reasoning models (LRMs), improving generalization by providing step-level supervision. Yet most approaches require ground-truth labels to seed or filter these traces - an expensive bottleneck in domains like biology where wet-lab data are scarce. We propose a label-free alternative: uncertainty-based filtering, which uses a model's own confidence - quantified through established uncertainty metrics like self-consistency and predictive perplexity - as a substitute for external labels. We sample multiple reasoning traces and retain only low-uncertainty subsets. Applied to biological perturbation prediction, a domain where wet-lab labels are especially costly, we show that the filtered subset has higher accuracy, and that supervised fine-tuning (SFT) on uncertainty-filtered data outperforms unfiltered synthetic data, narrows the gap to ground-truth training, and surpasses strong LRM baselines. Ablations show that per-class filtering corrects for class-specific uncertainty scales and that hybrid uncertainty metrics yield higher-quality datasets. Our results suggest that model-internal confidence is a powerful signal for efficient reasoning dataset creation, enabling LRMs in domains where supervision is expensive.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。