arXiv:2510.05871cs.AIcs.LG2025-10被引 3

用模型自信心过滤生成推理链,低成本构建生物推理数据集

Towards Label-Free Biological Reasoning Synthetic Dataset Creation via Uncertainty Filtering

  • 用模型置信度替代真实标签筛选推理路径
  • 过滤后数据使模型准确率提升,接近真实标签训练效果
  • 适合湿实验数据稀缺的生物医学领域研究者

合成思维链广泛用于训练大型推理模型(LRMs),通过提供逐步监督提升泛化能力。然而多数方法需依赖真实标签来生成或筛选这些链,而在生物学等湿实验数据稀缺的领域,这一过程成本高昂。本文提出无需标签的替代方案:基于不确定性的过滤机制,利用模型自身置信度(如自一致性、预测困惑度)作为外部标签的替代信号,采样多个推理路径并保留低不确定性子集。在生物扰动预测任务中,该方法显著提升模型准确率;使用过滤后数据进行有监督微调,性能优于未过滤的合成数据,缩小与真实标签训练的差距,并超越多个强基线模型。消融实验表明,按类别过滤可校正类别间不确定性差异,混合不确定性指标能生成更高质量数据集。结果表明,模型内部置信度是高效推理数据构建的强大信号,适用于标注成本高的领域。

原文摘要 · Abstract (English)

Synthetic chain-of-thought (CoT) traces are widely used to train large reasoning models (LRMs), improving generalization by providing step-level supervision. Yet most approaches require ground-truth labels to seed or filter these traces - an expensive bottleneck in domains like biology where wet-lab data are scarce. We propose a label-free alternative: uncertainty-based filtering, which uses a model's own confidence - quantified through established uncertainty metrics like self-consistency and predictive perplexity - as a substitute for external labels. We sample multiple reasoning traces and retain only low-uncertainty subsets. Applied to biological perturbation prediction, a domain where wet-lab labels are especially costly, we show that the filtered subset has higher accuracy, and that supervised fine-tuning (SFT) on uncertainty-filtered data outperforms unfiltered synthetic data, narrows the gap to ground-truth training, and surpasses strong LRM baselines. Ablations show that per-class filtering corrects for class-specific uncertainty scales and that hybrid uncertainty metrics yield higher-quality datasets. Our results suggest that model-internal confidence is a powerful signal for efficient reasoning dataset creation, enabling LRMs in domains where supervision is expensive.

推理生成无监督数据生物信息学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。