通过新方法降低实验误差,提升批量数据下的预测准确性。
When three experiments are better than two: Avoiding intractable correlated aleatoric uncertainty by leveraging a novel bias--variance tradeoff
- 基于偏差-方差权衡,设计减少实验间偏差的主动学习策略。
- 利用协偏差-协方差关系,批量处理历史数据性能超越BALD等方法。
- 适合需要高精度实验决策的科学建模与数据采集场景。
真实世界实验中普遍存在异方差随机不确定性,且在批量设置下可能存在相关性。通过偏差-方差分解,可将模型分布与真实变量间的期望均方误差表示为认知不确定性项、偏差平方项和随机不确定性项之和。本文据此提出新型主动学习策略,直接减小实验轮次间的偏差,适用于含噪与无噪模型系统。进一步研究通过新型协偏差-协方差关系,以二次方式利用历史数据,自然引出基于特征分解的批处理机制。当采用该基于差值的方法结合二次估计器时,在批量设置下表现优于多种经典方法,包括BALD与最小置信度策略。
原文摘要 · Abstract (English)
Real-world experimental scenarios are characterized by the presence of heteroskedastic aleatoric uncertainty, and this uncertainty can be correlated in batched settings. The bias--variance tradeoff can be used to write the expected mean squared error between a model distribution and a ground-truth random variable as the sum of an epistemic uncertainty term, the bias squared, and an aleatoric uncertainty term. We leverage this relationship to propose novel active learning strategies that directly reduce the bias between experimental rounds, considering model systems both with and without noise. Finally, we investigate methods to leverage historical data in a quadratic manner through the use of a novel cobias--covariance relationship, which naturally proposes a mechanism for batching through an eigendecomposition strategy. When our difference-based method leveraging the cobias--covariance relationship is utilized in a batched setting (with a quadratic estimator), we outperform a number of canonical methods including BALD and Least Confidence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。