3D重建模型的置信度常被误用,实际预测误差偏高2.4倍。
A Calibration Audit of Confidence in Feed-Forward 3D Reconstruction
- 用幂律模型校正置信度大小,保留排序能力
- 训练后模型在未见数据上仍严重高估可靠性
- 校正后中位误差降低至1.35倍,局部微调可到1.12倍
前馈式3D重建模型输出每个像素的置信度,下游系统常将其视为可靠性信号。该置信度作为损失权重训练,而非不确定性度量,其能否用于误差预测尚未验证。我们在十三个数据集上审计了七个公开骨干模型,从四个维度评估置信度:误差排序能力、平均水平准确性、全范围一致性及置信区间覆盖性。结果显示置信度能较好排序误差,但当测试条件与训练不一致时,预测不确定性普遍偏低,中位误差偏移达2.4倍;且越自信,误差越偏离。即使损失函数已收敛,问题依然存在。一个在自身损失下恢复的模型,在数百次更新内即可达到最优,但在未见帧上仍保持过自信。针对每个模型-数据集组合,我们拟合出双常数幂律校正,仅调整幅度而不改变排序。尽管如此,场景尺度缺失导致校正无法完全弥补误差,即使最佳校正仍约有三分之二的场景落在五点误差带外,因缺失的是形状信息而非简单偏移。我们发布了审计协议、结果及每对模型-数据集的校正常数。若用目标数据集外的常数,中位误差由2.4倍降至1.35倍;再用少量标注样本微调,可进一步降至1.12倍。
原文摘要 · Abstract (English)
Feed-forward 3D reconstruction models emit a per-pixel confidence that downstream systems read as a reliability signal. It is trained as a loss weight, not as an uncertainty magnitude, and whether it can be used as an error prediction has not been measured. We audit seven released backbones on thirteen datasets and score the confidence on four properties, how well it ranks error, whether its level is right on average, whether it holds across the confidence range, and whether its intervals cover the truth. The confidence ranks error well, but the predicted uncertainty is too low when it is read under conditions that are not exactly those of training. The median case is off by 2.4x across all seven models, and the error prediction is further off the more confident the model is. We show that this phenomenon can appear even though the loss's optimum is reached. A released model resumed under its own loss reaches that optimum on its training data within a few hundred updates and stays overconfident on unseen frames. A power law with two constants per backbone and dataset corrects the overall magnitude of the predicted uncertainty and leaves the ranking untouched. What no rescaling reaches is the scene, which we attribute to the model's missing knowledge of scale across predictions. Every correction we tried is close to right on average and still leaves two thirds of held-out scenes outside a five-point band, because what a scene is missing is a shape rather than a shift. We release the audit protocol, its results, and the fitted constants per model and dataset. Fitted with the target dataset held out, the constants bring the median case from 2.4x off to 1.35x, and a refit on a few labelled scenes of that dataset reaches 1.12x.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。