arXiv:2605.01346cs.CV2026-05

通过对比竞争假设,让模型在模糊时自动放弃判断,提升决策可靠性。

CHASE: Competing Hypotheses for Ambiguity-Aware Selective Prediction

  • 构建竞争假设框架,用多解释比对识别真实模糊性。
  • 在高模糊场景下三分类准确率提升8.8%,对齐率提高11.0%。
  • 适用于需要谨慎决策的生物视频分析等结构化模糊场景。

标准选择性预测方法通常从单一预测分支输出中估计不确定性。然而,在部分可观测场景下,局部时间证据可能相互矛盾,传统置信度分数会误导判断。我们提出CHASE(竞争假设的模糊感知选择性预测),一种显式比较结构化时间解释以决定是否做决策或放弃的框架。由于真实模糊会导致竞争假设之间的得分差距缩小,CHASE通过优化基于排序的筛选器来最大化假设间距,从而全局区分安全决策与根本不确定的情况。我们在隐含连接推断任务上评估该框架,使用受控的、基于物理的巨单层囊泡(GUV)模拟器,并实现零样本定性迁移至真实GUV视频。实验表明,显式推理竞争假设带来了更优的性能平衡。相比经典不确定性基线,CHASE在无弃权准确率、三分类准确率和整体模糊对齐弃权(覆盖率80%)上均取得统计显著提升。具体而言,在极高模糊场景下,其三分类准确率相对提升最高达8.8%,整体对齐率相对提升最高达11.0%。在80%覆盖率下保持与最优基线相当的选择风险边界,90%覆盖率下总风险降低9.9%,为结构化模糊下的决策提供了更可靠的方案。

原文摘要 · Abstract (English)

Standard selective prediction methods typically estimate uncertainty from the output of a single predictive branch. While effective for general uncertainty estimation, these approaches often struggle under partial observability, where local temporal evidence can be contradictory and standard confidence scores become misleading. We introduce CHASE (Competing Hypotheses for Ambiguity-Aware Selective Prediction), a selective prediction framework that explicitly compares structured temporal explanations to determine whether to commit to a decision or abstain. Because genuine ambiguity causes the score gap between competing hypotheses to collapse, CHASE optimizes a ranking-aware selector over these hypothesis margins to globally separate safe commitments from fundamentally uncertain ones. We evaluate this framework on the problem of hidden connectivity inference, utilizing a controlled, physically grounded simulator inspired by the dynamics of giant unilamellar vesicles (GUVs), alongside zero-shot qualitative transfer (without retraining or fine tuning) to representative real GUV videos. Our experiments demonstrate that explicitly reasoning over competing hypotheses provides a superior balance of metrics. Compared to canonical uncertainty baselines, CHASE achieves statistically significant gains in overall no-abstain accuracy, three-way accuracy, and overall ambiguity-aligned abstention (at 80% coverage). Specifically, it yields up to an 11.0% relative mean improvement in overall alignment, alongside up to an 8.8% relative boost in three-way accuracy in the very-high ambiguity regime. By maintaining a selective risk boundary strictly at par with the best baselines at 80% coverage, and reducing overall risk by 9.9% at 90% coverage, this framework offers a more reliable approach to decision-making under structured ambiguity.

选择性预测模糊识别生物影像决策可信

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。