用可校准的预测集量化场景图生成的不确定性,提升可靠性与可解释性。
Conformal Prediction and MLLM aided Uncertainty Quantification in Scene Graph Generation
- 基于置信度预测框架,为任意现有场景图生成方法提供不确定性量化
- 构建统计严格覆盖的预测集,确保真值在置信区间内概率达标
- 结合多模态大模型筛选最合理场景图,适合追求可靠性的视觉理解应用
场景图生成(SGG)旨在通过识别物体及其成对关系来结构化表达图像内容。然而,长尾分布和预测波动等固有挑战使得不确定性量化对其实用性至关重要。本文提出一种新型基于置信度预测(CP)的框架,可适配任意现有SGG方法,通过构建生成场景图的预测集来量化其预测不确定性,实现统计上严格的覆盖率保证。此外,为确保预测集中包含最具实际可解释性的场景图,设计了一种基于多模态大模型(MLLM)的后处理策略,从中筛选出视觉与语义上最合理的场景图。实验表明,该方法能从单张图像生成多样化的可能场景图,评估SGG方法的可靠性,并整体提升生成性能。
原文摘要 · Abstract (English)
Scene Graph Generation (SGG) aims to represent visual scenes by identifying objects and their pairwise relationships, providing a structured understanding of image content. However, inherent challenges like long-tailed class distributions and prediction variability necessitate uncertainty quantification in SGG for its practical viability. In this paper, we introduce a novel Conformal Prediction (CP) based framework, adaptive to any existing SGG method, for quantifying their predictive uncertainty by constructing well-calibrated prediction sets over their generated scene graphs. These scene graph prediction sets are designed to achieve statistically rigorous coverage guarantees. Additionally, to ensure these prediction sets contain the most practically interpretable scene graphs, we design an effective MLLM-based post-processing strategy for selecting the most visually and semantically plausible scene graphs within these prediction sets. We show that our proposed approach can produce diverse possible scene graphs from an image, assess the reliability of SGG methods, and improve overall SGG performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。