提出多候选显著性检测,让模型一次输出多个合理分割结果。
Pluralistic Salient Object Detection
- 用专家混合模型同时生成多个分割掩码,支持不同用户意图
- 构建含10万+图像-掩码对的DUTS-MQ数据集,标注人类偏好分数
- 新数据集解决标注不一致问题,适合研究主观性视觉任务
我们提出多候选显著性检测(PSOD),旨在为给定图像生成多个合理的显著性分割结果。与传统方法仅输出单一掩码不同,该任务承认真实图像中存在多个显著对象,且显著性定义受用户意图影响而具有模糊性。为此,我们构建了两个新数据集:DUTS-MM 和 DUTS-MQ。DUTS-MM 在 DUTS 基础上改进标注质量,尤其在边界和细粒度结构上更准确,缓解标注不一致问题,并为存在歧义的图像提供多个真值掩码;DUTS-MQ 包含约10万张图像-掩码对,附带人工标注的偏好评分,支持学习真实人类对分割质量的判断。基于这两个数据集,我们提出一个基于混合专家(MOE)设计的简单有效基线模型,配备双预测头,可同时生成多个掩码并预测其人类偏好分数。大量实验与分析验证了新数据集的价值及所提框架的有效性。
原文摘要 · Abstract (English)
We introduce pluralistic salient object detection (PSOD), a novel task aimed at generating multiple plausible salient segmentation results for a given input image. Unlike conventional SOD methods that produce a single segmentation mask for salient objects, this new setting recognizes the inherent complexity of real-world images, comprising multiple objects, and the ambiguity in defining salient objects due to different user intentions. To study this task, we present two new SOD datasets "DUTS-MM" and "DUS-MQ", along with newly designed evaluation metrics. DUTS-MM builds upon the DUTS dataset but enriches the ground-truth mask annotations from three aspects which 1) improves the mask quality especially for boundary and fine-grained structures; 2) alleviates the annotation inconsistency issue; and 3) provides multiple ground-truth masks for images with saliency ambiguity. DUTS-MQ consists of approximately 100K image-mask pairs with human-annotated preference scores, enabling the learning of real human preferences in measuring mask quality. Building upon these two datasets, we propose a simple yet effective pluralistic SOD baseline based on a Mixture-of-Experts (MOE) design. Equipped with two prediction heads, it simultaneously predicts multiple masks using different query prompts and predicts human preference scores for each mask candidate. Extensive experiments and analyses underscore the significance of our proposed datasets and affirm the effectiveness of our PSOD framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。