VISTA通过可视化分析提升大模型生成标签质量,助力开放词汇图像分割
VISTA: A Visual Analytics Framework to Enhance Foundation Model-Generated Data Labels
- 融合多阶段验证与人工专家判断,系统识别标签隐含错误
- 在两个基准数据集上显著提升模型性能,验证标签质量关键作用
- 适合关注数据质量、尤其是开放词汇分割的研究者使用
多模态基础模型(如CLIP和LLaVA)的进步推动了大规模数据集的自动标注,提升了开放词汇目标检测与分割等挑战性下游任务的模型表现。然而,现有方法更关注数据数量而非质量,缺乏对大模型生成标签质量的深入研究。这是因为无真实标签情况下验证海量数据极具挑战:现有方法仅依赖有限指标识别问题数据,或仅对少量数据进行人工验证,无法全面覆盖潜在问题。为此,我们提出VISTA——一个视觉分析框架,通过结合多阶段数据验证策略与人类专家经验,帮助用户发现、理解并修正大模型生成标签中的隐藏缺陷。在两个基准数据集上的详细用例及专家评审表明,VISTA从定量和定性两方面均有效提升了多模态模型性能。
原文摘要 · Abstract (English)
The advances in multi-modal foundation models (FMs) (e.g., CLIP and LLaVA) have facilitated the auto-labeling of large-scale datasets, enhancing model performance in challenging downstream tasks such as open-vocabulary object detection and segmentation. However, the quality of FM-generated labels is less studied as existing approaches focus more on data quantity over quality. This is because validating large volumes of data without ground truth presents a considerable challenge in practice. Existing methods typically rely on limited metrics to identify problematic data, lacking a comprehensive perspective, or apply human validation to only a small data fraction, failing to address the full spectrum of potential issues. To overcome these challenges, we introduce VISTA, a visual analytics framework that improves data quality to enhance the performance of multi-modal models. Targeting the complex and demanding domain of open-vocabulary image segmentation, VISTA integrates multi-phased data validation strategies with human expertise, enabling humans to identify, understand, and correct hidden issues within FM-generated labels. Through detailed use cases on two benchmark datasets and expert reviews, we demonstrate VISTA's effectiveness from both quantitative and qualitative perspectives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。