VISLIX用AI自动发现视觉模型缺陷数据片段,支持专家交互验证。
VISLIX: An XAI Framework for Validating Vision Models with Slice Discovery and Analysis
- 基于大模型自动识别图像数据中的性能短板片段,无需额外标注
- 可生成自然语言分析报告,帮助理解模型在特定场景的失效原因
- 支持专家交互式假设验证,适用于自动驾驶等高风险领域
真实世界的机器学习模型在部署前需严格评估,尤其在自动驾驶、监控等安全关键领域。模型评估常聚焦于数据片段——具有特定特征的数据子集。自动发现数据片段可定位模型表现不佳的条件,助力开发者改进性能。然而,现有视觉模型评估中的数据切片方法面临多重挑战:一、依赖图像元数据或视觉概念,在目标检测等任务中效果有限;二、理解数据片段耗时耗力,高度依赖专家经验;三、缺乏人机协同机制,无法支持专家主动提出并验证假设。为此,我们提出VISLIX,一种新型可视化分析框架,利用前沿基础模型辅助领域专家分析计算机视觉模型的数据片段。该方法无需图像元数据或视觉概念,能自动生成自然语言洞察,并支持用户交互式验证假设。通过专家研究与三个应用场景验证,结果表明VISLIX在提供目标检测模型全面验证见解方面具有显著有效性。
原文摘要 · Abstract (English)
Real-world machine learning models require rigorous evaluation before deployment, especially in safety-critical domains like autonomous driving and surveillance. The evaluation of machine learning models often focuses on data slices, which are subsets of the data that share a set of characteristics. Data slice finding automatically identifies conditions or data subgroups where models underperform, aiding developers in mitigating performance issues. Despite its popularity and effectiveness, data slicing for vision model validation faces several challenges. First, data slicing often needs additional image metadata or visual concepts, and falls short in certain computer vision tasks, such as object detection. Second, understanding data slices is a labor-intensive and mentally demanding process that heavily relies on the expert's domain knowledge. Third, data slicing lacks a human-in-the-loop solution that allows experts to form hypothesis and test them interactively. To overcome these limitations and better support the machine learning operations lifecycle, we introduce VISLIX, a novel visual analytics framework that employs state-of-the-art foundation models to help domain experts analyze slices in computer vision models. Our approach does not require image metadata or visual concepts, automatically generates natural language insights, and allows users to test data slice hypothesis interactively. We evaluate VISLIX with an expert study and three use cases, that demonstrate the effectiveness of our tool in providing comprehensive insights for validating object detection models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。