arXiv:2505.06030cs.AIcs.CV2025-05中稿 · IJCNN 2025被引 1

通过反事实解释,找出模型误判3D物体识别的原因。

Why Are You Wrong? Counterfactual Explanations for Language Grounding with 3D Objects

  • 生成与原描述结构相似的反事实文本,揭示错误原因。
  • 在ShapeTalk数据集上验证,能有效暴露模型偏差和描述缺陷。
  • 帮助工程师优化模型,提升人机交互理解力。

将自然语言与几何形状结合是机器人和语言辅助设计中的新兴研究方向。核心任务是目标物体识别,即根据文本描述选择对应的3D物体。由于语言描述和空间关系的多样性,该任务复杂度高,亟需深入理解神经网络模型的行为。然而,当前研究仍有限:当模型在提供看似正确的描述时仍出现误判,从业者往往无法知晓原因。本文提出一种方法,通过生成反事实示例来回答“为什么错了”。该方法针对一个包含两个物体和一段文本描述的误分类样本,生成一个结构相近、语义合理但会导致正确预测的替代描述。我们在ShapeTalk数据集上使用三种不同模型进行了评估,结果显示生成的反事实描述保持原结构、语义连贯且有意义,能揭示描述中的弱点、模型偏差,并增进对模型行为的理解。这些洞察有助于从业者更有效地与系统交互,也帮助工程师改进模型性能。

原文摘要 · Abstract (English)

Combining natural language and geometric shapes is an emerging research area with multiple applications in robotics and language-assisted design. A crucial task in this domain is object referent identification, which involves selecting a 3D object given a textual description of the target. Variability in language descriptions and spatial relationships of 3D objects makes this a complex task, increasing the need to better understand the behavior of neural network models in this domain. However, limited research has been conducted in this area. Specifically, when a model makes an incorrect prediction despite being provided with a seemingly correct object description, practitioners are left wondering: "Why is the model wrong?". In this work, we present a method answering this question by generating counterfactual examples. Our method takes a misclassified sample, which includes two objects and a text description, and generates an alternative yet similar formulation that would have resulted in a correct prediction by the model. We have evaluated our approach with data from the ShapeTalk dataset along with three distinct models. Our counterfactual examples maintain the structure of the original description, are semantically similar and meaningful. They reveal weaknesses in the description, model bias and enhance the understanding of the models behavior. Theses insights help practitioners to better interact with systems as well as engineers to improve models.

语言接地3D物体反事实解释模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。