用语义聚类解决动作识别评估中的动词歧义问题
Towards Robust Evaluation of Visual Activity Recognition: Resolving Verb Ambiguity with Sense Clustering
- 通过视觉语言聚类构建动词意义分组,捕捉同一图像的多重合理描述
- imSitu数据集上每张图平均对应4个语义簇,体现不同观察视角
- 评估结果更贴近人类判断,适合追求精准评估的动作识别研究
视觉活动识别系统的评估因动词语义和图像解释的固有歧义而困难。在描述图像中的动作时,同义动词可能指代同一事件(如 brushing vs. grooming),而不同视角可能导致同样合理但不同的动词选择(如 piloting vs. operating)。标准的精确匹配评估依赖单一标准答案,无法捕捉这些歧义,导致模型性能评估不完整。为此,我们提出一种视觉语言聚类框架,构建动词意义簇,实现更稳健的评估。对 imSitu 数据集的分析表明,每张图像平均映射到约四个意义簇,每个簇代表图像的一种不同视角。我们评估了多个活动识别模型,并将基于簇的评估与标准方法对比。此外,人类一致性分析表明,基于簇的评估更符合人类判断,提供更细致的模型性能评估。
原文摘要 · Abstract (English)
Evaluating visual activity recognition systems is challenging due to inherent ambiguities in verb semantics and image interpretation. When describing actions in images, synonymous verbs can refer to the same event (e.g., brushing vs. grooming), while different perspectives can lead to equally valid but distinct verb choices (e.g., piloting vs. operating). Standard exact-match evaluation, which relies on a single gold answer, fails to capture these ambiguities, resulting in an incomplete assessment of model performance. To address this, we propose a vision-language clustering framework that constructs verb sense clusters, providing a more robust evaluation. Our analysis of the imSitu dataset shows that each image maps to around four sense clusters, with each cluster representing a distinct perspective of the image. We evaluate multiple activity recognition models and compare our cluster-based evaluation with standard evaluation methods. Additionally, our human alignment analysis suggests that the cluster-based evaluation better aligns with human judgments, offering a more nuanced assessment of model performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。