arXiv:2411.01925cs.CV2024-11

利用视觉数据上下文不确定性,提升深度模型训练效率与泛化能力。

Exploiting Contextual Uncertainty of Visual Data for Efficient Training of Deep Models

  • 引入上下文多样性,用于主动学习,优化标注策略。
  • 提出数据修复算法,减少模型对典型场景的偏差。
  • 基于类别设计互补标注,增强模型在领域迁移下的表现。

真实世界中的物体很少孤立出现,其空间布局由功能需求和人机交互规律决定。例如椅子常靠近桌子,电脑通常放在上方。人类在识别模糊图像时会依赖此类上下文线索。深度神经网络同样可利用数据中的上下文信息学习表征。本文提出三方面贡献:(1) 引入上下文多样性(CDAL)于主动学习,在语义分割、目标检测和图像分类任务中均有效;(2) 提出数据修复算法,构建上下文公平的数据集,使模型能识别脱离典型上下文的物体;(3) 提出基于类别的标注方法,选择上下文相关且互补的类别,以应对领域偏移问题。强调人工参与的重要性,通过人机协同实现精准标注,并开发野生动物相机陷阱图像检索系统与乡村道路质量预警系统。大规模标注采用人类专家与零样本模型结合策略,并在各阶段集成人类反馈以持续优化。

原文摘要 · Abstract (English)

Objects, in the real world, rarely occur in isolation and exhibit typical arrangements governed by their independent utility, and their expected interaction with humans and other objects in the context. For example, a chair is expected near a table, and a computer is expected on top. Humans use this spatial context and relative placement as an important cue for visual recognition in case of ambiguities. Similar to human's, DNN's exploit contextual information from data to learn representations. Our research focuses on harnessing the contextual aspects of visual data to optimize data annotation and enhance the training of deep networks. Our contributions can be summarized as follows: (1) We introduce the notion of contextual diversity for active learning CDAL and show its applicability in three different visual tasks semantic segmentation, object detection and image classification, (2) We propose a data repair algorithm to curate contextually fair data to reduce model bias, enabling the model to detect objects out of their obvious context, (3) We propose Class-based annotation, where contextually relevant classes are selected that are complementary for model training under domain shift. Understanding the importance of well-curated data, we also emphasize the necessity of involving humans in the loop to achieve accurate annotations and to develop novel interaction strategies that allow humans to serve as fact-checkers. In line with this we are working on developing image retrieval system for wildlife camera trap images and reliable warning system for poor quality rural roads. For large-scale annotation, we are employing a strategic combination of human expertise and zero-shot models, while also integrating human input at various stages for continuous feedback.

主动学习上下文建模数据修复人机协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。