arXiv:2507.17079cs.CV2025-07综述被引 3

少样本学习让视频与3D目标检测模型仅凭少量标注就能识别新物体,大幅减少人工标注成本。

Few-Shot Learning in Video and 3D Object Detection: A Survey

  • 利用时序结构传播信息,通过管状提议和时间匹配网络实现跨帧知识迁移
  • 在仅有1~2个样本情况下,3D检测模型仍可保持较高准确率,解决点云稀疏性难题
  • 适合自动驾驶等需低标注成本的真实场景,尤其适用于视频与三维感知系统

少样本学习(FSL)使目标检测模型仅需少数标注样本即可识别新类别,从而降低昂贵的手动标注成本。本综述系统分析了视频与3D目标检测中近期的少样本学习进展。在视频领域,由于跨帧标注更费力,利用帧间信息传播的技术(如管状提议、时间匹配网络)可从少数样本中高效挖掘时空结构。3D检测面临点云稀疏、纹理缺失等挑战,现有方法结合专用点云网络与针对类别不平衡设计的损失函数,实现小样本下的稳定检测。该技术显著减少了3D标注需求,推动自动驾驶等实际应用落地。核心挑战包括泛化与过拟合平衡、原型匹配机制设计及多模态数据特性适配。综上,少样本学习通过有效融合特征、时序与多模态信息,在降低监督需求的同时,为视频与3D等真实世界应用提供了可行路径。

原文摘要 · Abstract (English)

Few-shot learning (FSL) enables object detection models to recognize novel classes given only a few annotated examples, thereby reducing expensive manual data labeling. This survey examines recent FSL advances for video and 3D object detection. For video, FSL is especially valuable since annotating objects across frames is more laborious than for static images. By propagating information across frames, techniques like tube proposals and temporal matching networks can detect new classes from a couple examples, efficiently leveraging spatiotemporal structure. FSL for 3D detection from LiDAR or depth data faces challenges like sparsity and lack of texture. Solutions integrate FSL with specialized point cloud networks and losses tailored for class imbalance. Few-shot 3D detection enables practical autonomous driving deployment by minimizing costly 3D annotation needs. Core issues in both domains include balancing generalization and overfitting, integrating prototype matching, and handling data modality properties. In summary, FSL shows promise for reducing annotation requirements and enabling real-world video, 3D, and other applications by efficiently leveraging information across feature, temporal, and data modalities. By comprehensively surveying recent advancements, this paper illuminates FSL's potential to minimize supervision needs and enable deployment across video, 3D, and other real-world applications.

少样本学习3D检测视频理解自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。