arXiv:2603.10538cs.CV2026-03中稿 · CVPR

DSFlash实现实时全景场景图生成,速度快且资源占用低。

DSFlash: Comprehensive Panoptic Scene Graph Generation in Realtime

  • 采用轻量设计实现全景场景图生成,兼顾全面关系与高速推理。
  • 在RTX 3090上达56帧/秒,性能媲美现有顶尖方法。
  • 仅需单块旧GTX 1080训练不足24小时,适合资源受限团队使用。

场景图生成(SGG)旨在从图像中提取详细的图结构,作为复杂下游任务(如具身智能体推理)的稳健中间表示。然而,在实际部署中——尤其是在计算资源受限的边缘设备上——速度与资源效率是关键挑战,而现有研究对此关注有限。为此,我们提出DSFlash,一种低延迟的全景场景图生成模型,以克服上述限制。DSFlash可在标准RTX 3090 GPU上以56帧/秒的速度处理视频流,且性能不逊于现有最先进方法。重要的是,与以往仅关注显著关系的方法不同,DSFlash生成全面的场景图,提供更丰富的上下文信息,同时保持优异的延迟表现。此外,DSFlash资源消耗低,仅需不到24小时即可在单块九年前的GTX 1080 GPU上完成训练。这一可及性使其特别适合计算资源有限的研究者和实践者,便于他们为特定应用定制和微调SGG模型。

原文摘要 · Abstract (English)

Scene Graph Generation (SGG) aims to extract a detailed graph structure from an image, a representation that holds significant promise as a robust intermediate step for complex downstream tasks like reasoning for embodied agents. However, practical deployment in real-world applications - especially on resource constrained edge devices - requires speed and resource efficiency, challenges that have received limited attention in existing research. To bridge this gap, we introduce DSFlash, a low-latency model for panoptic scene graph generation designed to overcome these limitations. DSFlash can process a video stream at 56 frames per second on a standard RTX 3090 GPU, without compromising performance against existing state-of-the-art methods. Crucially, unlike prior approaches that often restrict themselves to salient relationships, DSFlash computes comprehensive scene graphs, offering richer contextual information while maintaining its superior latency. Furthermore, DSFlash is light on resources, requiring less than 24 hours to train on a single, nine-year-old GTX 1080 GPU. This accessibility makes DSFlash particularly well-suited for researchers and practitioners operating with limited computational resources, empowering them to adapt and fine-tune SGG models for specialized applications.

场景图生成实时推理边缘计算轻量化模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。