arXiv:2506.20279cs.CV2025-06被引 3

用生成模型先验实现真实场景下高效统一的密集预测

From Ideal to Real: Unified and Data-Efficient Dense Prediction for Real-World Scenarios

  • 基于生成模型视觉先验,统一处理25类真实任务
  • 仅需0.01%数据量,性能超越现有方法
  • 参数增量低于0.1%,适合资源受限部署

密集预测在计算机视觉中具有重要意义,旨在为输入图像学习像素级标注。尽管该领域取得进展,现有方法多聚焦理想化条件,难以泛化到真实场景,且实际应用中真实数据极度稀缺。为此,我们首先构建DenseWorld基准,涵盖25类对应紧迫现实应用的密集预测任务,并实现跨任务统一评估。随后提出DenseDiT,利用生成模型的视觉先验,通过统一策略执行多样真实场景下的密集预测任务。DenseDiT结合参数复用机制与两个轻量级分支,自适应融合多尺度上下文,实现小于0.1%额外参数的高效微调,在激活视觉先验的同时有效适配多样化任务。在DenseWorld上的评估显示,现有通用与专用基线性能显著下降,凸显其真实世界泛化能力有限。相比之下,DenseDiT仅使用基线0.01%的训练数据即达到更优结果,验证了其在真实部署中的实用价值。

原文摘要 · Abstract (English)

Dense prediction tasks hold significant importance of computer vision, aiming to learn pixel-wise annotated labels for input images. Despite advances in this field, existing methods primarily focus on idealized conditions, exhibiting limited real-world generalization and struggling with the acute scarcity of real-world data in practical scenarios. To systematically study this problem, we first introduce DenseWorld, a benchmark spanning a broad set of 25 dense prediction tasks that correspond to urgent real-world applications, featuring unified evaluation across tasks. We then propose DenseDiT, which exploits generative models' visual priors to perform diverse real-world dense prediction tasks through a unified strategy. DenseDiT combines a parameter-reuse mechanism and two lightweight branches that adaptively integrate multi-scale context. This design enables DenseDiT to achieve efficient tuning with less than 0.1% additional parameters, activating the visual priors while effectively adapting to diverse real-world dense prediction tasks. Evaluations on DenseWorld reveal significant performance drops in existing general and specialized baselines, highlighting their limited real-world generalization. In contrast, DenseDiT achieves superior results using less than 0.01% training data of baselines, underscoring its practical value for real-world deployment.

密集预测生成模型真实场景数据高效

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。