用线性上下文学习从视频中提取像素级动态特征,提升多任务表现。
Featurising Pixels from Dynamic 3D Scenes with Linear In-Context Learners

- 基于线性上下文学习,融合深度与运动线索训练像素特征。
- 在视频分割、法向量估计等任务上显著优于基线方法。
- 适合需要时空一致像素表示的视觉任务研究者使用。
视觉模型在像素级推理方面展现出巨大潜力。尽管存在大量视觉基础模型,我们仍缺乏有效编码视觉场景像素级时空特性的表示方法。现有框架或基于图像预训练任务,忽略动态元素;或基于视频序列进行动作级推理,难以扩展至密集像素级预测。本文提出LILA框架,通过线性上下文学习从视频中学习像素级特征描述符。其核心是利用现成网络估算的时空线索图(深度与运动),即使这些线索存在噪声,仍可在未经清理的视频数据集上有效训练,实现语义与几何属性的时序一致性嵌入。我们在视频对象分割、表面法向量估计和语义分割等多个任务上验证了所学表示的优异性能。
原文摘要 · Abstract (English)
One of the most exciting applications of vision models involve pixel-level reasoning. Despite the abundance of vision foundation models, we still lack representations that effectively embed spatio-temporal properties of visual scenes at the pixel level. Existing frameworks either train on image-based pretext tasks, which do not account for dynamic elements, or on video sequences for action-level reasoning, which does not scale to dense pixel-level prediction. We present a framework that learns pixel-accurate feature descriptors from videos, LILA. The core element of our training framework is linear in-context learning. LILA leverages spatio-temporal cue maps -- depth and motion -- estimated with off-the-shelf networks. Despite the noisy nature of those cues, LILA trains effectively on uncurated video datasets, embedding semantic and geometric properties in a temporally consistent manner. We demonstrate compelling empirical benefits of the learned representation across a diverse suite of vision tasks: video object segmentation, surface normal estimation and semantic segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。