仅用相机运动轨迹就能感知视频内容,无需看到像素。
Seeing without Pixels: Perception from Camera Trajectories
- 用对比学习训练模型,将相机轨迹映射到语义空间。
- 轨迹信息能准确还原视频动作与场景,效果媲美视觉输入。
- 适合无摄像头图像或低算力场景,如移动设备、机器人感知。
能否仅凭相机运动轨迹(即其在空间中移动的路径)来感知视频内容,而无需看到像素?本文首次系统性地探究这一看似不可能的问题。为此,我们提出一种对比学习框架,训练出专用于编码相机姿态轨迹的CamFormer模型,将其投影至与自然语言对齐的联合嵌入空间。研究发现,尽管表面简单,相机轨迹却是揭示视频内容的极富信息量信号。换句话说,“你怎么动”确实能反映“你在做什么”(第一人称视角)或“你观察什么”(第三人称视角)。我们在多种下游任务上验证了所学的CamFormer嵌入表示的通用性,涵盖跨模态对齐、分类与时间分析等。重要的是,这些表示对不同相机姿态估计方法均保持鲁棒性,包括高精度多传感器与标准单目RGB估计算法。研究结果确立了相机轨迹作为轻量、鲁棒且多功能的视频感知新模态。
原文摘要 · Abstract (English)
Can one perceive a video's content without seeing its pixels, just from the camera trajectory-the path it carves through space? This paper is the first to systematically investigate this seemingly implausible question. Towards this end, we propose a contrastive learning framework to train CamFormer, a dedicated encoder that projects camera pose trajectories into a joint embedding space, aligning them with natural language. We find that, contrary to its apparent simplicity, the camera trajectory is a remarkably informative signal to uncover video content. In other words, "how you move" can indeed provide valuable cues about "what you are doing" (egocentric) or "observing" (exocentric). We demonstrate the versatility of our learned CamFormer embeddings on a diverse suite of downstream tasks, ranging from cross-modal alignment to classification and temporal analysis. Importantly, our representations are robust across diverse camera pose estimation methods, including both high-fidelity multi-sensored and standard RGB-only estimators. Our findings establish camera trajectory as a lightweight, robust, and versatile modality for perceiving video content.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。