让2D跨模态模型在线适应3D多物体场景的遮挡与特征区分。
Video and Language Alignment in 2D Systems for 3D Multi-object Scenes with Multi-Information Derivative-Free Control
- 用无导数优化最小化遗憾,提升多变量互信息估计
- 无需预训练/微调,实时适应遮挡与特征变化
- 适合需要动态相机控制的3D多物体视觉语言任务
在处理3D场景时,基于2D视觉输入的跨模态系统面临维度失配问题。通过引入场景内相机可缓解此问题,但需学习控制模块。本文提出一种新方法,通过无导数优化最小化遗憾来改进多变量互信息估计。该算法使现成的2D跨模态系统能够在线适应物体遮挡并区分特征。表达性度量与基于价值的优化相结合,辅助场景内相机直接从视觉语言模型的噪声输出中学习。所提流程在无需预训练或微调的前提下,显著提升了多物体3D场景中的跨模态任务性能。
原文摘要 · Abstract (English)
Cross-modal systems trained on 2D visual inputs are presented with a dimensional shift when processing 3D scenes. An in-scene camera bridges the dimensionality gap but requires learning a control module. We introduce a new method that improves multivariate mutual information estimates by regret minimisation with derivative-free optimisation. Our algorithm enables off-the-shelf cross-modal systems trained on 2D visual inputs to adapt online to object occlusions and differentiate features. The pairing of expressive measures and value-based optimisation assists control of an in-scene camera to learn directly from the noisy outputs of vision-language models. The resulting pipeline improves performance in cross-modal tasks on multi-object 3D scenes without resorting to pretraining or finetuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。