用语言和图片定位室内位置,模型比人还准。
The Wallpaper is Ugly: Indoor Localization using Vision and Language
- 用预训练视觉语言模型计算文本与图像的相似度
- 在未见环境、文本、图像上仍能准确定位
- 微调后的CLIP模型超越人类定位能力
我们研究如何利用自然语言查询和环境图像,在已知地图的室内环境中定位用户。基于近期预训练的视觉-语言模型,我们学习文本描述与环境图像之间的相似度评分,通过该评分识别最匹配语言查询的位置,从而估计用户位置。该方法可在训练中未见过的环境、文本和图像上实现定位。一个经过微调的CLIP模型在我们的评估中表现优于人类。
原文摘要 · Abstract (English)
We study the task of locating a user in a mapped indoor environment using natural language queries and images from the environment. Building on recent pretrained vision-language models, we learn a similarity score between text descriptions and images of locations in the environment. This score allows us to identify locations that best match the language query, estimating the user's location. Our approach is capable of localizing on environments, text, and images that were not seen during training. One model, finetuned CLIP, outperformed humans in our evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。