构建10语言多语种视觉指代理解数据集,实现高效跨语言物体定位。
Comprehension of Multilingual Expressions Referring to Target Objects in Visual Inputs
- 基于机器翻译与上下文增强构建10语言统一数据集
- 小模型下多语言表现稳定,跨语言泛化能力强
- 注意力锚点+残差优化,提升多语种视觉定位效率
指代表达理解(REC)要求模型根据自然语言描述在图像中定位目标物体。尽管已有显著进展,该领域研究仍以英语为主,难以满足全球化部署需求。本文通过两项贡献推进多语言REC:首先,通过机器翻译和上下文增强,系统扩展12个英文REC基准,构建覆盖10种语言的统一多语言数据集;其次,提出一种基于注意力锚点的高效神经架构,采用多语言SigLIP2编码器。该方法从注意力分布生成粗粒度空间锚点,并通过学习残差进行精炼。实验表明,即使使用较小模型,标准基准上仍具竞争力。多语言评估显示各语言间性能一致,验证了高效多语言视觉定位系统的可行性。
原文摘要 · Abstract (English)
Referring Expression Comprehension (REC) requires models to localize objects in images based on different types of natural language descriptions. Even with significant progress, research on the area remains predominantly English-centric, despite increasing global deployment demands. This work addresses multilingual REC through two main contributions. First, we construct a unified multilingual dataset spanning 10 languages, by systematically expanding 12 existing English REC benchmarks through machine translation and context-based translation enhancement. Second, we introduce an attention-anchored efficient neural architecture that uses a multilingual SigLIP2 encoder. Our attention-based approach generates coarse spatial anchors from attention distributions, which are subsequently refined through learned residuals. Experimental evaluation demonstrates competitive performance on standard benchmarks despite the use of a relatively small model. Multilingual evaluation shows consistent capabilities across languages, establishing the practical feasibility of efficient multilingual visual grounding systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。