用语言描述辅助拼图,提升破损缝隙下的拼接准确率。
VLHSA: Vision-Language Hierarchical Semantic Alignment for Jigsaw Puzzle Solving with Eroded Gaps
- 通过多层级语义对齐,将图像块与文本描述匹配。
- 在多种数据集上提升拼图准确率,最高达14.2个百分点。
- 适合研究多模态融合或视觉推理的学者参考。
拼图求解在计算机视觉中仍具挑战性,需同时理解局部碎片细节与全局空间关系。传统方法仅依赖边缘匹配和视觉连贯性等视觉线索,较少利用自然语言描述在复杂场景(尤其是破损缝隙拼图)中的语义引导。本文提出一种视觉-语言框架,借助文本上下文增强拼图组装性能。核心为视觉-语言分层语义对齐(VLHSA)模块,通过从局部标记到全局上下文的多层级语义匹配,实现图像块与文本描述的对齐。该模块整合双视觉编码器与语言特征,支持跨模态推理。实验表明,本方法在多个数据集上显著优于现有先进模型,拼图准确率提升达14.2个百分点。消融实验验证了VLHSA模块在超越纯视觉方法中的关键作用。本工作为拼图求解建立了融合多模态语义的新范式。
原文摘要 · Abstract (English)
Jigsaw puzzle solving remains challenging in computer vision, requiring an understanding of both local fragment details and global spatial relationships. While most traditional approaches only focus on visual cues like edge matching and visual coherence, few methods explore natural language descriptions for semantic guidance in challenging scenarios, especially for eroded gap puzzles. We propose a vision-language framework that leverages textual context to enhance puzzle assembly performance. Our approach centers on the Vision-Language Hierarchical Semantic Alignment (VLHSA) module, which aligns visual patches with textual descriptions through multi-level semantic matching from local tokens to global context. Also, a multimodal architecture that combines dual visual encoders with language features for cross-modal reasoning is integrated into this module. Experiments demonstrate that our method significantly outperforms state-of-the-art models across various datasets, achieving substantial improvements, including a 14.2 percentage point gain in piece accuracy. Ablation studies confirm the critical role of the VLHSA module in driving improvements over vision-only approaches. Our work establishes a new paradigm for jigsaw puzzle solving by incorporating multimodal semantic insights.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。