评测视觉语言模型在家具组装图与视频对齐中的表现,发现视觉编码是关键瓶颈。
Benchmarking and Mechanistic Analysis of Vision-Language Models for Cross-Depiction Assembly Instruction Alignment

- 构建包含29款宜家家具的1623个问题基准测试集
- 文本能帮助理解指令但削弱图-视频对齐效果
- 模型架构比参数量更能决定对齐准确率
二维组装图通常抽象难懂,亟需智能助手实时监测进度、识别错误并提供步骤指导。在混合现实场景中,系统需从摄像头画面中识别已完成和进行中的步骤,并与图纸指令对齐。视觉语言模型(VLMs)在此任务中展现潜力,但面临描绘差异难题——组装图与视频帧间共享视觉特征极少。为系统评估该差距,我们构建了IKEA-Bench,涵盖29款宜家家具的6类任务共1623个问题,并评估19个不同规模(2B-38B参数)的VLMs在三种对齐策略下的表现。主要发现:(1) 通过文本可恢复组装指令理解,但文本同时降低图-视频对齐性能;(2) 模型架构家族比参数量更显著影响对齐准确率;(3) 视频理解仍是核心瓶颈,策略改进无法突破。三层次机制分析揭示:图与视频占据ViT网络中不相交的子空间,添加文本促使模型从视觉驱动转向文本驱动推理。结果表明,视觉编码是提升跨描绘鲁棒性的首要优化目标。
原文摘要 · Abstract (English)
2D assembly diagrams are often abstract and hard to follow, creating a need for intelligent assistants that can monitor progress, detect errors, and provide step-by-step guidance. In mixed reality settings, such systems must recognize completed and ongoing steps from the camera feed and align them with the diagram instructions. Vision Language Models (VLMs) show promise for this task, but face a depiction gap because assembly diagrams and video frames share few visual features. To systematically assess this gap, we construct IKEA-Bench, a benchmark of 1,623 questions across 6 task types on 29 IKEA furniture products, and evaluate 19 VLMs (2B-38B) under three alignment strategies. Our key findings: (1) assembly instruction understanding is recoverable via text, but text simultaneously degrades diagram-to-video alignment; (2) architecture family predicts alignment accuracy more strongly than parameter count; (3) video understanding remains a hard bottleneck unaffected by strategy. A three-level mechanistic analysis further reveals that diagrams and video occupy disjoint ViT subspaces, and that adding text shifts models from visual to text-driven reasoning. These results identify visual encoding as the primary target for improving cross-depiction robustness. Project page: https://ryenhails.github.io/IKEA-Bench/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。