arXiv:2512.11167cs.CVcs.AI2025-12中稿 · AAAI

复现并改进高分辨率图像分块推理方法,验证局部细节恢复效果

Image Tiling for High-Resolution Reasoning: Balancing Local Detail with Global Context

  • 采用图像分块策略,在保持计算效率的同时还原细粒度视觉信息
  • 分块能有效提升局部细节识别,但全局上下文对结果影响显著
  • 适合关注高分辨率多模态模型设计的研究者参考

可复现性是科学进步的核心,但复杂的多模态模型常缺乏透明的实现细节和可访问的训练基础设施。本文详细复现并批判性分析了发表于CVPR24的猴子视觉语言模型(Monkey VLM)(Li et al. 2023b),该方法通过图像分块实现高分辨率图像理解。原论文提出将大图像分割为小块以恢复细粒度视觉细节,同时保持计算效率。本研究使用开源检查点重新实现训练流程,确认分块策略能有效恢复局部细节。进一步探究全局上下文引入的影响,获得对未来高分辨率多模态建模的实用洞见。然而,实验结果存在偏差,其幅度高度依赖任务类型与分块粒度。

原文摘要 · Abstract (English)

Reproducibility remains a cornerstone of scientific progress, yet complex multimodal models often lack transparent implementation details and accessible training infrastructure. In this work, we present a detailed reproduction and critical analysis of the Monkey Vision-Language Model (VLM) (Li et al. 2023b) published in CVPR24, a recent approach to high-resolution image understanding via image tiling. The original paper proposed splitting large images into tiles to recover fine-grained visual details while maintaining computational efficiency. Our study replicates this strategy using open checkpoints and reimplements the training pipeline. We confirm the key finding of the original Monkey VLM work, namely that tiling effectively recovers local details. We then extend this work further, by investigating the effect of the inclusion of the global context, which provide practical insights for future high-resolution multimodal modeling. However, we also report deviations in the results, with the magnitude of these effects depending heavily on task type and tile granularity.

图像分块多模态高分辨率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。