提升视觉语言模型对高分辨率图像的处理能力,减少幻觉并增强文本密集任务表现。
VisualRWKV-HD and UHD: Advancing High-Resolution Processing for Visual Language Models
- 采用无损下采样技术融合高/低分辨率视觉编码器,不增加序列长度。
- 将图像分四段重组,融合高低分辨率特征,支持4096×4096像素输入。
- 在文档与文本密集任务中显著提升性能,适合高精度视觉理解场景。
准确理解复杂视觉信息对视觉语言模型(VLMs)至关重要。提升图像分辨率可增强视觉感知能力,不仅减少幻觉,还能提升文本密集或文档分析等高分辨率任务的表现。本文提出VisualRWKV-HD和VisualRWKV-UHD,两款专为处理高分辨率视觉输入设计的模型。VisualRWKV-HD采用无损下采样方法,有效整合高分辨率视觉编码器与低分辨率编码器,无需延长输入序列长度。VisualRWKV-UHD通过将图像划分为四个区域并重新组合,融合高/低分辨率特征,平衡粗粒度与细粒度信息。模型支持最高4096×4096像素输入,实现更细致全面的视觉处理。两者在多个VLM基准测试中表现优异,并在文本密集任务中显著提升性能。
原文摘要 · Abstract (English)
Accurately understanding complex visual information is crucial for visual language models (VLMs). Enhancing image resolution can improve visual perception capabilities, not only reducing hallucinations but also boosting performance in tasks that demand high resolution, such as text-rich or document analysis. In this paper, we present VisualRWKV-HD and VisualRWKV-UHD, two advancements in the VisualRWKV model family, specifically designed to process high-resolution visual inputs. For VisualRWKV-HD, we developed a lossless downsampling method to effectively integrate a high-resolution vision encoder with low-resolution encoders, without extending the input sequence length. For the VisualRWKV-UHD model, we enhanced image representation by dividing the image into four segments, which are then recombined with the original image. This technique allows the model to incorporate both high-resolution and low-resolution features, effectively balancing coarse and fine-grained information. As a result, the model supports resolutions up to 4096 x 4096 pixels, offering a more detailed and comprehensive visual processing capability. Both VisualRWKV-HD and VisualRWKV-UHD not only achieve strong results on VLM benchmarks but also show marked improvements in performance for text-rich tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。