arXiv:2411.18764cs.CVcs.AI2024-11被引 1

CoVis通过双层分割+大模型,让图像理解更全面高效。

CoVis: A Collaborative Framework for Fine-grained Graphic Visual Understanding

  • 双层分割网络+大模型生成,协同提取图像细粒度信息。
  • 32人实验显示,比现有方法特征提取更准,描述更完整。
  • 适合需要深度图文分析的科研、设计与教育场景。

图形化视觉内容有助于信息传播与创意启发。然而当前视觉内容的理解主要依赖个人知识背景,影响信息获取的质量与效率。为提升视觉信息传递质量并突破观察者的信息茧房限制,我们提出CoVis,一种细粒度图形视觉理解的协作框架。通过设计并实现级联双层分割网络与基于大语言模型(LLM)的内容生成器,该框架从图像中尽可能提取知识,并生成可视化分析结果,帮助观察者从更整体视角理解图像。基于32名参与者的定量与定性实验表明,CoVis在特征提取方面优于现有方法,生成的视觉描述比通用大模型更全面、更详细。

原文摘要 · Abstract (English)

Graphic visual content helps in promoting information communication and inspiration divergence. However, the interpretation of visual content currently relies mainly on humans' personal knowledge background, thereby affecting the quality and efficiency of information acquisition and understanding. To improve the quality and efficiency of visual information transmission and avoid the limitation of the observer due to the information cocoon, we propose CoVis, a collaborative framework for fine-grained visual understanding. By designing and implementing a cascaded dual-layer segmentation network coupled with a large-language-model (LLM) based content generator, the framework extracts as much knowledge as possible from an image. Then, it generates visual analytics for images, assisting observers in comprehending imagery from a more holistic perspective. Quantitative experiments and qualitative experiments based on 32 human participants indicate that the CoVis has better performance than current methods in feature extraction and can generate more comprehensive and detailed visual descriptions than current general-purpose large models.

图像理解多模态大模型视觉分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。