arXiv:2409.13941cs.CVcs.AI2024-09

用汽车拼成动物图像,互动问答揭示环保车信息

TalkMosaic: Interactive PhotoMosaic with Multi-modal LLM Q&A Interactions

  • 用多模态大模型实现图片点击即查车信息
  • 通过稀疏注意力与量化加速模型推理,提升响应速度
  • 适合环保主题艺术创作与智能交互设计人群

我们使用多种车型图像拼合出鸟类或狮子等动物形象,以环保为主题,在单张合成图像中最大化呈现车辆信息,提升公众对环境挑战的认知。提出一种新型图像交互方式:通过简单点击,可在拼贴图的局部图像与原始汽车图像间实时切换,并自动保存至桌面。构建了融合汽车知识与背景的定制化多模态GPT模型TalkMosaic,用户上传汽车图像后,可高效获取如‘如何购买符合高环保标准的轮胎’等精准问答。深入分析并采用概率性FlashAttention(PrFlashAttention)和阶梯自适应量化(SAQ)技术,显著加速多模态大模型推理。原型系统验证了该方法的可行性与有效性。

原文摘要 · Abstract (English)

We use images of cars of a wide range of varieties to compose an image of an animal such as a bird or a lion for the theme of environmental protection to maximize the information about cars in a single composed image and to raise the awareness about environmental challenges. We present a novel way of image interaction with an artistically-composed photomosaic image, in which a simple operation of "click and display" is used to demonstrate the interactive switch between a tile image in a photomosaic image and the corresponding original car image, which will be automatically saved on the Desktop. We build a multimodal custom GPT named TalkMosaic by incorporating car images information and the related knowledge to ChatGPT. By uploading the original car image to TalkMosaic, we can ask questions about the given car image and get the corresponding answers efficiently and effectively such as where to buy the tire in the car image that satisfies high environmental standards. We give an in-depth analysis on how to speed up the inference of multimodal LLM using sparse attention and quantization techniques with presented probabilistic FlashAttention (PrFlashAttention) and Staircase Adaptive Quantization (SAQ) methods. The implemented prototype demonstrates the feasibility and effectiveness of the presented approach.

图像生成多模态交互设计环保应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。