用大模型让图像压缩自动理解用户需求,告别手动调参。
Comp-X: On Defining an Interactive Learned Image Compression Paradigm With Expert-driven LLM Agent
- 引入多功能编码框架,统一感知、可变码率等不同需求
- 通过专家反馈增强上下文学习,让大模型精准理解压缩指令
- 构建首个交互式压缩评测基准,适合非专业用户使用
我们提出Comp-X,首个由大语言模型代理驱动的智能交互式图像压缩范式。传统编码器通常编码模式有限且依赖工程师手动选择,对非专业用户不友好。为此,我们推动编码范式演进,提出三大创新:(i) 多功能编码框架,将人类-机器感知、可变编码、空间比特分配等多种目标/需求统一于一个框架;(ii) 交互式编码代理,采用增强的上下文学习方法结合编码专家反馈,教会大模型理解编码请求、模式选择与工具使用;(iii) IIC-bench,首个专用评测基准,包含多样化的用户请求及编码专家标注,系统化评估智能交互式图像压缩。大量实验表明,Comp-X能高效理解编码请求,具备出色的文本交互能力,同时在单一框架下仍保持可比压缩性能,为图像压缩领域的通用人工智能(AGI)提供了有前景的方向。
原文摘要 · Abstract (English)
We present Comp-X, the first intelligently interactive image compression paradigm empowered by the impressive reasoning capability of large language model (LLM) agent. Notably, commonly used image codecs usually suffer from limited coding modes and rely on manual mode selection by engineers, making them unfriendly for unprofessional users. To overcome this, we advance the evolution of image coding paradigm by introducing three key innovations: (i) multi-functional coding framework, which unifies different coding modes of various objective/requirements, including human-machine perception, variable coding, and spatial bit allocation, into one framework. (ii) interactive coding agent, where we propose an augmented in-context learning method with coding expert feedback to teach the LLM agent how to understand the coding request, mode selection, and the use of the coding tools. (iii) IIC-bench, the first dedicated benchmark comprising diverse user requests and the corresponding annotations from coding experts, which is systematically designed for intelligently interactive image compression evaluation. Extensive experimental results demonstrate that our proposed Comp-X can understand the coding requests efficiently and achieve impressive textual interaction capability. Meanwhile, it can maintain comparable compression performance even with a single coding framework, providing a promising avenue for artificial general intelligence (AGI) in image compression.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。