让普通用户也能生成专业级图像描述,自动理解并执行复杂指令。
From Simple to Professional: A Combinatorial Controllable Image Captioning Agent
- 用多模态大模型+外部工具,把简单指令转为专业描述。
- 支持情感、关键词、焦点等多维度精准控制,输出可定制化。
- 全程透明展示推理与工具使用,适合需要可控生成的场景。
可控图像描述代理(CapAgent)是一种创新系统,旨在弥合用户操作简易性与专业级输出之间的差距。CapAgent能自动将用户提供的简单指令转化为详细、专业的指令,实现精确且上下文感知的图像描述生成。该系统利用多模态大语言模型(MLLMs)和外部工具(如目标检测工具、搜索引擎),确保生成的描述符合指定的语感、关键词、关注点及格式要求。CapAgent在每一步都透明地控制生成过程,并展示其推理逻辑与工具调用,增强用户信任与参与感。项目代码已公开于 https://github.com/xin-ran-w/CapAgent。
原文摘要 · Abstract (English)
The Controllable Image Captioning Agent (CapAgent) is an innovative system designed to bridge the gap between user simplicity and professional-level outputs in image captioning tasks. CapAgent automatically transforms user-provided simple instructions into detailed, professional instructions, enabling precise and context-aware caption generation. By leveraging multimodal large language models (MLLMs) and external tools such as object detection tool and search engines, the system ensures that captions adhere to specified guidelines, including sentiment, keywords, focus, and formatting. CapAgent transparently controls each step of the captioning process, and showcases its reasoning and tool usage at every step, fostering user trust and engagement. The project code is available at https://github.com/xin-ran-w/CapAgent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。