arXiv:2507.12841cs.CV2025-07被引 5

提出统一框架、数据集与评测标准,实现多模态可控图文生成

AnyCap Project: A Unified Framework, Dataset, and Benchmark for Controllable Omni-modal Captioning

  • 轻量级插件框架,无需重训练即可增强基模型可控性
  • 在30万条数据上提升内容准确率45%、风格契合度12%
  • 适配视频/图像/文本多模态,适合需要精准指令控制的研究者

可控图文生成对精确的多模态对齐和指令遵循至关重要,但现有模型常缺乏细粒度控制和可靠评估。为此,我们提出AnyCap项目,涵盖模型、数据集与评测体系。引入AnyCapModel(ACM),一种轻量级即插即用框架,可在不重训练基础模型的前提下,通过融合用户指令与模态特征,提升多模态生成的可控性。为缓解可控多模态数据稀缺问题,构建AnyCapDataset(ACD),覆盖三种模态、28类用户指令,含30万条高质量数据。进一步提出AnyCapEval,通过解耦内容准确性和风格一致性,提供更可靠的评估指标。ACM在多种基模型上显著提升性能,在AnyCapEval中使GPT-4o的內容得分提高45%,风格得分提升12%,并在MIA-Bench和VidCapBench等主流基准上取得显著进步。

原文摘要 · Abstract (English)

Controllable captioning is essential for precise multimodal alignment and instruction following, yet existing models often lack fine-grained control and reliable evaluation protocols. To address this gap, we present the AnyCap Project, an integrated solution spanning model, dataset, and evaluation. We introduce AnyCapModel (ACM), a lightweight plug-and-play framework that enhances the controllability of existing foundation models for omni-modal captioning without retraining the base model. ACM reuses the original captions from base models while incorporating user instructions and modality features to generate improved captions. To remedy the data scarcity in controllable multimodal captioning, we build AnyCapDataset (ACD), covering three modalities, 28 user-instruction types, and 300\,k high-quality data entries. We further propose AnyCapEval, a new benchmark that provides more reliable evaluation metrics for controllable captioning by decoupling content accuracy and stylistic fidelity. ACM markedly improves caption quality across a diverse set of base models on AnyCapEval. Notably, ACM-8B raises GPT-4oś content scores by 45\% and style scores by 12\%, and it also achieves substantial gains on widely used benchmarks such as MIA-Bench and VidCapBench.

可控生成多模态评测基准指令遵循

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。