arXiv:2505.24211cs.CL2025-05ACL被引 1

对比通用与专用模型在跨模态转移中的一致性表现

Are Any-to-Any Models More Consistent Across Modality Transfers Than Specialists?

  • 构建包含1000张图像的ACON数据集,评估多模态转换一致性
  • 点对点测试显示通用模型一致性不优于专用模型
  • 通过潜在空间分析发现弱但可检测的结构一致性

任何到任何的生成模型旨在统一框架下实现多模态间的无缝理解与生成,但其跨模态关系保持能力仍存疑。为探究此问题,我们构建了ACON数据集,包含1000张图像(500张新贡献),配以标题、编辑指令和问答对,用于严格评估跨模态转移。采用循环一致性、前向等变性和共轭等变性三个标准进行实验,结果表明:在点对点评估中,任何到任何模型并未表现出比专用模型更强的跨模态一致性。然而,在等变性评估中,通过对中间潜在空间进行多步编辑操作的结构化分析,发现了弱但可观测的一致性。代码与数据已开源。

原文摘要 · Abstract (English)

Any-to-any generative models aim to enable seamless interpretation and generation across multiple modalities within a unified framework, yet their ability to preserve relationships across modalities remains uncertain. Do unified models truly achieve cross-modal coherence, or is this coherence merely perceived? To explore this, we introduce ACON, a dataset of 1,000 images (500 newly contributed) paired with captions, editing instructions, and Q&A pairs to evaluate cross-modal transfers rigorously. Using three consistency criteria-cyclic consistency, forward equivariance, and conjugated equivariance-our experiments reveal that any-to-any models do not consistently demonstrate greater cross-modal consistency than specialized models in pointwise evaluations such as cyclic consistency. However, equivariance evaluations uncover weak but observable consistency through structured analyses of the intermediate latent space enabled by multiple editing operations. We release our code and data at https://github.com/JiwanChung/ACON.

跨模态一致性生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。