arXiv:2412.10594cs.CVcs.LG2024-12被引 2

构建统一多模态感知评估基准,发现通用模型与专用模型各有优劣。

Towards Unified Benchmark and Models for Multi-Modal Perceptual Metrics

  • 提出7类25个数据集的统一评测基准UniSim-Bench。
  • 通用模型平均表现尚可,但专用模型在单项任务上更优。
  • 微调后通用模型性能提升,但仍难泛化到未见任务。

人类对单模态和多模态输入相似性的感知极为复杂,难以构建准确模拟该感知的自动化度量方法。通用视觉-语言模型(如CLIP)和大型多模态模型(LMMs)可作为零样本感知度量工具,近期也有研究开发了针对特定感知任务的专用模型。然而,现有感知度量与人类感知的契合程度仍不明确。为此,我们引入UniSim-Bench,一个涵盖7个多模态感知相似性任务、共25个数据集的基准。评估显示,尽管通用模型平均表现尚可,但在个别任务上常落后于专用模型;而针对特定任务微调的度量模型,在未见相关任务上泛化能力差。作为迈向统一多任务感知相似性度量的第一步,我们在部分UniSim-Bench任务上微调了编码器型和生成型视觉-语言模型,结果取得最高平均性能,某些任务甚至超越专用模型。然而,这些模型在未见任务上仍表现不佳,凸显学习鲁棒统一感知相似性度量以捕捉人类相似性概念的持续挑战。代码与模型见https://github.com/SaraGhazanfari/UniSim。

原文摘要 · Abstract (English)

Human perception of similarity across uni- and multimodal inputs is highly complex, making it challenging to develop automated metrics that accurately mimic it. General purpose vision-language models, such as CLIP and large multi-modal models (LMMs), can be applied as zero-shot perceptual metrics, and several recent works have developed models specialized in narrow perceptual tasks. However, the extent to which existing perceptual metrics align with human perception remains unclear. To investigate this question, we introduce UniSim-Bench, a benchmark encompassing 7 multi-modal perceptual similarity tasks, with a total of 25 datasets. Our evaluation reveals that while general-purpose models perform reasonably well on average, they often lag behind specialized models on individual tasks. Conversely, metrics fine-tuned for specific tasks fail to generalize well to unseen, though related, tasks. As a first step towards a unified multi-task perceptual similarity metric, we fine-tune both encoder-based and generative vision-language models on a subset of the UniSim-Bench tasks. This approach yields the highest average performance, and in some cases, even surpasses taskspecific models. Nevertheless, these models still struggle with generalization to unseen tasks, highlighting the ongoing challenge of learning a robust, unified perceptual similarity metric capable of capturing the human notion of similarity. The code and models are available at https://github.com/SaraGhazanfari/UniSim.

多模态感知评估基准测试模型泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。