arXiv:2410.19974cs.LGcs.CL2024-10被引 1

对比大模型在多模态搜索中的成本与准确率,发现小模型加视觉反而更差。

Evaluating Cost-Accuracy Trade-offs in Multimodal Search Relevance Judgements

  • 测试多种大模型和多模态模型在多场景下的判断一致性
  • 发现模型表现受场景影响大,小模型加视觉反而降低准确率
  • 为实际应用选型提供成本-精度权衡参考

大型语言模型(LLMs)展现出作为搜索相关性评估工具的潜力,但缺乏在不同情境下持续表现最优的模型选择指南。本文评估了多个LLM和多模态语言模型(MLLMs)在多种多模态搜索场景中与人类判断的一致性。分析揭示了成本与准确性之间的权衡,表明模型性能随上下文显著变化。有趣的是,在小型模型中,引入视觉组件可能反而损害性能而非提升。这些发现凸显了在实际应用中选择合适模型的复杂性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated potential as effective search relevance evaluators. However, there is a lack of comprehensive guidance on which models consistently perform optimally across various contexts or within specific use cases. In this paper, we assess several LLMs and Multimodal Language Models (MLLMs) in terms of their alignment with human judgments across multiple multimodal search scenarios. Our analysis investigates the trade-offs between cost and accuracy, highlighting that model performance varies significantly depending on the context. Interestingly, in smaller models, the inclusion of a visual component may hinder performance rather than enhance it. These findings highlight the complexities involved in selecting the most appropriate model for practical applications.

多模态搜索大模型评估成本效益

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。