arXiv:2509.04480cs.CLcs.LG2025-09

用离散提示调优提升个性化视觉情绪识别准确率

Discrete Prompt Tuning via Recursive Utilization of Black-box Multimodal Large Language Model for Personalized Visual Emotion Recognition

  • 通过模拟人类提示工程,选择最优自然语言表示更新提示
  • 在个性化情绪识别任务上优于传统MLLM方法,提升显著
  • 适合需要个体化情绪分析的场景,如广告定制与情感计算

视觉情绪识别(VER)因其在舆论挖掘和广告设计等领域的广泛应用而备受关注。将该能力扩展至个体层面,可进一步拓展其实际应用潜力。近年来,多模态大语言模型(MLLMs)受到越来越多关注,并展现出与传统VER方法相当的性能。然而,由于训练数据涵盖广泛且多样化的通用观点,这些模型倾向于偏好多数观点和常见模式,限制了其在个性化情绪识别中的表现,而这正是实际应用的关键需求。为解决此问题,本文提出一种受人类提示工程启发的离散提示调优方法,针对每位个体进行适配。该方法从生成的多个提示中筛选出最优自然语言表示,并用于更新提示,实现精准的个性化情绪识别。

原文摘要 · Abstract (English)

Visual Emotion Recognition (VER) is an important research topic due to its wide range of applications, including opinion mining and advertisement design. Extending this capability to recognize emotions at the individual level further broadens its potential applications. Recently, Multimodal Large Language Models (MLLMs) have attracted increasing attention and demonstrated performance comparable to that of conventional VER methods. However, MLLMs are trained on large and diverse datasets containing general opinions, which causes them to favor majority viewpoints and familiar patterns. This tendency limits their performance in a personalized VER, which is crucial for practical and real-world applications, and indicates a key area for improvement. To address this limitation, the proposed method employs discrete prompt tuning inspired by the process of humans' prompt engineering to adapt the VER task to each individual. Our method selects the best natural language representation from the generated prompts and uses it to update the prompt for the realization of accurate personalized VER.

视觉情绪识别提示调优个性化MLLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。