基于用户个性的多模态立场检测模型,让机器更懂人的态度表达
PRISM of Opinions: A Persona-Reasoned Multimodal Framework for User-centric Conversational Stance Detection
- 从历史发帖中构建用户长期人格画像,捕捉个体差异
- 在6个真实话题上实现94.2%准确率,优于现有方法
- 适合关注社交媒体情感分析与个性化对话系统的研究者
多模态社交内容的爆发推动了多模态对话立场检测(MCSD)研究,旨在解析用户在复杂讨论中对特定目标的态度。然而,现有研究受限于:1)伪多模态——视觉信息仅存在于源帖子,评论被视为纯文本,与真实多模态互动脱节;2)用户同质化——忽略影响立场表达的个人特质。为此,我们提出首个以用户为中心的MCSD数据集U-MStance,包含超过4万条标注评论,覆盖六个真实议题。进一步提出PRISM模型,通过分析历史帖子与评论生成用户长期人格画像,再利用思维链对齐对话上下文中的文本与视觉线索,弥合跨模态语义与语用差距。最后引入双向任务强化机制,联合优化立场检测与立场感知回复生成,实现知识双向迁移。在U-MStance上的实验表明,PRISM显著超越强基线,验证了以用户为中心、上下文驱动的多模态推理在真实立场理解中的有效性。
原文摘要 · Abstract (English)
The rapid proliferation of multimodal social media content has driven research in Multimodal Conversational Stance Detection (MCSD), which aims to interpret users' attitudes toward specific targets within complex discussions. However, existing studies remain limited by: **1) pseudo-multimodality**, where visual cues appear only in source posts while comments are treated as text-only, misaligning with real-world multimodal interactions; and **2) user homogeneity**, where diverse users are treated uniformly, neglecting personal traits that shape stance expression. To address these issues, we introduce **U-MStance**, the first user-centric MCSD dataset, containing over 40k annotated comments across six real-world targets. We further propose **PRISM**, a **P**ersona-**R**easoned mult**I**modal **S**tance **M**odel for MCSD. PRISM first derives longitudinal user personas from historical posts and comments to capture individual traits, then aligns textual and visual cues within conversational context via Chain-of-Thought to bridge semantic and pragmatic gaps across modalities. Finally, a mutual task reinforcement mechanism is employed to jointly optimize stance detection and stance-aware response generation for bidirectional knowledge transfer. Experiments on U-MStance demonstrate that PRISM yields significant gains over strong baselines, underscoring the effectiveness of user-centric and context-grounded multimodal reasoning for realistic stance understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。