通过多模态细粒度推理提升帖子质量评估精度
Multimodal Fine-grained Reasoning for Post Quality Evaluation
- 将质量评估转为排序任务,融合图文信息捕捉细微差异
- 在艺术史数据集上相比最佳单模态方法提升9.52% NDCG@3
- 适合关注内容质量分析与多模态理解的研究者
准确评估帖子质量需复杂关系推理以捕捉细微的主题-帖子关联。现有研究存在三大局限:(1) 将任务视为单模态分类,忽视多模态线索与细粒度质量差异;(2) 深层多模态融合引入噪声,导致误导性信号;(3) 缺乏捕捉相关性与全面性等复杂语义关系的能力。为此,我们提出多模态细粒度主题-帖子关系推理(MFTRR)框架,模拟人类认知过程。MFTRR将帖子质量评估重构为排序任务,融合多模态数据以更好捕捉质量变化。其包含两个核心模块:(1) 局部-全局语义相关性推理模块,在局部与全局层面建模帖子与主题的细粒度语义交互,结合最大信息融合机制抑制噪声;(2) 多层级证据关系推理模块,探索宏观与微观层面的关系线索,强化基于证据的推理。我们在三个新构建的多模态主题-帖子数据集及公开的Lazada-Home数据集上评估了MFTRR。实验结果表明,MFTRR显著优于当前最优基线,在艺术史数据集上相较最佳单模态方法提升高达9.52% NDCG@3。
原文摘要 · Abstract (English)
Accurately assessing post quality requires complex relational reasoning to capture nuanced topic-post relationships. However, existing studies face three major limitations: (1) treating the task as unimodal categorization, which fails to leverage multimodal cues and fine-grained quality distinctions; (2) introducing noise during deep multimodal fusion, leading to misleading signals; and (3) lacking the ability to capture complex semantic relationships like relevance and comprehensiveness. To address these issues, we propose the Multimodal Fine-grained Topic-post Relational Reasoning (MFTRR) framework, which mimics human cognitive processes. MFTRR reframes post-quality assessment as a ranking task and incorporates multimodal data to better capture quality variations. It consists of two key modules: (1) the Local-Global Semantic Correlation Reasoning Module, which models fine-grained semantic interactions between posts and topics at both local and global levels, enhanced by a maximum information fusion mechanism to suppress noise; and (2) the Multi-Level Evidential Relational Reasoning Module, which explores macro- and micro-level relational cues to strengthen evidence-based reasoning. We evaluate MFTRR on three newly constructed multimodal topic-post datasets and the public Lazada-Home dataset. Experimental results demonstrate that MFTRR significantly outperforms state-of-the-art baselines, achieving up to 9.52% NDCG@3 improvement over the best unimodal method on the Art History dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。