多智能体框架提升新闻图文分类准确率与可解释性
MultiPress: A Multi-Agent Framework for Interpretable Multimodal News Classification
- 分三阶段部署感知、检索推理、融合评分智能体
- 在新构建数据集上优于强基线,显著提升分类效果
- 适合需要可解释性的新闻内容分析场景
随着多模态新闻内容的普及,有效的新闻主题分类需要模型能够联合理解文本和图像等异构数据。现有方法通常独立处理各模态或采用简单融合策略,难以捕捉复杂的跨模态交互并利用外部知识。为此,我们提出 MultiPress,一种新颖的三阶段多智能体框架,包含多模态感知、检索增强推理和门控融合评分三个专用智能体,并引入奖励驱动的迭代优化机制。我们在新构建的大规模多模态新闻数据集上验证了 MultiPress,结果表明其显著优于强基线,凸显了模块化多智能体协作和检索增强推理在提升分类准确率与可解释性方面的有效性。
原文摘要 · Abstract (English)
With the growing prevalence of multimodal news content, effective news topic classification demands models capable of jointly understanding and reasoning over heterogeneous data such as text and images. Existing methods often process modalities independently or employ simplistic fusion strategies, limiting their ability to capture complex cross-modal interactions and leverage external knowledge. To overcome these limitations, we propose MultiPress, a novel three-stage multi-agent framework for multimodal news classification. MultiPress integrates specialized agents for multimodal perception, retrieval-augmented reasoning, and gated fusion scoring, followed by a reward-driven iterative optimization mechanism. We validate MultiPress on a newly constructed large-scale multimodal news dataset, demonstrating significant improvements over strong baselines and highlighting the effectiveness of modular multi-agent collaboration and retrieval-augmented reasoning in enhancing classification accuracy and interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。