系统梳理文档分类中多模态与多视角融合方法,量化其效果并提出可复现的实践指南。
Document Classification Pattern Recognition via Information Fusion: A Systematic Review of Multimodal and Multiview Representation Approaches
- 构建统一框架,整合多模态与多视角融合研究范式
- 多模态融合提升准确率5.28个百分点,多视角融合提升4.67%
- 揭示方法可复现性差问题,仅11.8%研究使用统计检验
信息融合通过整合多个数据源(多模态)或表示(多视角),广泛用于提升文档分类性能。然而,该领域缺乏统一框架、量化效果评估及对实践者的清晰指导。本文系统综述139篇原始研究,提出一个正式框架以结构化该领域,开展定性分析识别关键趋势,并首次针对文档分类进行随机效应元分析(据我们所知),量化性能增益。结果显示,多模态融合显著提升准确率(均值+5.28个百分点,p=0.0016),F1-score方向正向但统计不显著;多视角融合带来一致但适度的增益:准确率+4.67%,F1-score+3.08%,召回率均显著提升(所有p<0.05)。关键发现是方法可复现性严重不足:仅11.8%(多模态)和23.3%(多视角)的研究使用统计检验验证结果,削弱了多数结论的可靠性。本综述的主要贡献为统一框架、首个量化证据基础及数据驱动的实践建议。结论指出,成功的融合不依赖算法复杂度,而取决于融合方法与任务上下文的战略匹配,以及对更严格验证的承诺。
原文摘要 · Abstract (English)
Information fusion is used widely to improve document classification by the integration of multiple data sources (multimodal) or representations (multiview). However, the field lacks a unified framework, a quantitative synthesis of its effectiveness, and clear guidance for practitioners. This systematic review addresses these gaps by analysing 139 primary studies. It introduces a formal framework to structure the field, presents the results of a qualitative analysis to identify key trends, and performs a random-effects meta-analysis (to our knowledge, the first focused on document classification) to quantify performance gains. Our meta-analysis reveals that multimodal fusion improves accuracy (mean gain of +5.28 percentage points, $p=0.0016$) significantly -- the F1-score effect is directionally positive but statistically non-significant in our primary model. Multiview fusion provides consistent but modest gains for accuracy (+4.67\%), F1-score (+3.08\%), and recall (all $p<0.05$). Critically, our qualitative synthesis uncovers challenges in reproducibility in methodological rigour: only 11.8\% (multimodal) and 23.3\% (multiview) of the studies use statistical tests to validate their findings, which undermines the reliability of many of their results. This review's primary contributions are a unifying framework, the first quantitative evidence base, and data-driven guidelines. This review concludes that successful information fusion depends not on algorithmic complexity, but on the strategic alignment of the fusion method with the task context and a commitment to more rigorous validation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。