融合音视频与语音的三模态模型,提升长视频主题发现质量
MMTM: Tri-Modal Topic Modeling for Long-Form Video via Similarity-Gated Fusion

- 用相似性门控融合语音、音频和视觉特征,实现多模态主题建模
- 噪声降低至0.06,主题稳定性提升,熵值达0.92,聚类有效性提高5-12倍
- 适合需要高质量视频主题分析的研究者,尤其关注跨模态内容理解
我们提出MMTM,一种用于长视频主题发现的模块化流程,通过确定性相似性门控融合语音识别、音频与视觉嵌入,并结合BERTopic聚类。在德语(Tagesschau)和英语(NBC)新闻广播上跨语言评估,联合三模态建模显著提升主题质量:噪声从0.27降至0.06,转换率从0.70降至0.21,归一化熵从0.84升至0.92,表明主题更连贯且时间更稳定。聚类有效性(Calinski-Harabasz)在各嵌入空间提升5-12倍。德语数据集词汇连贯性(NPMI)从0.77升至0.86,但该效果依赖语料,未在较短的NBC广播中迁移。我们开源了代码及一个经人工验证的54小时多模态视频主题语料库,包含双标注视觉评估与大模型辅助标注。
原文摘要 · Abstract (English)
We introduce MMTM, a modular pipeline for topic discovery in long-form video that integrates speech recognition, audio and visual embeddings, and BERTopic clustering through a deterministic similarity-gated fusion. Evaluated cross-lingually on German (Tagesschau) and English (NBC) broadcast news, joint tri-modal modeling substantially improves topic quality: noise drops from 0.27 to 0.06, transition rate from 0.70 to 0.21, and normalized entropy rises from 0.84 to 0.92, indicating more coherent and temporally stable topics. Cluster validity (Calinski-Harabasz) improves by 5-12X across embedding spaces. Lexical coherence (NPMI) rises from 0.77 to 0.86 on German but is corpus-dependent and does not transfer to the shorter NBC broadcasts. We release the pipeline code and a human-validated 54-hour multimodal video topic corpus with dual-annotator visual evaluation and LLM-assisted labeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。