用重建驱动学习跨模态表示,提升广播内容自动化理解能力
Reconstruction-Driven Multimodal Representation Learning for Automated Media Understanding
- 通过联合重建多模态数据,学习统一语义表示
- 在LUMA数据集上显著提升聚类与对齐指标
- 适合广播媒体自动化标签与检索场景
广播机构日益依赖人工智能自动化内容索引、打标和元数据生成等繁重流程。现有AI系统多基于单一模态(如视频、音频或文本),难以捕捉广播内容中的复杂跨模态关系。本文提出多模态自编码器(MMAE),在文本、音频和视觉数据间学习统一表示,实现元数据提取与语义聚类的端到端自动化。模型在新提出的LUMA数据集上训练,该数据集为真实媒体内容的全对齐多模态三元组基准。通过最小化各模态间的联合重建损失,MMAE发现无需大规模配对或对比数据集的模态不变语义结构。实验表明,相比线性基线,模型在轮廓系数(Silhouette)、调整兰德指数(ARI)和标准化互信息(NMI)等指标上均有显著提升,验证了基于重建的多模态嵌入可作为可扩展元数据生成与跨模态检索的基础。结果表明,重建驱动的多模态学习有望提升现代广播工作流中的自动化水平、可搜索性与内容管理效率。
原文摘要 · Abstract (English)
Broadcast and media organizations increasingly rely on artificial intelligence to automate the labor-intensive processes of content indexing, tagging, and metadata generation. However, existing AI systems typically operate on a single modality-such as video, audio, or text-limiting their understanding of complex, cross-modal relationships in broadcast material. In this work, we propose a Multimodal Autoencoder (MMAE) that learns unified representations across text, audio, and visual data, enabling end-to-end automation of metadata extraction and semantic clustering. The model is trained on the recently introduced LUMA dataset, a fully aligned benchmark of multimodal triplets representative of real-world media content. By minimizing joint reconstruction losses across modalities, the MMAE discovers modality-invariant semantic structures without relying on large paired or contrastive datasets. We demonstrate significant improvements in clustering and alignment metrics (Silhouette, ARI, NMI) compared to linear baselines, indicating that reconstruction-based multimodal embeddings can serve as a foundation for scalable metadata generation and cross-modal retrieval in broadcast archives. These results highlight the potential of reconstruction-driven multimodal learning to enhance automation, searchability, and content management efficiency in modern broadcast workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。