BOOM实现课件音频与幻灯片的多模态同步翻译,让全球学生听懂、看懂、读懂讲座。
BOOM: Beyond Only One Modality KIT's Multimodal Multilingual Lecture Companion
- 联合翻译音频与幻灯片,保持三模态同步输出
- 生成本地化文本、保留视觉元素的幻灯片和合成语音
- 提升摘要与问答等下游任务表现,适合教育科技开发者
教育全球化与在线学习的快速发展使教学内容本地化成为关键挑战。讲座内容天然具有多模态特性,融合语音与视觉幻灯片,需系统支持多模态处理。为提供完整学习体验,翻译必须同时保留文字、幻灯片视觉信息与语音听觉内容。我们提出 extbf{BOOM},一个面向多语言的多模态讲座助手,可联合翻译讲座音频与幻灯片,生成三模态同步输出:翻译后的文本、保留视觉元素的本地化幻灯片、以及合成语音。该端到端方法使学生能以母语获取讲座内容,同时尽可能保留原始信息。实验表明,考虑幻灯片信息的转录文本在摘要生成与问答等下游任务中也展现出显著提升。演示视频与代码已公开于 https://ai4lt.github.io/boom/(MIT 许可)。
原文摘要 · Abstract (English)
The globalization of education and rapid growth of online learning have made localizing educational content a critical challenge. Lecture materials are inherently multimodal, combining spoken audio with visual slides, which requires systems capable of processing multiple input modalities. To provide an accessible and complete learning experience, translations must preserve all modalities: text for reading, slides for visual understanding, and speech for auditory learning. We present \textbf{BOOM}, a multimodal multilingual lecture companion that jointly translates lecture audio and slides to produce synchronized outputs across three modalities: translated text, localized slides with preserved visual elements, and synthesized speech. This end-to-end approach enables students to access lectures in their native language while aiming to preserve the original content in its entirety. Our experiments demonstrate that slide-aware transcripts also yield cascading benefits for downstream tasks such as summarization and question answering. The demo video and code can be found at https://ai4lt.github.io/boom/ \footnote{All released code and models are licensed under the MIT License}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。