arXiv:2510.26978cs.CVcs.CL2025-10被引 1

基于语义帧聚合的Transformer模型,让实时视频评论更贴合对话上下文。

Semantic Frame Aggregation-based Transformer for Live Video Comment Generation

  • 按语义相关性加权视频帧,突出关键画面
  • 在320万条评论数据上实现更高相关性生成
  • 适合直播互动、多模态生成研究者

Twitch等平台上的实时视频评论互动日益流行,但自动生成契合语境的评论仍具挑战。视频流包含大量数据和无关内容,现有方法常忽略对与观众互动最相关的视频帧进行优先处理。为此,我们提出语义帧聚合式Transformer(SFAT)模型,结合CLIP的跨模态知识生成评论,并根据帧的语义相关性动态加权,通过加权求和聚焦关键帧。其交叉注意力解码器同时关注聊天与视频模态,确保生成内容反映双重上下文。此外,为弥补现有数据集以中文为主、类别有限的问题,我们构建了大规模多模态英文评论数据集:从Twitch提取,覆盖11类视频,共438小时时长,含320万条评论。实验表明,该模型在实时视频与对话上下文中生成评论方面优于现有方法。

原文摘要 · Abstract (English)

Live commenting on video streams has surged in popularity on platforms like Twitch, enhancing viewer engagement through dynamic interactions. However, automatically generating contextually appropriate comments remains a challenging and exciting task. Video streams can contain a vast amount of data and extraneous content. Existing approaches tend to overlook an important aspect of prioritizing video frames that are most relevant to ongoing viewer interactions. This prioritization is crucial for producing contextually appropriate comments. To address this gap, we introduce a novel Semantic Frame Aggregation-based Transformer (SFAT) model for live video comment generation. This method not only leverages CLIP's visual-text multimodal knowledge to generate comments but also assigns weights to video frames based on their semantic relevance to ongoing viewer conversation. It employs an efficient weighted sum of frames technique to emphasize informative frames while focusing less on irrelevant ones. Finally, our comment decoder with a cross-attention mechanism that attends to each modality ensures that the generated comment reflects contextual cues from both chats and video. Furthermore, to address the limitations of existing datasets, which predominantly focus on Chinese-language content with limited video categories, we have constructed a large scale, diverse, multimodal English video comments dataset. Extracted from Twitch, this dataset covers 11 video categories, totaling 438 hours and 3.2 million comments. We demonstrate the effectiveness of our SFAT model by comparing it to existing methods for generating comments from live video and ongoing dialogue contexts.

实时评论多模态生成视频理解Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。