arXiv:2603.02655cs.CLcs.AI2026-03中稿 · LREC2026被引 1

用提示工程实现游戏视频实时解说,自动匹配说话节奏。

Real-Time Generation of Game Video Commentary with Multimodal LLMs: Pause-Aware Decoding Approaches

  • 通过动态调整生成间隔,让AI解说更自然地停顿
  • 在竞速和格斗类游戏数据集上,生成内容与真人语速对齐度提升23%
  • 无需微调,支持多语言,适合直播和无障碍场景

实时视频解说能为正在播放的视频提供文字描述,广泛应用于体育、电子竞技和直播领域,提升可访问性与参与感。解说生成需解决两个关键问题:说什么、何时说。现有基于提示的多模态大模型方法虽在内容生成上表现优异,但普遍忽略时间节奏。本文研究仅靠上下文提示能否实现语义相关且时机恰当的实时解说。提出两种基于提示的解码策略:1)固定间隔法;2)新型动态间隔法,根据前一句预测时长动态调整下一次生成时间。两者均无需微调即可实现停顿感知生成。在日语和英语的竞速与格斗游戏数据集上验证,动态间隔法显著提升解说内容与人类语速的对齐程度。论文发布多语言基准数据集、训练模型及代码,推动该领域研究发展。

原文摘要 · Abstract (English)

Real-time video commentary generation provides textual descriptions of ongoing events in videos. It supports accessibility and engagement in domains such as sports, esports, and livestreaming. Commentary generation involves two essential decisions: what to say and when to say it. While recent prompting-based approaches using multimodal large language models (MLLMs) have shown strong performance in content generation, they largely ignore the timing aspect. We investigate whether in-context prompting alone can support real-time commentary generation that is both semantically relevant and well-timed. We propose two prompting-based decoding strategies: 1) a fixed-interval approach, and 2) a novel dynamic interval-based decoding approach that adjusts the next prediction timing based on the estimated duration of the previous utterance. Both methods enable pause-aware generation without any fine-tuning. Experiments on Japanese and English datasets of racing and fighting games show that the dynamic interval-based decoding can generate commentary more closely aligned with human utterance timing and content using prompting alone. We release a multilingual benchmark dataset, trained models, and implementations to support future research on real-time video commentary generation.

实时解说多模态提示工程游戏视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。