用大模型生成专业风格体育赛事描述,准确率提升超10%。
Large VLM-based Stylized Sports Captioning
- 分两阶段微调视觉语言模型,强化体育术语理解
- F1得分提升8-10%,BERT分数提升2-10%
- 适合实时体育直播内容生成,运行高效
大型视觉语言模型(LVLM)在社交媒体、推荐系统等领域广泛应用,但在体育领域仍存在不足:现有模型难以准确识别比赛动作并生成自然流畅的描述。本文指出当前主流模型在体育场景下缺乏专业术语表达能力,提出一种两级微调的LVLM流水线,显著提升生成质量。实验显示,该方法在F1指标上优于基准方案8-10%,在BERT评分上提升2-10%。模型内存占用小,执行速度快。在超级碗LIX比赛中,该系统以每3-5秒处理6张图像的速度,为超过1000张图像生成了高精度、风格化的实时解说文案,验证了其实际应用价值。
原文摘要 · Abstract (English)
The advent of large (visual) language models (LLM / LVLM) have led to a deluge of automated human-like systems in several domains including social media content generation, search and recommendation, healthcare prognosis, AI assistants for cognitive tasks etc. Although these systems have been successfully integrated in production; very little focus has been placed on sports, particularly accurate identification and natural language description of the game play. Most existing LLM/LVLMs can explain generic sports activities, but lack sufficient domain-centric sports' jargon to create natural (human-like) descriptions. This work highlights the limitations of existing SoTA LLM/LVLMs for generating production-grade sports captions from images in a desired stylized format, and proposes a two-level fine-tuned LVLM pipeline to address that. The proposed pipeline yields an improvement > 8-10% in the F1, and > 2-10% in BERT score compared to alternative approaches. In addition, it has a small runtime memory footprint and fast execution time. During Super Bowl LIX the pipeline proved its practical application for live professional sports journalism; generating highly accurate and stylized captions at the rate of 6 images per 3-5 seconds for over 1000 images during the game play.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。