构建电视新闻多模态标注框架,分析内容与观众行为关系
From Content to Audience: A Multimodal Annotation Framework for Broadcast Television Analytics
- 设计四维标注体系,融合视觉、语音、元数据等多源信息
- 大模型能利用视频时序信息提升效果,小模型易受上下文过载影响
- 在14个完整节目上验证,可关联内容特征与分龄观众偏好
自动标注广播电视内容面临结构化音视频编排、领域特定编辑模式及严格运营约束的挑战。尽管多模态大语言模型(MLLM)在通用视频理解中表现优异,其在广播场景下不同管道架构与输入配置的相对效能仍缺乏实证研究。本文系统评估了应用于意大利电视新闻的多模态标注流程,构建了涵盖视觉环境分类、话题分类、敏感内容检测和命名实体识别四个维度的领域专用基准数据集。对比两种管道架构在九个前沿模型(包括Gemini 3.0 Pro、LLaMA 4 Maverick、Qwen-VL系列和Gemma 3)上的表现,采用逐步增强的输入策略:结合视觉信号、自动语音识别、说话人辨识与元数据。实验表明,视频输入带来的增益具有强模型依赖性:大模型能有效利用时间连续性,而小模型在扩展多模态上下文时出现性能下降,可能因令牌过载所致。除基准测试外,所选管道部署于14个完整广播节目,实现分钟级标注,并与意大利媒体公司提供的标准化收视数据集成,完成话题层面观众敏感度与代际参与差异的相关性分析,验证了该框架在内容驱动观众分析中的实际可行性。
原文摘要 · Abstract (English)
Automated semantic annotation of broadcast television content presents distinctive challenges, combining structured audiovisual composition, domain-specific editorial patterns, and strict operational constraints. While multimodal large language models (MLLMs) have demonstrated strong general-purpose video understanding capabilities, their comparative effectiveness across pipeline architectures and input configurations in broadcast-specific settings remains empirically undercharacterized. This paper presents a systematic evaluation of multimodal annotation pipelines applied to broadcast television news in the Italian setting. We construct a domain-specific benchmark of clips labeled across four semantic dimensions: visual environment classification, topic classification, sensitive content detection, and named entity recognition. Two different pipeline architectures are evaluated across nine frontier models, including Gemini 3.0 Pro, LLaMA 4 Maverick, Qwen-VL variants, and Gemma 3, under progressively enriched input strategies combining visual signals, automatic speech recognition, speaker diarization, and metadata. Experimental results demonstrate that gains from video input are strongly model-dependent: larger models effectively leverage temporal continuity, while smaller models show performance degradation under extended multimodal context, likely due to token overload. Beyond benchmarking, the selected pipeline is deployed on 14 full broadcast episodes, with minute-level annotations integrated with normalized audience measurement data provided by an Italian media company. This integration enables correlational analysis of topic-level audience sensitivity and generational engagement divergence, demonstrating the operational viability of the proposed framework for content-based audience analytics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。