用多模态模型为金融短视频生成精准标题,发现视频本身已具强识别力。
FinCap: Topic-Aligned Captions for Short-Form Financial YouTube Videos
- 融合音视频与字幕,测试七种模态组合的联合推理能力
- 仅用视频在四类任务中表现最优,证明视觉线索关键
- 部分双模态组合优于全模态,提示噪声干扰需警惕
我们评估多模态大语言模型(MLLMs)在金融短时视频(SVs)中进行主题对齐字幕生成的能力,通过联合分析字幕(T)、音频(A)和视频(V)实现。基于624个标注的YouTube短视頻,我们在五个主题上测试了所有七种模态组合(T, A, V, TA, TV, AV, TAV):主要建议、情感分析、视频目的、视觉分析及金融实体识别。结果显示,仅使用视频在四个主题上表现强劲,凸显其捕捉视觉上下文与有效线索(如情绪、手势、肢体语言)的价值。某些选择性组合如TV或AV常优于全模态(TAV),表明过多模态可能引入噪声。本研究建立了金融短视頻字幕生成的首个基准,揭示了该领域中复杂视觉线索建模的潜力与挑战。所有代码与数据可在GitHub公开获取,采用CC-BY-NC-SA 4.0许可。
原文摘要 · Abstract (English)
We evaluate multimodal large language models (MLLMs) for topic-aligned captioning in financial short-form videos (SVs) by testing joint reasoning over transcripts (T), audio (A), and video (V). Using 624 annotated YouTube SVs, we assess all seven modality combinations (T, A, V, TA, TV, AV, TAV) across five topics: main recommendation, sentiment analysis, video purpose, visual analysis, and financial entity recognition. Video alone performs strongly on four of five topics, underscoring its value for capturing visual context and effective cues such as emotions, gestures, and body language. Selective pairs such as TV or AV often surpass TAV, implying that too many modalities may introduce noise. These results establish the first baselines for financial short-form video captioning and illustrate the potential and challenges of grounding complex visual cues in this domain. All code and data can be found on our Github under the CC-BY-NC-SA 4.0 license.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。