arXiv:2507.08104cs.MMcs.AI2025-07KDD被引 3

构建多模态金融影响力数据集,评估模型识别投资建议可信度的能力

VideoConviction: A Multimodal Benchmark for Human Conviction and Stock Market Recommendations

  • 基于6000+人工标注,构建视频与文本结合的金融推荐评估基准
  • 模型虽能识别股票代码但难区分真实建议与普通评论,高可信度推荐仍逊于指数基金
  • 反向操作(做空推荐)年化收益超标普500 6.8%,但风险更高

社交媒体放大了被称为'金融网红'(finfluencers)的金融意见领袖影响力,他们通过YouTube等平台发布股票建议。理解其影响需分析语气、表达风格和面部表情等多模态信号,远超传统文本分析。我们提出VideoConviction,一个包含6000+专家标注的多模态数据集,耗时457小时人工标注,用于评测多模态大语言模型(MLLMs)和文本大语言模型(LLMs)在金融话语中的表现。结果表明,尽管多模态输入有助于提取股票代码(如苹果公司AAPL),但两类模型均难以区分投资行为与可信度——即通过自信表达和详细推理传达的信心程度,常将一般评论误判为明确建议。高可信度建议表现优于低可信度,但仍不及主流标普500指数基金。一种逆向策略(做空金融网红建议)年化收益率超出标普500 6.8%,但风险更大(夏普比率0.41对0.65)。该基准支持对完整视频与分段视频输入的多样化评估,推动多模态金融研究深入发展。代码、数据集与评测排行榜已开源,采用CC BY-NC 4.0许可。

原文摘要 · Abstract (English)

Social media has amplified the reach of financial influencers known as "finfluencers," who share stock recommendations on platforms like YouTube. Understanding their influence requires analyzing multimodal signals like tone, delivery style, and facial expressions, which extend beyond text-based financial analysis. We introduce VideoConviction, a multimodal dataset with 6,000+ expert annotations, produced through 457 hours of human effort, to benchmark multimodal large language models (MLLMs) and text-based large language models (LLMs) in financial discourse. Our results show that while multimodal inputs improve stock ticker extraction (e.g., extracting Apple's ticker AAPL), both MLLMs and LLMs struggle to distinguish investment actions and conviction--the strength of belief conveyed through confident delivery and detailed reasoning--often misclassifying general commentary as definitive recommendations. While high-conviction recommendations perform better than low-conviction ones, they still underperform the popular S\&P 500 index fund. An inverse strategy--betting against finfluencer recommendations--outperforms the S\&P 500 by 6.8\% in annual returns but carries greater risk (Sharpe ratio of 0.41 vs. 0.65). Our benchmark enables a diverse evaluation of multimodal tasks, comparing model performance on both full video and segmented video inputs. This enables deeper advancements in multimodal financial research. Our code, dataset, and evaluation leaderboard are available under the CC BY-NC 4.0 license.

多模态金融分析视频理解大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。