arXiv:2410.02637cs.AIcs.CV2024-10被引 11

用图表让多模态模型读懂时间序列数据,省钱又高效。

Plots Unlock Time-Series Understanding in Multimodal Models

  • 用现有视觉编码器解析时间序列图表,无需额外训练。
  • 图表输入比原始数据文本提升120%性能(合成数据),实测提升150%。
  • 适合医疗、金融等领域处理复杂噪声数据,降低模型调用成本90%。

尽管多模态基础模型已能处理文本以外的数据,但在医疗、金融和社会科学等领域的多维时间序列数据分析中仍被严重低估,错失了更丰富、数据驱动的洞察机会。本文提出一种简单但有效的方法,利用现有模型的视觉编码器通过图表‘看懂’时间序列数据,避免了额外且可能昂贵的模型训练。实证评估表明,该方法优于将原始时间序列作为文本输入,且视觉表示可使模型API成本降低高达90%。通过逐步增加复杂度的合成数据任务验证假设,涵盖从干净数据的功能形式识别到含噪散点图的趋势提取。为证明从合成任务到真实场景的泛化能力,将方法应用于消费者健康任务——跌倒检测、活动识别和状态评估,这些任务涉及异构、噪声数据及多步推理。在GPT与Gemini模型族上,图表输入在零样本合成任务上性能最高提升120%,真实任务上最高提升150%,充分展现了其对基础模型原生能力的高效利用潜力。

原文摘要 · Abstract (English)

While multimodal foundation models can now natively work with data beyond text, they remain underutilized in analyzing the considerable amounts of multi-dimensional time-series data in fields like healthcare, finance, and social sciences, representing a missed opportunity for richer, data-driven insights. This paper proposes a simple but effective method that leverages the existing vision encoders of these models to "see" time-series data via plots, avoiding the need for additional, potentially costly, model training. Our empirical evaluations show that this approach outperforms providing the raw time-series data as text, with the additional benefit that visual time-series representations demonstrate up to a 90% reduction in model API costs. We validate our hypothesis through synthetic data tasks of increasing complexity, progressing from simple functional form identification on clean data, to extracting trends from noisy scatter plots. To demonstrate generalizability from synthetic tasks with clear reasoning steps to more complex, real-world scenarios, we apply our approach to consumer health tasks - specifically fall detection, activity recognition, and readiness assessment - which involve heterogeneous, noisy data and multi-step reasoning. The overall success in plot performance over text performance (up to an 120% performance increase on zero-shot synthetic tasks, and up to 150% performance increase on real-world tasks), across both GPT and Gemini model families, highlights our approach's potential for making the best use of the native capabilities of foundation models.

时间序列多模态图表理解低耗推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。