用图表+原始数据融合提升时间序列分类可解释性与性能
VTBench: A Multimodal Framework for Time-Series Classification with Chart-Based Representations

- 将时间序列转为线图、面积图等图表,与原始数据联合建模
- 31个数据集实验表明多图表融合可提升准确率,但冗余会降低性能
- 提供图表选型和融合策略的实用指南,适合关注可解释性的研究者
时间序列分类(TSC)虽在深度学习推动下取得显著进展,但多数模型仅依赖原始数值输入,忽视了其他表示形式。尽管基于纹理的编码(如GAF、RP)能将时间序列转为二维图像,但常需复杂预处理且可解释性差。相比之下,图表可视化更直观,在特定领域表现良好,但其有效性尚未系统评估,尤其缺乏对不同图表类型、视觉编码方式及数据集的全面分析。本文提出VTBench,一个系统化且可扩展的多模态框架,通过融合原始序列与图表可视化重新审视TSC。该框架生成轻量级、人类可读的线图、面积图、柱状图和散点图,提供同一信号的不同视角。我们设计模块化架构,支持单图表视觉-数值融合、多图表视觉融合及完整多模态融合。在31个UCR数据集上的实验表明:(1) 仅使用图表的模型在部分场景下表现竞争力,尤其在小数据集上;(2) 多图表融合可通过捕捉互补视觉线索提升准确率;(3) 多模态模型在视觉特征非冗余时性能提升或保持稳定,但在引入冗余时可能下降。我们进一步提炼出选择图表类型、融合策略与配置的实用建议。VTBench为可解释且高效的多模态时间序列分类建立了统一基础。
原文摘要 · Abstract (English)
Time-series classification (TSC) has advanced significantly with deep learning, yet most models rely solely on raw numerical inputs, overlooking alternative representations. While texture-based encodings such as Gramian Angular Fields (GAF) and Recurrence Plots (RP) convert time series into 2D images, they often require heavy preprocessing and yield less intuitive representations. In contrast, chart-based visualizations offer more interpretable alternatives and show promise in specific domains; however, their effectiveness remains underexplored, with limited systematic evaluation across chart types, visual encoding choices, and datasets. In this work, we introduce VTBench, a systematic and extensible framework that re-examines TSC through multimodal fusion of raw sequences and chart-based visualizations. VTBench generates lightweight, human-interpretable plots -- line, area, bar, and scatter, providing complementary views of the same signal. We develop a modular architecture supporting multiple fusion strategies, including single-chart visual-numerical fusion, multi-chart visual fusion, and full multimodal fusion with raw inputs. Through experiments across 31 UCR datasets, we show that: (1) chart-only models are competitive in selected settings, particularly on smaller datasets; (2) combining multiple chart types can improve accuracy by capturing complementary visual cues; and (3) multimodal models improve or maintain performance when visual features provide non-redundant information, but may degrade accuracy when they introduce redundancy. We further distill practical guidelines for selecting chart types, fusion strategies, and configurations. VTBench establishes a unified foundation for interpretable and effective multimodal time-series classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。