用逆向生成法构建跨领域时间序列描述数据集
Domain-Independent Automatic Generation of Descriptive Texts for Time-Series Data
- 提出反向生成方法构建描述文本对
- 新数据集TACO支持跨领域文本生成
- 对比学习模型可泛化到未见领域
由于缺乏带描述文本的时间序列数据标注,训练生成描述性文本的模型面临挑战。本研究提出一种系统化方法,从时间序列数据中生成跨领域描述文本。识别出两种构建时间序列与描述文本配对的方法:正向与反向。通过实现新颖的反向方法,构建了时间序列观测自动描述数据集(Temporal Automated Captions for Observations, TACO)。实验表明,基于对比学习的模型在TACO数据集上训练后,能够为新领域的时序数据生成描述性文本。
原文摘要 · Abstract (English)
Due to scarcity of time-series data annotated with descriptive texts, training a model to generate descriptive texts for time-series data is challenging. In this study, we propose a method to systematically generate domain-independent descriptive texts from time-series data. We identify two distinct approaches for creating pairs of time-series data and descriptive texts: the forward approach and the backward approach. By implementing the novel backward approach, we create the Temporal Automated Captions for Observations (TACO) dataset. Experimental results demonstrate that a contrastive learning based model trained using the TACO dataset is capable of generating descriptive texts for time-series data in novel domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。