用上下文微调实现轻量级时间序列数据估值,兼顾效率与时序依赖。
Lightweight Time Series Data Valuation on Time Series Foundation Models via In-Context Finetuning
- 基于上下文微调近似影响函数,评估样本贡献。
- 在多数据集上表现可靠,计算开销可控。
- 适合关注数据质量与模型泛化的研究者。
时间序列基础模型(TSFMs)因在大规模多样化时间序列数据上进行预训练而展现出日益增强的能力。因此,时间序列数据的质量对TSFM性能至关重要,准确高效的时序数据估值不可或缺。然而,传统数据估值方法(如影响函数)由于难以扩展到日益庞大的TSFM模型规模,存在严重计算瓶颈,且常无法保留时序依赖性。本文提出LTSV:通过上下文微调实现对时间序列基础模型的轻量级数据估值。基于理论证据——上下文微调可近似影响函数,LTSV通过测量上下文微调后损失的变化来估计样本贡献,利用TSFMs强大的泛化能力,生成鲁棒且可迁移的数据估值。为捕捉时序依赖,引入时间块聚合机制,整合重叠时间窗口内的每块影响得分。在多个时间序列数据集和模型上的实验表明,LTSV始终提供可靠且强劲的估值性能,同时保持可管理的计算开销。结果表明,对时间序列基础模型进行上下文微调,为数据归因与模型泛化之间提供了实用而有效的桥梁。
原文摘要 · Abstract (English)
Time series foundation models (TSFMs) have demonstrated increasing capabilities due to their extensive pretraining on large volumes of diverse time series data. Consequently, the quality of time series data is crucial to TSFM performance, rendering an accurate and efficient data valuation of time series for TSFMs indispensable. However, traditional data valuation methods, such as influence functions, face severe computational bottlenecks due to their poor scalability with growing TSFM model sizes and often fail to preserve temporal dependencies. In this paper, we propose LTSV, a Lightweight Time Series Valuation on TSFMS via in-context finetuning. Grounded in the theoretical evidence that in-context finetuning approximates the influence function, LTSV estimates a sample's contribution by measuring the change in context loss after in-context finetuning, leveraging the strong generalization capabilities of TSFMs to produce robust and transferable data valuations. To capture temporal dependencies, we introduce temporal block aggregation, which integrates per-block influence scores across overlapping time windows. Experiments across multiple time series datasets and models demonstrate that LTSV consistently provides reliable and strong valuation performance, while maintaining manageable computational requirements. Our results suggest that in-context finetuning on time series foundation models provides a practical and effective bridge between data attribution and model generalization in time series learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。