arXiv:2605.30344cs.AI2026-05

构建带解释的时序异常检测数据集,训练出高效可解释的视觉语言模型。

Tiny but Trusted: Efficient Vision-Language Reasoning for Time-Series Anomaly Detection

论文配图:Tiny but Trusted: Efficient Vision-Language Reasoning for Time-Series Anomaly Detection
图 1 · 摘自论文原文
  • 用多模型生成细粒度解释,构建高质量时序异常数据集
  • 在新数据集上微调后,精度和F1提升超20个百分点
  • 模型参数少、推理快,适合工业场景实时异常检测

视觉语言模型在诸多任务中表现优异,但在时序数据异常检测中效果不佳。现有公开基准仅提供区间标注,缺乏自然语言解释,难以微调模型生成可信决策。为此,研究者构建了VisAnomBench,基于公开时序数据集并融合多个大模型生成的高质异常解释,采用细粒度任务奖励筛选。在此基础上,通过微调得到参数高效的VisAnomReasoner模型。在VisAnomBench上,该模型实现更精准的异常定位,精度与F1分别提升至少21.23和23.87个百分点;在TSB-AD-U基准上也展现强泛化能力,精度和F1分别提升9.57和13.39个百分点。

原文摘要 · Abstract (English)

Recent advances in Vision-Language Models (VLMs) have achieved impressive performance across many tasks, yet prior studies report unsatisfactory performance when applying large language or multimodal models to finding abnormal patterns in sequential data. Public anomaly detection benchmarks typically provide interval annotations but not natural-language rationales, making it difficult to fine-tune VLMs to produce grounded, interpretable decisions. To address this gap, we construct VisAnomBench, a curated benchmark built from public time-series datasets and augmented with high-quality anomaly explanations selected from multiple large VLMs using fine-grained, task-specific rewards. Through fine-tuning on this benchmark, we develop VisAnomReasoner, a parameter-efficient VLM for time-series anomaly detection. Experimental results on VisAnomBench show that VisAnomReasoner achieves more accurate anomaly localization and consistently outperforms all baselines, with improvements of at least 21.23 and 23.87 percentage points in precision and F1, respectively. Additional experiments on the TSB-AD-U benchmark demonstrate strong cross-benchmark generalization, with VisAnomReasoner improving precision and F1 by 9.57 and 13.39 percentage points, respectively.

时序异常检测视觉语言模型可解释性高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。