用智能体框架让大模型自动识别时间序列数据质量维度并精准评分
TSQAgent: Rating Time Series Data Quality via Dedicated Agentic Reasoning

- 设计三角色智能体:感知者选关键维度,检查者定量分析,裁决者综合判断
- 在11个真实数据集上,质量评分准确率提升23%,下游任务性能显著改善
- 首次构建专用评测基准TSQBench,揭示当前大模型在维度识别上的严重不足
评估时间序列(TS)数据质量虽基础却极具挑战,因质量维度多样复杂。近期大语言模型(LLMs)通过成对比较和分维度评估展现潜力,但现有方法依赖人工预设维度且仅基于文本推理,难以识别真正相关维度或进行有证据支撑的量化对比。为此,我们构建了专用于评估LLMs能力的TSQBench基准,测试其在(i)理解并识别相关质量维度,以及(ii)特定维度下进行质量比较两方面的能力。分析表明,当前LLMs在维度识别与基于证据的质量比较上均表现不佳。为此,我们提出TSQAgent——一种新型代理式推理框架,包含三个协同角色:感知者(聚焦维度选择)、检查者(维度内定量分析)与裁决者(聚合并优化最终判断)。特别地,引入能识别并优先处理最相关维度的代理推理策略,并结合外部分析工具的工作流,实现选定维度上的精确量化比较。在自建基准及11个真实数据集上的实验表明,该框架不仅显著提升LLMs在质量理解与量化比较方面的能力,更有效转化为更优的质量感知数据选择,从而提升下游性能与数据效率。
原文摘要 · Abstract (English)
Assessing the quality of time series (TS) data is fundamental yet inherently challenging due to the multifaceted nature of quality dimensions. Recently, large language models (LLMs) have emerged as a promising paradigm for TS quality assessment via pairwise comparison and per-dimension evaluation. However, existing approaches rely on manually predefined quality dimensions and purely text-based reasoning, leaving it unknown whether LLMs can identify truly relevant quality dimensions or perform grounded and quantitative quality comparisons. To investigate this, we construct TSQBench, a dedicated benchmark for evaluating LLMs on two progressive capabilities: (i) understanding and identifying relevant quality dimensions, and (ii) performing quality comparison under specific dimensions. Our analysis reveals that current LLMs consistently struggle with both dimension identification and evidence-grounded quality comparison. To address these limitations, we propose TSQAgent, a novel agentic reasoning framework for TS quality rating consisting of three collaborative roles: Perceiver for focused dimension selection, Inspector for dimension-wise quantitative analysis, and Adjudicator that aggregates and refines the final judgment. In particular, we introduce an agentic reasoning strategy that instills the ability to identify and prioritize the most relevant quality dimensions, and further propose an agent workflow equipped with external analytical tools to enable precise quantitative comparisons over selected dimensions. Experiments on both the proposed benchmark and eleven real-world datasets demonstrate that our framework not only substantially improves LLMs' capabilities in quality understanding and quantitative comparison but also effectively translates these improvements into better quality-aware data selection, leading to enhanced downstream performance and data efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。