构建首个联合评估时间约束与规划能力的中文基准,测试大模型在复杂对话中的任务调度能力。
TCP: a Benchmark for Temporal Constraint-Based Planning
- 基于真实对话生成带多维时间约束的协作项目场景
- 强模型在该基准上仍表现不佳,暴露时序推理短板
- 适合研究大模型时序规划、智能调度与对话理解的学者
时间推理与规划是大语言模型(LLMs)的关键能力,但现有基准大多孤立评估且复杂度有限。为此,我们提出时间约束规划(TCP)基准,联合评估这两项能力。每个实例包含围绕协作项目的自然对话,其中显性或隐性表达多种相互依赖的时间约束,模型需推导满足所有约束的最优调度方案。我们通过生成抽象问题原型,并与多个领域的现实场景结合,利用大模型将其扩展为对话,再对抽样样本进行人工质量检查以确保可靠性。评估结果显示,即使最强的现有模型在TCP上也表现受限,凸显其难度并揭示了大模型在时间约束规划上的不足。我们分析了失败案例,开源了该基准,期望推动未来研究。
原文摘要 · Abstract (English)
Temporal reasoning and planning are essential capabilities for large language models (LLMs), yet most existing benchmarks evaluate them in isolation and under limited forms of complexity. To address this gap, we introduce the Temporal Constraint-based Planning (TCP) benchmark that jointly assesses both capabilities. Each instance in TCP features a naturalistic dialogue around a collaborative project, where diverse and interdependent temporal constraints are explicitly or implicitly expressed, and models must infer an optimal schedule that satisfies all constraints. To construct TCP, we generate abstract problem prototypes that are then paired with realistic scenarios from various domains and enriched into dialogues using an LLM. A human quality check is performed on a sampled subset to confirm the reliability of our benchmark. We evaluate state-of-the-art LLMs and find that even the strongest models may struggle with TCP, highlighting its difficulty and revealing limitations in LLMs' temporal constraint-based planning abilities. We analyze underlying failure cases, open source our benchmark, and hope our findings can inspire future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。