构建多认知维度的旅行规划评测基准,揭示大模型协同处理复杂任务的能力瓶颈。
ItinBench: Benchmarking Planning Across Multiple Cognitive Dimensions with Large Language Models
- 融合空间推理与语言推理,设计跨认知域的旅行路线评测任务
- 多模型对比显示:在多任务并行时性能显著下降,难以保持稳定表现
- 适合研究通用智能、多任务推理与评估方法的学者参考
具备高级认知能力的大语言模型正作为各类推理与规划任务的智能体涌现。传统评估多聚焦于受控环境中的特定推理或规划问题。近期研究将旅行规划作为载体,将多种语言推理任务整合到真实世界场景中。然而,推理不仅限于语言层面,全面评估大模型需涵盖多个认知领域。为此,我们提出 ItinBench 基准,将空间推理任务(如路线优化)融入行程规划,同时保留传统语言推理任务。该基准对 Llama 3.1 8B、Mistral Large、Gemini 1.5 Pro 及 GPT 系列等多个模型进行多任务联合评估。结果表明,当需同时处理多个认知维度时,大模型难以维持高水平且一致的表现。通过引入人类级认知领域的多样化任务,ItinBench 为构建更贴近真实挑战的综合推理评测体系提供了新视角。代码与数据集见:https://ethanwtl.github.io/IBweb/
原文摘要 · Abstract (English)
Large language models (LLMs) with advanced cognitive capabilities are emerging as agents for various reasoning and planning tasks. Traditional evaluations often focus on specific reasoning or planning questions within controlled environments. Recent studies have explored travel planning as a medium to integrate various verbal reasoning tasks into real-world contexts. However, reasoning tasks extend beyond verbal reasoning alone, and a comprehensive evaluation of LLMs requires a testbed that incorporates tasks from multiple cognitive domains. To address this gap, we introduce ItinBench, a benchmark that features one task of spatial reasoning, i.e., route optimization, into trip itinerary planning while keeping the traditional verbal reasoning tasks. ItinBench evaluates various LLMs across diverse tasks simultaneously, including Llama 3.1 8B, Mistral Large, Gemini 1.5 Pro, and GPT family. Our findings reveal that LLMs struggle to maintain high and consistent performance when concurrently handling multiple cognitive dimensions. By incorporating tasks from distinct human-level cognitive domains, ItinBench provides new insights into building more comprehensive reasoning testbeds that better reflect real-world challenges. The code and dataset: https://ethanwtl.github.io/IBweb/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。