构建移动数据问答基准,检验大模型理解人类行为语义的能力
MobQA: A Benchmark Dataset for Semantic Understanding of Human Mobility Data through Question Answering
- 设计三类问题:事实提取、语义推理、自由解释,需时空语义综合推理
- 5800个高质量问答对,长轨迹下模型表现显著下降
- 主流大模型在事实回答上优秀,但解释和推理能力仍严重不足
本文提出MobQA,一个通过自然语言问答评估大语言模型(LLMs)对人类移动数据语义理解能力的基准数据集。现有模型虽能精准预测移动模式,却难以理解其背后原因或语义含义。MobQA涵盖日到周粒度的多样化人类GPS轨迹,包含5800个高质量问答对,分为三类互补问题:事实检索(精确数据提取)、多选推理(语义推断)和自由解释(解释性描述),均需空间、时间与语义综合推理。对主流大模型的评估显示,其在事实检索上表现良好,但在语义推理和解释任务中存在明显局限,且轨迹长度显著影响模型效果。结果揭示了当前顶尖大模型在语义移动理解上的成就与瓶颈。
原文摘要 · Abstract (English)
This paper presents MobQA, a benchmark dataset designed to evaluate the semantic understanding capabilities of large language models (LLMs) for human mobility data through natural language question answering. While existing models excel at predicting human movement patterns, it remains unobvious how much they can interpret the underlying reasons or semantic meaning of those patterns. MobQA provides a comprehensive evaluation framework for LLMs to answer questions about diverse human GPS trajectories spanning daily to weekly granularities. It comprises 5,800 high-quality question-answer pairs across three complementary question types: factual retrieval (precise data extraction), multiple-choice reasoning (semantic inference), and free-form explanation (interpretive description), which all require spatial, temporal, and semantic reasoning. Our evaluation of major LLMs reveals strong performance on factual retrieval but significant limitations in semantic reasoning and explanation question answering, with trajectory length substantially impacting model effectiveness. These findings demonstrate the achievements and limitations of state-of-the-art LLMs for semantic mobility understanding.\footnote{MobQA dataset is available at https://github.com/CyberAgentAILab/mobqa.}
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。