arXiv:2504.16505cs.CVcs.MM2025-04AAAI

专为旅游设计的多模态模型,能看图识景、推理路线,比通用模型更懂城市旅行。

TraveLLaMA: A Multimodal Travel Assistant with Large-Scale Dataset and Structured Reasoning

  • 构建26.5万条旅游问答数据集,含图文与专家思维链标注
  • 通过结构化推理使回答准确率提升10.8%,支持可解释决策
  • 用户测试评分82.5,适合开发智能导游与出行助手

旅游规划日益依赖数字助手,但现有多模态AI系统缺乏对城市环境的专项知识与上下文理解。本文提出TraveLLaMA,一个面向综合旅行辅助的专用多模态语言模型。核心贡献包括:(1) TravelQA数据集,包含26.5万条问答对,涵盖16万条文本问答(来自真实旅行资料)、10万条图文问答(含地图与位置图像)以及5000条专家标注的思维链推理样本;(2) Travel-CoT结构化推理框架,将旅行问题分解为空间、时间与实用维度,使回答准确率提升10.8%,并提供可解释的决策路径;(3) 交互式代理系统,经500名用户参与的广泛测评验证。在先进视觉语言模型(LLaVA、Qwen-VL、Shikra)上微调后,基础性能提升6.2–9.4%,结合Travel-CoT进一步优化。模型在情境化推荐、地图解读与场景理解方面表现优异,可提供营业时间、文化背景等实用信息。用户研究显示其系统可用性得分达82.5,显著优于通用模型,为多模态旅行助手树立新标准。

原文摘要 · Abstract (English)

Tourism and travel planning increasingly rely on digital assistance, yet existing multimodal AI systems often lack specialized knowledge and contextual understanding of urban environments. We present TraveLLaMA, a specialized multimodal language model designed for comprehensive travel assistance. Our work addresses the fundamental challenge of developing practical AI travel assistants through three key contributions: (1) TravelQA, a novel dataset of 265k question-answer pairs combining 160k text QA from authentic travel sources, 100k vision-language QA featuring maps and location imagery, and 5k expert-annotated Chain-of-Thought reasoning examples; (2) Travel-CoT, a structured reasoning framework that decomposes travel queries into spatial, temporal, and practical dimensions, improving answer accuracy by 10.8\% while providing interpretable decision paths; and (3) an interactive agent system validated through extensive user studies. Through fine-tuning experiments on state-of-the-art vision-language models (LLaVA, Qwen-VL, Shikra), we achieve 6.2-9.4\% base improvements, further enhanced by Travel-CoT reasoning. Our model demonstrates superior capabilities in contextual travel recommendations, map interpretation, and scene understanding while providing practical information such as operating hours and cultural insights. User studies with 500 participants show TraveLLaMA achieves a System Usability Scale score of 82.5, significantly outperforming general-purpose models and establishing new standards for multimodal travel assistance systems.

旅行助手多模态推理框架数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。