让旅行搜索更自然:用大小模型协同实现图文酒店精准匹配
HotelMatch-LLM: Joint Multi-Task Training of Small and Large Language Models for Efficient Multimodal Hotel Retrieval
- 大小模型分工协作:小模型处理查询,大模型生成酒店嵌入
- 多任务优化使检索效果提升,最高达0.681(基准0.603)
- 支持全量房源图片处理,适合实际旅游场景应用
我们提出HotelMatch-LLM,一种面向旅行领域的多模态稠密检索模型,支持以自然语言进行房源搜索,突破传统搜索需先选目的地、再调参数的限制。该模型具备三大创新:(1)领域特化的多任务优化,引入三项新目标——检索、视觉与语言建模;(2)非对称稠密检索架构,使用小型语言模型(SLM)高效处理在线查询,大型语言模型(LLM)生成酒店嵌入;(3)全面的图像处理能力,可处理全部房源图片图集。在四个多样化测试集上的实验表明,HotelMatch-LLM显著优于现有先进模型,包括VISTA和MARVEL。尤其在主查询类型测试集中,其得分达0.681,远超最有效基线MARVEL的0.603。分析显示,多任务优化影响显著,模型在不同LLM架构下具有强泛化能力,并可高效处理大规模图像图集。
原文摘要 · Abstract (English)
We present HotelMatch-LLM, a multimodal dense retrieval model for the travel domain that enables natural language property search, addressing the limitations of traditional travel search engines which require users to start with a destination and editing search parameters. HotelMatch-LLM features three key innovations: (1) Domain-specific multi-task optimization with three novel retrieval, visual, and language modeling objectives; (2) Asymmetrical dense retrieval architecture combining a small language model (SLM) for efficient online query processing and a large language model (LLM) for embedding hotel data; and (3) Extensive image processing to handle all property image galleries. Experiments on four diverse test sets show HotelMatch-LLM significantly outperforms state-of-the-art models, including VISTA and MARVEL. Specifically, on the test set -- main query type -- we achieve 0.681 for HotelMatch-LLM compared to 0.603 for the most effective baseline, MARVEL. Our analysis highlights the impact of our multi-task optimization, the generalizability of HotelMatch-LLM across LLM architectures, and its scalability for processing large image galleries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。