arXiv:2605.07677cs.IRcs.AI2026-05

让旅游推荐可解释、可验证、能纠错,提升真实决策可信度。

TRACE: Tourism Recommendation with Accountable Citation Evidence

论文配图:TRACE: Tourism Recommendation with Accountable Citation Evidence
图 1 · 摘自论文原文
  • 引入带真实评论引用的多轮对话数据集,支持可追溯推荐。
  • 发现大模型在召回和纠错上强,但引用密度低;检索器引用准但准确率差。
  • 适合关注可解释性与鲁棒性的旅游推荐研究者。

旅游推荐是对话式推荐系统(CRS)的高风险场景:看似合理的建议可能浪费真实金钱与行程时间。现有基准主要以单一召回率评估实体提及,虽有空间或知识图谱上下文,但均未结合多轮对话、原文评论证据与拒绝后恢复能力。为此,我们提出TRACE,包含10,000条多轮旅游推荐对话,覆盖2,400个Yelp景点与34,208条评论,覆盖美国八大城市,配套14种检索、规划与大模型基线,以及25项按准确性、可溯源性、恢复能力分类的指标。实验揭示‘三重能力差距’:大模型零样本在闭集召回@1与拒绝恢复上领先,但引用密度低于检索器;非大模型检索器实现表面字面溯源,但准确率低;多评论合成方法在恢复阶段表现不佳。可溯源性评分与人工引用精确度高度相关(斯皮尔曼ρ=+0.80,p<10⁻²⁰),配对t检验复现了基线排名差异(主导对比p<0.01)。TRACE将可问责旅游推荐重构为兼顾正确景点、可验证证据与自适应修复的联合目标,而非单一指标排行榜。

原文摘要 · Abstract (English)

Tourism is a high-stakes setting for conversational recommender systems (CRS): a plausible-sounding suggestion can waste real money and trip time once a traveler acts on it. Existing CRS benchmarks primarily evaluate systems with a single Recall@k score over entity mentions, and tourism-specific resources add spatial or knowledge-graph context, yet none of them couple multi-turn recommendation with verbatim review-span evidence and rejection recovery. This leaves an evaluation gap for tourism recommendation that is simultaneously trustworthy, verifiable, and adaptive: recommend the right point of interest (POI) for multi-aspect preferences (such as cuisine, price, atmosphere, walking distance), justify each suggestion with verifiable evidence from prior visitors so the traveler can act without trial and error, and recover when the first recommendation is rejected mid-dialogue. We introduce TRACE, where each item is a multi-turn tourism recommendation dialogue with review-span citations and explicit rejection turns: 10,000 dialogues over 2,400 Yelp POIs and 34,208 reviews across eight U.S. cities, paired with 14 retrieval, planning, and LLM baselines, along with 25 metrics organized under Accuracy, Grounding, and Recovery. Across these baselines, TRACE reveals the Three-Competency Gap: LLM Zero-Shot leads in closed-set Recall@1 and rejection recovery but cites less densely than retrievers; non-LLM retrievers achieve surface-verbatim grounding but with low accuracy; Multi-Review Synthesis fails at recovery. The Grounding Score agrees with human citation precision (Spearman rho=+0.80, p<10^-20), and paired t-tests reproduce the per-baseline ranking (p<0.01 on the dominant contrasts). TRACE reframes accountable tourism recommendation as a joint target (right POI, verifiable evidence, adaptive repair) rather than a single-axis leaderboard.

旅游推荐可解释性对话系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。