用大模型评估城市旅行推荐,兼顾实用、多样与环保。
Multi-Dimensional Evaluation of Sustainable City Trips with LLM-as-a-Judge and Human-in-the-Loop

- 让多个大模型当裁判,分四维度打分:相关性、多样性、可持续性、人气平衡
- 经专家校准后,模型对环保的判断仍不一致,暴露评价标准分歧
- 适合做智能旅行推荐系统评测的研究者和开发者参考
评估复杂的对话式旅行推荐极具挑战,因人工标注成本高,且标准指标忽略利益相关者目标。本文研究使用大模型作为裁判,从四个维度——相关性、多样性、可持续性与人气平衡——评估城市旅行推荐列表,并提出三阶段校准框架:(1) 多个大模型进行基线评分,(2) 专家评估识别系统性偏差,(3) 通过规则与少量示例进行维度专项校准。在两种推荐场景中,我们发现模型存在特定偏差,且各维度间差异显著,即使整体排名一致。校准虽使各维度判断更清晰,但仍暴露出对可持续性的理解分歧,凸显透明、有偏见意识的大模型评估的重要性。相关提示与代码已开源:https://github.com/ashmibanerjee/trs-llm-calibration。
原文摘要 · Abstract (English)
Evaluating nuanced conversational travel recommendations is challenging when human annotations are costly and standard metrics ignore stakeholder-centric goals. We study LLMs-as-Judges for sustainable city-trip lists across four dimensions -- relevance, diversity, sustainability, and popularity balance, and propose a three-phase calibration framework: (1) baseline judging with multiple LLMs, (2) expert evaluation to identify systematic misalignment, and (3) dimension-specific calibration via rules and few-shot examples. Across two recommendation settings, we observe model-specific biases and high dimension-level variance, even when judges agree on overall rankings. Calibration clarifies reasoning per dimension but exposes divergent interpretations of sustainability, highlighting the need for transparent, bias-aware LLM evaluation. Prompts and code are released for reproducibility: https://github.com/ashmibanerjee/trs-llm-calibration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。