不用标注数据,用专家思维链评估旅游领域大模型表现。
LETToT: Label-Free Evaluation of Large Language Models On Tourism Using Expert Tree-of-Thought
- 用专家构建的思维链结构替代标注数据进行评估
- 优化后比基线提升4.99%-14.15%质量,小模型性能显著增强
- 适合关注领域模型评估、减少标注依赖的研究者
在旅游等特定领域评估大语言模型(LLM)仍具挑战,源于标注基准成本高昂及幻觉问题。本文提出无需标签的旅游领域大模型评估框架LETToT,利用专家生成的思维链结构代替标注数据。首先,通过与通用质量维度对齐及专家反馈,迭代优化层级化思维链组件,结果表明系统优化后的专家思维链相较基线提升4.99%-14.15%相对质量。其次,将优化后的思维链应用于不同规模模型(32B-671B参数)的评估:(1) 规模定律在专业领域依然存在(DeepSeek-V3表现最优),但推理增强的小模型(如DeepSeek-R1-Distill-Llama-70B)可缩小差距;(2) 对于参数量小于72B的模型,显式推理架构在准确性和简洁性上均显著优于对照组(p<0.05)。本工作建立了一种可扩展、无需标签的领域专用模型评估范式,为传统标注基准提供了有力替代方案。
原文摘要 · Abstract (English)
Evaluating large language models (LLMs) in specific domain like tourism remains challenging due to the prohibitive cost of annotated benchmarks and persistent issues like hallucinations. We propose $\textbf{L}$able-Free $\textbf{E}$valuation of LLM on $\textbf{T}$ourism using Expert $\textbf{T}$ree-$\textbf{o}$f-$\textbf{T}$hought (LETToT), a framework that leverages expert-derived reasoning structures-instead of labeled data-to access LLMs in tourism. First, we iteratively refine and validate hierarchical ToT components through alignment with generic quality dimensions and expert feedback. Results demonstrate the effectiveness of our systematically optimized expert ToT with 4.99-14.15\% relative quality gains over baselines. Second, we apply LETToT's optimized expert ToT to evaluate models of varying scales (32B-671B parameters), revealing: (1) Scaling laws persist in specialized domains (DeepSeek-V3 leads), yet reasoning-enhanced smaller models (e.g., DeepSeek-R1-Distill-Llama-70B) close this gap; (2) For sub-72B models, explicit reasoning architectures outperform counterparts in accuracy and conciseness ($p<0.05$). Our work established a scalable, label-free paradigm for domain-specific LLM evaluation, offering a robust alternative to conventional annotated benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。