构建旅行领域低资源任务基准,揭示大模型在真实场景下的性能瓶颈
TravelBench : Exploring LLM Performance in Low-Resource Domains
- 基于真实数据构建14个旅行领域NLP任务基准
- 小模型推理能力提升显著,大模型仍存性能天花板
- 适合关注低资源场景落地的大模型研究者
现有大模型评估基准难以反映模型在低资源任务中的真实能力,导致该领域解决方案发展受限。为此,我们基于真实场景的匿名数据,构建了涵盖7类常见NLP任务的14个旅行领域数据集,并分析了多种大模型在其中的表现。结果表明,通用基准评估无法充分揭示模型在低资源任务中的实际表现。即使经过大量训练浮点运算(FLOPs),现成大模型在复杂、领域特定的任务中仍面临性能瓶颈。此外,推理能力对小型模型在特定任务上的表现提升更为显著,使其能更准确地判断任务结果。
原文摘要 · Abstract (English)
Results on existing LLM benchmarks capture little information over the model capabilities in low-resource tasks, making it difficult to develop effective solutions in these domains. To address these challenges, we curated 14 travel-domain datasets spanning 7 common NLP tasks using anonymised data from real-world scenarios, and analysed the performance across LLMs. We report on the accuracy, scaling behaviour, and reasoning capabilities of LLMs in a variety of tasks. Our results confirm that general benchmarking results are insufficient for understanding model performance in low-resource tasks. Despite the amount of training FLOPs, out-of-the-box LLMs hit performance bottlenecks in complex, domain-specific scenarios. Furthermore, reasoning provides a more significant boost for smaller LLMs by making the model a better judge on certain tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。