现有大模型评测与真实用户需求严重脱节,研究提出新框架揭示差距。
LLM Benchmark-User Need Misalignment for Climate Change
- 构建人类与AI知识获取行为框架,识别不同交互模式。
- 发现现有评测与真实用户需求存在显著不匹配。
- 适合关注模型评测、RAG系统和气候知识服务的研究者。
气候变化是影响公共决策与政策讨论的重大社会科学议题。随着大语言模型(LLMs)日益成为获取气候知识的入口,现有评测是否反映真实用户需求,对评估模型在实际场景中的表现至关重要。本文提出一种主动知识行为框架,捕捉人类与人类、人类与AI之间知识寻求与提供行为的差异。进一步构建了主题-意图-形式分类体系,并应用于分析代表不同知识行为的气候相关数据。结果表明,当前评测与真实用户需求存在显著偏差,而人类与大模型之间的知识互动模式与人际互动高度相似。这些发现为评测设计、检索增强生成(RAG)系统开发及大模型训练提供了可操作指导。代码已开源:https://github.com/OuchengLiu/LLM-Misalign-Climate-Change。
原文摘要 · Abstract (English)
Climate change is a major socio-scientific issue shapes public decision-making and policy discussions. As large language models (LLMs) increasingly serve as an interface for accessing climate knowledge, whether existing benchmarks reflect user needs is critical for evaluating LLM in real-world settings. We propose a Proactive Knowledge Behaviors Framework that captures the different human-human and human-AI knowledge seeking and provision behaviors. We further develop a Topic-Intent-Form taxonomy and apply it to analyze climate-related data representing different knowledge behaviors. Our results reveal a substantial mismatch between current benchmarks and real-world user needs, while knowledge interaction patterns between humans and LLMs closely resemble those in human-human interactions. These findings provide actionable guidance for benchmark design, RAG system development, and LLM training. Code is available at https://github.com/OuchengLiu/LLM-Misalign-Climate-Change.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。