首个评估大模型房产服务能力的基准测试,揭示其仍有较大提升空间。
REAL: Benchmarking Abilities of Large Language Models for Housing Transactions and Services
- 构建涵盖5316条数据的房产服务评测集,分4大主题14类任务
- 现有大模型在房产推理与记忆上表现不足,错误率仍高
- 适合研究智能客服、房产AI助手的开发者与评估者参考
大语言模型(LLMs)的发展推动了多个领域聊天机器人的进步。亟需评估这些模型是否能在房产交易与服务中扮演与人类相当的代理角色。本文提出房地产代理大语言模型评估(REAL),是首个专为评估大模型在房产交易与服务领域能力而设计的评测套件。REAL包含5,316条高质量评估样本,覆盖记忆、理解、推理和幻觉四个主题,划分为14个类别,用于检验大模型在真实房产场景中的知识与能力。同时,该评测被用于评估当前最先进的大模型性能。实验结果表明,大模型在实际房产应用中仍有显著改进空间。
原文摘要 · Abstract (English)
The development of large language models (LLMs) has greatly promoted the progress of chatbot in multiple fields. There is an urgent need to evaluate whether LLMs can play the role of agent in housing transactions and services as well as humans. We present Real Estate Agent Large Language Model Evaluation (REAL), the first evaluation suite designed to assess the abilities of LLMs in the field of housing transactions and services. REAL comprises 5,316 high-quality evaluation entries across 4 topics: memory, comprehension, reasoning and hallucination. All these entries are organized as 14 categories to assess whether LLMs have the knowledge and ability in housing transactions and services scenario. Additionally, the REAL is used to evaluate the performance of most advanced LLMs. The experiment results indicate that LLMs still have significant room for improvement to be applied in the real estate field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。