构建用户中心的金融能力评测基准,评估大模型真实金融场景表现。
UCFE: A User-Centric Financial Expertise Benchmark for Large Language Models
- 结合专家评分与动态任务交互,模拟真实金融变化情境。
- 11个大模型在该基准上得分与人类偏好相关性达0.78。
- 适合研究金融大模型性能与用户体验的学者和开发者。
本文提出UCFE:一种以用户为中心的金融专业知识评测基准,旨在评估大语言模型(LLMs)处理复杂现实金融任务的能力。UCFE采用混合方法,结合人工专家评估与动态、任务特定的交互,模拟不断演变的金融场景。首先,我们对804名参与者进行了用户研究,收集其在金融任务中的反馈;其次,基于这些反馈构建了涵盖广泛用户意图与交互的数据集。该数据集作为基准,用于评估11个LLM服务,采用LLM-as-Judge方法。结果表明,基准得分与人类偏好间存在显著一致性,皮尔逊相关系数为0.78,验证了UCFE数据集及评估方法的有效性。UCFE不仅揭示了大模型在金融领域的潜力,还提供了评估其性能与用户满意度的稳健框架。
原文摘要 · Abstract (English)
This paper introduces the UCFE: User-Centric Financial Expertise benchmark, an innovative framework designed to evaluate the ability of large language models (LLMs) to handle complex real-world financial tasks. UCFE benchmark adopts a hybrid approach that combines human expert evaluations with dynamic, task-specific interactions to simulate the complexities of evolving financial scenarios. Firstly, we conducted a user study involving 804 participants, collecting their feedback on financial tasks. Secondly, based on this feedback, we created our dataset that encompasses a wide range of user intents and interactions. This dataset serves as the foundation for benchmarking 11 LLMs services using the LLM-as-Judge methodology. Our results show a significant alignment between benchmark scores and human preferences, with a Pearson correlation coefficient of 0.78, confirming the effectiveness of the UCFE dataset and our evaluation approach. UCFE benchmark not only reveals the potential of LLMs in the financial domain but also provides a robust framework for assessing their performance and user satisfaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。