自动构建垂直领域评测集,动态测试大模型真实表现
TestAgent: Automatic Benchmarking and Exploratory Interaction for Evaluating LLMs in Vertical Domains
- 用检索增强生成自动构造领域问题,实现评测集自动生成
- 通过强化学习动态调整提问策略,揭示模型知识边界与稳定性
- 在医疗、法律等多领域验证,适合需深度评估的专业场景
随着大语言模型(LLMs)在高度专业化的垂直领域广泛应用,其领域性能评估变得至关重要。然而,现有评估方法通常依赖人工构建静态单轮数据集,存在两大局限:(i) 手动构建成本高,需为每个新领域重复进行;(ii) 静态单轮评估与真实应用中的动态多轮交互不匹配,难以评估专业性与稳定性。为此,我们提出TestAgent框架,实现垂直领域的自动基准测试与探索性动态评估。TestAgent利用检索增强生成从用户提供的知识源中生成领域特定问题,并结合两阶段标准生成过程,实现可扩展的自动化基准创建。此外,引入强化学习引导的多轮交互策略,根据实时模型响应自适应决定问题类型,动态探测知识边界与稳定性。在医疗、法律和政府领域的大量实验表明,TestAgent能高效实现跨领域基准生成,并通过动态探索性评估提供更深入的模型行为洞察。本工作建立了一种大语言模型在垂直领域中自动化、深度评估的新范式。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) are increasingly deployed in highly specialized vertical domains, the evaluation of their domain-specific performance becomes critical. However, existing evaluations for vertical domains typically rely on the labor-intensive construction of static single-turn datasets, which present two key limitations: (i) manual data construction is costly and must be repeated for each new domain, and (ii) static single-turn evaluations are misaligned with the dynamic multi-turn interactions in real-world applications, limiting the assessment of professionalism and stability. To address these, we propose TestAgent, a framework for automatic benchmarking and exploratory dynamic evaluation in vertical domains. TestAgent leverages retrieval-augmented generation to create domain-specific questions from user-provided knowledge sources, combined with a two-stage criteria generation process, thereby enabling scalable and automated benchmark creation. Furthermore, it introduces a reinforcement learning-guided multi-turn interaction strategy that adaptively determines question types based on real-time model responses, dynamically probing knowledge boundaries and stability. Extensive experiments across medical, legal, and governmental domains demonstrate that TestAgent enables efficient cross-domain benchmark generation and yields deeper insights into model behavior through dynamic exploratory evaluation. This work establishes a new paradigm for automated and in-depth evaluation of LLMs in vertical domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。