arXiv:2502.04349cs.CLcs.AI2025-02被引 1

用虚拟用户动态评估大模型对话数据采集能力

Dynamic benchmarking framework for LLM-based conversational data capture

  • 构建仿真用户模拟多轮对话,动态测试模型表现
  • 少样本条件下自适应策略提升数据提取准确率
  • 适合评估真实场景中对话系统的实用性能

大语言模型的快速发展推动了对话系统的发展,但现有评估框架多聚焦单一任务,难以捕捉多轮对话的动态特性。本文提出一种动态基准测试框架,通过与合成用户交互来评估基于大语言模型的对话代理。该框架结合生成式智能体模拟,从信息抽取、上下文感知和自适应互动三个维度进行评估。通过模拟多样化的用户行为,实现了可扩展、自动化且灵活的评估方法。在贷款申请场景下的实验表明,在单次和少样本提取条件下,该框架能有效验证模型性能。结果发现,自适应策略显著提升了对模糊回应的处理能力,提高了数据提取准确性。未来工作将拓展至更广泛领域,并引入对话连贯性、用户参与度等新指标。本研究为评估基于大语言模型的对话系统提供了结构化、可扩展的方法,有助于实际应用部署。

原文摘要 · Abstract (English)

The rapid evolution of large language models (LLMs) has transformed conversational agents, enabling complex human-machine interactions. However, evaluation frameworks often focus on single tasks, failing to capture the dynamic nature of multi-turn dialogues. This paper introduces a dynamic benchmarking framework to assess LLM-based conversational agents through interactions with synthetic users. The framework integrates generative agent simulation to evaluate performance on key dimensions: information extraction, context awareness, and adaptive engagement. By simulating various aspects of user behavior, our work provides a scalable, automated, and flexible benchmarking approach. Experimental evaluation - within a loan application use case - demonstrates the framework's effectiveness under one-shot and few-shot extraction conditions. Results show that adaptive strategies improve data extraction accuracy, especially when handling ambiguous responses. Future work will extend its applicability to broader domains and incorporate additional metrics (e.g., conversational coherence, user engagement). This study contributes a structured, scalable approach to evaluating LLM-based conversational agents, facilitating real-world deployment.

对话系统评估框架LLM动态测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。