用真实对话自动生成代码任务评估标准,省时高效且准确
Conv-to-Bench: Evaluating Language Models Via User-Assistant Dialogues In Code Tasks

- 从用户与助手的真实多轮对话中自动提取任务指令和验证标准
- 生成的评估集与人工标准高度一致,相关性达ρ=1.000,计算开销低
- 适合需要快速构建高质量评估基准的研究者和开发者
大型语言模型(LLMs)的快速发展已超出传统评估基准的可扩展性,后者仍严重依赖人力专家标注。为此,我们提出Conv-to-Bench,一种多阶段框架,可将真实的多轮用户-助手对话自动转化为结构化、可验证的需求清单。通过利用真实对话日志中的“指令演化”现象,该方法将零散的用户意图拆解为整合后的指令和二元评价标准。应用于编程领域时,Conv-to-Bench生成的评估集与人工编写的基准(如BigCodeBench)近乎完全对齐,斯皮尔曼相关系数最高达ρ = 1.000,且计算开销显著降低。对LLM作为裁判者的验证进一步确认其可靠性,主评估者与人工验证真值间达到κ = 0.705的一致性。全面消融研究显示,尽管多轮交互能捕捉用户意图的迭代演化,但以指令为中心的提取方式更具鲁棒性。最终,Conv-to-Bench提供了一种可扩展、低成本的范式,可随用户导向型AI应用的多样化持续维持高保真评估标准。
原文摘要 · Abstract (English)
The rapid advancement of Large Language Models (LLMs) has outpaced the scalability of traditional evaluation benchmarks, which remain heavily dependent on labor-intensive expert curation. We address this bottleneck with Conv-to-Bench, a multi-stage framework that automatically transforms authentic multi-turn user-assistant dialogues into structured, verifiable requirement checklists. By leveraging the "instructional evolution" found in real-world conversational logs, our approach deconstructs fragmented user intent into consolidated instructions and binary evaluation criteria. Applied to the programming domain, Conv-to-Bench produces evaluation sets that demonstrate near-perfect alignment with human-authored standards like BigCodeBench, achieving Spearman correlations of up to $ρ$ = 1.000 with significantly lower computational overhead. Validation of the LLM-as-a-judge framework further confirms its reliability, with the primary evaluator achieving substantial agreement with human-verified ground truth ($κ$ = 0.705). Our comprehensive ablation studies reveal that while multi-turn interactions capture the iterative evolution of user intent, instruction-centric extraction provides a more robust foundation. Ultimately, Conv-to-Bench provides a scalable, cost-effective paradigm for maintaining high-fidelity evaluation standards as user-centric AI applications continue to diversify.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。