构建韩国民众公开接口基准,提升大模型多步调用能力
Multi-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe

- 通过真实调用验证工具链,构建可执行的多步任务数据
- 9B模型经训练后接近27B模型性能,跨基准显著提升
- 适合关注政务类AI Agent落地与工具调用优化的研究者
数据主权法规要求公共机构部署开源、本地化的大型语言模型代理,以在实时政府API间执行多步工具调用。然而,现有开源模型在此类任务中表现持续不佳,且缺乏有效评估基准。我们提出韩国开放公共API基准(KOPA-Bench),包含145个真实任务。为缩小性能差距,我们提出EDGE——一种基于实际执行的动态图数据合成方法。EDGE构建工具输出到输入的依赖图,仅保留真实调用成功后的连接,并遍历这些验证链接生成可执行的多步轨迹。使用该数据集通过GRPO微调后,我们的9B模型几乎达到同系列未微调27B模型的性能,在KOPA-Bench和BFCL基准上均有显著提升。
原文摘要 · Abstract (English)
Data-sovereignty regulations increasingly require public institutions to deploy open-source, on-premise LLM agents that chain multiple tool-calls across live government APIs. However, open-source models consistently underperform in this multi-step setting, and no existing benchmark measures the gap. We introduce the Korean Open Public API Benchmark (KOPA-Bench), comprising 145 real-world tasks. To close this gap, we present EDGE, an Execution-grounded Dynamic Graph for tool-calling data synthEsis driven by live execution. EDGE builds a graph of how each tool's output can feed another's input, keeps only the links that succeed when actually called against the live APIs, and traverses these verified links to synthesize executable multi-step trajectories. Fine-tuned via GRPO on the resulting dataset, our 9B model nearly matches the untuned 27B model from the same family, improving substantially not only on KOPA-Bench but also on the BFCL benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。