arXiv:2506.10622cs.CLcs.AI2025-06Conference of the …被引 3

一站式对话智能体开发工具,支持生成、评估与可解释性分析

SDialog: A Python Toolkit for End-to-End Agent Building, User Simulation, Dialog Generation, and Evaluation

论文配图:SDialog: A Python Toolkit for End-to-End Agent Building, User Simulation, Dialog Generation, and Evaluation
图 1 · 摘自论文原文
  • 统一框架整合对话生成、评估与可解释性分析
  • 支持多智能体仿真与语音声学建模,实现端到端对话构建
  • 适用于对话系统研究者,提升实验效率与系统理解

我们提出 SDialog,一个 MIT 许可的开源 Python 工具包,将对话生成、评估与机制可解释性整合为单一端到端框架,用于构建和分析基于大语言模型(LLM)的对话智能体。该工具围绕标准化对话表示构建,提供:(1) 基于角色的多智能体仿真,支持可组合编排,实现受控的合成对话生成;(2) 综合评估体系,包含语言学指标、大模型作为评判者及功能正确性验证器;(3) 机制可解释性工具,支持激活检查、特征消融与诱导控制;(4) 全流程音频生成,含 3D 房间建模与麦克风效果模拟。工具兼容所有主流大模型后端,可在统一 API 下实现混合后端实验。通过对话为中心的架构耦合生成、评估与可解释性,SDialog 使研究人员能够更系统地构建、基准测试与理解对话系统。

原文摘要 · Abstract (English)

We present SDialog, an MIT-licensed open-source Python toolkit that unifies dialog generation, evaluation and mechanistic interpretability into a single end-to-end framework for building and analyzing LLM-based conversational agents. Built around a standardized Dialog representation, SDialog provides: (1) persona-driven multi-agent simulation with composable orchestration for controlled, synthetic dialog generation, (2) comprehensive evaluation combining linguistic metrics, LLM-as-a-judge and functional correctness validators, (3) mechanistic interpretability tools for activation inspection and steering via feature ablation and induction, and (4) audio generation with full acoustic simulation including 3D room modeling and microphone effects. The toolkit integrates with all major LLM backends, enabling mixed-backend experiments under a unified API. By coupling generation, evaluation, and interpretability in a dialog-centric architecture, SDialog enables researchers to build, benchmark and understand conversational systems more systematically.

对话系统工具包可解释性多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。