一站式对话智能体开发工具,支持生成、评估与可解释性分析
SDialog: A Python Toolkit for End-to-End Agent Building, User Simulation, Dialog Generation, and Evaluation
- 统一对话表示框架,支持角色驱动的多智能体模拟
- 集成语言指标、LLM评判与功能正确性验证的综合评估体系
- 适合对话系统研究者,尤其关注可解释性与跨后端实验
我们提出 SDialog,一个 MIT 许可的开源 Python 工具包,将对话生成、评估与机制可解释性整合到一个端到端框架中,用于构建和分析基于大语言模型(LLM)的对话智能体。该工具包以标准化的 exttt{Dialog} 表示为核心,提供:(1) 可组合编排的角色驱动多智能体模拟,实现受控的合成对话生成;(2) 结合语言学指标、LLM-as-a-judge 与功能正确性验证的综合性评估;(3) 激活检查与特征扰动/诱导的机制可解释性工具;(4) 支持 3D 房间建模与麦克风效应的完整音频生成。工具包兼容所有主流 LLM 后端,通过统一 API 支持混合后端实验。通过在对话为中心的架构中融合生成、评估与可解释性,SDialog 使研究人员能够更系统地构建、评测与理解对话系统。
原文摘要 · Abstract (English)
We present SDialog, an MIT-licensed open-source Python toolkit that unifies dialog generation, evaluation and mechanistic interpretability into a single end-to-end framework for building and analyzing LLM-based conversational agents. Built around a standardized \texttt{Dialog} representation, SDialog provides: (1) persona-driven multi-agent simulation with composable orchestration for controlled, synthetic dialog generation, (2) comprehensive evaluation combining linguistic metrics, LLM-as-a-judge and functional correctness validators, (3) mechanistic interpretability tools for activation inspection and steering via feature ablation and induction, and (4) audio generation with full acoustic simulation including 3D room modeling and microphone effects. The toolkit integrates with all major LLM backends, enabling mixed-backend experiments under a unified API. By coupling generation, evaluation, and interpretability in a dialog-centric architecture, SDialog enables researchers to build, benchmark and understand conversational systems more systematically.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。