评测对话系统的多维度能力,关注语言、文化与安全的全面评估。
Overview of Dialog System Evaluation Track: Dimensionality, Language, Culture and Safety at DSTC 12
- 构建多维度自动评价框架,覆盖10类对话属性
- 跨语言安全检测最高准确率0.9648,文化敏感性仍需提升
- 适合研究对话评估、AI安全与跨文化AI的学者参考
大型语言模型的快速发展加剧了对话系统评估的需求,但现有评估手段仍显不足,尤其在安全性方面常存在定义狭窄或文化偏见。DSTC12 Track 1“对话系统评估:维度、语言、文化与安全”旨在填补这一空白,包含两个子任务:(1)对话级多维度自动评估指标,(2)多语言多文化安全检测。任务1针对10个对话维度,以Llama-3-8B为基线,平均斯皮尔曼相关系数达0.1681,显示改进空间巨大。任务2中,参赛模型在多语言安全子集上显著超越Llama-Guard-3-1B基线(最高ROC-AUC 0.9648),但在文化敏感性子集上基线表现更优(0.5126 ROC-AUC),凸显文化感知安全的挑战。本文介绍所提供数据集、基线模型及各任务提交结果。
原文摘要 · Abstract (English)
The rapid advancement of Large Language Models (LLMs) has intensified the need for robust dialogue system evaluation, yet comprehensive assessment remains challenging. Traditional metrics often prove insufficient, and safety considerations are frequently narrowly defined or culturally biased. The DSTC12 Track 1, "Dialog System Evaluation: Dimensionality, Language, Culture and Safety," is part of the ongoing effort to address these critical gaps. The track comprised two subtasks: (1) Dialogue-level, Multi-dimensional Automatic Evaluation Metrics, and (2) Multilingual and Multicultural Safety Detection. For Task 1, focused on 10 dialogue dimensions, a Llama-3-8B baseline achieved the highest average Spearman's correlation (0.1681), indicating substantial room for improvement. In Task 2, while participating teams significantly outperformed a Llama-Guard-3-1B baseline on the multilingual safety subset (top ROC-AUC 0.9648), the baseline proved superior on the cultural subset (0.5126 ROC-AUC), highlighting critical needs in culturally-aware safety. This paper describes the datasets and baselines provided to participants, as well as submission evaluation results for each of the two proposed subtasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。