arXiv:2506.02449cs.CLcs.HC2025-06EMNLP

用合成数据评测对话系统隐式个性化能力

IP-Dialog: Evaluating Implicit Personalization in Dialogue Systems with Synthetic Data

  • 自动生成含12类用户属性的合成对话数据
  • 涵盖10项任务,验证模型对背景信息的理解与推理
  • 提供因果图分析模型决策路径,适合研究个性化对话者

现代对话系统需从对话中隐式推断用户背景并提供个性化服务,但高质量数据稀缺。传统数据构建方法耗时耗力且存在隐私风险。为此,我们提出一种自动合成数据生成方法,构建了包含10项任务和12种用户属性的隐式个性化对话(IP-Dialog)基准数据集。同时设计了一套包含四项指标的评估框架,用于衡量模型对用户属性的感知与推理能力。我们还提出了五种因果图,揭示模型在隐式个性化中的推理路径。大量实验验证了数据集的可靠性与有效性。

原文摘要 · Abstract (English)

In modern dialogue systems, the ability to implicitly infer user backgrounds from conversations and leverage this information for personalized assistance is crucial. However, the scarcity of high-quality data remains a fundamental challenge to evaluating and improving this capability. Traditional dataset construction methods are labor-intensive, resource-demanding, and raise privacy concerns. To address these issues, we propose a novel approach for automatic synthetic data generation and introduce the Implicit Personalized Dialogue (IP-Dialog) benchmark along with a training dataset, covering 10 tasks and 12 user attribute types. Additionally, we develop a systematic evaluation framework with four metrics to assess both attribute awareness and reasoning capabilities. We further propose five causal graphs to elucidate models' reasoning pathways during implicit personalization. Extensive experiments yield insightful observations and prove the reliability of our dataset.

对话系统个性化合成数据评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。