构建新标注体系,让对话系统更懂个人事实。
An Annotation Scheme and Classifier for Personal Facts in Dialogue

- 设计包含人口属性与持有物的新分类体系,支持结构化存储。
- 基于Gemma-300M的多头分类器达81.6%宏F1,比最强少样本LLM高9个百分点。
- 适用于需精准理解用户个人背景的智能对话系统研发者。
大型语言模型的发展推动了个性化对话系统的应用。本文提出一种扩展的个人事实分类标注方案,解决了现有方法(如PeaCoK)的局限性。新方案引入人口统计、拥有物等新类别,以及持续时间、有效性、后续性等属性,支持事实的结构化存储、质量筛选及对话延续性判断。我们从多会话聊天数据中人工标注了2,779条事实,并基于Transformer编码器训练了多头分类器。结合Gemma-300M编码器,该分类器在宏F1上达到81.6±2.6%,显著优于所有少样本LLM基线(最佳为GPT-5.4-mini,72.92%),且计算开销更低。错误分析显示,在语义边界区分、时间维度理解及后续行为推断方面仍存挑战。数据集和分类器已公开。
原文摘要 · Abstract (English)
The advancement of Large Language Models (LLMs) has enabled their application in personalized dialogue systems. We present an extended annotation scheme for personal fact classification that addresses limitations in existing approaches, particularly PeaCoK. Our scheme introduces new categories (Demographics, Possessions) and attributes (Duration, Validity, Followup) that enable structured storage, quality filtering, and identification of facts suitable for dialogue continuation. We manually annotated 2,779 facts from Multi-Session Chat and trained a multi-head classifier based on transformer encoders. Combined with the Gemma-300M encoder, the classifier achieves $81.6 \pm 2.6$\% macro F1, outperforming all few-shot LLM baselines (best: GPT-5.4-mini, 72.92\%) by nearly 9 percentage points while requiring substantially fewer computational resources. Error analysis reveals persistent challenges in semantic boundary disambiguation, temporal aspect interpretation, and pragmatic reasoning for followup assessment. The dataset\footnotemark[1] and classifier\footnotemark[2] are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。