用合成对话自动发现社交规范,提升模型对人际互动规则的理解能力。
"Hiding in Plain Sight": Designing Synthetic Dialog Generation for Uncovering Socially Situated Norms
- 通过自评估与规范发现机制生成对话,不依赖预设标签。
- 构建了包含多种人物属性与对话轨迹的高质量合成数据集。
- 适合研究社交规范、对话理解与伦理对齐的学者与工程师。
自然语境中的对话蕴含着深层的社会规范,体现对话者间的关系与沟通意图。本文提出一种多阶段框架,通过自我评估与规范发现机制,从丰富上下文的交互中自动挖掘社会规范,而非依赖预定义标签。基于该框架,我们构建了 NormHint——一个涵盖广泛对话者属性(如年龄、职业、性格)、关系类型、话题及对话发展路径的合成对话数据集。该数据集在每一轮对话中精细标注了规范违反情况、参与者详细描述及修复建议,包括早期干预带来的替代路径。人工验证与自动化分析表明,该数据集在多样话题上具备高度自然性与真实性。此外,实验发现,使用我们的规范违反数据微调模型后,其识别和理解对话中潜在规范违反的能力显著提升。
原文摘要 · Abstract (English)
Naturally situated conversations encapsulate the social norms inherent to their context, reflecting both the relationships between interlocutors and the underlying communicative intent. In this paper, we propose a novel, multi-step framework for generating dialogues that automatically uncovers social norms from rich, context-laden interactions through a process of self-assessment and norm discovery, rather than relying on predefined norm labels. Leveraging this framework, we construct NormHint, a comprehensive synthetic dialogue dataset spanning a wide range of interlocutor attributes (e.g., age, profession, personality), relationship types, conversation topics, and conversational trajectories. NormHint is meticulously annotated with turn-level norm violation information, detailed participant descriptions, and remediation suggestions-including alternative trajectories achieved through early intervention. Human validation and automated analysis demonstrate that our dataset captures diverse conversational topics with high naturalness and realism. Moreover, we discovered that fine-tuning a model with our norm violation data significantly enhances its ability to detect and understand potential norm violations in conversations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。