构建跨文化对话数据集,提升多轮对话中社会规范识别能力
MINDS: A Cross-cultural Dialogue Corpus for Social Norm Classification and Adherence Detection
- 用检索增强框架结合语义分块,实现多轮对话中的规范推理
- 在31组跨语言对话上验证,规范遵循判断准确率显著提升
- 适合研究跨文化对话与社交智能系统的学者和开发者
社会规范是指导人际交流的隐性文化准则,不同于客观常识,其具有主观性、情境依赖性和文化差异性,给计算模型带来挑战。现有工作虽提供有价值的规范标注,但多针对孤立语句或合成对话,难以捕捉真实对话的动态多轮特性。本文提出Norm-RAG——一种基于检索增强的代理式框架,用于多轮对话中的精细社会规范推理。该框架建模话语层面属性,包括交际意图、说话人角色、人际框架和语言线索,并通过新型语义分块技术从结构化规范文档中检索依据,实现可解释且上下文敏感的规范遵从与违反判断。我们进一步构建MINDS(Multilingual Interactions with Norm-Driven Speech)——一个包含31组多轮中英、西英对话的双语数据集,每轮均通过多标注者共识标注规范类别与遵从状态,反映跨文化真实规范表达。实验表明,Norm-RAG在规范检测与泛化能力上均有提升,为文化自适应和社会智能对话系统提供了新范式。
原文摘要 · Abstract (English)
Social norms are implicit, culturally grounded expectations that guide interpersonal communication. Unlike factual commonsense, norm reasoning is subjective, context-dependent, and varies across cultures, posing challenges for computational models. Prior works provide valuable normative annotations but mostly target isolated utterances or synthetic dialogues, limiting their ability to capture the fluid, multi-turn nature of real-world conversations. In this work, we present Norm-RAG, a retrieval-augmented, agentic framework for nuanced social norm inference in multi-turn dialogues. Norm-RAG models utterance-level attributes including communicative intent, speaker roles, interpersonal framing, and linguistic cues and grounds them in structured normative documentation retrieved via a novel Semantic Chunking approach. This enables interpretable and context-aware reasoning about norm adherence and violation across multilingual dialogues. We further introduce MINDS (Multilingual Interactions with Norm-Driven Speech), a bilingual dataset comprising 31 multi-turn Mandarin-English and Spanish-English conversations. Each turn is annotated for norm category and adherence status using multi-annotator consensus, reflecting cross-cultural and realistic norm expression. Our experiments show that Norm-RAG improves norm detection and generalization, demonstrates improved performance for culturally adaptive and socially intelligent dialogue systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。