无需标注偏好,通过内容接受度自动对齐语言模型与社区规范。
Density-Guided Response Optimization: Community-Grounded Alignment via Implicit Acceptance Signals
- 利用社区对内容的接受行为构建隐式偏好信号,捕捉其在表征空间中的密度结构。
- 在无标注场景下,模型生成响应被人类与专家更广泛认可,优于传统方法。
- 适合缺乏标注资源或敏感话题的在线社区,具实用价值与伦理优势。
部署于在线社区的语言模型需适应不同社会、文化及领域语境下的规范。以往对齐方法依赖显式偏好监督或预设原则,在资源充足场景有效,却难以覆盖多数在线社区——尤其是缺乏机构支持、标注基础设施或涉及敏感话题的群体,因偏好收集成本高、伦理风险大或文化错位而受限。我们观察到,社区已通过内容接受、互动和留存等行为隐式表达偏好。这些行为在表征空间中形成可测量的几何结构:被接受的内容占据连贯且高密度区域,反映社区特定规范;被拒绝内容则位于稀疏或错位区域。我们据此提出密度引导响应优化(DGRO),无需显式偏好标签即可对齐语言模型。基于标注偏好数据,我们证明局部密度可恢复成对社区判断,表明几何结构蕴含有意义的偏好信号。随后在跨平台、跨主题、跨语言的多种低标注环境中应用DGRO,结果表明,经DGRO对齐的模型生成内容在人类标注者、领域专家及模型评估中均优于监督与提示基线。我们主张将DGRO作为显式偏好监督不可行或与实际实践脱节时的可行替代方案,并讨论从涌现接受行为中学习的潜在影响与风险。
原文摘要 · Abstract (English)
Language models deployed in online communities must adapt to norms that vary across social, cultural, and domain-specific contexts. Prior alignment approaches rely on explicit preference supervision or predefined principles, which are effective for well-resourced settings but exclude most online communities -- particularly those without institutional backing, annotation infrastructure, or organized around sensitive topics -- where preference elicitation is costly, ethically fraught, or culturally misaligned. We observe that communities already express preferences implicitly through what content they accept, engage with, and allow to persist. We show that this acceptance behavior induces measurable geometric structure in representation space: accepted responses occupy coherent, high-density regions that reflect community-specific norms, while rejected content falls in sparser or misaligned areas. We operationalize this structure as an implicit preference signal for alignment and introduce density-guided response optimization (DGRO), a method that aligns language models to community norms without requiring explicit preference labels. Using labeled preference data, we demonstrate that local density recovers pairwise community judgments, indicating that geometric structure encodes meaningful preference signal. We then apply DGRO in annotation-scarce settings across diverse communities spanning platform, topic, and language. DGRO-aligned models consistently produce responses preferred by human annotators, domain experts, and model-based judges over supervised and prompt-based baselines. We position DGRO as a practical alignment alternative for communities where explicit preference supervision is unavailable or misaligned with situated practices, and discuss the implications and risks of learning from emergent acceptance behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。