为人文社科设计高效偏好对齐方法,提升大模型理解力与判断力。
BridgeAlign: Bridging Preference Alignment for Humanities and Social Sciences

- 分三阶段合成人文社科偏好数据:种子文档筛选、基于角色的指令逆向生成、质量退化构建精细对比对。
- 用21万条合成数据对齐模型,在17个基准上超越11个基线,同时提升人类偏好与知识能力。
- 首次实现人文社科领域无权衡的双优表现,适合需要深度理解的学术与写作场景。
尽管大语言模型的数据合成已广泛开展,但主要集中于可验证答案的领域,忽视了人文社科(HSS)中更依赖微妙质量判断的开放性任务。因此,偏好对齐成为覆盖广义HSS任务的自然范式。然而现有方法或成本过高,或不适用于广泛的人文社科领域。为此,我们提出BridgeAlign,首个面向广义人文社科领域的偏好对齐流程,包含三个阶段:i) 种子文档遴选:通过启发式/基于LLM的过滤与文本优化,从网络语料中收集HSS种子文档;ii) 偏好数据合成:通过基于角色的指令逆向生成偏好三元组,并结合问答一致性检查;iii) 偏好优化:不再依赖简单的真人-模型启发式规则,而是先以人文社科质量量规为基础锚定偏好,再通过受控的质量退化生成近边界偏好对,实现更细粒度的质量区分。在超过21万条合成偏好样本的对齐下,Qwen3-8B在17个基准上平均表现优于11个强基线,且在人类偏好与知识能力上均领先,二者间无权衡,经大量实验验证并结合现有理论进行情境化分析。
原文摘要 · Abstract (English)
While data synthesis for large language models (LLMs) is prevalent, it primarily targets domains with verifiable answers, overlooking open-ended humanities and social sciences (HSS), where nuanced quality judgments matter more than objective correctness. This makes preference alignment a natural paradigm for broad HSS tasks. Yet existing methods are either costly or not tailored to broad HSS disciplines. We thus propose BridgeAlign, among the first preference-alignment pipelines for broad HSS disciplines, with three phases: i) Seed Curation: curating HSS seed documents from web corpora via heuristic/LLM-based filtering and text refinement; ii) Preference Data Synthesis: generating preference triplets via persona-based instruction inversion with Q&A consistency checks; iii) Preference Optimization: moving beyond naive human-vs-model heuristics by first grounding preferences in HSS quality rubric, then generating transitional responses via controlled quality degradation to form near-boundary preference pairs for finer-grained quality discrimination. Aligning over 210k synthetic preference samples, BridgeAlign enables Qwen3-8B to achieve the best average across 17 benchmarks against 11 strong baselines; importantly, leading on both human-preference and knowledge-based capabilities at once, with no trade-off between them, as supported by extensive experiments and contextualized by existing theories.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。