用对抗性训练让大模型学会真正理解社交情境
Social-R1: Towards Human-like Social Reasoning in LLMs
- 设计难例基准ToMBench-Hard,逼模型深度推理
- 40亿参数模型超越更大模型,跨8个任务泛化强
- 通过过程奖励对齐人类思维,不只看结果
尽管大型语言模型在多个领域表现出色,但社会智能——即感知社会线索、推断心理状态并生成适当回应的能力——仍是关键挑战,尤其在实现人机协作和开发真正服务人类需求的AI方面。当前模型常依赖表面模式而非真正的社会推理。我们主张,培养类人社会智能需在难以走捷径的挑战性案例上进行训练。为此,我们提出ToMBench-Hard,一个对抗性基准,用于提供困难的社交推理训练样本。在此基础上,我们提出Social-R1,一种强化学习框架,通过多维奖励将模型推理与人类认知对齐。不同于基于结果的强化学习,Social-R1监督整个推理过程,强制结构一致性、逻辑完整性和信息密度。结果表明,该方法使一个40亿参数模型超越更大规模的模型,并在八个不同基准上实现稳健泛化。这些发现表明,通过轨迹级对齐的挑战性训练案例,可为高效且可靠的社交智能提供可行路径。
原文摘要 · Abstract (English)
While large language models demonstrate remarkable capabilities across numerous domains, social intelligence - the capacity to perceive social cues, infer mental states, and generate appropriate responses - remains a critical challenge, particularly for enabling effective human-AI collaboration and developing AI that truly serves human needs. Current models often rely on superficial patterns rather than genuine social reasoning. We argue that cultivating human-like social intelligence requires training with challenging cases that resist shortcut solutions. To this end, we introduce ToMBench-Hard, an adversarial benchmark designed to provide hard training examples for social reasoning. Building on this, we propose Social-R1, a reinforcement learning framework that aligns model reasoning with human cognition through multi-dimensional rewards. Unlike outcome-based RL, Social-R1 supervises the entire reasoning process, enforcing structural alignment, logical integrity, and information density. Results show that our approach enables a 4B parameter model to surpass much larger counterparts and generalize robustly across eight diverse benchmarks. These findings demonstrate that challenging training cases with trajectory-level alignment offer a path toward efficient and reliable social intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。