构建闭环框架,让对话系统更像真人聊天。
Toward Anthropomorphic Dialogue: A Closed-Loop Framework for Human-Like Chat Generation, Evaluation, and Preference Alignment

- 用角色设定、记忆和虚拟时间实现拟人化对话运行时
- 在55种人格下测试,最强模型准确率达39.00%
- 奖励机制聚焦薄弱技能,适合研究对话智能的团队
真实私人对话不仅需要流畅回复,还需保持人格、关系、长期记忆、知识边界、媒介特定时机及连贯多轮结构。我们提出AnthroDial闭环框架,将拟人化对话建模为系统架构、可执行评估与诊断对齐的联合问题。该框架包含:(1) 基于角色条件的调度对话运行时,融合人格卡、情景卡、长期记忆、虚拟时间与单次决策;(2) 可执行基准测试,含L0有效性门控、每轮五个维度与对话级五个维度;(3) 后训练流水线,筛选16,436个调度决策样本用于SFT,采用GRPO并引入认知诊断、ZPD感知奖励。奖励通过卡尔曼滤波维护各行为维度的能力估计,优先提升能力缺口大的维度,并以滚动回放得分作为任务级ZPD匹配,聚焦优化可学弱项。在包含55个人格、50种情景、50个角色-情景绑定及每模型100个角色条件案例的基准上,评估了16个系统,涵盖前沿基线、开源模型、思考/无思考变体及SFT/强化学习消融实验。最强非训练基线达32.00%严格准确率,而Qwen3.6-27B-SFT+RL达到39.00%严格准确率与98.5总体得分。在9B无思考设置中,SFT与强化学习使严格准确率从0.00%提升至13.00%与18.37%。结果表明,当生成、评估与奖励设计共享同一行为维度时,拟人化对话性能显著提升。
原文摘要 · Abstract (English)
Human-like private chat requires more than fluent response generation: a system must preserve persona, relationship, memory, bounded knowledge, medium-specific timing, and a coherent multi-turn arc. We present AnthroDial, a closed-loop framework that formulates anthropomorphic dialogue as a joint problem of system architecture, executable evaluation, and diagnostic alignment. It combines (1) a role-conditioned scheduled dialogue runtime with persona and scenario cards, long-term memory, virtual time, and single-draft message decisions; (2) an executable benchmark with an L0 validity gate, five per-turn dimensions, and five dialogue-level dimensions; and (3) a post-training pipeline that filters 16,436 scheduled-decision examples for SFT and applies GRPO with a cognitive-diagnostic, ZPD-aware reward. The reward maintains Kalman-filtered capability estimates for each behavioral dimension, upweights dimensions with larger capability deficits, and uses rollout scores as task-level ZPD matches to focus optimization on learnable weak skills. On a benchmark with 55 personas, 50 scenarios, 50 persona-scenario bindings, and 100 role-conditioned cases per model, we evaluate 16 systems spanning frontier baselines, open models, thinking/no-think variants, and SFT/RL ablations. The strongest non-trained baseline reaches 32.00% strict ACC, while Qwen3.6-27B-SFT+RL reaches 39.00% strict ACC and a 98.5 overall score. In the 9B no-think setting, SFT and RL improve strict ACC from 0.00% to 13.00% and 18.37%. These results show that anthropomorphic dialogue benefits when generation, evaluation, and reward shaping share the same behavioral dimensions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。