通过反馈循环提升对话生成的意图与情感一致性
PragAlign: Feedback-Guided Pragmatic Alignment for Controlled Synthetic Dialogue Generation

- 用生成-评估-修正循环,逐轮优化对话质量
- 99.5%对话通过评估,显著优于单次生成的72.25%
- 特别改善情感对齐,适合需精准控制的对话场景
合成对话生成可支持隐私受限服务场景的研究,但生成对话必须保持交际意图、情感含义和自然对话流。我们提出PragAlign,一种基于反馈的可控合成对话生成框架,可依据服务上下文、目标意图和目标情绪进行生成,并支持辅助性格特征控制。该框架采用生成-评估-修正循环,由基于大模型的评估器对意图对齐、情感对齐、连贯性、流畅性和综合质量进行评分,并提供针对性反馈,最多进行三轮修正。在800组匹配的对话规范测试中,PragAlign达到99.50%的评估者定义通过率,远高于单次生成的72.25%和无结构反馈重复生成的95.88%。结果表明,多次尝试带来主要提升,而结构化反馈则主要改善最后一环的多约束满足,而非整体平均质量。修正收益集中在情感对齐,也是消融实验中的主要失败模式。另对1,200条生成对话的人工评估显示,意图表达和对话流高度可识别,但情感适当性更不稳定且主观性强。结果支持PragAlign作为提升评估者定义交际约束满足的质量控制框架,同时指出情感实现与独立人类感知质量仍是开放挑战。
原文摘要 · Abstract (English)
Synthetic dialogue generation can support research in privacy-restricted service settings, but generated conversations must preserve communicative intent, affective meaning, and natural dialogue flow. We introduce PragAlign, a feedback-guided framework for controlled synthetic dialogue generation conditioned on service context, target intent, and target emotion, with auxiliary trait-style controls. PragAlign uses a generate--evaluate--revise loop in which an LLM-based evaluator scores intent alignment, emotion alignment, coherence, fluency, and aggregate quality, then provides criterion-specific feedback for up to three refinement rounds. On 800 matched dialogue specifications, PragAlign achieves 99.50\% evaluator-defined acceptance, compared with 72.25\% for one-shot generation and 95.88\% for repeated generation without structured feedback. This indicates that repeated attempts account for much of the gain over one-shot generation, while structured feedback primarily improves last-mile multi-constraint satisfaction rather than broad average quality. Refinement gains are concentrated in emotion alignment, which is also the dominant failure mode in ablations. A separate human evaluation of 1,200 generated dialogues shows that intent expression and dialogue flow are highly recognizable to annotators, while emotion appropriateness is less stable and more subjective. These results support PragAlign as a quality-control framework for improving evaluator-defined communicative constraint satisfaction, while showing that affective realization and independent human-perceived quality remain open challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。