用任务示例增强语义表示,提升对话生成质量,尤其在小数据和复杂任务中效果显著。
Impact of enriched meaning representations for language generation in dialogue tasks: A comprehensive exploration of the relevance of tasks, corpora and metrics
- 输入中加入任务示例(MR-句子对)以丰富语义表示,提升生成效果。
- 小数据、高变异性场景下,增强输入使生成质量显著提升,零样本也有效。
- 基于人工评分的语义评估指标更易发现遗漏等细微问题,优于纯嵌入式指标。
对话系统需生成多样语言形式以实现流畅准确交互。自然语言生成(NLG)引擎将语义表示(MRs)转化为句子,直接影响用户感知。传统MRs通过对话动作(DA)编码交流功能(如告知、请求、确认),并通过槽位-值对列举语义内容。本文研究在训练与推理时引入原始数据中提取的MR-句子对作为任务示例,是否能提升微调模型的生成质量。分析涵盖五种关注不同语言维度的指标,以及四个在领域、规模、词汇、MR变异性和采集方式上各不相同的语料库。据我们所知,这是首个在多领域、多语料特征及评估指标下系统比较MR影响的对话NLG研究。关键发现:增强输入在复杂任务和小数据集、高MR与句子变异性场景中有效;在零样本设置中对任意领域均有帮助。语义类指标比词汇类指标更能准确反映生成质量;其中基于人工评分的语义指标可检测嵌入式指标常忽略的遗漏等问题。此外,指标得分演变及槽位准确率、对话动作准确率的优异表现表明生成模型在语义与沟通意图层面具备快速适应能力与鲁棒性。
原文摘要 · Abstract (English)
Conversational systems should generate diverse language forms to interact fluently and accurately with users. In this context, Natural Language Generation (NLG) engines convert Meaning Representations (MRs) into sentences, directly influencing user perception. These MRs usually encode the communicative function (e.g., inform, request, confirm) via DAs and enumerate the semantic content with slot-value pairs. In this work, our objective is to analyse whether providing a task demonstrator to the generator enhances the generations of a fine-tuned model. This demonstrator is an MR-sentence pair extracted from the original dataset that enriches the input at training and inference time. The analysis involves five metrics that focus on different linguistic aspects, and four datasets that differ in multiple features, such as domain, size, lexicon, MR variability, and acquisition process. To the best of our knowledge, this is the first study on dialogue NLG implementing a comparative analysis of the impact of MRs on generation quality across domains, corpus characteristics, and the metrics used to evaluate these generations. Our key insight is that the proposed enriched inputs are effective for complex tasks and small datasets with high variability in MRs and sentences. They are also beneficial in zero-shot settings for any domain. Moreover, the analysis of the metrics shows that semantic metrics capture generation quality more accurately than lexical metrics. In addition, among these semantic metrics, those trained with human ratings can detect omissions and other subtle semantic issues that embedding-based metrics often miss. Finally, the evolution of the metric scores and the excellent results for Slot Accuracy and Dialogue Act Accuracy demonstrate that the generative models present fast adaptability to different tasks and robustness at semantic and communicative intention levels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。