arXiv:2410.03839cs.CL2024-10

构建更忠实的广告文本生成评估数据集,解决现有数据失真问题。

FaithCAMERA: Construction of a Faithful Dataset for Ad Text Generation

  • 联合广告创作者精修参考文本,确保与输入文档一致。
  • 剔除不忠实体训练数据后,实体级忠实度与信息量提升,句级下降。
  • 适合关注广告生成真实性与数据质量的研究者使用。

在广告文本生成(ATG)中,优质广告文本需兼具忠实性与信息量:既要忠实于输入文档,又要包含吸引潜在客户的重点信息。现有评估数据集CAMERA(arXiv:2309.12030)虽适合评估信息量,但其参考文本常包含与输入不符的信息,阻碍了ATG研究进展。本文与内部广告创作者合作,精修CAMERA参考文本,构建新评估数据集FaithCAMERA,确保参考文本的忠实性。利用FaithCAMERA,可评估现有提升忠实性的方法在保持信息量的同时生成高质量广告文本的能力。实验表明,移除包含不忠实体的训练数据后,实体级别忠实度与信息量均提升,但句级表现下降。这表明未来ATG研究不仅需扩大数据规模,更需保障数据忠实性。本数据集将公开发布。

原文摘要 · Abstract (English)

In ad text generation (ATG), desirable ad text is both faithful and informative. That is, it should be faithful to the input document, while at the same time containing important information that appeals to potential customers. The existing evaluation data, CAMERA (arXiv:2309.12030), is suitable for evaluating informativeness, as it consists of reference ad texts created by ad creators. However, these references often include information unfaithful to the input, which is a notable obstacle in promoting ATG research. In this study, we collaborate with in-house ad creators to refine the CAMERA references and develop an alternative ATG evaluation dataset called FaithCAMERA, in which the faithfulness of references is guaranteed. Using FaithCAMERA, we can evaluate how well existing methods for improving faithfulness can generate informative ad text while maintaining faithfulness. Our experiments show that removing training data that contains unfaithful entities improves the faithfulness and informativeness at the entity level, but decreases both at the sentence level. This result suggests that for future ATG research, it is essential not only to scale the training data but also to ensure their faithfulness. Our dataset will be publicly available.

广告生成数据集忠实性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。