arXiv:2509.24792cs.CL2025-09EMNLP

提出新评估方法,更准判断自动生成缝纫步骤的时空连贯性。

Evaluating Spatiotemporal Consistency in Automatically Generated Sewing Instructions

  • 基于树结构的自动评估法,捕捉步骤间时空关系。
  • 与人工标注错误数和评分相关性更高,优于传统指标。
  • 对刻意设计的干扰案例更鲁棒,适合质检类应用。

本文提出一种新型树状结构自动评估指标,用于衡量大语言模型生成的分步装配指令在时空一致性上的表现,相比传统的BLEU和BERT相似度评分更具准确性。我们将该指标应用于缝纫指令领域,实验表明其与人工标注的错误数量及人类质量评分的相关性更强,证明了该方法在评估缝纫指令时空合理性方面的优越性。进一步实验显示,该指标在面对专门设计以误导依赖文本相似性的评估方法的反事实案例时,具有更强的鲁棒性。

原文摘要 · Abstract (English)

In this paper, we propose a novel, automatic tree-based evaluation metric for LLM-generated step-by-step assembly instructions, that more accurately reflects spatiotemporal aspects of construction than traditional metrics such as BLEU and BERT similarity scores. We apply our proposed metric to the domain of sewing instructions, and show that our metric better correlates with manually-annotated error counts as well as human quality ratings, demonstrating our metric's superiority for evaluating the spatiotemporal soundness of sewing instructions. Further experiments show that our metric is more robust than traditional approaches against artificially-constructed counterfactual examples that are specifically constructed to confound metrics that rely on textual similarity.

评估指标自然语言生成缝纫指令

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。