arXiv:2505.24683cs.CLcs.AI2025-05EMNLP被引 6

用反馈提升用户对机器翻译的判断力,尤其问答表效果最佳

Should I Share this Translation? Evaluating Quality Feedback for User Reliance on Machine Translation

  • 通过显式(错误标注、大模型解释)和隐式(反译、问答表)反馈提升用户决策
  • 除错误标注外,所有反馈均显著提高判断准确性和合理使用率
  • 问答表效果最优,帮助高、负担低,适合非母语者使用

随着人工智能在工作与日常生活中的广泛应用,用户需要可靠的反馈机制来负责任地使用AI,尤其是在自身无法评估AI输出质量的场景中。本文研究了单语用户在是否分享机器翻译结果时的决策行为,先无反馈,再引入四种质量反馈:(1) 错误标注的显式反馈;(2) 大语言模型解释的显式反馈;(3) 反向翻译的隐式反馈;(4) 问答表的隐式反馈。结果表明,除错误标注外,其余三种反馈均显著提升决策准确性和合理依赖度。特别地,隐式反馈(尤其是问答表)在决策准确率、合理依赖、用户感知方面均优于显式反馈,获得最高帮助性与信任度评分,同时心理负担最低。

原文摘要 · Abstract (English)

As people increasingly use AI systems in work and daily life, feedback mechanisms that help them use AI responsibly are urgently needed, particularly in settings where users are not equipped to assess the quality of AI predictions. We study a realistic Machine Translation (MT) scenario where monolingual users decide whether to share an MT output, first without and then with quality feedback. We compare four types of quality feedback: explicit feedback that directly give users an assessment of translation quality using (1) error highlights and (2) LLM explanations, and implicit feedback that helps users compare MT inputs and outputs through (3) backtranslation and (4) question-answer (QA) tables. We find that all feedback types, except error highlights, significantly improve both decision accuracy and appropriate reliance. Notably, implicit feedback, especially QA tables, yields significantly greater gains than explicit feedback in terms of decision accuracy, appropriate reliance, and user perceptions, receiving the highest ratings for helpfulness and trust, and the lowest for mental burden.

机器翻译人机交互反馈机制用户依赖

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。