arXiv:2608.19083cs.HCcs.CL2026-08

AI翻译可读性高却难评估,源文可见也不代表质量判断准确。

When Readability and Source Retention Diverge: An Evaluability Gap in AI Translation

论文配图:When Readability and Source Retention Diverge: An Evaluability Gap in AI Translation
图 1 · 摘自论文原文
  • 对比可读性与忠实度两种输出风格,检验源文条件对评价的影响。
  • 复杂文本中,忠实度输出保留更多信息,但可读性输出仍获更高评分。
  • 源文可见不等于能有效评估内容保留,影响用户信任与数据披露意愿。

可读性高的AI生成结果可能造成评价空白:即使提供源文,整体质量判断未必反映输出对源文的保留程度。本研究通过2×2实验设计(N=306)考察源文类型与输出呈现方式对感知翻译质量的影响,并分析输出评价与系统信任、声明披露意愿的关系。使用TransLingo平台,比较简单叙事与复杂哲思文学文本在大模型生成的可读性输出与研究者修订的忠实度输出之间的差异。描述性刺激审计显示,在两种源文条件下,忠实度输出均具备更高的源文保留率。因子分析揭示呈现方式与源文条件存在显著交互作用:对于简单文本,忠实度输出评分更高;而对于复杂文本,两类输出评分无显著差异。感知智能、拟人化归因及任务信任也呈现类似依赖源文条件的模式。结构方程模型进一步表明,任务信任是声明披露意愿的近端预测因子,且各评价维度间存在并发关联。结果说明,源文可见并不等同于可评价,尤其在复杂文本中,展示源文无法确保质量评分反映内容保留差异。这区分了翻译输出评估支持与个人文本委托决策的数据处理支持。

原文摘要 · Abstract (English)

Readable AI output can leave an evaluability gap: even when the source is shown, an overall-quality judgment may not reflect what an output preserves. We investigated how source-text condition and output rendering relate to perceived translation quality, and how output and system appraisals relate to trust and stated disclosure willingness in a plain-text interface. A focal 2 * 2 comparison (N=306) using TransLingo examined simple generated narratives and complex literary-philosophical prose alongside LLM-generated readability-oriented outputs and researcher-revised fidelity-oriented outputs. A descriptive stimulus audit indicated greater source retention in fidelity-oriented outputs in both source-text conditions. Factorial analyses showed a significant rendering-by-source-text-condition interaction in perceived quality. Participants rated fidelity-oriented outputs higher than readability-oriented outputs for the simple narratives, whereas no reliable rendering difference emerged for the complex prose. A corresponding source-condition-dependent pattern was observed for perceived intelligence, agency-oriented anthropomorphic attribution, and task-performance trust. A separate theory-ordered appraisal-structure SEM characterized concurrent associations among perceived quality, perceived intelligence, agency-oriented anthropomorphic attribution, task-performance trust, and stated disclosure willingness across six domains, with task-performance trust as the proximal correlate of stated willingness. The observed rating pattern distinguishes source access from source evaluability: for the complex stimuli, displaying the source did not ensure that one overall-quality rating reflected differences in retained content. It also separates support for evaluating translation output from data-handling support for decisions about what personal text to entrust to a system.

AI翻译可读性源文保留可信度评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。