arXiv:2502.09532cs.CLcs.AI2025-02被引 15

用户用AI写多语言文案时,会因一语言表现差而连带降低对另一语言的信任。

Mind the Gap! Choice Independence in Using Multilingual LLMs for Persuasive Co-Writing Tasks in Different Languages

  • 测试不同语言的LLM表现如何影响用户后续选择
  • 西班牙语LLM表现差导致英语任务使用率下降32%
  • 女性用户误以为广告是AI生成会减少捐款,适合人机协作设计者参考

生成式AI的兴起催生了新型写作助手,这些系统通常依赖多语言大模型(LLMs),使全球工作者能在多种语言中修改或创作内容。然而,现有研究表明多语言LLM在不同语言上的表现存在显著差异,使用者因此面临输出质量不一致的风险。值得注意的是,近期研究发现人们倾向于将算法错误泛化到独立任务中,违背了选择独立性原则。本文分析了用户在撰写慈善广告任务中使用多语言写作助手时,是否受第二语言性能的影响。结果表明,曾接触西班牙语LLM的用户,其对英语LLM的使用率下降32%。尽管这一行为未影响广告整体说服力,但用户对广告来源的认知却起关键作用:西班牙语女性参与者若认为广告由AI生成,捐赠意愿显著下降。此外,多数人无法准确区分人类与LLM生成的广告。本研究对多语言LLM作为辅助工具的设计、开发与采纳具有重要启示。

原文摘要 · Abstract (English)

Recent advances in generative AI have precipitated a proliferation of novel writing assistants. These systems typically rely on multilingual large language models (LLMs), providing globalized workers the ability to revise or create diverse forms of content in different languages. However, there is substantial evidence indicating that the performance of multilingual LLMs varies between languages. Users who employ writing assistance for multiple languages are therefore susceptible to disparate output quality. Importantly, recent research has shown that people tend to generalize algorithmic errors across independent tasks, violating the behavioral axiom of choice independence. In this paper, we analyze whether user utilization of novel writing assistants in a charity advertisement writing task is affected by the AI's performance in a second language. Furthermore, we quantify the extent to which these patterns translate into the persuasiveness of generated charity advertisements, as well as the role of peoples' beliefs about LLM utilization in their donation choices. Our results provide evidence that writers who engage with an LLM-based writing assistant violate choice independence, as prior exposure to a Spanish LLM reduces subsequent utilization of an English LLM. While these patterns do not affect the aggregate persuasiveness of the generated advertisements, people's beliefs about the source of an advertisement (human versus AI) do. In particular, Spanish-speaking female participants who believed that they read an AI-generated advertisement strongly adjusted their donation behavior downwards. Furthermore, people are generally not able to adequately differentiate between human-generated and LLM-generated ads. Our work has important implications for the design, development, integration, and adoption of multilingual LLMs as assistive agents -- particularly in writing tasks.

多语言LLM人机协作认知偏差说服力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。