arXiv:2412.19610cs.CL2024-12被引 2

对比AI与人类写商品广告,发现GPT-4表现最佳。

Machine Generated Product Advertisements: Benchmarking LLMs Against Human Performance

  • 用四种AI模型生成100个商品描述,对比人类写作
  • GPT-4在情感、说服力、清晰度等指标上领先
  • 其他模型常出现逻辑混乱、脱离产品主题的问题

本研究采用多维度评估模型,对比四款AI模型(Gemma 2B、LLAMA、GPT2、ChatGPT 4)在有无样例条件下生成的100个商品描述与人工撰写版本。评估涵盖情感倾向、可读性、说服力、搜索引擎优化(SEO)、清晰度、情感吸引力及行动号召有效性。结果显示,ChatGPT 4表现最优;其余模型存在严重缺陷,输出常逻辑不清、结构松散,缺乏上下文相关性,难以聚焦产品核心信息,导致语句断裂且信息无效。研究揭示了当前AI在电商内容生成中的能力边界与不足。

原文摘要 · Abstract (English)

This study compares the performance of AI-generated and human-written product descriptions using a multifaceted evaluation model. We analyze descriptions for 100 products generated by four AI models (Gemma 2B, LLAMA, GPT2, and ChatGPT 4) with and without sample descriptions, against human-written descriptions. Our evaluation metrics include sentiment, readability, persuasiveness, Search Engine Optimization(SEO), clarity, emotional appeal, and call-to-action effectiveness. The results indicate that ChatGPT 4 performs the best. In contrast, other models demonstrate significant shortcomings, producing incoherent and illogical output that lacks logical structure and contextual relevance. These models struggle to maintain focus on the product being described, resulting in disjointed sentences that do not convey meaningful information. This research provides insights into the current capabilities and limitations of AI in the creation of content for e-Commerce.

AI生成电商文案大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。