用真实决策效果评估文本生成质量,发现分析性内容更利于人机协作。
Decision-Oriented Text Evaluation
- 通过人类和大模型在金融交易中的实际表现来评估文本影响。
- 仅靠摘要时人机表现均不如随机,但加入分析后人机协同显著提升。
- 适合关注人机协作与生成文本实用价值的研究者。
自然语言生成正越来越多地应用于高风险领域,但常见的内在评估方法(如n-gram重叠或句子合理性)与实际决策效果的相关性较弱。本文提出一种面向决策的评估框架,直接衡量生成文本对人类和大语言模型(LLM)决策结果的影响。以市场摘要文本——包括客观的早间总结和主观的收盘分析——为测试案例,评估基于这些文本做出交易决策的人类投资者和自主的LLM代理的财务表现。研究发现,仅依赖摘要时,人类和LLM的表现均无法持续超越随机水平。然而,更丰富的分析性评论使人类与LLM协作团队显著优于单独的人类或模型基线。该方法强调应以生成文本促进人机协同决策的能力作为核心评估标准,揭示了传统内在指标的关键局限。
原文摘要 · Abstract (English)
Natural language generation (NLG) is increasingly deployed in high-stakes domains, yet common intrinsic evaluation methods, such as n-gram overlap or sentence plausibility, weakly correlate with actual decision-making efficacy. We propose a decision-oriented framework for evaluating generated text by directly measuring its influence on human and large language model (LLM) decision outcomes. Using market digest texts--including objective morning summaries and subjective closing-bell analyses--as test cases, we assess decision quality based on the financial performance of trades executed by human investors and autonomous LLM agents informed exclusively by these texts. Our findings reveal that neither humans nor LLM agents consistently surpass random performance when relying solely on summaries. However, richer analytical commentaries enable collaborative human-LLM teams to outperform individual human or agent baselines significantly. Our approach underscores the importance of evaluating generated text by its ability to facilitate synergistic decision-making between humans and LLMs, highlighting critical limitations of traditional intrinsic metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。