为大模型求解能力设计信任度量指标,区分逻辑与自然语言问题。
From Logic to Language: A Trust Index for Problem Solving with LLMs
- 提出向量化的信任指数Q,量化大模型在模糊问题中的解答质量。
- 用语义熵衡量答案对问题表述变化的鲁棒性,情绪值反映主观满意度。
- 适用于评估大模型在开放、主观任务中的表现,如客服、创意生成。
传统计算依赖形式化逻辑系统,在规则明确的问题上表现优异,但难以处理含糊、动态和主观的问题。大语言模型(LLMs)的出现使系统能够通过自然语言解决这类问题。本文提出一个统一框架,区分形式语言与自然语言问题的求解空间。形式问题可用二元正确性评估,而自然语言问题需考虑语义模糊性与主观性,因此引入向量化的信任指数Q,反映解答的连续适切性。文中提出两个统计质量维度:归一化双语义熵衡量答案在问题表述变化下的鲁棒性与概念多样性;情感值将主观评价转化为可优化的量化指标。这些概念有助于更严谨地理解大模型在复杂任务中的能力边界与本质特征。
原文摘要 · Abstract (English)
Classical computation, grounded in formal, logical systems, has been the engine of technological progress for decades, excelling at problems that can be described with unambiguous rules. This paradigm, however, leaves a vast ocean of human problems -- those characterized by ambiguity, dynamic environments, and subjective context -- largely untouched. The advent of Large Language Models (LLMs) represents a fundamental shift, enabling computational systems to engage with this previously inaccessible domain using natural language. This paper introduces a unified framework to understand and contrast these problem-solving paradigms. We define and delineate the problem spaces addressable by formal languages versus natural language. While solutions to the former problem class can be evaluated using binary quality measures, the latter requires a much more nuanced definition of approximate solution space taking into account the vagueness, subjectivity and ambiguity inherent to natural language. We therefore introduce a vector-valued trust index Q, which reflects solution quality and distinguishes the binary correctness of formal solutions from the continuous adequacy spectrum characteristic of natural language solutions. Within this framework, we propose two statistical quality dimensions. Normalized bi-semantic entropy measures robustness and conceptual diversity of LLM answers given semantic variation in problem formulations. Emotional valence maps subjective valuation of a solution to a quantifiable metric that can be maximized by invoking statistical measures. The concepts introduced in this work will provide a more rigorous understanding of the capabilities, limitations, and inherent nature of problem-solving in the age of LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。