跳转至

weather_interaction.py — verl/interactions/weather_interaction.py

文件路径

verl/interactions/weather_interaction.py

文件概述

WeatherInteraction 是一个面向天气查询的演示交互环境,用于训练 LLM 学会使用工具来回答问题。它评判 LLM 是否正确使用了天气查询工具,而不是仅凭自己的知识回答。

训练目标:教会 LLM 在遇到天气问题时,应该调用天气工具而不是直接编造答案。

关键代码讲解

1. 生成响应(核心逻辑)

async def generate_response(self, instance_id, messages, **kwargs):
    content = "no tool call"
    for i in range(len(messages) - 1, -1, -1):
        item = messages[i]
        if item.get("role") == "tool":
            content = item.get("content")
            break
    self._instance_dict[instance_id]["response"] = content

    reward = await self.calculate_score(instance_id)
    if reward == 1.0:
        response = "Thank you for your weather query!"
        should_terminate_sequence = True
    else:
        response = "Please use the weather tool to get the weather information."
        should_terminate_sequence = True
    return should_terminate_sequence, response, reward, {}

关键区别(与 Gsm8kInteraction 对比): 1. 查找的是工具调用结果:role == "tool"(而非 role == "assistant")。 2. 总是终止对话:无论成功失败,should_terminate_sequence 都是 True(天气查询是一次性操作)。 3. 如果没找到工具调用结果,content 保持为 "no tool call"。

2. 计算分数

async def calculate_score(self, instance_id, **kwargs):
    if self._instance_dict[instance_id]["response"] == "no tool call":
        return 0.0
    return 1.0

评分逻辑极其简单: - 如果 LLM 使用了工具(有 role="tool" 的消息),得 1.0 分。 - 如果 LLM 没有使用工具,得 0.0 分。

这意味着只要 LLM 学会了"遇到天气问题就调用工具",就能得满分。

与 Gsm8kInteraction 的对比

特性 Gsm8kInteraction WeatherInteraction
检查目标 LLM 的文字回答 LLM 是否调用了工具
查找消息 role="assistant" role="tool"
多轮交互 答错继续 一轮即终止
评分逻辑 与标准答案比对 是否使用了工具

核心类/函数列表

类/方法 作用
WeatherInteraction 天气查询交互环境
start_interaction 初始化交互会话
generate_response 检查 LLM 是否使用了工具
calculate_score 简单二值评分
finalize_interaction 清理实例

与其他模块的关系

  • 继承自 BaseInteraction(base.py)。
  • 不依赖任何评分工具:评分逻辑硬编码在类中。
  • 与工具系统配合:需要与天气查询工具一起使用,交互环境负责评判,工具负责执行。

小结

WeatherInteraction 是一个简单但有意义的演示——它展示了如何训练 LLM 学会在合适的时候使用工具。这种"工具使用意识"的训练是 Tool-augmented RL 的基础。