weather_interaction.py — verl/interactions/weather_interaction.py¶
文件路径¶
verl/interactions/weather_interaction.py
文件概述¶
WeatherInteraction 是一个面向天气查询的演示交互环境,用于训练 LLM 学会使用工具来回答问题。它评判 LLM 是否正确使用了天气查询工具,而不是仅凭自己的知识回答。
训练目标:教会 LLM 在遇到天气问题时,应该调用天气工具而不是直接编造答案。
关键代码讲解¶
1. 生成响应(核心逻辑)¶
async def generate_response(self, instance_id, messages, **kwargs):
content = "no tool call"
for i in range(len(messages) - 1, -1, -1):
item = messages[i]
if item.get("role") == "tool":
content = item.get("content")
break
self._instance_dict[instance_id]["response"] = content
reward = await self.calculate_score(instance_id)
if reward == 1.0:
response = "Thank you for your weather query!"
should_terminate_sequence = True
else:
response = "Please use the weather tool to get the weather information."
should_terminate_sequence = True
return should_terminate_sequence, response, reward, {}
关键区别(与 Gsm8kInteraction 对比):
1. 查找的是工具调用结果:role == "tool"(而非 role == "assistant")。
2. 总是终止对话:无论成功失败,should_terminate_sequence 都是 True(天气查询是一次性操作)。
3. 如果没找到工具调用结果,content 保持为 "no tool call"。
2. 计算分数¶
async def calculate_score(self, instance_id, **kwargs):
if self._instance_dict[instance_id]["response"] == "no tool call":
return 0.0
return 1.0
评分逻辑极其简单:
- 如果 LLM 使用了工具(有 role="tool" 的消息),得 1.0 分。
- 如果 LLM 没有使用工具,得 0.0 分。
这意味着只要 LLM 学会了"遇到天气问题就调用工具",就能得满分。
与 Gsm8kInteraction 的对比¶
| 特性 | Gsm8kInteraction | WeatherInteraction |
|---|---|---|
| 检查目标 | LLM 的文字回答 | LLM 是否调用了工具 |
| 查找消息 | role="assistant" |
role="tool" |
| 多轮交互 | 答错继续 | 一轮即终止 |
| 评分逻辑 | 与标准答案比对 | 是否使用了工具 |
核心类/函数列表¶
| 类/方法 | 作用 |
|---|---|
WeatherInteraction |
天气查询交互环境 |
start_interaction |
初始化交互会话 |
generate_response |
检查 LLM 是否使用了工具 |
calculate_score |
简单二值评分 |
finalize_interaction |
清理实例 |
与其他模块的关系¶
- 继承自
BaseInteraction(base.py)。 - 不依赖任何评分工具:评分逻辑硬编码在类中。
- 与工具系统配合:需要与天气查询工具一起使用,交互环境负责评判,工具负责执行。
小结¶
WeatherInteraction 是一个简单但有意义的演示——它展示了如何训练 LLM 学会在合适的时候使用工具。这种"工具使用意识"的训练是 Tool-augmented RL 的基础。