image_zoom_in_tool.py — verl/tools/image_zoom_in_tool.py¶
文件路径¶
verl/tools/image_zoom_in_tool.py
文件概述¶
ImageZoomInTool 是一个视觉工具,允许 LLM 对图片的指定区域进行放大裁剪。这在多模态 RL 训练中非常有用——LLM 可以通过调用这个工具来"仔细看"图片中的某个区域,从而更好地理解图片内容。
应用场景:LLM 在解答视觉问题时,可能需要放大图片中的文字、物体等细节。
关键代码讲解¶
1. 并发框架¶
与 sandbox_fusion_tools.py 和 search_tool.py 一样,这个文件包含了自己的 TokenBucketWorker + VisualExecutionWorker + init_visual_execution_pool 并发框架。结构完全一致,这里不再赘述。
2. ImageZoomInTool 配置¶
class ImageZoomInTool(BaseTool):
MIN_DIMENSION = 28 # 裁剪区域的最小尺寸(像素)
def __init__(self, config, tool_schema):
super().__init__(config, tool_schema)
self._instance_dict = {}
self.num_workers = config.get("num_workers", 20)
self.rate_limit = config.get("rate_limit", 50)
self.timeout = config.get("timeout", 30)
MIN_DIMENSION = 28 确保裁剪区域不会太小(小于 28x28 像素的裁剪没有实际意义)。
3. 创建实例(加载图片)¶
async def create(self, instance_id=None, **kwargs):
if instance_id is None:
instance_id = str(uuid4())
create_kwargs = kwargs.get("create_kwargs", {})
if create_kwargs:
kwargs.update(create_kwargs)
image = kwargs.get("image")
if image is None:
raise ValueError("Missing required 'image' parameter in kwargs")
img = fetch_image({"image": image})
self._instance_dict[instance_id] = {
"image": img,
"response": "",
"reward": 0.0,
}
return instance_id, ToolResponse()
创建实例时需要传入图片。支持多种图片来源:PIL Image 对象、URL、本地路径、base64 编码等。使用 qwen_vl_utils.fetch_image 统一加载。
4. 边界框验证¶
def _validate_bbox(self, left, top, right, bottom):
if not (left < right and top < bottom):
return False
height = bottom - top
width = right - left
if min(height, width) == 0:
return False
if max(height, width) / min(height, width) > 100:
return False
return True
验证边界框的合法性: - 左上角必须在右下角的左上方。 - 宽度和高度不能为 0。 - 宽高比不能超过 100(防止极端窄条形裁剪)。
5. 边界框调整(核心算法)¶
def _maybe_resize_bbox(self, bbox_2d, image_width, image_height):
left, top, right, bottom = bbox_2d
# 1. 将边界框限制在图片范围内
left = max(0.0, float(left))
top = max(0.0, float(top))
right = min(float(image_width), float(right))
bottom = min(float(image_height), float(bottom))
# 2. 如果裁剪区域太小,以中心点为基准扩大
if height < self.MIN_DIMENSION or width < self.MIN_DIMENSION:
center_x = (left + right) / 2.0
center_y = (top + bottom) / 2.0
ratio = self.MIN_DIMENSION / min(height, width)
target_width = width * ratio
target_height = height * ratio
# 如果扩大后超出图片边界,按比例缩小
# ...
return current_bbox
这个方法处理各种边界情况:
1. 将边界框裁剪到图片范围内。
2. 如果裁剪区域太小,自动扩大到 MIN_DIMENSION。
3. 保持宽高比不变。
4. 确保扩大后不超出图片边界。
6. 执行裁剪¶
async def execute(self, instance_id, parameters, **kwargs):
bbox_2d = parameters.get("bbox_2d")
label = parameters.get("label", "")
if not bbox_2d or len(bbox_2d) != 4:
return ToolResponse(text="Error: bbox_2d parameter is missing..."), -0.05, {"success": False}
image = self._instance_dict[instance_id]["image"]
resized_bbox = self._maybe_resize_bbox(bbox_2d, image.size[0], image.size[1])
if resized_bbox is None:
return ToolResponse(text="Error: bounding box is invalid..."), -0.05, {"success": False}
cropped_image = image.crop(resized_bbox)
return ToolResponse(image=[cropped_image], text=f"Zoomed in on the region {bbox_2d}."), 0.0, {"success": True}
执行流程:
1. 获取 bbox_2d(边界框坐标 [x1, y1, x2, y2])和可选的 label。
2. 验证和调整边界框。
3. 使用 PIL 的 crop 方法裁剪图片。
4. 返回 ToolResponse,其中包含裁剪后的图片和文字描述。
5. 失败时返回 -0.05 的惩罚分。
核心类/函数列表¶
| 类/函数 | 作用 |
|---|---|
VisualExecutionWorker |
视觉处理任务执行器 |
init_visual_execution_pool |
初始化视觉处理执行池 |
ImageZoomInTool |
图片放大裁剪工具 |
_validate_bbox |
验证边界框合法性 |
_maybe_resize_bbox |
调整边界框(裁剪、扩大) |
create |
加载图片创建实例 |
execute |
执行裁剪操作 |
与其他模块的关系¶
- 继承自
BaseTool。 - 依赖
qwen_vl_utils.fetch_image加载图片。 - 并发框架与
sandbox_fusion_tools.py的设计模式相同。 - 是多模态 RL 的关键组件,支持视觉问答等任务。
小结¶
ImageZoomInTool 展示了 verl 工具系统如何支持多模态交互。LLM 可以指定图片中的一个区域(通过边界框坐标),工具会返回裁剪放大后的图片。边界框的智能调整算法确保了裁剪结果的可用性。