跳转至

image_zoom_in_tool.py — verl/tools/image_zoom_in_tool.py

文件路径

verl/tools/image_zoom_in_tool.py

文件概述

ImageZoomInTool 是一个视觉工具,允许 LLM 对图片的指定区域进行放大裁剪。这在多模态 RL 训练中非常有用——LLM 可以通过调用这个工具来"仔细看"图片中的某个区域,从而更好地理解图片内容。

应用场景:LLM 在解答视觉问题时,可能需要放大图片中的文字、物体等细节。

关键代码讲解

1. 并发框架

与 sandbox_fusion_tools.py 和 search_tool.py 一样,这个文件包含了自己的 TokenBucketWorker + VisualExecutionWorker + init_visual_execution_pool 并发框架。结构完全一致,这里不再赘述。

2. ImageZoomInTool 配置

class ImageZoomInTool(BaseTool):
    MIN_DIMENSION = 28  # 裁剪区域的最小尺寸(像素)

    def __init__(self, config, tool_schema):
        super().__init__(config, tool_schema)
        self._instance_dict = {}
        self.num_workers = config.get("num_workers", 20)
        self.rate_limit = config.get("rate_limit", 50)
        self.timeout = config.get("timeout", 30)

MIN_DIMENSION = 28 确保裁剪区域不会太小(小于 28x28 像素的裁剪没有实际意义)。

3. 创建实例(加载图片)

async def create(self, instance_id=None, **kwargs):
    if instance_id is None:
        instance_id = str(uuid4())
    create_kwargs = kwargs.get("create_kwargs", {})
    if create_kwargs:
        kwargs.update(create_kwargs)
    image = kwargs.get("image")
    if image is None:
        raise ValueError("Missing required 'image' parameter in kwargs")
    img = fetch_image({"image": image})
    self._instance_dict[instance_id] = {
        "image": img,
        "response": "",
        "reward": 0.0,
    }
    return instance_id, ToolResponse()

创建实例时需要传入图片。支持多种图片来源:PIL Image 对象、URL、本地路径、base64 编码等。使用 qwen_vl_utils.fetch_image 统一加载。

4. 边界框验证

def _validate_bbox(self, left, top, right, bottom):
    if not (left < right and top < bottom):
        return False
    height = bottom - top
    width = right - left
    if min(height, width) == 0:
        return False
    if max(height, width) / min(height, width) > 100:
        return False
    return True

验证边界框的合法性: - 左上角必须在右下角的左上方。 - 宽度和高度不能为 0。 - 宽高比不能超过 100(防止极端窄条形裁剪)。

5. 边界框调整(核心算法)

def _maybe_resize_bbox(self, bbox_2d, image_width, image_height):
    left, top, right, bottom = bbox_2d
    # 1. 将边界框限制在图片范围内
    left = max(0.0, float(left))
    top = max(0.0, float(top))
    right = min(float(image_width), float(right))
    bottom = min(float(image_height), float(bottom))

    # 2. 如果裁剪区域太小,以中心点为基准扩大
    if height < self.MIN_DIMENSION or width < self.MIN_DIMENSION:
        center_x = (left + right) / 2.0
        center_y = (top + bottom) / 2.0
        ratio = self.MIN_DIMENSION / min(height, width)
        target_width = width * ratio
        target_height = height * ratio
        # 如果扩大后超出图片边界,按比例缩小
        # ...
    return current_bbox

这个方法处理各种边界情况: 1. 将边界框裁剪到图片范围内。 2. 如果裁剪区域太小,自动扩大到 MIN_DIMENSION。 3. 保持宽高比不变。 4. 确保扩大后不超出图片边界。

6. 执行裁剪

async def execute(self, instance_id, parameters, **kwargs):
    bbox_2d = parameters.get("bbox_2d")
    label = parameters.get("label", "")

    if not bbox_2d or len(bbox_2d) != 4:
        return ToolResponse(text="Error: bbox_2d parameter is missing..."), -0.05, {"success": False}

    image = self._instance_dict[instance_id]["image"]
    resized_bbox = self._maybe_resize_bbox(bbox_2d, image.size[0], image.size[1])

    if resized_bbox is None:
        return ToolResponse(text="Error: bounding box is invalid..."), -0.05, {"success": False}

    cropped_image = image.crop(resized_bbox)
    return ToolResponse(image=[cropped_image], text=f"Zoomed in on the region {bbox_2d}."), 0.0, {"success": True}

执行流程: 1. 获取 bbox_2d(边界框坐标 [x1, y1, x2, y2])和可选的 label。 2. 验证和调整边界框。 3. 使用 PIL 的 crop 方法裁剪图片。 4. 返回 ToolResponse,其中包含裁剪后的图片和文字描述。 5. 失败时返回 -0.05 的惩罚分。

核心类/函数列表

类/函数 作用
VisualExecutionWorker 视觉处理任务执行器
init_visual_execution_pool 初始化视觉处理执行池
ImageZoomInTool 图片放大裁剪工具
_validate_bbox 验证边界框合法性
_maybe_resize_bbox 调整边界框(裁剪、扩大)
create 加载图片创建实例
execute 执行裁剪操作

与其他模块的关系

  • 继承自 BaseTool。
  • 依赖 qwen_vl_utils.fetch_image 加载图片。
  • 并发框架与 sandbox_fusion_tools.py 的设计模式相同。
  • 是多模态 RL 的关键组件,支持视觉问答等任务。

小结

ImageZoomInTool 展示了 verl 工具系统如何支持多模态交互。LLM 可以指定图片中的一个区域(通过边界框坐标),工具会返回裁剪放大后的图片。边界框的智能调整算法确保了裁剪结果的可用性。