跳转至

naive_rollout_rob.py — NaiveRolloutRob 实现了 OpenVLA 模型的推理逻辑

文件路径: verl/experimental/vla/naive_rollout_rob.py 模块路径: verl.experimental.vla.naive_rollout_rob

文件概述

NaiveRolloutRob 实现了 OpenVLA 模型的推理逻辑,包括图像预处理(process_input)和动作生成(generate_sequences)。这是模型从观测到动作的完整推理流水线。

一、图像预处理 process_input

def process_input(task_descriptions, images_and_states, processor):
    """将原始观测转换为模型输入

    Args:
        task_descriptions: 任务描述文本列表
        images_and_states: 包含 full_image 的字典
        processor: PrismaticProcessor
    """
    batchdata = {"input_ids": [], "attention_mask": [], "pixel_values": []}

    for i in range(len(task_descriptions)):
        # 1. 图像预处理: resize -> RGB -> center crop
        image = resize_image(images_and_states["full_image"][i].cpu().numpy(), (224, 224))
        image = Image.fromarray(image).convert("RGB")
        image = center_crop_image(image)

        # 2. 构造提示文本
        prompt = f"In: What action should the robot take to {task_description.lower()}?\nOut:"

        # 3. 使用 processor 编码
        batch_feature = processor(prompt, image)

        # 4. 确保末尾有特殊 token 29871
        if not torch.all(input_ids[:, -1] == 29871):
            input_ids = torch.cat(
                (input_ids, torch.Tensor([29871]).long().unsqueeze(0)), dim=1
            )
    ...

左填充处理

VLA 模型使用左填充(left padding),确保有效 token 在右侧对齐:

    # 对不同长度的 input_ids 做 padding
    batchdata["input_ids"] = pad_sequence(
        batchdata["input_ids"],
        batch_first=True,
        padding_value=processor.tokenizer.pad_token_id
    )

    # 将 padding 移到左侧(排序方式)
    padding_mask = ~batchdata["input_ids"].ne(processor.tokenizer.pad_token_id)
    sorted_indices = torch.argsort(padding_mask.int(), dim=1, descending=True, stable=True)
    batchdata["input_ids"] = torch.gather(batchdata["input_ids"], 1, sorted_indices)

二、NaiveRolloutRob 类

初始化

class NaiveRolloutRob(BaseRollout):
    def __init__(self, model_config, module=None):
        if module is not None:
            self.module = module
        else:
            self.module = OpenVLAForActionPrediction.from_pretrained(
                model_config["path"], trust_remote_code=True
            )
        # 设置单图输入
        self.module.vision_backbone.set_num_images_in_input(1)
        # 加载 processor
        self.processor = PrismaticProcessor.from_pretrained(model_config["path"])
        # 加载归一化统计量
        self.module.norm_stats = json.load(open(dataset_statistics_path))

推理方法

@torch.no_grad()
def generate_sequences(self, prompts: DataProto) -> DataProto:
    """完整推理流程"""
    # 1. 提取参数
    do_sample = prompts.meta_info["do_sample"]
    temperature = prompts.meta_info["temperature"]

    # 2. 预处理输入
    vla_input = process_input(
        task_descriptions, images_and_states, self.processor
    )

    # 3. 模型推理(一步生成)
    vla_output = self._generate_one_step(
        vla_input, do_sample, temperature, max_prompt_length
    )

    # 4. 封装为 DataProto
    return DataProto.from_dict(tensors=vla_output)

权重更新(异步)

async def update_weights(self, weights_iterator, **kwargs):
    """从外部迭代器更新模型权重

    用于将训练后的新权重加载到推理模型中。
    """
    for name, param in weights_iterator:
        cleaned_name = name.replace("_fsdp_wrapped_module.", "")
        if cleaned_name in target_state_dict:
            target_state_dict[cleaned_name].copy_(param, non_blocking=True)

核心类/函数列表

名称 类型 说明
pad_sequence_to_length 函数 将 tensor 填充到指定长度
process_input 函数 图像+文本 -> 模型输入
NaiveRolloutRob 类 OpenVLA 推理封装
generate_sequences 方法 从观测生成动作
_generate_one_step 方法 单步模型推理
update_weights 方法 异步更新模型权重
release / resume 方法 GPU/CPU 迁移

与其他模块的关系

  • 使用 OpenVLAForActionPrediction(models/openvla_oft/modeling_prismatic.py)
  • 使用 PrismaticProcessor(models/openvla_oft/processing_prismatic.py)
  • 使用 action_utils.py 的图像预处理函数
  • 被 RobActorRolloutRefWorker(fsdp_workers.py)调用

小结

NaiveRolloutRob 是 OpenVLA 模型推理的"最后一公里"。它负责将原始的环境观测(图像+任务描述)转换为模型输入,调用模型生成动作,然后将结果封装返回。"Naive"意味着它使用直接的单步推理,没有复杂的采样策略或 beam search。