naive_rollout_rob.py — NaiveRolloutRob 实现了 OpenVLA 模型的推理逻辑¶
文件路径:
verl/experimental/vla/naive_rollout_rob.py模块路径:verl.experimental.vla.naive_rollout_rob
文件概述¶
NaiveRolloutRob 实现了 OpenVLA 模型的推理逻辑,包括图像预处理(process_input)和动作生成(generate_sequences)。这是模型从观测到动作的完整推理流水线。
一、图像预处理 process_input¶
def process_input(task_descriptions, images_and_states, processor):
"""将原始观测转换为模型输入
Args:
task_descriptions: 任务描述文本列表
images_and_states: 包含 full_image 的字典
processor: PrismaticProcessor
"""
batchdata = {"input_ids": [], "attention_mask": [], "pixel_values": []}
for i in range(len(task_descriptions)):
# 1. 图像预处理: resize -> RGB -> center crop
image = resize_image(images_and_states["full_image"][i].cpu().numpy(), (224, 224))
image = Image.fromarray(image).convert("RGB")
image = center_crop_image(image)
# 2. 构造提示文本
prompt = f"In: What action should the robot take to {task_description.lower()}?\nOut:"
# 3. 使用 processor 编码
batch_feature = processor(prompt, image)
# 4. 确保末尾有特殊 token 29871
if not torch.all(input_ids[:, -1] == 29871):
input_ids = torch.cat(
(input_ids, torch.Tensor([29871]).long().unsqueeze(0)), dim=1
)
...
左填充处理¶
VLA 模型使用左填充(left padding),确保有效 token 在右侧对齐:
# 对不同长度的 input_ids 做 padding
batchdata["input_ids"] = pad_sequence(
batchdata["input_ids"],
batch_first=True,
padding_value=processor.tokenizer.pad_token_id
)
# 将 padding 移到左侧(排序方式)
padding_mask = ~batchdata["input_ids"].ne(processor.tokenizer.pad_token_id)
sorted_indices = torch.argsort(padding_mask.int(), dim=1, descending=True, stable=True)
batchdata["input_ids"] = torch.gather(batchdata["input_ids"], 1, sorted_indices)
二、NaiveRolloutRob 类¶
初始化¶
class NaiveRolloutRob(BaseRollout):
def __init__(self, model_config, module=None):
if module is not None:
self.module = module
else:
self.module = OpenVLAForActionPrediction.from_pretrained(
model_config["path"], trust_remote_code=True
)
# 设置单图输入
self.module.vision_backbone.set_num_images_in_input(1)
# 加载 processor
self.processor = PrismaticProcessor.from_pretrained(model_config["path"])
# 加载归一化统计量
self.module.norm_stats = json.load(open(dataset_statistics_path))
推理方法¶
@torch.no_grad()
def generate_sequences(self, prompts: DataProto) -> DataProto:
"""完整推理流程"""
# 1. 提取参数
do_sample = prompts.meta_info["do_sample"]
temperature = prompts.meta_info["temperature"]
# 2. 预处理输入
vla_input = process_input(
task_descriptions, images_and_states, self.processor
)
# 3. 模型推理(一步生成)
vla_output = self._generate_one_step(
vla_input, do_sample, temperature, max_prompt_length
)
# 4. 封装为 DataProto
return DataProto.from_dict(tensors=vla_output)
权重更新(异步)¶
async def update_weights(self, weights_iterator, **kwargs):
"""从外部迭代器更新模型权重
用于将训练后的新权重加载到推理模型中。
"""
for name, param in weights_iterator:
cleaned_name = name.replace("_fsdp_wrapped_module.", "")
if cleaned_name in target_state_dict:
target_state_dict[cleaned_name].copy_(param, non_blocking=True)
核心类/函数列表¶
| 名称 | 类型 | 说明 |
|---|---|---|
pad_sequence_to_length |
函数 | 将 tensor 填充到指定长度 |
process_input |
函数 | 图像+文本 -> 模型输入 |
NaiveRolloutRob |
类 | OpenVLA 推理封装 |
generate_sequences |
方法 | 从观测生成动作 |
_generate_one_step |
方法 | 单步模型推理 |
update_weights |
方法 | 异步更新模型权重 |
release / resume |
方法 | GPU/CPU 迁移 |
与其他模块的关系¶
- 使用
OpenVLAForActionPrediction(models/openvla_oft/modeling_prismatic.py) - 使用
PrismaticProcessor(models/openvla_oft/processing_prismatic.py) - 使用
action_utils.py的图像预处理函数 - 被
RobActorRolloutRefWorker(fsdp_workers.py)调用
小结¶
NaiveRolloutRob 是 OpenVLA 模型推理的"最后一公里"。它负责将原始的环境观测(图像+任务描述)转换为模型输入,调用模型生成动作,然后将结果封装返回。"Naive"意味着它使用直接的单步推理,没有复杂的采样策略或 beam search。