跳转至

megatron_actor.py — Megatron 并行 Actor

文件概述

基于 Megatron-LM 的 MegatronPPOActor 实现(约 834 行)。支持张量并行(TP)、流水线并行(PP)和 Router Replay(路由重放),适合超大规模模型训练。

核心类

MegatronPPOActor

class MegatronPPOActor(BasePPOActor):
    def __init__(self, config, megatron_module, megatron_optimizer, ...):
        self.megatron_module = megatron_module  # Megatron 模型(可能跨多个 PP stage)
        self.megatron_optimizer = megatron_optimizer

与 FSDP Actor 的关键差异

流水线并行前向传播

def forward_backward_batch(self, data):
    """使用 Megatron 的流水线调度器执行前向+反向"""
    forward_backward_func = get_forward_backward_func()

    # Megatron 的 PP 调度器会自动处理:
    # 1. 将微批次分配到不同 PP stage
    # 2. 在 stage 之间传递激活值
    # 3. 反向传播时传递梯度
    losses = forward_backward_func(
        forward_step_func=self.forward_step,
        data_iterator=batch_iter,
        model=self.megatron_module,
        num_microbatches=n_micro_batch,
    )

Router Replay(路由重放)

# 在 MoE 模型训练中,确保训练时使用与推理时相同的路由决策
if enable_routing_replay:
    # 设置路由重放动作为 REPLAY_FORWARD
    RouterReplay.set_global_router_replay_action(RouterReplayAction.REPLAY_FORWARD)
    # 传入推理时记录的路由索引
    set_router_replay_data(layers_topk_idx, ...)

与其他模块的关系

  • 继承自 base.py 的 BasePPOActor
  • 被 megatron_workers.py 创建和使用
  • 使用 Megatron Core 的并行策略模块

小结

MegatronPPOActor 在 DataParallelPPOActor 的基础上增加了流水线并行和路由重放支持,适合需要极致并行的大规模训练场景。