megatron_actor.py — Megatron 并行 Actor¶
文件概述¶
基于 Megatron-LM 的 MegatronPPOActor 实现(约 834 行)。支持张量并行(TP)、流水线并行(PP)和 Router Replay(路由重放),适合超大规模模型训练。
核心类¶
MegatronPPOActor¶
class MegatronPPOActor(BasePPOActor):
def __init__(self, config, megatron_module, megatron_optimizer, ...):
self.megatron_module = megatron_module # Megatron 模型(可能跨多个 PP stage)
self.megatron_optimizer = megatron_optimizer
与 FSDP Actor 的关键差异¶
流水线并行前向传播¶
def forward_backward_batch(self, data):
"""使用 Megatron 的流水线调度器执行前向+反向"""
forward_backward_func = get_forward_backward_func()
# Megatron 的 PP 调度器会自动处理:
# 1. 将微批次分配到不同 PP stage
# 2. 在 stage 之间传递激活值
# 3. 反向传播时传递梯度
losses = forward_backward_func(
forward_step_func=self.forward_step,
data_iterator=batch_iter,
model=self.megatron_module,
num_microbatches=n_micro_batch,
)
Router Replay(路由重放)¶
# 在 MoE 模型训练中,确保训练时使用与推理时相同的路由决策
if enable_routing_replay:
# 设置路由重放动作为 REPLAY_FORWARD
RouterReplay.set_global_router_replay_action(RouterReplayAction.REPLAY_FORWARD)
# 传入推理时记录的路由索引
set_router_replay_data(layers_topk_idx, ...)
与其他模块的关系¶
- 继承自
base.py的BasePPOActor - 被
megatron_workers.py创建和使用 - 使用 Megatron Core 的并行策略模块
小结¶
MegatronPPOActor 在 DataParallelPPOActor 的基础上增加了流水线并行和路由重放支持,适合需要极致并行的大规模训练场景。