transformer_impl.py — VeOmni 训练引擎实现¶
文件概述¶
基于 VeOmni 框架的训练引擎实现(约 588 行)。VeOmni 是字节跳动开发的分布式训练框架,支持 FSDP2 + EP + Ulysses SP。
核心类¶
VeOmniEngine(FSDPEngine)¶
继承自 FSDPEngine,使用 VeOmni 的并行状态管理和模型构建。
class VeOmniEngine(FSDPEngine):
def __init__(self, ...):
# 使用 VeOmni 初始化并行状态
parallel_state.init_parallel_state(
dp_size=dp_size,
dp_replicate_size=...,
dp_shard_size=...,
ep_size=self.engine_config.expert_parallel_size,
ulysses_size=self.engine_config.ulysses_parallel_size,
)
def _build_model_optimizer(self):
"""使用 VeOmni 的模型构建和并行化"""
module = build_foundation_model(
config_path=..., weights_path=...,
attn_implementation=..., moe_implementation=...,
)
module = build_parallelize_model(module, ...)
VeOmniEngineWithLMHead¶
@EngineRegistry.register(model_type="language_model", backend=["veomni"], device=["cuda", "npu"])
class VeOmniEngineWithLMHead(VeOmniEngine, FSDPEngineWithLMHead):
"""同时支持 CUDA 和 NPU 的语言模型引擎"""
OmniSequenceShardCollator¶
class OmniSequenceShardCollator:
"""VeOmni 序列并行的数据切分整理器
功能:
- 沿序列维度切分输入(sp_slice)
- 填充到序列并行对齐大小(sp_padding)
"""
EP 权重收集¶
def get_per_tensor_param(self):
"""支持 Expert Parallel 的权重导出
专家层权重通过 all_gather_into_tensor 跨 EP rank 收集
"""
if is_expert_layer and is_proj and ps.ep_enabled:
torch.distributed.all_gather_into_tensor(stacked_tensor, unsharded_tensor, group=ps.ep_group)
yield from process_func(name, stacked_tensor)
与其他模块的关系¶
- 继承
engine/fsdp/transformer_impl.py的FSDPEngine - 使用 VeOmni 框架的 parallel_state、build_foundation_model 等
- 使用
engine/veomni/utils.py的 offload 和 MoE 参数处理
小结¶
VeOmni 引擎将字节跳动的 VeOmni 训练框架集成到 verl 中,特别优化了 MoE 模型和多模态模型的训练。