跳转至

transformer_impl.py — VeOmni 训练引擎实现

文件概述

基于 VeOmni 框架的训练引擎实现(约 588 行)。VeOmni 是字节跳动开发的分布式训练框架,支持 FSDP2 + EP + Ulysses SP。

核心类

VeOmniEngine(FSDPEngine)

继承自 FSDPEngine,使用 VeOmni 的并行状态管理和模型构建。

class VeOmniEngine(FSDPEngine):
    def __init__(self, ...):
        # 使用 VeOmni 初始化并行状态
        parallel_state.init_parallel_state(
            dp_size=dp_size,
            dp_replicate_size=...,
            dp_shard_size=...,
            ep_size=self.engine_config.expert_parallel_size,
            ulysses_size=self.engine_config.ulysses_parallel_size,
        )

    def _build_model_optimizer(self):
        """使用 VeOmni 的模型构建和并行化"""
        module = build_foundation_model(
            config_path=..., weights_path=...,
            attn_implementation=..., moe_implementation=...,
        )
        module = build_parallelize_model(module, ...)

VeOmniEngineWithLMHead

@EngineRegistry.register(model_type="language_model", backend=["veomni"], device=["cuda", "npu"])
class VeOmniEngineWithLMHead(VeOmniEngine, FSDPEngineWithLMHead):
    """同时支持 CUDA 和 NPU 的语言模型引擎"""

OmniSequenceShardCollator

class OmniSequenceShardCollator:
    """VeOmni 序列并行的数据切分整理器

    功能:
    - 沿序列维度切分输入(sp_slice)
    - 填充到序列并行对齐大小(sp_padding)
    """

EP 权重收集

def get_per_tensor_param(self):
    """支持 Expert Parallel 的权重导出

    专家层权重通过 all_gather_into_tensor 跨 EP rank 收集
    """
    if is_expert_layer and is_proj and ps.ep_enabled:
        torch.distributed.all_gather_into_tensor(stacked_tensor, unsharded_tensor, group=ps.ep_group)
        yield from process_func(name, stacked_tensor)

与其他模块的关系

  • 继承 engine/fsdp/transformer_impl.py 的 FSDPEngine
  • 使用 VeOmni 框架的 parallel_state、build_foundation_model 等
  • 使用 engine/veomni/utils.py 的 offload 和 MoE 参数处理

小结

VeOmni 引擎将字节跳动的 VeOmni 训练框架集成到 verl 中,特别优化了 MoE 模型和多模态模型的训练。