跳转至

weight_converter.py — 权重名称转换器

文件路径

verl/models/mcore/weight_converter.py

文件概述

提供 Megatron-Core 权重名称到 HuggingFace 权重名称的映射转换。在 mcore 训练中,权重使用 mcore 的命名格式存储;保存 checkpoint 时需要转换回 HF 格式。

关键代码讲解

基类

class McoreToHFWeightConverterBase:
    def __init__(self, hf_config, mcore_config):
        self.hf_config = hf_config
        self.mcore_config = mcore_config

    def convert_param(self, name, params_one_group):
        """将 mcore 参数名和分片参数列表转换为 HF 格式"""
        raise NotImplementedError

Dense 模型的名称映射

class McoreToHFWeightConverterDense(McoreToHFWeightConverterBase):
    def _convert_attention_param(self, name, params):
        layer_number = name.split(".")[2]
        if "self_attention.linear_qkv.weight" in name:
            # mcore: decoder.layers.0.self_attention.linear_qkv.weight
            # HF:    model.layers.0.self_attn.q_proj.weight
            #        model.layers.0.self_attn.k_proj.weight
            #        model.layers.0.self_attn.v_proj.weight
            convert_names = [
                f"model.layers.{layer_number}.self_attn.q_proj.weight",
                f"model.layers.{layer_number}.self_attn.k_proj.weight",
                f"model.layers.{layer_number}.self_attn.v_proj.weight",
            ]
        elif "self_attention.linear_proj.weight" in name:
            # mcore 的 linear_proj 对应 HF 的 o_proj
            convert_names = [f"model.layers.{layer_number}.self_attn.o_proj.weight"]

    def _convert_mlp_param(self, name, params):
        if "mlp.linear_fc1.weight" in name:
            # mcore: linear_fc1 = gate_proj + up_proj(合并)
            # HF:    gate_proj 和 up_proj(分开)
            convert_names = [
                f"model.layers.{layer_number}.mlp.gate_proj.weight",
                f"model.layers.{layer_number}.mlp.up_proj.weight",
            ]
        elif "mlp.linear_fc2.weight" in name:
            # mcore: linear_fc2 = down_proj
            convert_names = [f"model.layers.{layer_number}.mlp.down_proj.weight"]

关键映射: - mcore 将 Q/K/V 合并为 linear_qkv,HF 是分开的 q_proj/k_proj/v_proj - mcore 将 gate_proj 和 up_proj 合并为 linear_fc1,HF 是分开的

不同模型的专用转换器

转换器类 适用模型 特殊处理
McoreToHFWeightConverterDense LLaMA, Qwen2, Qwen3 标准 QKV/MLP 拆分
McoreToHFWeightConverterQwen2Moe Qwen2 MoE 专家权重 + 共享专家
McoreToHFWeightConverterMixtral Mixtral 专家权重(无共享专家)
McoreToHFWeightConverterDpskv3 DeepSeek-V3 MLA 权重 + MoE
McoreToHFWeightConverterQwen2_5_VL Qwen2.5-VL 视觉编码器权重
McoreToHFWeightConverterQwen3Moe Qwen3 MoE 与 Qwen2 MoE 类似

与其他模块的关系

  • 被 registry.py 的 get_mcore_weight_converter() 调用
  • 被 checkpoint 保存流程调用

小结

权重转换器处理 mcore 和 HF 之间的命名差异。最关键的转换是 QKV 合并/拆分和 gate_up 合并/拆分。MoE 模型还需要处理专家权重的命名映射。