weight_converter.py — 权重名称转换器¶
文件路径¶
verl/models/mcore/weight_converter.py
文件概述¶
提供 Megatron-Core 权重名称到 HuggingFace 权重名称的映射转换。在 mcore 训练中,权重使用 mcore 的命名格式存储;保存 checkpoint 时需要转换回 HF 格式。
关键代码讲解¶
基类¶
class McoreToHFWeightConverterBase:
def __init__(self, hf_config, mcore_config):
self.hf_config = hf_config
self.mcore_config = mcore_config
def convert_param(self, name, params_one_group):
"""将 mcore 参数名和分片参数列表转换为 HF 格式"""
raise NotImplementedError
Dense 模型的名称映射¶
class McoreToHFWeightConverterDense(McoreToHFWeightConverterBase):
def _convert_attention_param(self, name, params):
layer_number = name.split(".")[2]
if "self_attention.linear_qkv.weight" in name:
# mcore: decoder.layers.0.self_attention.linear_qkv.weight
# HF: model.layers.0.self_attn.q_proj.weight
# model.layers.0.self_attn.k_proj.weight
# model.layers.0.self_attn.v_proj.weight
convert_names = [
f"model.layers.{layer_number}.self_attn.q_proj.weight",
f"model.layers.{layer_number}.self_attn.k_proj.weight",
f"model.layers.{layer_number}.self_attn.v_proj.weight",
]
elif "self_attention.linear_proj.weight" in name:
# mcore 的 linear_proj 对应 HF 的 o_proj
convert_names = [f"model.layers.{layer_number}.self_attn.o_proj.weight"]
def _convert_mlp_param(self, name, params):
if "mlp.linear_fc1.weight" in name:
# mcore: linear_fc1 = gate_proj + up_proj(合并)
# HF: gate_proj 和 up_proj(分开)
convert_names = [
f"model.layers.{layer_number}.mlp.gate_proj.weight",
f"model.layers.{layer_number}.mlp.up_proj.weight",
]
elif "mlp.linear_fc2.weight" in name:
# mcore: linear_fc2 = down_proj
convert_names = [f"model.layers.{layer_number}.mlp.down_proj.weight"]
关键映射:
- mcore 将 Q/K/V 合并为 linear_qkv,HF 是分开的 q_proj/k_proj/v_proj
- mcore 将 gate_proj 和 up_proj 合并为 linear_fc1,HF 是分开的
不同模型的专用转换器¶
| 转换器类 | 适用模型 | 特殊处理 |
|---|---|---|
McoreToHFWeightConverterDense |
LLaMA, Qwen2, Qwen3 | 标准 QKV/MLP 拆分 |
McoreToHFWeightConverterQwen2Moe |
Qwen2 MoE | 专家权重 + 共享专家 |
McoreToHFWeightConverterMixtral |
Mixtral | 专家权重(无共享专家) |
McoreToHFWeightConverterDpskv3 |
DeepSeek-V3 | MLA 权重 + MoE |
McoreToHFWeightConverterQwen2_5_VL |
Qwen2.5-VL | 视觉编码器权重 |
McoreToHFWeightConverterQwen3Moe |
Qwen3 MoE | 与 Qwen2 MoE 类似 |
与其他模块的关系¶
- 被
registry.py的get_mcore_weight_converter()调用 - 被 checkpoint 保存流程调用
小结¶
权重转换器处理 mcore 和 HF 之间的命名差异。最关键的转换是 QKV 合并/拆分和 gate_up 合并/拆分。MoE 模型还需要处理专家权重的命名映射。