patch.py — 这个文件为 vLLM 中的 MoE(Mixture of Experts)模型提供权重加载补...¶
模块路径: verl.utils.vllm.patch
文件概述¶
这个文件为 vLLM 中的 MoE(Mixture of Experts)模型提供权重加载补丁。问题背景:vLLM 的 MoE 模型权重(如 w13_weight、w2_weight)缺少 weight_loader 属性,导致 verl 的权重同步机制无法正常工作。
关键代码讲解¶
1. 支持的 MoE 模型¶
SUPPORTED_MOE_MODELS = []
try:
from vllm.model_executor.models.deepseek_v2 import DeepseekV2ForCausalLM, DeepseekV3ForCausalLM
SUPPORTED_MOE_MODELS.append(DeepseekV2ForCausalLM)
SUPPORTED_MOE_MODELS.append(DeepseekV3ForCausalLM)
except ImportError:
pass
# 类似地添加 Mixtral, Qwen2Moe, Qwen3Moe, KimiVL 等
使用 try-except 逐个尝试导入,兼容不同版本的 vLLM。
2. 补丁函数¶
def patch_vllm_moe_model_weight_loader(model):
"""为 MoE 模型的 expert 权重添加 weight_loader"""
# 获取内部模型
inner_model = getattr(model, "model", None) or getattr(model, "language_model", None)
# 检查是否为 MoE 模型
if not isinstance(model, tuple(SUPPORTED_MOE_MODELS)):
return
# 遍历所有层,为 expert 权重添加 weight_loader
for layer_idx, layer in enumerate(inner_model.layers):
mlp = getattr(layer, mlp_attr, None)
experts = getattr(mlp, "experts", None)
if not experts or not hasattr(experts, "weight_loader"):
continue
for name, param in mlp.named_parameters():
if "w13_weight" in name or "w2_weight" in name:
param.weight_loader = experts.weight_loader
核心修复:将 experts.weight_loader 赋给每个 expert 权重参数。这是 vLLM 0.8.2 中的一个 bug 的 workaround。
核心类/函数列表¶
| 函数名 | 作用 |
|---|---|
patch_vllm_moe_model_weight_loader |
为 MoE 权重添加 weight_loader |
与其他模块的关系¶
- 不在
__init__.py中导入(需要在 vLLM 实例创建后才能使用) - 用于 verl 的 weight sync 流程
小结¶
一个针对 vLLM MoE 模型 bug 的 workaround。在 vLLM 修复此问题前,这个补丁确保 verl 能正确地将更新后的 expert 权重同步到推理引擎。