跳转至

math_dapo.py — 这个文件实现了 DAPO 风格数学题的奖励评分逻辑

模块路径: verl.utils.reward_score.math_dapo

文件概述

这个文件实现了 DAPO 风格数学题的奖励评分逻辑,适用于 math_dapo、math、aime 等数据集。评分结合了正确性检验和格式奖励两个维度:模型不仅需要算出正确答案,还需要按照 <think>...</think>...\boxed{...} 的格式输出。

评分依赖外部库 mathruler 来提取和评判答案。

关键代码讲解

1. 格式奖励 format_reward

def format_reward(predict_str: str) -> float:
    pattern = re.compile(r"<think>.*</think>.*\\boxed\{.*\}.*", re.DOTALL)
    match_result = re.fullmatch(pattern, predict_str)
    return 1.0 if match_result else 0.0

检查模型输出是否遵循了 <think>思考过程</think>...\boxed{答案} 的格式。re.DOTALL 使 . 匹配换行符。

2. 正确性奖励 acc_reward

def acc_reward(predict_str: str, ground_truth: str, use_boxed: bool = True) -> float:
    if use_boxed:
        answer = extract_boxed_content(predict_str)  # 从 \boxed{} 中提取答案
    else:
        answer = predict_str
    return 1.0 if grade_answer(answer, ground_truth) else 0.0

使用 mathruler 库的 grade_answer 来判断答案是否正确,支持更复杂的数学等价判断。

3. 综合评分 compute_score

def compute_score(predict_str, ground_truth, use_boxed=True, format_score=0.1):
    return (1.0 - format_score) * acc_reward(predict_str, ground_truth, use_boxed) \
         + format_score * format_reward(predict_str)

最终得分 = 90% 正确性分数 + 10% 格式分数(默认比例)。这种设计鼓励模型在学习正确答案的同时也学习输出格式。

核心类/函数列表

函数名 作用
format_reward 检查输出格式是否符合 <think>...\boxed{}
acc_reward 判断答案正确性
compute_score 综合格式和正确性的加权评分

与其他模块的关系

  • 被 __init__.py 调用,处理 math_dapo、math、aime 系列数据集
  • 依赖外部库 mathruler 的 extract_boxed_content 和 grade_answer
  • 与 geo3k.py 代码结构完全一致(geo3k 是几何题的评分,但评分逻辑相同)

小结

DAPO 评分的特色是加权混合奖励:正确性占主导,但格式也有少量权重。这种设计在 RL 训练早期有助于模型先学会按格式输出思考过程和答案,然后再提高正确率。format_score 参数可调节两者的比重。