跳转至

geo3k.py — 这个文件实现了 Geometry3K 几何题数据集的奖励评分

模块路径: verl.utils.reward_score.geo3k

文件概述

这个文件实现了 Geometry3K 几何题数据集的奖励评分。代码结构与 math_dapo.py 完全一致,同样使用格式奖励 + 正确性奖励的加权方式。

关键代码讲解

def format_reward(predict_str: str) -> float:
    pattern = re.compile(r"<think>.*</think>.*\\boxed\{.*\}.*", re.DOTALL)
    return 1.0 if re.fullmatch(pattern, predict_str) else 0.0

def acc_reward(predict_str: str, ground_truth: str, use_boxed: bool = True) -> float:
    if use_boxed:
        answer = extract_boxed_content(predict_str)
    else:
        answer = predict_str
    return 1.0 if grade_answer(answer, ground_truth) else 0.0

def compute_score(predict_str, ground_truth, use_boxed=True, format_score=0.1):
    return (1.0 - format_score) * acc_reward(predict_str, ground_truth, use_boxed) \
         + format_score * format_reward(predict_str)

逻辑与 math_dapo.py 完全相同,区别仅在于它被 __init__.py 的 hiyouga/geometry3k 数据集路由调用。

核心类/函数列表

函数名 作用
format_reward 检查 <think>...\boxed{} 格式
acc_reward 判断答案正确性
compute_score 加权综合评分

与其他模块的关系

  • 被 __init__.py 调用,处理 hiyouga/geometry3k 数据集
  • 与 math_dapo.py 代码结构完全一致
  • 依赖 mathruler 库

小结

专为几何题数据集设计的评分模块,但评分逻辑与 DAPO 数学评分完全相同。