跳转至

nvtx_profile.py — NVIDIA NVTX 分析

文件路径: verl/utils/profiler/nvtx_profile.py

文件概述

基于 NVIDIA NVTX (NVIDIA Tools Extension) 的性能标记实现。配合 Nsight Systems 使用,可以在时间线上看到各个函数的执行区间。

核心函数

def mark_start_range(message=None, color=None, domain=None, category=None):
    return nvtx.start_range(message=message, color=color, domain=domain, category=category)

def mark_end_range(range_id):
    return nvtx.end_range(range_id)

def mark_annotate(message=None, color=None, ...):
    """函数装饰器,自动标记函数执行区间"""
    def decorator(func):
        return nvtx.annotate(profile_message, ...)(func)
    return decorator

NsightSystemsProfiler

class NsightSystemsProfiler(DistProfiler):
    def start(self, **kwargs):
        if not self.discrete:
            torch.cuda.profiler.start()

    def annotate(self, message=None, ...):
        """为 worker 函数添加分析区间"""
        def wrapper(*args, **kwargs_inner):
            if self.discrete:
                torch.cuda.profiler.start()
            mark_range = mark_start_range(...)
            result = func(...)
            mark_end_range(mark_range)
            if self.discrete:
                torch.cuda.profiler.stop()

支持两种模式: - 连续模式: 整个训练 step 是一个分析会话 - 离散模式: 每个 task(如 actor update、critic update)独立分析

与其他模块的关系

  • 被 profile.py 的 DistProfiler 在 NVIDIA GPU 上使用
  • 需要 nvtx Python 包

小结

NVIDIA GPU 上的高性能分析标记实现。