prometheus_utils.py — 提供 Prometheus 监控配置的自动更新功能¶
文件路径: verl/experimental/agent_loop/prometheus_utils.py
文件概述¶
提供 Prometheus 监控配置的自动更新功能。Prometheus 是一个流行的开源监控系统,verl 使用它来收集 rollout 服务器(vLLM/SGLang)的性能指标。该文件的核心功能是在所有 Ray 节点上自动写入 Prometheus 配置文件并触发热重载。
关键代码讲解¶
update_prometheus_config - 核心函数¶
def update_prometheus_config(config: PrometheusConfig, server_addresses: list[str], rollout_name: str = None):
"""
更新 Prometheus 配置文件并在所有节点上重载。
Args:
config: Prometheus 配置(包含文件路径和端口)
server_addresses: vLLM 或 SGLang 服务器地址列表
rollout_name: rollout 后端名称(如 "vllm"、"sglang")
"""
工作流程¶
- 构建配置 JSON:
prometheus_config_json = {
"global": {"scrape_interval": "10s", "evaluation_interval": "10s"},
"scrape_configs": [
{
"job_name": "ray",
"file_sd_configs": [{"files": ["/tmp/ray/prom_metrics_service_discovery.json"]}],
},
{
"job_name": "rollout",
"static_configs": [{"targets": server_addresses}]
},
],
}
包含两个采集任务:
- ray:从 Ray 的服务发现文件采集 Ray 集群指标
- rollout:从 rollout 服务器(vLLM/SGLang)地址采集推理指标
- 分布式写入配置文件(使用 Ray Remote):
@ray.remote(num_cpus=0)
def write_config_file(config_data, config_path):
os.makedirs(os.path.dirname(config_path), exist_ok=True)
with open(config_path, "w") as f:
yaml.dump(config_data, f, default_flow_style=False, indent=2)
# 在所有存活节点上调度写入任务
for node in alive_nodes:
node_ip = node["NodeManagerAddress"]
task = write_config_file.options(
resources={"node:" + node_ip: 0.001} # 调度到特定节点
).remote(prometheus_config_json, config.file)
- 触发 Prometheus 热重载:
@ray.remote(num_cpus=0)
def reload_prometheus(port):
reload_url = f"http://{ip_address}:{port}/-/reload"
subprocess.run(["curl", "-X", "POST", reload_url], ...)
核心类/函数列表¶
| 名称 | 类型 | 说明 |
|---|---|---|
update_prometheus_config() |
函数 | 更新 Prometheus 配置并重载 |
与其他模块的关系¶
- 被
AgentLoopManager._initialize_llm_servers()调用,在 rollout 服务器启动后自动配置监控 - 依赖
verl.workers.config.RolloutConfig中的PrometheusConfig获取端口和文件路径
小结¶
prometheus_utils.py 是一个运维辅助模块,实现了 Prometheus 监控配置的自动化管理。在分布式训练中,rollout 服务器可能分布在多个节点上,该模块确保每个节点都能正确配置 Prometheus 采集目标,实现对推理服务器性能的实时监控。