CVE-2026-94093

Summary

A security vulnerability has been detected in DLR-RM stable-baselines3 up to 2.9.0. This affects the function PPO.load/load_replay_buffer/VecNormalize.load of the file save_util.py. Such manipulation leads to deserialization. It is possible to launch the attack remotely. The exploit has been disclosed publicly and may be used. In v2.9.0 the PyTorch tensor load path is hardened (weights_only=True), but that hardening was later reverted on master via PR #1913 "Hotfix: revert loading with weights_only=True" [blocked] to fix PyTorch 1.13 compat - so even the one "safe" path is inconsistent across versions. #2281 was closed as a duplicate of #1831 since both are unsafe pickle deserialization - but #1831's fix (PR #41) only gated the Hugging Face Hub loader in the separate huggingface_sb3 package. This finding covers the core stable_baselines3 load APIs (PPO.load, load_replay_buffer, VecNormalize.load), which have no safe mode or gate and remained exploitable in v2.9.0 until the outstanding hardening (PR #2264) ships.

Affected Software

VendorProductVersion RangeStatus
DLR-RMstable-baselines32.0affected
DLR-RMstable-baselines32.1affected
DLR-RMstable-baselines32.2affected
DLR-RMstable-baselines32.3affected
DLR-RMstable-baselines32.4affected
DLR-RMstable-baselines32.5affected
DLR-RMstable-baselines32.6affected
DLR-RMstable-baselines32.7affected
DLR-RMstable-baselines32.8affected
DLR-RMstable-baselines32.9.0affected

Weaknesses

  • CWE-502: Deserialization
  • CWE-20: Improper Input Validation

References