Skip to main content

Training with Ray/RLlib

SuperDex Gym provides a Ray/RLlib integration for RL training on SuperDex Physics using Proximal Policy Optimization (PPO) or Soft Actor-Critic (SAC), and for running inference with trained checkpoints.

caution

Algorithm support: PPO is the recommended workflow. SAC is available experimentally, but its bundled settings and stopping thresholds have not been tuned or validated for the included environments. When SAC is selected, the training script ignores any per-environment settings under the recipe's ppo key and uses the same shared SAC algorithm configuration for every environment.

Quickstart​

Dependency requirements​

After uv sync --extra core, install the supported training dependencies from the project_superdex root:

uv pip install torch==2.7.1 --extra-index-url https://download.pytorch.org/whl/cpu
uv pip install "ray[rllib]==2.49.0" "moviepy" "pillow>=10.1" "tensorboard"

After installing the dependencies above, run the commands below from superdex_lab/apps/rllib with plain uv run. Because superdex_lab is an unmanaged uv project, uv run reuses the existing environment without synchronizing dependencies or replacing installed packages with local source builds. See Dependencies for apps.

Working directory​

All paths in this guide are relative to the superdex_lab project root. Change to the RLlib application directory before running the commands:

cd superdex_lab/apps/rllib

train_samples.py and run_inference.py support both package-relative imports and direct-script sibling imports. Invoking either script by path also works because Python adds the script directory to sys.path. The commands in this guide use direct-script execution from apps/rllib/. visualize_training_history.py has no sibling imports and runs from any directory.

Minimum training run​

Run one CartPole PPO iteration with one environment runner:

uv run python train_samples.py \
--pattern "superdex_gym/CartPole-v0" --num_env_runners 1 --max_iterations 1

A successful run creates a Tune trial named cart_pole_<trial_id> under ~/ray_results/PPO/, writes params.json and training logs in that trial directory, and writes a final checkpoint_* directory. The console reports training progress and exits after the first iteration. The final checkpoint is always written even when the configured checkpoint frequency has not elapsed.

Minimum inference run​

Pass the absolute path to the generated checkpoint_* directory, not its parent trial directory:

uv run python run_inference.py /path/to/checkpoint

The checkpoint must be a valid RLlib checkpoint with policy weights under learner_group/learner/rl_module/<DEFAULT_MODULE_ID>/. Its parent trial directory must contain params.json; pointing the command at the trial directory or another level fails. The script uses Ray's DEFAULT_MODULE_ID constant instead of a literal module name.

A successful inference run runs 10 episodes, prints each episode's return and completion reason, and reports the completed episode count. It does not initialize Ray. Pass --video to render offscreen and record MP4 files. Without --video, inference runs without rendering, because interactive viewing is not available.

Script locations and runtime behavior​

ScriptLocationPurpose
train_samples.pyapps/rllib/train_samples.pyTrain one or more environments
run_inference.pyapps/rllib/run_inference.pyRoll out a trained checkpoint
visualize_training_history.pyapps/rllib/visualize_training_history.pyStitch checkpoint videos into a montage

train_samples.py calls a bare ray.init(). It has no --ray-address option. To attach it to an existing cluster, set the RAY_ADDRESS environment variable; Ray reads that variable directly. run_inference.py does not initialize Ray.

Training a custom environment​

train_samples.py trains exact canonical IDs listed in its recipe manifests. To train an environment of your own from your own script, register it with Gymnasium and then expose it to Ray Tune:

from utils import register_envs

register_envs()

Register your custom Gymnasium spec in the superdex_gym namespace before calling register_envs(). register_envs() first registers the public SuperDex specs, then exposes every spec in the superdex_gym namespace to Tune. A spec registered under any other namespace is left untouched and must be registered with Tune by its owner.

The indirection matters because RLlib supplies env_config separately from Gymnasium constructor arguments. The Tune creator forwards env_config to env_spec.make(**env_config): for a SuperDex spec the shared factory recursively merges those values over the registered cfg defaults on an isolated deep copy (explicit caller values win while untouched nested defaults survive), and for any other spec they pass straight through as constructor keyword arguments.

Scripts Overview​

1. train_samples.py​

Train multiple sample environments simultaneously.

This script trains one Ray Tune experiment per selected environment, so several tasks can be trained side by side with a shared configuration.

Which environments are trainable. Explicit manifests map exact canonical IDs to a filesystem-safe output slug and recipe path. Paths are manifest data, not values inferred from environment modules, classes, variants, or IDs. Missing IDs have no recipe: variants never inherit base recipes and bases never inherit variant recipes. In an open-source build the trainable set is:

Gymnasium IDOutput slugRecipe
superdex_gym/AntNoContact-v0ant_no_contactsuperdex/lab/rllib/recipes/ant_no_contact/train.json
superdex_gym/CartPole-v0cart_polesuperdex/lab/rllib/recipes/cart_pole/train.json
superdex_gym/HalfCheetah-v0half_cheetahsuperdex/lab/rllib/recipes/half_cheetah/train.json

The registered base superdex_gym/Ant-v0 is not trainable because the manifest has no entry for that exact ID. The slug affects only the Tune trial directory name.

All three recipes use assets included in the checkout.

Features:

  • Simultaneous training of multiple environments
  • Algorithm selection: PPO (recommended) or SAC (experimental)
  • Pattern-based environment selection
  • Per-environment PPO hyperparameters from the recipe files
  • Centralized experiment management

Usage Examples:

# Train all trainable environments with PPO (default).
uv run python train_samples.py

# Train specific environments using canonical-ID patterns.
uv run python train_samples.py --pattern "superdex_gym/CartPole-v0"
uv run python train_samples.py --pattern "superdex_gym/*Cheetah*"

# Train with a custom configuration.
uv run python train_samples.py --num_env_runners 64 --checkpoint_freq 5

# Experimental SAC run on CartPole, limited to one training iteration.
# This checks the training path, not learning or convergence.
uv run python train_samples.py \
--algorithm SAC --pattern "superdex_gym/CartPole-v0" --max_iterations 1

# Train selected environments with video recording enabled (off by default)
uv run python train_samples.py --pattern "superdex_gym/Ant*" --video_on_checkpoint --output_path ./benchmark_results

# High-throughput training for benchmarking
uv run python train_samples.py --num_env_runners 128 --checkpoint_freq 20
caution
--pattern matches canonical IDs

Patterns use fnmatch against exact canonical Gymnasium IDs, not recipe slugs or discovery short names. For example, use "superdex_gym/CartPole-v0" or "superdex_gym/*Cheetah*". A pattern that matches no trainable canonical ID raises an error that lists the available IDs.

Command-line Options:

OptionDefaultNotes
--algorithm, -aPPOPPO (recommended) or SAC (experimental). SAC uses shared defaults and ignores per-environment ppo overrides.
--num_env_runners, -n32Parallel experience collectors. Capped to max(1, num_cpus - num_learners), with a printed WARNING: line, when num_env_runners + num_learners >= num_cpus.
--checkpoint_freq, -cf10Checkpoint every N training iterations. A final checkpoint is always written at the end of training.
--pattern, -p*Selects canonical Gymnasium IDs with fnmatch
--num_learners, -nl1Parallel policy-update processes. Unlike --num_env_runners, a value at or above the CPU count raises rather than being capped.
--output_path, -o~/ray_results/Root for Tune results
--video_on_checkpoint, -vidoffEncoding adds per-checkpoint overhead, so pass --video_on_checkpoint to enable. It is a BooleanOptionalAction, so --no-video_on_checkpoint is also accepted. When the renderer is unavailable it downgrades to disabled, emitting a warning and continuing rather than failing.
--profileoffEnables the environment profiler and dump_timings_to_info. Applies to every environment — profiling lives in the MochiEnv base class.
--max_iterations, -mi00 means no iteration limit; training stops on the recipe's reward / env-steps criteria

2. run_inference.py​

Run inference using trained RLlib policy checkpoints.

Usage Examples:

# Basic inference with a trained checkpoint.
uv run python run_inference.py /path/to/checkpoint

# Use the exploration forward pass.
uv run python run_inference.py /path/to/checkpoint --explore_during_inference

# Record videos of inference episodes.
uv run python run_inference.py /path/to/checkpoint --video

# Custom number of episodes and video path.
uv run python run_inference.py /path/to/checkpoint --num_episodes 20 --video_path ./inference_videos

Command-line Options:

OptionDefaultNotes
checkpoint_path—Required. Path to the checkpoint directory.
--num_episodes10
--explore_during_inferenceoffSamples stochastically from the exploration distribution instead of taking the greedy action.
--videooffRenders offscreen and records MP4 files
--video_paththe checkpoint directoryImplies --video

Without --video, inference runs without rendering, because interactive viewing is not available. Pass --video to render offscreen instead.

How actions are selected

The action distribution class is taken from the restored module via get_inference_action_dist_cls() / get_exploration_action_dist_cls(), so PPO gets a diagonal Gaussian, SAC a squashed Gaussian, and discrete action spaces a categorical. Without --explore_during_inference the distribution is collapsed with to_deterministic() before sampling, yielding the greedy action; with the flag, the action is drawn stochastically.

Videos are written as inference_000.mp4, inference_001.mp4, … at a hardcoded 30 fps — unlike run_sample.py, which uses the environment's control frequency.

Checkpoint requirements:

  • A valid RLlib checkpoint directory.
  • params.json in the checkpoint's parent (trial) directory — not inside the checkpoint directory itself. Pointing at the wrong level is the most common failure here.
  • Policy weights under learner_group/learner/rl_module/<DEFAULT_MODULE_ID>/. The script uses Ray's DEFAULT_MODULE_ID constant rather than a literal name.

Training recipe manifests​

Per-environment training settings are data files, not code. A manifest maps an exact canonical Gymnasium ID to a filesystem-safe output slug and an explicit recipe path. Public entries ship with the superdex.lab.rllib package under superdex/lab/rllib/recipes/manifest.json and load through the installable superdex.lab.rllib.recipe_manifest module. Add an exact manifest entry to make an environment trainable. No filename, class, module, variant, or related ID is used as fallback identity.

Schema​

{
"description": "Training recipe for the HalfCheetah benchmark env. PPO is the recommended workflow. SAC uses the experimental shared baseline.",
"normalize_observations": true,
"stop_criteria": {
"episode_return_mean": 9800,
"num_env_steps_sampled_lifetime": 50000000
},
"ppo": {
"env_runners": {
"batch_mode": "truncate_episodes"
},
"training": {
"clip_param": 0.2,
"gamma": 0.99,
"grad_clip": 0.5,
"kl_coeff": 1.0,
"lambda_": 0.95,
"lr": 0.0003,
"minibatch_size": 4096,
"num_epochs": 32,
"train_batch_size": 65536,
"vf_loss_coeff": 0.5
}
}
}

A recipe holds training settings only. Top-level observation normalization is algorithm-independent and applies to PPO and SAC. The bundled stopping thresholds were selected for PPO and have not been validated for SAC.

Unknown top-level keys and unknown keys inside stop_criteria are rejected on every run with a ValueError naming the offending key and supported set. Validation inside the ppo block occurs only under --algorithm PPO, because SAC ignores that entire section. Consequently, an unknown ppo key raises under PPO but is a silent no-op under SAC. An env_config section is a hard error under either algorithm.

Environment defaults come from the exact registered Gymnasium EnvSpec, never from the recipe. train_samples.py serializes only the runtime profile and dump_timings_to_info overrides; the Ray creator recursively merges them over the registered defaults.

KeyMeaning
descriptionHuman-readable only. The training script never reads it.
normalize_observationsWhen true, installs an RLlib MeanStdFilter environment-to-module connector for each environment runner. This is algorithm-independent and applies to both PPO and SAC.
stop_criteriaTwo keys are recognised: episode_return_mean maps to RLlib's env_runners/episode_return_mean, and num_env_steps_sampled_lifetime maps to the metric of the same name. Any other key raises ValueError. --max_iterations adds a training_iteration stop on top.
ppoPPO overrides layered onto default_ppo_config. ppo["env_runners"] and ppo["training"] are passed to the corresponding RLlib configuration methods. The additional train_batch_size_per_runner setting is multiplied by the effective runtime --num_env_runners and written to train_batch_size. Setting it together with ppo.training.train_batch_size raises under PPO because both configure the same value. SAC ignores and does not validate the entire ppo section.
Recipe scope and algorithm-specific validation
  • A recipe is scoped to exactly one canonical ID. A variant does not inherit its base's recipe, and a base does not pick up a variant's. Add the exact ID, output slug, and explicit recipe path to the appropriate manifest.
  • The ppo block is ignored under SAC. --algorithm SAC uses default_sac_config for every environment without per-environment PPO overrides. Because that block is not read, a misspelled key or conflicting batch-size settings inside it are errors under PPO and silent no-ops under SAC.

Environment configuration and RLlib recipe manifests​

Environment implementation and variant files remain separate from RLlib recipe data:

ArtifactMeaning
<name>_env.pyEnv module. Its MochiEnv subclass becomes a Gymnasium ID (AntEnv → superdex_gym/Ant-v0, short name ant)
<module>_<variant>.jsonGym config variant — the only place env configuration may live. Registered as its own Gymnasium ID and short name (ant_env_no_contact.json → superdex_gym/AntNoContact-v0, short name ant_no_contact)
superdex/lab/rllib/recipes/manifest.jsonPublic exact-ID mapping to slugs and explicit recipe paths, loaded through superdex.lab.rllib.recipe_manifest

<variant> must be a snake_case token, with three rules on its segments:

  • No segment may be env. Every env module ends in the _env segment, so this is what keeps a longer sibling module's files (foo_env_extra_env.json) from reading as a variant of the shorter one (foo_env.py).
  • No segment may be train or benchmark. These usage categories remain reserved even though RLlib recipe paths are selected explicitly by the manifest above.
  • A test segment marks the variant test-only. It is discovered and smoke-tested, but never registered with Gymnasium and never listed by a CLI. Use this for degenerate configurations (no gravity, no damping) that are worth crash-checking but are not shippable tasks.

Because training recipes may not configure the environment, every configuration that gets trained is also a named, runnable, smoke-tested environment. Training fails loudly if a train recipe contains an env_config section.

benchmark recipes are the exception to the training-only schema: they may carry an env_cfg measurement baseline. No public environment ships a benchmark recipe. The separate worker-sweep script apps/envs/benchmark.py owns its baseline independently.

General Usage Notes​

Training Configuration​

PPO (default_ppo_config):

SettingValue
Network2-layer fully connected, [64, 64], tanh activation
Value functionShares layers with the policy (vf_share_layers: True)
Learning rate0.0003, fixed — there is no schedule
train_batch_size16384 (32 × 512)
minibatch_size4096
num_epochs15
lambda_0.95
vf_loss_coeff0.01
rollout_fragment_length512
CallbacksLogRewardAndInfoCallbacks
EvaluationRequests one complete episode every iteration, parallel to training, explore=False

Experimental SAC baseline (default_sac_config):

These shared settings are provided to exercise the RLlib integration. They have not been tuned or validated for the included environments, and the recipe's ppo overrides are not applied. The recipe stopping thresholds are likewise not evidence that SAC will reach them.

SettingValue
ModelRLlib's default SAC RLModule; no model overrides are applied
Actor / critic / alpha learning rates3e-5 / 3e-4 / 1e-4
initial_alpha0.1, with target_entropy="auto"
train_batch_size256
tau0.005, target_network_update_freq 1
n_step1
Replay bufferEpisodeReplayBuffer, capacity 1e5, batch 256 × 1
Learning starts after1024 sampled steps
rollout_fragment_length1
CallbacksRLlib's stock DefaultCallbacks
EvaluationEvery iteration, parallel to training, explore=False

Both run on CPU — num_gpus_per_learner=0.

Distributed Training​

  • Environment runners: parallel processes collecting experience.
  • Learners: parallel processes updating policy parameters.
  • Evaluation: a dedicated evaluation runner, in parallel with training. PPO evaluation remains subject to Ray's evaluation_sample_timeout_s (120 seconds by default), and its reported returns use the configured five-episode smoothing window.

num_env_runners is capped to the available CPU count minus the learner count, with a printed WARNING: line, once num_env_runners + num_learners reaches the CPU count; num_learners at or above the CPU count raises instead. An example such as --num_env_runners 128 on a smaller machine is therefore reduced, and says so.

Checkpointing and Monitoring​

  • Checkpoints every --checkpoint_freq iterations, plus one at the end of training.
  • TensorBoard logging for training metrics.
  • Optional per-checkpoint video generation.

Performance Optimization​

  • CPU usage: worker counts are capped against the detected CPU count.
  • Memory: batch sizes are set per algorithm in the tables above.

Output Structure​

Training outputs​

Results are rooted at --output_path (default ~/ray_results/). Ray Tune creates an algorithm directory, then one trial directory per environment named <recipe_slug>_<trial_id>:

<output_path>/
|-- PPO/
| |-- ant_no_contact_a1b2c3d4/
| |-- cart_pole_e5f6a7b8/
| |-- half_cheetah_c9d0e1f2/
| |-- checkpoint_000000/
| | |-- learner_group/
| | |-- video_000.mp4 # if --video_on_checkpoint
| | |-- video_labels.json
| |-- params.json # read by run_inference.py
| |-- progress.csv
| |-- events.out.tfevents.*
|-- SAC/
|-- ...

Note that params.json sits in the trial directory, one level above each checkpoint_*, and that the per-checkpoint videos live inside the checkpoint directory.

Inference outputs​

  • Console output with episode returns and completion reasons
  • MP4 files if recording is enabled
  • Completed episode count

Debug options​

  • Ray dashboard at http://localhost:8265 for cluster monitoring.
  • TensorBoard: uv run tensorboard --logdir ~/ray_results (or your --output_path). Nothing installs the tensorboard command — ray[rllib] brings tensorboardx, which writes the event files but does not read them — so uv pip install tensorboard first.
  • Tune verbosity is fixed at V1_EXPERIMENT; there is no flag to change it.

Callbacks Directory​

apps/rllib/callbacks/ contains two callbacks. They are wired in different ways: LogRewardAndInfoCallbacks is an RLlib DefaultCallbacks subclass attached to the algorithm config, while CheckpointVideoGeneratorCallback is a Ray Tune callback attached to the tuner.

LogRewardAndInfoCallbacks​

Detailed episode statistics logging for PPO only. This RLlib DefaultCallbacks subclass is attached to the PPO algorithm config and hooks on_episode_end only. SAC uses RLlib's stock DefaultCallbacks, so the metrics below do not appear in a SAC run.

At the end of each episode, it emits four metric families. Each reports a mean or standard deviation over the steps of that episode:

MetricMeaning
info_<key>_meanMean of info[<key>] across the episode's steps
info_<key>_stdevStandard deviation of the same
reward_meanMean per-step reward. Not the episode return.
reward_stdevStandard deviation of the per-step reward

Because MochiEnv surfaces every reward term as info["reward_<term>"], a term named upright_reward arrives as info_reward_upright_reward_mean. String-valued info entries — terminated_reason, for instance — are skipped, since they cannot be averaged.

Enabling --profile also switches on dump_timings_to_info, which adds a timings/* family to info and therefore a corresponding set of info_timings/*_mean series. That is a lot of extra metrics; turn it off once you have the numbers you wanted.

CheckpointVideoGeneratorCallback​

Video generation at training checkpoints. A Ray Tune callback, attached to the tuner rather than to the algorithm config.

CheckpointVideoGeneratorCallback(
count=1, # sequential rollouts, one video each
fmt="mp4",
fps=30,
pattern=None, # fnmatch pattern(s) over trial IDs; None means all
log_to_tensorboard=True,
tb_log_every=1,
)

count and tb_log_every must both be >= 1. The TensorBoard counter is zero-based per trial, so the first checkpoint — the pre-training baseline — is always logged, then every tb_log_every-th one after it.

Output, written into each checkpoint_*/ directory:

FileContents
video_000.mp4, video_001.mp4, …One per requested rollout (count)
video_labels.json{"training_iteration": int, "num_env_steps_sampled_lifetime": int}, consumed by visualize_training_history.py

The TensorBoard summary is written to the trial directory rather than the checkpoint directory, so the clips share a run with the scalar metrics. The tags are rollout/video_<index>.

The callback runs the requested rollouts sequentially through one environment. That environment is still created through AsyncVectorEnv, but purely for process isolation: the renderer does not support multiple contexts per process, and the Ray Tune driver cannot initialize it in-process. Increasing count therefore increases checkpoint video-generation time approximately linearly; it does not add concurrent render workers.

To keep checkpointing responsive, this TensorBoard preview is capped at 300 frames and a 480-pixel longest side while preserving approximately the same playback duration. The video_*.mp4 in the checkpoint remains full-resolution and contains every rendered frame; use that file for behavior review and archival.

caution
moviepy is required for the TensorBoard videos

TensorBoardX encodes video summaries through moviepy. When it is missing, TensorBoardX logs moviepy is not installed. at ERROR level once per checkpoint and writes no video payload, but nothing is raised: the MP4 files still land in the checkpoint directory and training continues. Nothing appears in TensorBoard, and that logged error is the only warning you get. moviepy is not a declared dependency of superdex-lab either; install it alongside Ray — see Installation and Setup.

Rendering happens in spawned subprocesses

The renderer does not support more than one GL context per process, and the Ray Tune driver is multi-threaded. The callback therefore runs its rollouts in a Gymnasium AsyncVectorEnv with the "spawn" start method — forking the driver and then initializing OpenGL/EGL in the child segfaults. This also means the rollout environments are constructed fresh, with render_mode forced to "rgb_array".

Only the sidecar-writing and TensorBoard-logging helpers are best-effort. Ray runs on_checkpoint before the checkpoint is registered, so those two helpers log and swallow failures instead of aborting training. The rest of on_checkpoint has no try/except: failures while loading the module, rendering the rollout, or writing the animations propagate. If either metric is unavailable, the sidecar is skipped instead of being written with placeholder zeros; visualize_training_history.py can then fall back to timestamp matching.

Usage: both callbacks are wired up by train_samples.py:

  • LogRewardAndInfoCallbacks for PPO runs (SAC uses RLlib's DefaultCallbacks)
  • CheckpointVideoGeneratorCallback when --video_on_checkpoint is passed (it is off by default) and the renderer is available