Training with Ray/RLlib
SuperDex Gym provides a Ray/RLlib integration for RL training on SuperDex Physics using Proximal Policy Optimization (PPO) or Soft Actor-Critic (SAC), and for running inference with trained checkpoints.
Algorithm support: PPO is the recommended workflow. SAC is available
experimentally, but its bundled settings and stopping thresholds have not been tuned
or validated for the included environments. When SAC is selected, the training script
ignores any per-environment settings under the recipe's ppo key and uses the same
shared SAC algorithm configuration for every environment.
Quickstart
Dependency requirements
After uv sync --extra core, install the supported training dependencies from the
project_superdex root:
uv pip install torch==2.7.1 --extra-index-url https://download.pytorch.org/whl/cpu
uv pip install "ray[rllib]==2.49.0" "moviepy" "pillow>=10.1" "tensorboard"
After installing the dependencies above, run the commands below from
superdex_lab/apps/rllib with plain uv run. Because superdex_lab is an
unmanaged uv project, uv run reuses the existing environment without synchronizing
dependencies or replacing installed packages with local source builds. See
Dependencies for apps.
Working directory
All paths in this guide are relative to the superdex_lab project root. Change to the
RLlib application directory before running the commands:
cd superdex_lab/apps/rllib
train_samples.py and run_inference.py support both package-relative imports and
direct-script sibling imports. Invoking either script by path also works because Python
adds the script directory to sys.path. The commands in this guide use direct-script
execution from apps/rllib/. visualize_training_history.py has no sibling imports and
runs from any directory.
Minimum training run
Run one CartPole PPO iteration with one environment runner:
uv run python train_samples.py \
--pattern "superdex_gym/CartPole-v0" --num_env_runners 1 --max_iterations 1
A successful run creates a Tune trial named cart_pole_<trial_id> under
~/ray_results/PPO/, writes params.json and training logs in that trial directory,
and writes a final checkpoint_* directory. The console reports training progress
and exits after the first iteration. The final checkpoint is always written even when
the configured checkpoint frequency has not elapsed.
Minimum inference run
Pass the absolute path to the generated checkpoint_* directory, not its parent trial
directory:
uv run python run_inference.py /path/to/checkpoint
The checkpoint must be a valid RLlib checkpoint with policy weights under
learner_group/learner/rl_module/<DEFAULT_MODULE_ID>/. Its parent trial directory
must contain params.json; pointing the command at the trial directory or another
level fails. The script uses Ray's DEFAULT_MODULE_ID constant instead of a literal
module name.
A successful inference run runs 10 episodes, prints each episode's return and
completion reason, and reports the completed episode count. It does not initialize
Ray. Pass --video to render offscreen and record MP4 files. Without --video,
inference runs without rendering, because interactive viewing is not available.
Script locations and runtime behavior
| Script | Location | Purpose |
|---|---|---|
train_samples.py | apps/rllib/train_samples.py | Train one or more environments |
run_inference.py | apps/rllib/run_inference.py | Roll out a trained checkpoint |
visualize_training_history.py | apps/rllib/visualize_training_history.py | Stitch checkpoint videos into a montage |
train_samples.py calls a bare ray.init(). It has no --ray-address option. To
attach it to an existing cluster, set the RAY_ADDRESS environment variable; Ray
reads that variable directly. run_inference.py does not initialize Ray.
Training a custom environment
train_samples.py trains exact canonical IDs listed in its recipe manifests. To train
an environment of your own from your own script, register it with Gymnasium and then
expose it to Ray Tune:
from utils import register_envs
register_envs()
Register your custom Gymnasium spec in the superdex_gym namespace before calling
register_envs(). register_envs() first registers the public SuperDex specs, then
exposes every spec in the superdex_gym namespace to Tune. A spec registered under any
other namespace is left untouched and must be registered with Tune by its owner.
The indirection matters because RLlib supplies env_config separately from Gymnasium
constructor arguments. The Tune creator forwards env_config to
env_spec.make(**env_config): for a SuperDex spec the shared factory recursively merges
those values over the registered cfg defaults on an isolated deep copy (explicit caller
values win while untouched nested defaults survive), and for any other spec they pass
straight through as constructor keyword arguments.
Scripts Overview
1. train_samples.py
Train multiple sample environments simultaneously.
This script trains one Ray Tune experiment per selected environment, so several tasks can be trained side by side with a shared configuration.
Which environments are trainable. Explicit manifests map exact canonical IDs to a filesystem-safe output slug and recipe path. Paths are manifest data, not values inferred from environment modules, classes, variants, or IDs. Missing IDs have no recipe: variants never inherit base recipes and bases never inherit variant recipes. In an open-source build the trainable set is:
| Gymnasium ID | Output slug | Recipe |
|---|---|---|
superdex_gym/AntNoContact-v0 | ant_no_contact | superdex/lab/rllib/recipes/ant_no_contact/train.json |
superdex_gym/CartPole-v0 | cart_pole | superdex/lab/rllib/recipes/cart_pole/train.json |
superdex_gym/HalfCheetah-v0 | half_cheetah | superdex/lab/rllib/recipes/half_cheetah/train.json |
The registered base superdex_gym/Ant-v0 is not trainable because the manifest has no
entry for that exact ID. The slug affects only the Tune trial directory name.
All three recipes use assets included in the checkout.
Features:
- Simultaneous training of multiple environments
- Algorithm selection: PPO (recommended) or SAC (experimental)
- Pattern-based environment selection
- Per-environment PPO hyperparameters from the recipe files
- Centralized experiment management
Usage Examples:
# Train all trainable environments with PPO (default).
uv run python train_samples.py
# Train specific environments using canonical-ID patterns.
uv run python train_samples.py --pattern "superdex_gym/CartPole-v0"
uv run python train_samples.py --pattern "superdex_gym/*Cheetah*"
# Train with a custom configuration.
uv run python train_samples.py --num_env_runners 64 --checkpoint_freq 5
# Experimental SAC run on CartPole, limited to one training iteration.
# This checks the training path, not learning or convergence.
uv run python train_samples.py \
--algorithm SAC --pattern "superdex_gym/CartPole-v0" --max_iterations 1
# Train selected environments with video recording enabled (off by default)
uv run python train_samples.py --pattern "superdex_gym/Ant*" --video_on_checkpoint --output_path ./benchmark_results
# High-throughput training for benchmarking
uv run python train_samples.py --num_env_runners 128 --checkpoint_freq 20
--pattern matches canonical IDsPatterns use fnmatch against exact canonical Gymnasium IDs, not recipe slugs or
discovery short names. For example, use "superdex_gym/CartPole-v0" or
"superdex_gym/*Cheetah*". A pattern that matches no trainable canonical ID raises an
error that lists the available IDs.
Command-line Options:
| Option | Default | Notes |
|---|---|---|
--algorithm, -a | PPO | PPO (recommended) or SAC (experimental). SAC uses shared defaults and ignores per-environment ppo overrides. |
--num_env_runners, -n | 32 | Parallel experience collectors. Capped to max(1, num_cpus - num_learners), with a printed WARNING: line, when num_env_runners + num_learners >= num_cpus. |
--checkpoint_freq, -cf | 10 | Checkpoint every N training iterations. A final checkpoint is always written at the end of training. |
--pattern, -p | * | Selects canonical Gymnasium IDs with fnmatch |
--num_learners, -nl | 1 | Parallel policy-update processes. Unlike --num_env_runners, a value at or above the CPU count raises rather than being capped. |
--output_path, -o | ~/ray_results/ | Root for Tune results |
--video_on_checkpoint, -vid | off | Encoding adds per-checkpoint overhead, so pass --video_on_checkpoint to enable. It is a BooleanOptionalAction, so --no-video_on_checkpoint is also accepted. When the renderer is unavailable it downgrades to disabled, emitting a warning and continuing rather than failing. |
--profile | off | Enables the environment profiler and dump_timings_to_info. Applies to every environment — profiling lives in the MochiEnv base class. |
--max_iterations, -mi | 0 | 0 means no iteration limit; training stops on the recipe's reward / env-steps criteria |
2. run_inference.py
Run inference using trained RLlib policy checkpoints.
Usage Examples:
# Basic inference with a trained checkpoint.
uv run python run_inference.py /path/to/checkpoint
# Use the exploration forward pass.
uv run python run_inference.py /path/to/checkpoint --explore_during_inference
# Record videos of inference episodes.
uv run python run_inference.py /path/to/checkpoint --video
# Custom number of episodes and video path.
uv run python run_inference.py /path/to/checkpoint --num_episodes 20 --video_path ./inference_videos
Command-line Options:
| Option | Default | Notes |
|---|---|---|
checkpoint_path | — | Required. Path to the checkpoint directory. |
--num_episodes | 10 | |
--explore_during_inference | off | Samples stochastically from the exploration distribution instead of taking the greedy action. |
--video | off | Renders offscreen and records MP4 files |
--video_path | the checkpoint directory | Implies --video |
Without --video, inference runs without rendering, because interactive viewing is not
available. Pass --video to render offscreen instead.
The action distribution class is taken from the restored module via
get_inference_action_dist_cls() / get_exploration_action_dist_cls(), so PPO
gets a diagonal Gaussian, SAC a squashed Gaussian, and discrete action spaces a
categorical. Without --explore_during_inference the distribution is collapsed
with to_deterministic() before sampling, yielding the greedy action; with the
flag, the action is drawn stochastically.
Videos are written as inference_000.mp4, inference_001.mp4, … at a hardcoded
30 fps — unlike run_sample.py, which uses the environment's control frequency.
Checkpoint requirements:
- A valid RLlib checkpoint directory.
params.jsonin the checkpoint's parent (trial) directory — not inside the checkpoint directory itself. Pointing at the wrong level is the most common failure here.- Policy weights under
learner_group/learner/rl_module/<DEFAULT_MODULE_ID>/. The script uses Ray'sDEFAULT_MODULE_IDconstant rather than a literal name.
Training recipe manifests
Per-environment training settings are data files, not code. A manifest maps an exact
canonical Gymnasium ID to a filesystem-safe output slug and an explicit recipe path.
Public entries ship with the superdex.lab.rllib package under
superdex/lab/rllib/recipes/manifest.json and load through the installable
superdex.lab.rllib.recipe_manifest module. Add an exact manifest entry to make an
environment trainable. No filename, class, module, variant, or related ID is used as
fallback identity.
Schema
{
"description": "Training recipe for the HalfCheetah benchmark env. PPO is the recommended workflow. SAC uses the experimental shared baseline.",
"normalize_observations": true,
"stop_criteria": {
"episode_return_mean": 9800,
"num_env_steps_sampled_lifetime": 50000000
},
"ppo": {
"env_runners": {
"batch_mode": "truncate_episodes"
},
"training": {
"clip_param": 0.2,
"gamma": 0.99,
"grad_clip": 0.5,
"kl_coeff": 1.0,
"lambda_": 0.95,
"lr": 0.0003,
"minibatch_size": 4096,
"num_epochs": 32,
"train_batch_size": 65536,
"vf_loss_coeff": 0.5
}
}
}
A recipe holds training settings only. Top-level observation normalization is algorithm-independent and applies to PPO and SAC. The bundled stopping thresholds were selected for PPO and have not been validated for SAC.
Unknown top-level keys and unknown keys inside stop_criteria are rejected on every run
with a ValueError naming the offending key and supported set. Validation inside the
ppo block occurs only under --algorithm PPO, because SAC ignores that entire section.
Consequently, an unknown ppo key raises under PPO but is a silent no-op under SAC. An
env_config section is a hard error under either algorithm.
Environment defaults come from the exact registered Gymnasium EnvSpec, never from the
recipe. train_samples.py serializes only the runtime profile and
dump_timings_to_info overrides; the Ray creator recursively merges them over the
registered defaults.
| Key | Meaning |
|---|---|
description | Human-readable only. The training script never reads it. |
normalize_observations | When true, installs an RLlib MeanStdFilter environment-to-module connector for each environment runner. This is algorithm-independent and applies to both PPO and SAC. |
stop_criteria | Two keys are recognised: episode_return_mean maps to RLlib's env_runners/episode_return_mean, and num_env_steps_sampled_lifetime maps to the metric of the same name. Any other key raises ValueError. --max_iterations adds a training_iteration stop on top. |
ppo | PPO overrides layered onto default_ppo_config. ppo["env_runners"] and ppo["training"] are passed to the corresponding RLlib configuration methods. The additional train_batch_size_per_runner setting is multiplied by the effective runtime --num_env_runners and written to train_batch_size. Setting it together with ppo.training.train_batch_size raises under PPO because both configure the same value. SAC ignores and does not validate the entire ppo section. |
- A recipe is scoped to exactly one canonical ID. A variant does not inherit its base's recipe, and a base does not pick up a variant's. Add the exact ID, output slug, and explicit recipe path to the appropriate manifest.
- The
ppoblock is ignored under SAC.--algorithm SACusesdefault_sac_configfor every environment without per-environment PPO overrides. Because that block is not read, a misspelled key or conflicting batch-size settings inside it are errors under PPO and silent no-ops under SAC.
Environment configuration and RLlib recipe manifests
Environment implementation and variant files remain separate from RLlib recipe data:
| Artifact | Meaning |
|---|---|
<name>_env.py | Env module. Its MochiEnv subclass becomes a Gymnasium ID (AntEnv → superdex_gym/Ant-v0, short name ant) |
<module>_<variant>.json | Gym config variant — the only place env configuration may live. Registered as its own Gymnasium ID and short name (ant_env_no_contact.json → superdex_gym/AntNoContact-v0, short name ant_no_contact) |
superdex/lab/rllib/recipes/manifest.json | Public exact-ID mapping to slugs and explicit recipe paths, loaded through superdex.lab.rllib.recipe_manifest |
<variant> must be a snake_case token, with three rules on its segments:
- No segment may be
env. Every env module ends in the_envsegment, so this is what keeps a longer sibling module's files (foo_env_extra_env.json) from reading as a variant of the shorter one (foo_env.py). - No segment may be
trainorbenchmark. These usage categories remain reserved even though RLlib recipe paths are selected explicitly by the manifest above. - A
testsegment marks the variant test-only. It is discovered and smoke-tested, but never registered with Gymnasium and never listed by a CLI. Use this for degenerate configurations (no gravity, no damping) that are worth crash-checking but are not shippable tasks.
Because training recipes may not configure the environment, every configuration that gets
trained is also a named, runnable, smoke-tested environment. Training fails loudly if a
train recipe contains an env_config section.
benchmark recipes are the exception to the training-only schema: they may carry an
env_cfg measurement baseline. No public environment ships a benchmark recipe. The
separate worker-sweep script apps/envs/benchmark.py owns its baseline independently.
General Usage Notes
Training Configuration
PPO (default_ppo_config):
| Setting | Value |
|---|---|
| Network | 2-layer fully connected, [64, 64], tanh activation |
| Value function | Shares layers with the policy (vf_share_layers: True) |
| Learning rate | 0.0003, fixed — there is no schedule |
train_batch_size | 16384 (32 × 512) |
minibatch_size | 4096 |
num_epochs | 15 |
lambda_ | 0.95 |
vf_loss_coeff | 0.01 |
rollout_fragment_length | 512 |
| Callbacks | LogRewardAndInfoCallbacks |
| Evaluation | Requests one complete episode every iteration, parallel to training, explore=False |
Experimental SAC baseline (default_sac_config):
These shared settings are provided to exercise the RLlib integration. They have not
been tuned or validated for the included environments, and the recipe's ppo
overrides are not applied. The recipe stopping thresholds are likewise not evidence
that SAC will reach them.
| Setting | Value |
|---|---|
| Model | RLlib's default SAC RLModule; no model overrides are applied |
| Actor / critic / alpha learning rates | 3e-5 / 3e-4 / 1e-4 |
initial_alpha | 0.1, with target_entropy="auto" |
train_batch_size | 256 |
tau | 0.005, target_network_update_freq 1 |
n_step | 1 |
| Replay buffer | EpisodeReplayBuffer, capacity 1e5, batch 256 × 1 |
| Learning starts after | 1024 sampled steps |
rollout_fragment_length | 1 |
| Callbacks | RLlib's stock DefaultCallbacks |
| Evaluation | Every iteration, parallel to training, explore=False |
Both run on CPU — num_gpus_per_learner=0.
Distributed Training
- Environment runners: parallel processes collecting experience.
- Learners: parallel processes updating policy parameters.
- Evaluation: a dedicated evaluation runner, in parallel with training.
PPO evaluation remains subject to Ray's
evaluation_sample_timeout_s(120 seconds by default), and its reported returns use the configured five-episode smoothing window.
num_env_runners is capped to the available CPU count minus the learner count, with a
printed WARNING: line, once num_env_runners + num_learners reaches the CPU count;
num_learners at or above the CPU count raises instead. An example such as
--num_env_runners 128 on a smaller machine is therefore reduced, and says so.
Checkpointing and Monitoring
- Checkpoints every
--checkpoint_freqiterations, plus one at the end of training. - TensorBoard logging for training metrics.
- Optional per-checkpoint video generation.
Performance Optimization
- CPU usage: worker counts are capped against the detected CPU count.
- Memory: batch sizes are set per algorithm in the tables above.
Output Structure
Training outputs
Results are rooted at --output_path (default ~/ray_results/). Ray Tune creates
an algorithm directory, then one trial directory per environment named
<recipe_slug>_<trial_id>:
<output_path>/
|-- PPO/
| |-- ant_no_contact_a1b2c3d4/
| |-- cart_pole_e5f6a7b8/
| |-- half_cheetah_c9d0e1f2/
| |-- checkpoint_000000/
| | |-- learner_group/
| | |-- video_000.mp4 # if --video_on_checkpoint
| | |-- video_labels.json
| |-- params.json # read by run_inference.py
| |-- progress.csv
| |-- events.out.tfevents.*
|-- SAC/
|-- ...
Note that params.json sits in the trial directory, one level above each
checkpoint_*, and that the per-checkpoint videos live inside the checkpoint
directory.
Inference outputs
- Console output with episode returns and completion reasons
- MP4 files if recording is enabled
- Completed episode count
Debug options
- Ray dashboard at
http://localhost:8265for cluster monitoring. - TensorBoard:
uv run tensorboard --logdir ~/ray_results(or your--output_path). Nothing installs thetensorboardcommand —ray[rllib]bringstensorboardx, which writes the event files but does not read them — souv pip install tensorboardfirst. - Tune verbosity is fixed at
V1_EXPERIMENT; there is no flag to change it.