Skip to content

Policy Tuning and Deployment

Policy Tuning and Deployment

rlbotics turns a training run into a shippable policy - read the diagnostics, tune what matters, and export one self-contained ONNX file.

TIP

Install with pip install telekinesis-rlbotics — see Install Support for Reinforcement Learning for the simulator backends.

Tuning a Run

A reward curve tells you a run improved; the KL divergence and the action standard deviation tell you whether the next thousand iterations will improve or collapse.

Deploying a Policy

The export carries the observation normalization and, where the environment has action bounds, the action scaling - so a robot needs nothing but the file.

Where to Go Next?

Continue to the next tutorial.

Custom Environments and Algorithms

Train against your own simulator or robot, and plug in an algorithm other than PPO.

Next tutorial →

Support