Batched GPU simulation#
The optional mjorbit_warp backend uses MuJoCo Warp. The supported Pixi GPU
workflow targets Linux x86-64 with an NVIDIA GPU. The CPU package remains
available in the same environment; backend selection is explicit by import.
pixi install -e warp
pixi run -e warp example-batched
Complete example#
"""Step 256 worlds with explicit GPU upload and pull.
Run on Linux with ``pixi run -e warp example-batched``.
"""
import numpy as np
from mjorbit.constants import GM_EARTH, R_EARTH
from mjorbit.testdata import FREE_BODY_XML
from mjorbit_warp import MjoModel, OrbitInit, mjo_forward, mjo_pull, mjo_step, mjo_upload
def main() -> None:
radius_km = R_EARTH + 400.0
model = MjoModel.from_xml_path(FREE_BODY_XML, mj_timestep=0.01)
data = model.make_data(
orbit=OrbitInit(
R_eci=[radius_km, 0.0, 0.0],
V_eci=[0.0, np.sqrt(GM_EARTH / radius_km), 0.0],
),
nworld=256,
)
# Each world starts at a different chief-relative offset, in meters.
data.qpos[:, 0] = np.linspace(0.0, 10.0, data.nworld)
mjo_upload(model, data, fields="qpos")
mjo_forward(model, data)
for _ in range(100):
mjo_step(model, data)
mjo_pull(model, data, fields=["time", "qpos", "orbit"])
print("qpos shape:", data.qpos.shape)
print("Time (s):", data.time)
print("First/last spacecraft offsets (m):", data.qpos[[0, -1], :3])
if __name__ == "__main__":
main()
This runs 256 copies of the bundled free-body model with different initial offsets. Initial kernel compilation can take longer than later runs.
Shapes and synchronization#
Use model.make_data(orbit=..., nworld=N). With N > 1, host state arrays
have a leading world dimension: qpos is (N, nq), qvel is (N, nv), and
ctrl is (N, nu). With nworld=1, host arrays use the unbatched shape.
The device state is authoritative. The public NumPy arrays are host mirrors:
After editing host
qpos,qvel,ctrl, or orbit, callmjo_uploadwith the fields to transfer.Call
mjo_forwardafter changing initial conditions, thenmjo_stepin the simulation loop.Call
mjo_pullbefore reading current state on the host.
Spacecraft command buffers (rw_torque_cmd, mtq_dipole_cmd, and
thr_force_cmd) are uploaded automatically by ordinary stepping. CUDA graph
capture uses device-side control; do not assume a host edit will be captured.
For simple host-driven code, mjo_step(model, data, sync=True) uploads and
pulls the full supported state each step, at an additional transfer cost.
Precision and current scope#
MuJoCo Warp’s multibody state, device chief position/velocity, and reference orbit propagation use single precision. The absolute epoch anchor, solar and magnetic angle accumulation, and the cancellation-sensitive J2 difference use double precision. The chief-centered formulation keeps robot coordinates near the origin. CPU/GPU results should be compared with precision-appropriate tolerances, especially around contacts and eclipse boundaries.
The gravity kernels currently use fixed Earth constants. Use the CPU backend
for custom central-body gravity; supplying different gm, j2, or gravity
radius values does not change the GPU gravity kernels. The environment kernels
do consume the configured atmosphere, magnetic, and eclipse parameters. See
forces and disturbances
for model equations and the shared articulated-body drag limitation.
Reaction wheels, thrusters, and magnetorquers are supported. CMGs raise
NotImplementedError during GPU model construction. Native sensor data is
available; the CPU noisy sensor namespace is not implemented on GPU. The
standard task viewer and mjorbit.planning planner use the CPU backend.
GPU/RL demonstrations have their own entry points:
pixi run -e warp banner-viewer-gpu
pixi run -e rl ppo-truss-train --help
See the source PPO guide for training and checkpoint playback. A trained policy is not bundled.