Robot checkpoint archive

Public research archive of multiple pi0.5 experiment families. PPO+CSD is intentionally excluded. Checkpoint names denote local run iterations; BC iteration counts are not necessarily optimizer-step-equivalent to PPO.

Each checkpoint contains actor/model_state_dict/full_weights.pt for the RLinf model loader and distributed actor/dcp_checkpoint training state. These are not standalone Transformers or LeRobot exports. LoRA checkpoints include the full model state, not adapter-only exports. Loading requires the matching RLinf/OpenPI code, configuration, base assets and normalization statistics. Resume support depends on the runner; these files do not imply exact environment/RNG replay.

See manifest.json for the frozen checkpoint inventory. Local files are not deleted.

Main PPO training runs

Main run Checkpoint directory Saved steps Contents
Main LIBERO-Goal PPO libero-goal-ppo 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150 Full weights and distributed optimizer/resume state
Main LIBERO-Spatial PPO libero-spatial-ppo-ns5 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150 Full weights and distributed optimizer/resume state

Both main runs use five denoising steps and standard action-expert PPO with the VLM frozen. “Main” identifies the original training runs; Goal weights under ablations/ are separate post-training parameter replacements. Spatial 7/3 PPO, LoRA, AdaRMS-only, and AdaRMS-frozen experiments are separate families.

All 30 main checkpoint directories were checked for full weights, distributed checkpoint metadata, and four training-state shards against the family manifests on 2026-09-20. Remote file sizes matched. See each family manifest for the complete inventory; the root manifest is the earlier archive snapshot. Existing checkpoint paths are unchanged.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support