You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Q-Planning โ€” Q-function (LIBERO)

Off-policy Q-function over action chunks for the LIBERO suites, from Beyond Imitation: Self-Improving Robot Policies via Off-Policy Q-Planning (CoRL 2026).

At inference the Q-function scores action chunks proposed by a frozen behaviour-cloning policy and selects among them. Only the Q-function is updated on deployment rollouts โ€” including failed ones, which imitation learning cannot use. The BC policy's weights are never touched.

This checkpoint does not act on its own

It is a critic: it scores another policy's proposals and emits no actions. To run it you need the frozen FastWAM BC policy, which is published by its own authors:

# 1. Official FastWAM weights (~12 GB) + normalisation stats
hf download yuanty/fastwam libero_uncond_2cam224.pt \
    libero_uncond_2cam224_dataset_stats.json --local-dir $QPLANNING_ROOT/fastwam

# 2. Convert to the pinned LeRobot fork's config schema.
#    Pulls Wan-AI/Wan2.2-TI2V-5B (~20-24 GB) on first run โ€” FastWAM deliberately does not
#    store its text encoder in the checkpoint.
python scripts/convert_fastwam_checkpoint.py \
    --checkpoint        $QPLANNING_ROOT/fastwam/libero_uncond_2cam224.pt \
    --dataset-stats     $QPLANNING_ROOT/fastwam/libero_uncond_2cam224_dataset_stats.json \
    --output-dir        $QPLANNING_ROOT/fastwam/hf_checkpoint \
    --wan22-weights-dir $QPLANNING_ROOT/wan22

# 3. Point QPLANNING_BC_CKPT_LIBERO at the converted directory.

ZibinDong/fastwam_libero_uncond_2cam224 is the same weights already converted, and additionally bundles the Wan2.2 VAE and UMT5 text encoder โ€” but it targets upstream LeRobot's nested config schema, which the pinned fork does not read. In particular its n_action_steps: 32 would silently override the fork's 10, the paper's replan interval. Converting the official .pt yields a correct config for free.

Contents

path what
checkpoints/{005000..040000}/pretrained_model/ eight training steps, 5k apart
checkpoints/040000/pretrained_model/ final โ€” use this one
checkpoints/040000/training_state/ optimizer, scheduler, RNG state; for resuming training only
resolved.yaml the fully resolved run config
# Just the final checkpoint (~5 GB instead of 44 GB)
hf download Tianjiao-Yu/qplanning-qfunction-libero \
    --include "checkpoints/040000/pretrained_model/*" --local-dir ./q_libero

Model

Vision / text DINOv2-large + T5-v1.1-base
Trunk 18 decoder layers, d_model 1024, 16 heads, FFN 4096
Value head HL-Gauss, 101 bins over [-0.01, 1.01], ฯƒ = 0.0075
Chunk / discount h = 32, ฮณ = 0.99, target ฯ„ = 0.005
Cameras observation.images.image, observation.images.image2, 224ร—224
Action / state dim 7 / 8
Replan interval n_action_steps: 10

Trained 40k steps on 8ร—H100: batch 24, AdamW lr 3e-4 (backbone 9e-5), cosine decay to 1e-6, 2k warmup, weight decay 1e-4, sparse reward, on HuggingFaceVLA/libero.

Known caveat. FastWAM's checkpoint sets toggle_action_dimensions: [-1], a gripper sign convention. A probe battery found this Q-function is indifferent to gripper polarity (grip_flip 0.51 vs chance 0.50), so verify that convention on a short run before committing to a long evaluation.

Citation

@inproceedings{giridhar2026qplanning,
  title     = {Beyond Imitation: Self-Improving Robot Policies via Off-Policy Q-Planning},
  author    = {Giridhar, Varun and Khandelwal, Anant and Collins, Jeremy A.
               and Georgiev, Ignat and Garg, Animesh},
  booktitle = {Conference on Robot Learning (CoRL)},
  year      = {2026}
}

Apache-2.0. Builds on LeRobot (Apache-2.0). FastWAM (yuantianyuan01/FastWAM, arXiv 2603.16666) and LIBERO are the work of their respective authors and are not redistributed here.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading