Skip to content

Evaluate an OpenPI Policy

armnet-openpi-eval is the OpenPI counterpart of armnet-lerobot-eval: it evaluates an OpenPI policy (pi0 / pi0.5) on a managed real-robot SO-101 cell instead of a LeRobot policy. Each episode is recorded as a LeRobot Dataset and (optionally) pushed to the Hugging Face Hub so you can inspect the rollout afterwards.

Prerequisites

  • An OpenPI checkpoint — an Orbax step folder containing params/ and assets/.
  • That checkpoint already on your Armnet Volume (upload it out-of-band, or let the CLI upload a local copy for you — see below).
  • The OpenPI training config name registered in the installed openpi (for example pi05_so101_stacking_rings).
  • A Hugging Face token stored as a Armnet Secret if you want the recorded eval dataset pushed to the Hub.

Evaluate a Checkpoint Already on Your Volume

Point --openpi.checkpoint_dir at the Orbax step folder on your volume with a volume:// path:

armnet-openpi-eval \
  --openpi.checkpoint_dir=volume://openpi/checkpoints/5000 \
  --openpi.config_name=pi05_so101_stacking_rings \
  --armnet.n_episodes=3 \
  --armnet.episode_time_s=30 \
  --armnet.fps=20 \
  --armnet.task=ring_insert \
  --armnet.secrets="{HF_TOKEN: huggingface-token}"

The runtime resolves volume://openpi/checkpoints/5000 to ctx.volume.path("openpi/checkpoints/5000") inside the cell container.

Upload a Local Checkpoint While Evaluating

If the checkpoint only exists on your machine, upload it as part of the run with --armnet.volume_checkpoint_path. The CLI uploads the local --openpi.checkpoint_dir to that volume path first, then evaluates from there:

armnet-openpi-eval \
  --openpi.checkpoint_dir=./checkpoints/pi05_rings/5000 \
  --armnet.volume_checkpoint_path=openpi/checkpoints/5000 \
  --openpi.config_name=pi05_so101_stacking_rings \
  --armnet.n_episodes=3 \
  --armnet.task=ring_insert \
  --armnet.secrets="{HF_TOKEN: huggingface-token}"

Pass --armnet.volume_overwrite=True to replace files that already exist on the volume.

Recorded Eval Datasets

Every episode is recorded into a LeRobotDataset. By default the dataset is pushed to the Hub under an auto-generated repo id:

<hf_user>/eval_<task>_<timestamp>
  • The HF_TOKEN you pass through --armnet.secrets is used both to push the dataset and to resolve <hf_user> from the token.
  • Override the namespace with --armnet.hf_user, or set an explicit repo id with --armnet.record_dataset_repo_id.
  • Disable the upload entirely with --armnet.push_to_hub=False.

Shared leaderboard (opt in)

Leaderboard submission is disabled by default, so private evaluations remain private. To publish a cell-scored result, name the Hugging Face model repository the run should be attributed to:

armnet-openpi-eval \
  --openpi.checkpoint_dir=volume://openpi/checkpoints/5000 \
  --openpi.config_name=pi05_so101_stacking_rings \
  --armnet.task=ring_insert \
  --armnet.submit_to_leaderboard=True \
  --armnet.leaderboard_policy_repo_id=my-user/my-policy \
  --armnet.secrets="{HF_TOKEN: huggingface-token}"

The runtime records the BusyBox/cell success count, not a policy-reported metric. The supplied HF_TOKEN must have write access to the Armnet leaderboard dataset; ask Armnet for access before enabling the flag. By default the runtime resolves the model revision and model type from the repository. Pin them with --armnet.leaderboard_revision and --armnet.leaderboard_model_type when evaluating a local or otherwise non-resolvable checkpoint.

Common Flags

armnet-openpi-eval \
  --openpi.checkpoint_dir=volume://openpi/checkpoints/5000 \
  --openpi.config_name=pi05_so101_stacking_rings \
  --openpi.actions_to_execute=25 \
  --openpi.device=cuda \
  --armnet.embodiment=lerobot/so-101 \
  --armnet.task=ring_insert \
  --armnet.language_instruction="Insert the colourful ring into the central wooden peg" \
  --armnet.n_episodes=10 \
  --armnet.episode_time_s=30 \
  --armnet.fps=20 \
  --armnet.secrets="{HF_TOKEN: huggingface-token}" \
  --seed=1000 \
  --armnet.detach=True
Flag Default Purpose
--openpi.checkpoint_dir (required) volume:// path to the Orbax step folder (or a local path when uploading).
--openpi.config_name (required) OpenPI training config registered in the installed openpi.
--openpi.actions_to_execute 25 Actions consumed from each predicted action chunk before re-querying the policy.
--openpi.device cuda Inference device on the cell (OpenPI requires a CUDA-capable cell).
--armnet.embodiment lerobot/so-101 Target cell type, for example lerobot/so-101 or lerobot/bimanual_yam. Must match the cell.
--armnet.task (none) Cell task slug; omit to let any cell of the embodiment pick the job up.
--armnet.n_episodes 1 Number of eval episodes.
--armnet.episode_time_s 30 Per-episode cap in seconds.
--armnet.fps 20 Policy/control rate on the cell.
--armnet.language_instruction (none) Instruction passed to the policy; overrides the task's own instruction for this job.
--armnet.submit_to_leaderboard False Opt in to publishing the cell-scored aggregate to the shared leaderboard.
--armnet.leaderboard_policy_repo_id (none) Hugging Face model repo attributed on the leaderboard; required when submission is enabled.
--armnet.leaderboard_revision (auto) Override the model commit attributed to the result.
--armnet.leaderboard_model_type (auto) Override the inferred model family (pi0, pi0.5, etc.).
--armnet.detach False Return immediately instead of streaming logs and blocking until done.

See Also