Evaluate an OpenPI Policy¶
armnet-openpi-eval is the OpenPI counterpart of
armnet-lerobot-eval: it evaluates an
OpenPI policy (pi0 / pi0.5) on a managed real-robot SO-101 cell instead of a
LeRobot policy. Each episode is recorded as a LeRobot Dataset and (optionally)
pushed to the Hugging Face Hub so you can inspect the rollout afterwards.
Prerequisites¶
- An OpenPI checkpoint — an Orbax step folder containing
params/andassets/. - That checkpoint already on your Armnet Volume (upload it out-of-band, or let the CLI upload a local copy for you — see below).
- The OpenPI training config name registered in the installed
openpi(for examplepi05_so101_stacking_rings). - A Hugging Face token stored as a Armnet Secret if you want the recorded eval dataset pushed to the Hub.
Evaluate a Checkpoint Already on Your Volume¶
Point --openpi.checkpoint_dir at the Orbax step folder on your volume with a
volume:// path:
armnet-openpi-eval \
--openpi.checkpoint_dir=volume://openpi/checkpoints/5000 \
--openpi.config_name=pi05_so101_stacking_rings \
--armnet.n_episodes=3 \
--armnet.episode_time_s=30 \
--armnet.fps=20 \
--armnet.task=ring_insert \
--armnet.secrets="{HF_TOKEN: huggingface-token}"
The runtime resolves volume://openpi/checkpoints/5000 to
ctx.volume.path("openpi/checkpoints/5000") inside the cell container.
Upload a Local Checkpoint While Evaluating¶
If the checkpoint only exists on your machine, upload it as part of the run with
--armnet.volume_checkpoint_path. The CLI uploads the local
--openpi.checkpoint_dir to that volume path first, then evaluates from there:
armnet-openpi-eval \
--openpi.checkpoint_dir=./checkpoints/pi05_rings/5000 \
--armnet.volume_checkpoint_path=openpi/checkpoints/5000 \
--openpi.config_name=pi05_so101_stacking_rings \
--armnet.n_episodes=3 \
--armnet.task=ring_insert \
--armnet.secrets="{HF_TOKEN: huggingface-token}"
Pass --armnet.volume_overwrite=True to replace files that already exist on
the volume.
Recorded Eval Datasets¶
Every episode is recorded into a LeRobotDataset. By default the dataset is
pushed to the Hub under an auto-generated repo id:
<hf_user>/eval_<task>_<timestamp>
- The
HF_TOKENyou pass through--armnet.secretsis used both to push the dataset and to resolve<hf_user>from the token. - Override the namespace with
--armnet.hf_user, or set an explicit repo id with--armnet.record_dataset_repo_id. - Disable the upload entirely with
--armnet.push_to_hub=False.
Shared leaderboard (opt in)¶
Leaderboard submission is disabled by default, so private evaluations remain private. To publish a cell-scored result, name the Hugging Face model repository the run should be attributed to:
armnet-openpi-eval \
--openpi.checkpoint_dir=volume://openpi/checkpoints/5000 \
--openpi.config_name=pi05_so101_stacking_rings \
--armnet.task=ring_insert \
--armnet.submit_to_leaderboard=True \
--armnet.leaderboard_policy_repo_id=my-user/my-policy \
--armnet.secrets="{HF_TOKEN: huggingface-token}"
The runtime records the BusyBox/cell success count, not a policy-reported
metric. The supplied HF_TOKEN must have write access to the Armnet leaderboard
dataset; ask Armnet for access before enabling the flag. By default the runtime
resolves the model revision and model type from the repository. Pin them with
--armnet.leaderboard_revision and --armnet.leaderboard_model_type when
evaluating a local or otherwise non-resolvable checkpoint.
Common Flags¶
armnet-openpi-eval \
--openpi.checkpoint_dir=volume://openpi/checkpoints/5000 \
--openpi.config_name=pi05_so101_stacking_rings \
--openpi.actions_to_execute=25 \
--openpi.device=cuda \
--armnet.embodiment=lerobot/so-101 \
--armnet.task=ring_insert \
--armnet.language_instruction="Insert the colourful ring into the central wooden peg" \
--armnet.n_episodes=10 \
--armnet.episode_time_s=30 \
--armnet.fps=20 \
--armnet.secrets="{HF_TOKEN: huggingface-token}" \
--seed=1000 \
--armnet.detach=True
| Flag | Default | Purpose |
|---|---|---|
--openpi.checkpoint_dir |
(required) | volume:// path to the Orbax step folder (or a local path when uploading). |
--openpi.config_name |
(required) | OpenPI training config registered in the installed openpi. |
--openpi.actions_to_execute |
25 |
Actions consumed from each predicted action chunk before re-querying the policy. |
--openpi.device |
cuda |
Inference device on the cell (OpenPI requires a CUDA-capable cell). |
--armnet.embodiment |
lerobot/so-101 |
Target cell type, for example lerobot/so-101 or lerobot/bimanual_yam. Must match the cell. |
--armnet.task |
(none) | Cell task slug; omit to let any cell of the embodiment pick the job up. |
--armnet.n_episodes |
1 |
Number of eval episodes. |
--armnet.episode_time_s |
30 |
Per-episode cap in seconds. |
--armnet.fps |
20 |
Policy/control rate on the cell. |
--armnet.language_instruction |
(none) | Instruction passed to the policy; overrides the task's own instruction for this job. |
--armnet.submit_to_leaderboard |
False |
Opt in to publishing the cell-scored aggregate to the shared leaderboard. |
--armnet.leaderboard_policy_repo_id |
(none) | Hugging Face model repo attributed on the leaderboard; required when submission is enabled. |
--armnet.leaderboard_revision |
(auto) | Override the model commit attributed to the result. |
--armnet.leaderboard_model_type |
(auto) | Override the inferred model family (pi0, pi0.5, etc.). |
--armnet.detach |
False |
Return immediately instead of streaming logs and blocking until done. |
See Also¶
- Train and Deploy a LeRobot Policy — the LeRobot
policy equivalent (
armnet-lerobot-eval/armnet-lerobot-train). - Teleoperate and Record Datasets — collect the demonstration datasets you train these policies on.