Skip to content

Evaluate Policies

armnet-policy-eval is the generic Armnet policy evaluation command. It keeps one user-facing CLI while dispatching to policy-specific runtime backends. Each backend still builds its own Docker image, so LeRobot, OpenPI, GR00T N1.7, CRA GR00T N2, and future policy stacks can keep separate dependencies.

LeRobot Policies

For LeRobot-compatible policies, use the same flags as armnet-lerobot-eval; only the command name changes:

armnet-policy-eval \
  --policy.path=nikodembartnik/act_armnet_button_test \
  --armnet.embodiment=lerobot/so-101 \
  --armnet.task=push_green_button \
  --eval.n_episodes=1 \
  --eval.batch_size=1 \
  --armnet.episode_time_s=30 \
  --armnet.fps=20

If --policy.type is omitted, the command defaults to the LeRobot backend.

OpenPI Policies

Select OpenPI with --policy.type=openpi. The generic command maps --policy.path and --policy.config_name onto the existing OpenPI runtime backend:

armnet-policy-eval \
  --policy.type=openpi \
  --policy.path=volume://openpi/checkpoints/5000 \
  --policy.config_name=pi05_so101_stacking_rings \
  --policy.actions_to_execute=25 \
  --armnet.embodiment=lerobot/so-101 \
  --armnet.task=ring_insert \
  --armnet.n_episodes=3 \
  --armnet.episode_time_s=30 \
  --armnet.secrets="{HF_TOKEN: huggingface-token}"

OpenPI checkpoints are Orbax step folders containing params/ and assets/. Use a volume:// path, or upload a local checkpoint with the backend-specific volume upload flags.

GR00T Policies

Select NVIDIA GR00T N1.7 with --policy.type=groot:

armnet-policy-eval \
  --policy.type=groot \
  --policy.path=https://huggingface.co/pravsels/groot1.7_insert_candle \
  --armnet.embodiment=lerobot/so-101 \
  --armnet.task=insert_candle \
  --armnet.n_episodes=1 \
  --armnet.episode_time_s=30 \
  --armnet.secrets="{HF_TOKEN: huggingface-token}"

The GR00T backend accepts root-level Hugging Face model repos and nested checkpoint paths, for example:

pravsels/groot1.7_insert_candle
pravsels/groot1.7_insert_candle/checkpoints/8000/pretrained_model
https://huggingface.co/pravsels/groot1.7_insert_candle/tree/main/checkpoints/8000/pretrained_model

CRA GR00T N2 Policies

--policy.type=groot is GR00T N1.7. Use --policy.type=cra for CRA GR00T N2. CRA evals are SO-101 BusyBox only; YAM is rejected before submit.

Each image build fetches origin/main in the external/cra checkout and exports that tree into the build context. The working tree is not checked out, reset, or pulled. The fetched SHA is logged. The laptop that submits the job needs Git access to NVIDIA/Isaac-GR00T-CRA (the submodule remote). Missing source, network, or GitHub authorization fails before the job is queued.

armnet-policy-eval \
  --policy.type=cra \
  --policy.path=pravsels/cra_busybox_multitask_h48_cosine_20k \
  --policy.num_inference_timesteps=4 \
  --policy.actions_to_execute=16 \
  --armnet.embodiment=lerobot/so-101 \
  --armnet.task=busybox_left_toggle_down \
  --armnet.n_episodes=1 \
  --armnet.language_instruction='Turn the left toggle switch down to "Off" position' \
  --armnet.secrets="{HF_TOKEN: huggingface-token}"

--policy.path is a Hugging Face repo id or a volume:// path. Hub downloads reuse the cell's per-user HF_HOME cache. Volume paths never trigger a Hub download. Pass HF_TOKEN through --armnet.secrets so Hub checkpoints and recorded datasets can be read and pushed.

For a durable checkpoint that survives cache pruning, download once and upload to your volume:

huggingface-cli download pravsels/cra_busybox_multitask_h48_cosine_20k \
  --local-dir ./cra-busybox-h48-20k
armnet volume upload ./cra-busybox-h48-20k \
  cra/checkpoints/cra_busybox_multitask_h48_cosine_20k

Later runs use the volume path. Unchanged files are skipped on upload and on the cell-local mirror:

armnet-policy-eval \
  --policy.type=cra \
  --policy.path=volume://cra/checkpoints/cra_busybox_multitask_h48_cosine_20k \
  --armnet.embodiment=lerobot/so-101 \
  --armnet.task=busybox_left_toggle_down \
  --armnet.n_episodes=1 \
  --armnet.language_instruction='Turn the left toggle switch down to "Off" position' \
  --armnet.secrets="{HF_TOKEN: huggingface-token}"

Neither Hub cache nor the cell-local volume mirror is kept forever. Operators may prune HF_HOME and cell volume mirrors to control multi-user disk growth. A pruned Hub cache re-downloads; a pruned volume mirror is restored from the cloud volume, which remains the source of truth. Ancillary CRA/GR00T tokenizer artifacts go to ctx.cache_home/groot_hf_artifacts, which is also best-effort cache rather than durable storage.

Proven policy flags: --policy.path, --policy.revision, --policy.embodiment_tag (default so100_3rgb), --policy.num_inference_timesteps, --policy.actions_to_execute, --policy.device, --policy.use_torch_compile, --policy.action_ticket_seed, and --policy.action_ticket_path. State stays enabled; do not pass include_state or v21_to_v3_joint_fix.

Action tickets

With no ticket, every action chunk builds a new generator, seeds it with the policy default 42, and draws again. Joint denoising draws video noise from that generator first, then action noise. The draw matches across chunks and across runs because the seed matches. Nothing stores the tensor.

An action ticket replaces only that action noise. Video noise stays on seed 42. The same tensor is cloned into every inference and every episode in the job. policy.reset() does not resample it. Pass one of:

  • --policy.action_ticket_seed=<int> samples one N(0, I) tensor once, using the checkpoint action_cfg horizon and action dim, then reuses it.
  • --policy.action_ticket_path=volume://... loads an explicit .npz. Loading uses allow_pickle=False. The file must be schema cra-action-ticket-v1 and kind: action.

The two flags are mutually exclusive. A ticket rejects --policy.use_torch_compile, CUDA graph, and TensorRT. The shape is whatever the checkpoint declares. For cra_busybox_multitask_h48_cosine_20k that is (48, 32), not the six SO-101 joints. A mismatch fails before the robot connects.

The seed is a NumPy PCG64 recipe. The portable identity is the saved float32 tensor. Job progress and the final result include an action_ticket object with schema, kind, mode, shape, and SHA-256 of the canonical little-endian float32 bytes, plus seed or path. Raw latent values are not logged.

Sample locally, upload once, then evaluate one ticket on one task:

cd armnet
uv run python -m armnet_client.runtime_cra_eval.cra_ticket sample \
  --seed 123 --horizon 48 --action-dim 32 \
  --output /tmp/cra-ticket-123.npz
armnet volume upload /tmp/cra-ticket-123.npz cra/tickets/task-ticket-001.npz

Do not run either hardware command below until a human explicitly authorizes real hardware.

cd armnet
uv run armnet-policy-eval \
  --policy.type=cra \
  --policy.path=volume://cra/checkpoints/cra_busybox_multitask_h48_cosine_20k \
  --policy.action_ticket_seed=123 \
  --policy.num_inference_timesteps=4 \
  --policy.actions_to_execute=40 \
  --armnet.embodiment=lerobot/so-101 \
  --armnet.task=push_red_button \
  --armnet.n_episodes=5 \
  --armnet.language_instruction='Push the red button' \
  --armnet.secrets="{HF_TOKEN: huggingface-token}"
cd armnet
uv run armnet-policy-eval \
  --policy.type=cra \
  --policy.path=volume://cra/checkpoints/cra_busybox_multitask_h48_cosine_20k \
  --policy.action_ticket_path=volume://cra/tickets/task-ticket-001.npz \
  --policy.num_inference_timesteps=4 \
  --policy.actions_to_execute=40 \
  --armnet.embodiment=lerobot/so-101 \
  --armnet.task=push_red_button \
  --armnet.n_episodes=5 \
  --armnet.language_instruction='Push the red button' \
  --armnet.secrets="{HF_TOKEN: huggingface-token}"

Random search, ZOS, CEM, and sequential halving are consumers of this substrate. They are not implemented here. A later AutoResearch agent owns the search and its own evidence.

The six baseline commands, one task and one seed ticket each, are review text only. Do not submit them until a human authorizes hardware:

cd armnet
uv run armnet-policy-eval \
  --policy.type=cra \
  --policy.path=volume://cra/checkpoints/cra_busybox_multitask_h48_cosine_20k \
  --policy.action_ticket_seed=123 \
  --policy.num_inference_timesteps=4 \
  --policy.actions_to_execute=40 \
  --armnet.embodiment=lerobot/so-101 \
  --armnet.task=push_green_button \
  --armnet.n_episodes=5 \
  --armnet.language_instruction='Push the green button' \
  --armnet.secrets="{HF_TOKEN: huggingface-token}"
cd armnet
uv run armnet-policy-eval \
  --policy.type=cra \
  --policy.path=volume://cra/checkpoints/cra_busybox_multitask_h48_cosine_20k \
  --policy.action_ticket_seed=123 \
  --policy.num_inference_timesteps=4 \
  --policy.actions_to_execute=40 \
  --armnet.embodiment=lerobot/so-101 \
  --armnet.task=push_red_button \
  --armnet.n_episodes=5 \
  --armnet.language_instruction='Push the red button' \
  --armnet.secrets="{HF_TOKEN: huggingface-token}"
cd armnet
uv run armnet-policy-eval \
  --policy.type=cra \
  --policy.path=volume://cra/checkpoints/cra_busybox_multitask_h48_cosine_20k \
  --policy.action_ticket_seed=123 \
  --policy.num_inference_timesteps=4 \
  --policy.actions_to_execute=40 \
  --armnet.embodiment=lerobot/so-101 \
  --armnet.task=push_white_button \
  --armnet.n_episodes=5 \
  --armnet.language_instruction='Push the white button' \
  --armnet.secrets="{HF_TOKEN: huggingface-token}"
cd armnet
uv run armnet-policy-eval \
  --policy.type=cra \
  --policy.path=volume://cra/checkpoints/cra_busybox_multitask_h48_cosine_20k \
  --policy.action_ticket_seed=123 \
  --policy.num_inference_timesteps=4 \
  --policy.actions_to_execute=40 \
  --armnet.embodiment=lerobot/so-101 \
  --armnet.task=push_yellow_button \
  --armnet.n_episodes=5 \
  --armnet.language_instruction='Push the yellow button' \
  --armnet.secrets="{HF_TOKEN: huggingface-token}"
cd armnet
uv run armnet-policy-eval \
  --policy.type=cra \
  --policy.path=volume://cra/checkpoints/cra_busybox_multitask_h48_cosine_20k \
  --policy.action_ticket_seed=123 \
  --policy.num_inference_timesteps=4 \
  --policy.actions_to_execute=40 \
  --armnet.embodiment=lerobot/so-101 \
  --armnet.task=busybox_left_toggle_up \
  --armnet.n_episodes=5 \
  --armnet.language_instruction='Turn the left toggle switch up to "On" position' \
  --armnet.secrets="{HF_TOKEN: huggingface-token}"
cd armnet
uv run armnet-policy-eval \
  --policy.type=cra \
  --policy.path=volume://cra/checkpoints/cra_busybox_multitask_h48_cosine_20k \
  --policy.action_ticket_seed=123 \
  --policy.num_inference_timesteps=4 \
  --policy.actions_to_execute=40 \
  --armnet.embodiment=lerobot/so-101 \
  --armnet.task=busybox_left_toggle_down \
  --armnet.n_episodes=5 \
  --armnet.language_instruction='Turn the left toggle switch down to "Off" position' \
  --armnet.secrets="{HF_TOKEN: huggingface-token}"

Live BusyBox cameras are squashed to 320×240 with cv2.INTER_AREA on capture, before N2's codec roundtrip and before recording, so the full field of view is kept.

Structured Variation

An evaluation against one staged pose measures the policy on that pose. Add --armnet.variation=true and each episode runs in a slightly different scene: the rail carrying the arm shifts, the pan/tilt cameras aim a few degrees off their calibrated pose, and the lightbox brightens or dims.

armnet-policy-eval \
  --policy.path=pravsels/smolvla_open_lamp_door \
  --armnet.embodiment=lerobot/so-101 \
  --armnet.task=open_lamp_door \
  --armnet.n_episodes=20 \
  --armnet.variation=true \
  --armnet.variation-seed=42

Available on the LeRobot, OpenPI, GR00T, and CRA eval backends. Deliberately not on armnet-lerobot-record or armnet-lerobot-teleop: moving the scene under a human demonstrator changes what is taught, not what is measured.

Isolating one kind of generalisation

--armnet.variation=true enables every device family. Override them independently with:

  • --armnet.variation-rail — arm-base position on the rail
  • --armnet.variation-cameras — pan and tilt of each motorised camera
  • --armnet.variation-lighting — lightbox brightness

For example, test positional generalisation alone:

armnet-policy-eval \
  --policy.path=pravsels/smolvla_open_lamp_door \
  --armnet.task=open_lamp_door \
  --armnet.variation-rail=true \
  --armnet.variation-seed=42

An unspecified dimension inherits --armnet.variation (which defaults to false). Each physical axis draws from a random stream keyed by (seed, episode, axis name), rather than consuming from one shared sequence. The rail positions for a seed are therefore bit-for-bit identical whether lighting is enabled or disabled.

What varies, and by how much

Each value is drawn from a normal centred on that cell's own calibrated default, clipped at 2.5 standard deviations and then clamped to the device's physical travel. The spread is configured per device in the cell's JSON, so a cell whose rail parks near an end can use a narrower one.

Axis Default spread Typical envelope
Rail position 10 percent of travel ±25 percent around the cell's park position
Camera pan and tilt 5 degrees ±12.5 degrees around the calibrated aim, inside the mount's stops
Lightbox brightness 11 percent 2.5 to 57.5 percent, around a default of 30

Operators override the spread in robot_cell_*.json:

{
  "robot": {
    "rail": { "default_position_percent": 90, "variation_std_percent": 4 }
  },
  "lightbox": { "default_brightness": 30, "variation_std": 11 }
}

Camera mounts come in two kinds and are commanded in different units, though both swing the picture by the same ±12.5 degrees by default. An arducam_mount drives hobby servos over I2C in degrees, bounded by the bracket's published 0–180 pan and 15–145 tilt range, and takes variation_std_degrees. A servo_mount drives STS3215 serial servos in encoder steps over a 4096-step turn, and takes variation_std_steps or variation_std_degrees. Either axis can be tightened on its own with pan_variation_std_degrees or tilt_variation_std_degrees (or the same names ending in _steps). To let one direction travel further than the other, set pan_variation_below_degrees and pan_variation_above_degrees (tilt has the same pair). "Below" is toward smaller encoder steps, "above" toward larger ones, and both are still clipped to the measured stops. A serial bracket's reach depends on how it was assembled, so each one is bounded by the stops that cell measured with the range-test script on its Pi rather than by a shared constant:

{
  "robot": {
    "camera_configs": {
      "front": {
        "servo_mount": {
          "port": "/dev/ttyAMA0",
          "pan_id": 1, "tilt_id": 2,
          "pan": 679, "tilt": 3033,
          "pan_range": [37, 2185], "tilt_range": [1721, 3971],
          "pan_variation_std_degrees": 5.0,
          "tilt_variation_std_degrees": 5.0
        }
      }
    }
  }
}

A mount with no measured range falls back to the servo's own full turn, which varies more widely than anyone checked, so measure the stops before enabling camera variation on a newly built mount.

Lower the figure on a cell that parks a device near a physical stop. Cell-01 parks its rail at 70 percent and cell-03 at 90, so a wide spread on cell-03 would put much of the distribution against the end; it uses 4 instead. Zero pins an axis to its default while leaving the others varying.

Reproducibility

The seed fixes the sequence of offsets, not absolute device positions. Absolute positions are the cell's default plus the offset, and defaults differ between cells: cell-01 parks its rail at 70 percent, cell-03 at 90.

  • Same seed, same cell: identical scenes. This is the case that matters when comparing two policies.
  • Same seed, different cell: the same perturbation pattern in a different place.

Routing takes whichever cell for the task is free, so a re-run can land elsewhere. Every episode's metadata records the scenario it ran in, including cell_id, so a comparison can check it is like-for-like rather than assume so:

{
  "seed": 42,
  "episode": 3,
  "cell_id": "cell-01",
  "values": {
    "rail": 0.715,
    "brightness": 21.55
  }
}

A clamped key appears only when the device's travel limits, rather than the intended spread, decided a value — that axis varied less, and less symmetrically, than the configuration suggests. No cell currently ships a spread wide enough for it to fire, so seeing one means a config was widened past what that device can reach.

A not_applied key lists axes whose device would not move. A lightbox that does not answer costs that axis of variation and is reported in the job log; the episode runs on with that axis at whatever the reset left it, which is what the cell already does with the lightbox at job start. A rail that will not move is the exception and fails the episode, because it leaves the arm somewhere other than where the record claims.

The variation draw uses its own generator, so it neither consumes from nor disturbs the --seed that controls the policy's own randomness.

Setting a value by hand

--armnet.rail and --armnet.lightbox-brightness are held to the same envelope as a sampled value, and are checked before the job is queued. Because routing can pick any free cell for the task, a value has to be valid on all of them:

cell-03: --armnet.rail on rail=0.5 is outside 0.8 to 1

Recorded Eval Datasets

Each backend records eval episodes as a LeRobot dataset. By default, recorded datasets are pushed to Hugging Face under:

<hf_user>/eval_<task>_<timestamp>

Pass HF_TOKEN through --armnet.secrets to push private datasets or to resolve the target Hub username. Disable uploads with --armnet.push_to_hub=False.