Evaluate Policies¶
armnet-policy-eval is the generic Armnet policy evaluation command.
It keeps one user-facing CLI while dispatching to policy-specific runtime
backends. Each backend still builds its own Docker image, so LeRobot, OpenPI,
GR00T N1.7, CRA GR00T N2, and future policy stacks can keep separate
dependencies.
LeRobot Policies¶
For LeRobot-compatible policies, use the same flags as
armnet-lerobot-eval; only the command name changes:
armnet-policy-eval \
--policy.path=nikodembartnik/act_armnet_button_test \
--armnet.embodiment=lerobot/so-101 \
--armnet.task=push_green_button \
--eval.n_episodes=1 \
--eval.batch_size=1 \
--armnet.episode_time_s=30 \
--armnet.fps=20
If --policy.type is omitted, the command defaults to the LeRobot backend.
OpenPI Policies¶
Select OpenPI with --policy.type=openpi. The generic command maps
--policy.path and --policy.config_name onto the existing OpenPI runtime
backend:
armnet-policy-eval \
--policy.type=openpi \
--policy.path=volume://openpi/checkpoints/5000 \
--policy.config_name=pi05_so101_stacking_rings \
--policy.actions_to_execute=25 \
--armnet.embodiment=lerobot/so-101 \
--armnet.task=ring_insert \
--armnet.n_episodes=3 \
--armnet.episode_time_s=30 \
--armnet.secrets="{HF_TOKEN: huggingface-token}"
OpenPI checkpoints are Orbax step folders containing params/ and assets/.
Use a volume:// path, or upload a local checkpoint with the backend-specific
volume upload flags.
GR00T Policies¶
Select NVIDIA GR00T N1.7 with --policy.type=groot:
armnet-policy-eval \
--policy.type=groot \
--policy.path=https://huggingface.co/pravsels/groot1.7_insert_candle \
--armnet.embodiment=lerobot/so-101 \
--armnet.task=insert_candle \
--armnet.n_episodes=1 \
--armnet.episode_time_s=30 \
--armnet.secrets="{HF_TOKEN: huggingface-token}"
The GR00T backend accepts root-level Hugging Face model repos and nested checkpoint paths, for example:
pravsels/groot1.7_insert_candle
pravsels/groot1.7_insert_candle/checkpoints/8000/pretrained_model
https://huggingface.co/pravsels/groot1.7_insert_candle/tree/main/checkpoints/8000/pretrained_model
CRA GR00T N2 Policies¶
--policy.type=groot is GR00T N1.7. Use --policy.type=cra for CRA GR00T N2.
CRA evals are SO-101 BusyBox only; YAM is rejected before submit.
Each image build fetches origin/main in the external/cra checkout and
exports that tree into the build context. The working tree is not checked out,
reset, or pulled. The fetched SHA is logged. The laptop that submits the job
needs Git access to NVIDIA/Isaac-GR00T-CRA (the submodule remote). Missing
source, network, or GitHub authorization fails before the job is queued.
armnet-policy-eval \
--policy.type=cra \
--policy.path=pravsels/cra_busybox_multitask_h48_cosine_20k \
--policy.num_inference_timesteps=4 \
--policy.actions_to_execute=16 \
--armnet.embodiment=lerobot/so-101 \
--armnet.task=busybox_left_toggle_down \
--armnet.n_episodes=1 \
--armnet.language_instruction='Turn the left toggle switch down to "Off" position' \
--armnet.secrets="{HF_TOKEN: huggingface-token}"
--policy.path is a Hugging Face repo id or a volume:// path. Hub downloads
reuse the cell's per-user HF_HOME cache. Volume paths never trigger a Hub
download. Pass HF_TOKEN through --armnet.secrets so Hub checkpoints and
recorded datasets can be read and pushed.
For a durable checkpoint that survives cache pruning, download once and upload to your volume:
huggingface-cli download pravsels/cra_busybox_multitask_h48_cosine_20k \
--local-dir ./cra-busybox-h48-20k
armnet volume upload ./cra-busybox-h48-20k \
cra/checkpoints/cra_busybox_multitask_h48_cosine_20k
Later runs use the volume path. Unchanged files are skipped on upload and on the cell-local mirror:
armnet-policy-eval \
--policy.type=cra \
--policy.path=volume://cra/checkpoints/cra_busybox_multitask_h48_cosine_20k \
--armnet.embodiment=lerobot/so-101 \
--armnet.task=busybox_left_toggle_down \
--armnet.n_episodes=1 \
--armnet.language_instruction='Turn the left toggle switch down to "Off" position' \
--armnet.secrets="{HF_TOKEN: huggingface-token}"
Neither Hub cache nor the cell-local volume mirror is kept forever. Operators
may prune HF_HOME and cell volume mirrors to control multi-user disk growth.
A pruned Hub cache re-downloads; a pruned volume mirror is restored from the
cloud volume, which remains the source of truth. Ancillary CRA/GR00T tokenizer
artifacts go to ctx.cache_home/groot_hf_artifacts, which is also best-effort
cache rather than durable storage.
Proven policy flags: --policy.path, --policy.revision,
--policy.embodiment_tag (default so100_3rgb),
--policy.num_inference_timesteps, --policy.actions_to_execute,
--policy.device, --policy.use_torch_compile, --policy.action_ticket_seed,
and --policy.action_ticket_path. State stays enabled; do not pass
include_state or v21_to_v3_joint_fix.
Action tickets¶
With no ticket, every action chunk builds a new generator, seeds it with the
policy default 42, and draws again. Joint denoising draws video noise from
that generator first, then action noise. The draw matches across chunks and
across runs because the seed matches. Nothing stores the tensor.
An action ticket replaces only that action noise. Video noise stays on seed
42. The same tensor is cloned into every inference and every episode in the
job. policy.reset() does not resample it. Pass one of:
--policy.action_ticket_seed=<int>samples oneN(0, I)tensor once, using the checkpointaction_cfghorizon and action dim, then reuses it.--policy.action_ticket_path=volume://...loads an explicit.npz. Loading usesallow_pickle=False. The file must be schemacra-action-ticket-v1andkind: action.
The two flags are mutually exclusive. A ticket rejects
--policy.use_torch_compile, CUDA graph, and TensorRT. The shape is whatever
the checkpoint declares. For
cra_busybox_multitask_h48_cosine_20k that is (48, 32), not the six SO-101
joints. A mismatch fails before the robot connects.
The seed is a NumPy PCG64 recipe. The portable identity is the saved float32
tensor. Job progress and the final result include an action_ticket object
with schema, kind, mode, shape, and SHA-256 of the canonical little-endian
float32 bytes, plus seed or path. Raw latent values are not logged.
Sample locally, upload once, then evaluate one ticket on one task:
cd armnet
uv run python -m armnet_client.runtime_cra_eval.cra_ticket sample \
--seed 123 --horizon 48 --action-dim 32 \
--output /tmp/cra-ticket-123.npz
armnet volume upload /tmp/cra-ticket-123.npz cra/tickets/task-ticket-001.npz
Do not run either hardware command below until a human explicitly authorizes real hardware.
cd armnet
uv run armnet-policy-eval \
--policy.type=cra \
--policy.path=volume://cra/checkpoints/cra_busybox_multitask_h48_cosine_20k \
--policy.action_ticket_seed=123 \
--policy.num_inference_timesteps=4 \
--policy.actions_to_execute=40 \
--armnet.embodiment=lerobot/so-101 \
--armnet.task=push_red_button \
--armnet.n_episodes=5 \
--armnet.language_instruction='Push the red button' \
--armnet.secrets="{HF_TOKEN: huggingface-token}"
cd armnet
uv run armnet-policy-eval \
--policy.type=cra \
--policy.path=volume://cra/checkpoints/cra_busybox_multitask_h48_cosine_20k \
--policy.action_ticket_path=volume://cra/tickets/task-ticket-001.npz \
--policy.num_inference_timesteps=4 \
--policy.actions_to_execute=40 \
--armnet.embodiment=lerobot/so-101 \
--armnet.task=push_red_button \
--armnet.n_episodes=5 \
--armnet.language_instruction='Push the red button' \
--armnet.secrets="{HF_TOKEN: huggingface-token}"
Random search, ZOS, CEM, and sequential halving are consumers of this substrate. They are not implemented here. A later AutoResearch agent owns the search and its own evidence.
The six baseline commands, one task and one seed ticket each, are review text only. Do not submit them until a human authorizes hardware:
cd armnet
uv run armnet-policy-eval \
--policy.type=cra \
--policy.path=volume://cra/checkpoints/cra_busybox_multitask_h48_cosine_20k \
--policy.action_ticket_seed=123 \
--policy.num_inference_timesteps=4 \
--policy.actions_to_execute=40 \
--armnet.embodiment=lerobot/so-101 \
--armnet.task=push_green_button \
--armnet.n_episodes=5 \
--armnet.language_instruction='Push the green button' \
--armnet.secrets="{HF_TOKEN: huggingface-token}"
cd armnet
uv run armnet-policy-eval \
--policy.type=cra \
--policy.path=volume://cra/checkpoints/cra_busybox_multitask_h48_cosine_20k \
--policy.action_ticket_seed=123 \
--policy.num_inference_timesteps=4 \
--policy.actions_to_execute=40 \
--armnet.embodiment=lerobot/so-101 \
--armnet.task=push_red_button \
--armnet.n_episodes=5 \
--armnet.language_instruction='Push the red button' \
--armnet.secrets="{HF_TOKEN: huggingface-token}"
cd armnet
uv run armnet-policy-eval \
--policy.type=cra \
--policy.path=volume://cra/checkpoints/cra_busybox_multitask_h48_cosine_20k \
--policy.action_ticket_seed=123 \
--policy.num_inference_timesteps=4 \
--policy.actions_to_execute=40 \
--armnet.embodiment=lerobot/so-101 \
--armnet.task=push_white_button \
--armnet.n_episodes=5 \
--armnet.language_instruction='Push the white button' \
--armnet.secrets="{HF_TOKEN: huggingface-token}"
cd armnet
uv run armnet-policy-eval \
--policy.type=cra \
--policy.path=volume://cra/checkpoints/cra_busybox_multitask_h48_cosine_20k \
--policy.action_ticket_seed=123 \
--policy.num_inference_timesteps=4 \
--policy.actions_to_execute=40 \
--armnet.embodiment=lerobot/so-101 \
--armnet.task=push_yellow_button \
--armnet.n_episodes=5 \
--armnet.language_instruction='Push the yellow button' \
--armnet.secrets="{HF_TOKEN: huggingface-token}"
cd armnet
uv run armnet-policy-eval \
--policy.type=cra \
--policy.path=volume://cra/checkpoints/cra_busybox_multitask_h48_cosine_20k \
--policy.action_ticket_seed=123 \
--policy.num_inference_timesteps=4 \
--policy.actions_to_execute=40 \
--armnet.embodiment=lerobot/so-101 \
--armnet.task=busybox_left_toggle_up \
--armnet.n_episodes=5 \
--armnet.language_instruction='Turn the left toggle switch up to "On" position' \
--armnet.secrets="{HF_TOKEN: huggingface-token}"
cd armnet
uv run armnet-policy-eval \
--policy.type=cra \
--policy.path=volume://cra/checkpoints/cra_busybox_multitask_h48_cosine_20k \
--policy.action_ticket_seed=123 \
--policy.num_inference_timesteps=4 \
--policy.actions_to_execute=40 \
--armnet.embodiment=lerobot/so-101 \
--armnet.task=busybox_left_toggle_down \
--armnet.n_episodes=5 \
--armnet.language_instruction='Turn the left toggle switch down to "Off" position' \
--armnet.secrets="{HF_TOKEN: huggingface-token}"
Live BusyBox cameras are squashed to 320×240 with cv2.INTER_AREA on capture,
before N2's codec roundtrip and before recording, so the full field of view is
kept.
Structured Variation¶
An evaluation against one staged pose measures the policy on that pose. Add
--armnet.variation=true and each episode runs in a slightly different scene:
the rail carrying the arm shifts, the pan/tilt cameras aim a few degrees off
their calibrated pose, and the lightbox brightens or dims.
armnet-policy-eval \
--policy.path=pravsels/smolvla_open_lamp_door \
--armnet.embodiment=lerobot/so-101 \
--armnet.task=open_lamp_door \
--armnet.n_episodes=20 \
--armnet.variation=true \
--armnet.variation-seed=42
Available on the LeRobot, OpenPI, GR00T, and CRA eval backends. Deliberately not on armnet-lerobot-record
or armnet-lerobot-teleop: moving the scene under a human demonstrator changes
what is taught, not what is measured.
Isolating one kind of generalisation¶
--armnet.variation=true enables every device family. Override them independently with:
--armnet.variation-rail— arm-base position on the rail--armnet.variation-cameras— pan and tilt of each motorised camera--armnet.variation-lighting— lightbox brightness
For example, test positional generalisation alone:
armnet-policy-eval \
--policy.path=pravsels/smolvla_open_lamp_door \
--armnet.task=open_lamp_door \
--armnet.variation-rail=true \
--armnet.variation-seed=42
An unspecified dimension inherits --armnet.variation (which defaults to
false). Each physical axis draws from a random stream keyed by
(seed, episode, axis name), rather than consuming from one shared sequence.
The rail positions for a seed are therefore bit-for-bit identical whether
lighting is enabled or disabled.
What varies, and by how much¶
Each value is drawn from a normal centred on that cell's own calibrated default, clipped at 2.5 standard deviations and then clamped to the device's physical travel. The spread is configured per device in the cell's JSON, so a cell whose rail parks near an end can use a narrower one.
| Axis | Default spread | Typical envelope |
|---|---|---|
| Rail position | 10 percent of travel | ±25 percent around the cell's park position |
| Camera pan and tilt | 5 degrees | ±12.5 degrees around the calibrated aim, inside the mount's stops |
| Lightbox brightness | 11 percent | 2.5 to 57.5 percent, around a default of 30 |
Operators override the spread in robot_cell_*.json:
{
"robot": {
"rail": { "default_position_percent": 90, "variation_std_percent": 4 }
},
"lightbox": { "default_brightness": 30, "variation_std": 11 }
}
Camera mounts come in two kinds and are commanded in different units, though
both swing the picture by the same ±12.5 degrees by default. An arducam_mount
drives hobby servos over I2C in degrees, bounded by the bracket's published
0–180 pan and 15–145 tilt range, and takes variation_std_degrees. A
servo_mount drives STS3215 serial servos in encoder steps over a 4096-step
turn, and takes variation_std_steps or variation_std_degrees. Either axis
can be tightened on its own with pan_variation_std_degrees or
tilt_variation_std_degrees (or the same names ending in _steps). To let
one direction travel further than the other, set
pan_variation_below_degrees and pan_variation_above_degrees (tilt has the
same pair). "Below" is toward smaller encoder steps, "above" toward larger
ones, and both are still clipped to the measured stops. A serial
bracket's reach depends on how
it was assembled, so each one is bounded by the stops that cell measured with
the range-test script on its Pi rather than by a shared constant:
{
"robot": {
"camera_configs": {
"front": {
"servo_mount": {
"port": "/dev/ttyAMA0",
"pan_id": 1, "tilt_id": 2,
"pan": 679, "tilt": 3033,
"pan_range": [37, 2185], "tilt_range": [1721, 3971],
"pan_variation_std_degrees": 5.0,
"tilt_variation_std_degrees": 5.0
}
}
}
}
}
A mount with no measured range falls back to the servo's own full turn, which varies more widely than anyone checked, so measure the stops before enabling camera variation on a newly built mount.
Lower the figure on a cell that parks a device near a physical stop. Cell-01 parks its rail at 70 percent and cell-03 at 90, so a wide spread on cell-03 would put much of the distribution against the end; it uses 4 instead. Zero pins an axis to its default while leaving the others varying.
Reproducibility¶
The seed fixes the sequence of offsets, not absolute device positions. Absolute positions are the cell's default plus the offset, and defaults differ between cells: cell-01 parks its rail at 70 percent, cell-03 at 90.
- Same seed, same cell: identical scenes. This is the case that matters when comparing two policies.
- Same seed, different cell: the same perturbation pattern in a different place.
Routing takes whichever cell for the task is free, so a re-run can land
elsewhere. Every episode's metadata records the scenario it ran in, including
cell_id, so a comparison can check it is like-for-like rather than assume so:
{
"seed": 42,
"episode": 3,
"cell_id": "cell-01",
"values": {
"rail": 0.715,
"brightness": 21.55
}
}
A clamped key appears only when the device's travel limits, rather than the
intended spread, decided a value — that axis varied less, and less
symmetrically, than the configuration suggests. No cell currently ships a
spread wide enough for it to fire, so seeing one means a config was widened
past what that device can reach.
A not_applied key lists axes whose device would not move. A lightbox that
does not answer costs that axis of variation and is reported in the job log;
the episode runs on with that axis at whatever the reset left it, which is what the cell already
does with the lightbox at job start. A rail that will not move is the
exception and fails the episode, because it leaves the arm somewhere other than
where the record claims.
The variation draw uses its own generator, so it neither consumes from nor
disturbs the --seed that controls the policy's own randomness.
Setting a value by hand¶
--armnet.rail and --armnet.lightbox-brightness are held to the same
envelope as a sampled value, and are checked before the job is queued. Because
routing can pick any free cell for the task, a value has to be valid on all of
them:
cell-03: --armnet.rail on rail=0.5 is outside 0.8 to 1
Recorded Eval Datasets¶
Each backend records eval episodes as a LeRobot dataset. By default, recorded datasets are pushed to Hugging Face under:
<hf_user>/eval_<task>_<timestamp>
Pass HF_TOKEN through --armnet.secrets to push private datasets or to
resolve the target Hub username. Disable uploads with
--armnet.push_to_hub=False.