experiment · evergreen

On the arm: two traps

The arm that would not start and the jerky motion: what the logs showed, and the one flag that fixed it.

The first time I ran the model on the arm, the arm stayed in its rest pose. The second time it waited 6.8 seconds, picked the X, and ran out of time above the cell. The third time it worked, but in jerks. From the second run on I logged every command and joint angle, and the logs showed two traps.

Trap 1: a setting I copied

SmolVLA outputs a chunk at a time: the next 50 joint targets, about 1.7 seconds of motion. I had copied an inference setting from a recipe without measuring it here. It computed each new chunk from an observation 0.6–1 second old and appended it to the old one, so every 50 steps the arm was told to jump back. In the two logged runs with that setting, 17 of the 18 command jumps above 6° sit exactly on a chunk boundary. The largest is 29°, where a normal step is about 1°.

The fix was one flag, --inference.type=sync: finish the chunk, then compute the next. I first compared three settings without the arm, on a fixed image:

Inference settingLargest jumpPause
Queued, as first used31°none
sync6.5°0.4 s every 1.7 s
RTC on13°0.17 s

On the arm with sync, the run started at second 1.3 and the largest jump was 4.6°. It has been the default since.

Per-step change of the commanded pose over time for runs 2, 3 and 4, with chunk boundaries marked, and a table of jump counts

How far the command moved at every step in runs 2 and 3 (queued) and run 4 (sync). The grey lines are chunk boundaries.

Trap 2: my own waiting

In the recordings I waited 2.4 seconds on average before I moved, and the model copied that too. At rest it is undecided between "wait" and "start". This one is in the data, so no flag fixes it. Next round I will move the moment recording starts.

#inference #action-chunk #sync

See this note on the whiteboard →