experiment · evergreen
Training
SmolVLA fine-tuned for 20,000 steps on a Colab A100, then three checks on the Mac before the arm.
A policy is the function that turns what a robot sees into what it does
next. SmolVLA is a small one that takes camera images and a sentence. I did
not train it from scratch. I started from lerobot/smolvla_base and
fine-tuned it on my 45 episodes with LeRobot's default setting: only the part
that produces actions is trained, the vision-language model stays frozen.
The numbers: batch size 64, 20,000 steps, about 3 hours 45 minutes on a Colab A100. The loss went from 0.479 at step 100 down to 0.077. The notebook does all of it and uploads a checkpoint every 5,000 steps, so a dropped connection does not cost the whole run.
Before the model touched the arm I checked three things on the Mac:
- It loads onto the Mac's GPU and computes 50 steps of motion in 0.35 seconds.
- On recorded frames its prediction differs from my recorded motion by about 0.4–1° on average. Those are training frames, so this only says it learned its data.
- On the same frame, swapping the sentence for another cell changes the prediction by up to 50°. The model listens to the sentence.
None of this says the arm will succeed. It rules out the cheap mistakes before the motors are on.