note · growing

Giving it a brain

The model knows nine sentences, so Claude gets one tool with one argument, reads the board from the camera, and never drives the arm.

Three top-camera frames of the first game, one per arm move

The first game with Claude, near the end of each of the arm's three moves: middle center, bottom right, middle left.

The model that moves the arm is not smart. It imitates, and in training it saw exactly nine sentences:

put the red X in the {top|middle|bottom} {left|center|right} cell

"Put it in the corner" means nothing to it. So the thinking goes to Claude, and Claude never drives the arm. It gets one tool with one argument, place_x(cell). The code refuses a taken or unknown cell and writes the sentence, letter for letter as in the training data.

top camera
Claude
place_x("top left")
code
"put the red X in the top left cell"
top + wrist camera images, joint angles
SmolVLA policy
arm

Seeing the board. The frame from the top camera is rotated, cropped to the grid and shown to Claude, which reports the nine cells. On eight frames from real runs, each read twice, all 16 reads were correct.

No typing. The code checks the pick slot for a red X, notices when the arm is back at rest, and compares frames locally until a new piece shows up. Only then does Claude read the board again.

Dashboard. A local page with both cameras, the board and a stop button. The games so far were played from the terminal, so I have only seen it replay a recorded game.

The dashboard with the camera views, the drawn board and the status

The dashboard replaying a recorded game with scripted text. This is not a live game.

The split is not my idea: an earlier open SO-101 tic-tac-toe project was built the same way, and I used it as a reference. The code is in xox/.

#claude #board-reading #dashboard

See this note on the whiteboard →