Hardware
To prove our thesis, we mounted a piezoresistive tactile sensor on the moving jaw of an SO-101 gripper. The sensor is a row-column matrix with up to 32 rows and 32 columns. We currently downsample to 6×6 (36 taxels), because the readout board does not have enough pins for more. An Arduino mounted on the arm scans this region with its 10-bit ADC at roughly 200 Hz and sends each scan over USB. On the host, a small driver presents the stream as a camera, so each reading enters the training pipeline as a 6×6 single-channel image alongside the RGB views.
Task
The task is to pick up a purple origami cube. We chose this object because it can fail in two opposite ways: with too little grip force it slips out, and with too much it crushes. Each trial uses one of three cubes of different sizes (4.0, 4.5, and 5.0 cm on a side). In the camera images the three sizes are hard to tell apart. On the tactile sensor, however, they produce clearly different contact patterns, and each size requires a different gripper closure. We collected 40 teleoperated demonstrations, split roughly evenly across the three sizes. The task shares the main difficulties of fruit picking: objects vary in size, deform under load, and leave a narrow margin between slipping and crushing. This was the simplest experiment to stand-up based on our available resources.
Touch as a camera
We fine-tuned SmolVLA on a 40-episode dataset (pick_static_object) that records a front camera, a wrist camera, and the tactile sensor, all at 30 fps. The tactile readings were stored as a third camera stream. As a control, we fine-tuned a second policy on the same episodes with the tactile stream removed (pick_static_object_notactile). Adding touch as a camera made the policy worse: across our runs, success was up to 30% lower than for the policy without tactile input. This agrees with T-Rex (UC Berkeley and NVIDIA), where adding tactile input directly to π0.5 lowered its average success from 17% to 6%. Our working explanation is that the policy does not simply ignore the tactile image. With only 40 demonstrations, it fits to the tactile image and generalizes worse.
This clip isn’t available here. Watch episode 2 in the LeRobot visualizer.
With tactile streamOpen in visualizer
This clip isn’t available here. Watch episode 2 in the LeRobot visualizer.
Tactile stream removedOpen in visualizer
A separate tactile model
Our approach is to give touch its own model. A vision-only policy, the fine-tuned SmolVLA, moves the arm to the cube. When the tactile sensor registers first contact, control passes to a tactile supervisor: a linear policy that takes only the tactile readings as input and closes the gripper. The two models share no inputs, so the vision policy cannot fit to the tactile stream and the tactile policy never sees camera images. Splitting the task into two phases also gives each model a narrower problem: the vision policy only has to reach, and the tactile policy only has to grasp.
Results
We evaluated both configurations over 20 episodes each. The vision-only baseline (eval log) runs in the same hand-off framework with the tactile supervisor disabled, so SmolVLA controls the gripper for the whole episode. With vision alone, the arm missed the cube in 2 episodes, reached it but failed to secure it in 11, and succeeded in 7. With the tactile supervisor (eval log), it missed in 2, failed to secure the cube in 3, and succeeded in 15. Both configurations missed the cube equally often because they share the reaching policy, so the entire difference comes from the grasp. Of the 18 episodes in which the arm reached the cube, the vision-only policy secured the grasp in 7 (39%) and the tactile supervisor in 15 (83%). This difference is statistically significant (Fisher's exact test, p ≈ 0.015).
| Outcome | Vision only | Vision + tactile supervisor |
|---|---|---|
| Missed the cube | 2 | 2 |
| Reached, failed to secure | 11 | 3 |
| Succeeded | 7 | 15 |
This clip isn’t available here. Watch episode 2 in the LeRobot visualizer.
Vision onlyOpen in visualizer
This clip isn’t available here. Watch episode 2 in the LeRobot visualizer.
Vision + tactile supervisorOpen in visualizer