Dead Reckoning home
All postsDatasets on Hugging Face

Proving Tactile is Necessary for Dexterous Manipulation

Vision is sufficient for reaching an object but not for securing the grasp. A fine-tuned SmolVLA policy completes the grasp in 39% of trials. When the policy instead hands off to a tactile supervisor, a linear policy network, once the object is within reach, success rises to 83%.

Hardware

To prove our thesis, we mounted a piezoresistive tactile sensor on the moving jaw of an SO-101 gripper. The sensor is a row-column matrix with up to 32 rows and 32 columns. We currently downsample to 6×6 (36 taxels), because the readout board does not have enough pins for more. An Arduino mounted on the arm scans this region with its 10-bit ADC at roughly 200 Hz and sends each scan over USB. On the host, a small driver presents the stream as a camera, so each reading enters the training pipeline as a 6×6 single-channel image alongside the RGB views.

The SO-101 arm and gripper, with the tactile sensor on the moving jaw.
Figure 1. The SO-101 arm and gripper. The piezoresistive tactile sensor is mounted on the moving jaw.

Task

The task is to pick up a purple origami cube. We chose this object because it can fail in two opposite ways: with too little grip force it slips out, and with too much it crushes. Each trial uses one of three cubes of different sizes (4.0, 4.5, and 5.0 cm on a side). In the camera images the three sizes are hard to tell apart. On the tactile sensor, however, they produce clearly different contact patterns, and each size requires a different gripper closure. We collected 40 teleoperated demonstrations, split roughly evenly across the three sizes. The task shares the main difficulties of fruit picking: objects vary in size, deform under load, and leave a narrow margin between slipping and crushing. This was the simplest experiment to stand-up based on our available resources.

The three purple origami cubes used in the task.
Figure 2. The three origami cubes used in the task, 4.0, 4.5, and 5.0 cm on a side.

Touch as a camera

We fine-tuned SmolVLA on a 40-episode dataset (pick_static_object) that records a front camera, a wrist camera, and the tactile sensor, all at 30 fps. The tactile readings were stored as a third camera stream. As a control, we fine-tuned a second policy on the same episodes with the tactile stream removed (pick_static_object_notactile). Adding touch as a camera made the policy worse: across our runs, success was up to 30% lower than for the policy without tactile input. This agrees with T-Rex (UC Berkeley and NVIDIA), where adding tactile input directly to π0.5 lowered its average success from 17% to 6%. Our working explanation is that the policy does not simply ignore the tactile image. With only 40 demonstrations, it fits to the tactile image and generalizes worse.

This clip isn’t available here. Watch episode 2 in the LeRobot visualizer.

With tactile streamOpen in visualizer

This clip isn’t available here. Watch episode 2 in the LeRobot visualizer.

Tactile stream removedOpen in visualizer

Video 1. Episode 2 of the training data from the front camera, with and without the tactile stream. The inset shows the tactile stream as the policy receives it.

A separate tactile model

Our approach is to give touch its own model. A vision-only policy, the fine-tuned SmolVLA, moves the arm to the cube. When the tactile sensor registers first contact, control passes to a tactile supervisor: a linear policy that takes only the tactile readings as input and closes the gripper. The two models share no inputs, so the vision policy cannot fit to the tactile stream and the tactile policy never sees camera images. Splitting the task into two phases also gives each model a narrower problem: the vision policy only has to reach, and the tactile policy only has to grasp.

Results

We evaluated both configurations over 20 episodes each. The vision-only baseline (eval log) runs in the same hand-off framework with the tactile supervisor disabled, so SmolVLA controls the gripper for the whole episode. With vision alone, the arm missed the cube in 2 episodes, reached it but failed to secure it in 11, and succeeded in 7. With the tactile supervisor (eval log), it missed in 2, failed to secure the cube in 3, and succeeded in 15. Both configurations missed the cube equally often because they share the reaching policy, so the entire difference comes from the grasp. Of the 18 episodes in which the arm reached the cube, the vision-only policy secured the grasp in 7 (39%) and the tactile supervisor in 15 (83%). This difference is statistically significant (Fisher's exact test, p ≈ 0.015).

OutcomeVision onlyVision + tactile supervisor
Missed the cube22
Reached, failed to secure113
Succeeded715
Table 1. Outcomes over 20 evaluation episodes per configuration. Each square is one episode, grouped by outcome.

This clip isn’t available here. Watch episode 2 in the LeRobot visualizer.

Vision onlyOpen in visualizer

This clip isn’t available here. Watch episode 2 in the LeRobot visualizer.

Vision + tactile supervisorOpen in visualizer

Video 2. Evaluation episode 2 for each configuration, played side by side. The inset shows the tactile stream the supervisor reads.