Surgical Intent Recognition
The distinction between reaction and anticipation is the difference between a tool and an assistant. Current surgical robots react — they execute the surgeon's explicit commands with high fidelity. That's valuable. It's not enough.
The goal we set: predict the surgeon's next action 400ms before it begins, with > 90% accuracy, using only camera and EMG inputs from the surgical field.
Dataset
We collected 800 hours of laparoscopic cholecystectomy procedures from 12 hospitals, annotated by expert surgeons with 23 action classes (grasp, cut, irrigate, retract, clip, etc.). Each action was annotated with onset timestamps at 10ms resolution.
Model
Human Signal's intent graph (see our earlier note) was fine-tuned on this dataset. Input modalities:
Results
| Prediction horizon | Top-1 accuracy | Top-3 accuracy |
|---|---|---|
| 100ms ahead | 97.2% | 99.4% |
| 200ms ahead | 94.1% | 98.7% |
| 400ms ahead | 91.3% | 97.2% |
| 600ms ahead | 84.6% | 93.1% |
The 400ms horizon is the practical sweet spot — enough lead time for the robot to pre-position, not so far ahead that prediction becomes unreliable.