ASL fingerspelling

Communicate in Sign Language with 26 letters, plus an experimental numbers mode :)

The alphabet

Drawn from the landmarks the model was trained on.

Live readout

NO_HANDstate
--fps
nonearmed
-last
0.00v_bar
0.000sigma
0.00gate J
0.00gate Z
--arm≥
0.000rigid
--veto>
--smooth
-track
not runself-check
-emissions

How it works

This does fingerspelling only. One letter at a time, no word signs and no grammar. MediaPipe finds 21 hand landmarks, a random forest reads the handshape, and a state machine decides when you actually signed a letter. J and Z move, so you hold the starting pose still and then make the stroke. Drop your hand for a second to end a word. If one real word clearly fits the letters I show it underneath, and if several fit or none do I leave it alone. Your video never leaves your device.

I trained this on my own hands across four sessions, plus 8,561 frames of other people's hands from three public landmark sets. If you are not me, the numbers worth trusting are the ones measured by holding other people out. Holding out one signer at a time across ten named signers, it gets 0.9406 across three seeds (0.9379–0.9453) over all 23,984 of their held-out frames. Holding out a whole second set it never saw, 1,874 frames from several people recorded with this page's own landmarker, it gets 0.8292 across three seeds (0.8212–0.8362). A third set of five signers that nothing here ever trains on reads 0.889. Run end to end rather than frame by frame, 266 clips of other people fingerspelling gave the right letter on 94, the wrong letter on none, and silence on the other 172.

My own sessions score higher, and that number is not yours. With each of my sessions held out in turn it reads 0.926 of 4,878 held-out frames (95% interval 0.877–0.965 over the 57 recorded holds), and 0.895 on the one session I recorded on a different day (0.8848–0.8949 over three seeds). The release before this one measured 0.910 and 0.870 on my sessions and had no by-signer number at all. J and Z: 0.95 over 102 events (74 J/Z gestures and 28 movements) in 80 prompted items, cross-validated by item. The labels changed in an earlier release, so that replaces rather than improves on the earlier 0.864. Accuracy moves with hand shape and lighting. Across days my own M is read on about one frame in eight, which is a problem with my 2024 recording rather than with the model, and on other people's hands U and R are read as each other, with C, S and G next least reliable.

Numbers mode is experimental. It reads the ASL digits 0–9 with a separate forest trained on 218 signers from a public dataset (Sign Language Digits Dataset, Ankara Ayranci Anadolu High School, Apache-2.0), and never on me. I have not checked it on my own hand or on live video. It scores 0.986 leave-signer-out on that dataset's photos, and 96% of my own O, V, W, F and B frames read as 0, 2, 6, 9 and 4. J and Z are off in this mode, and a relaxed hand can read as 0 or 1. On my logged idle holds a relaxed hand reads as a digit, usually 0 or 1, about one time in six. No word is suggested here.