Upload a BSL video for English text. Speak or type English for BSL signing. Built for Deaf communities, researchers, and accessibility.
Vision, language, speech, and animation unified into a single pipeline. Every output is shown as text—nothing relies on audio.
Upload signing video. Video-SWIN-T recognises signs across 5,203 classes and translates to natural English.
Type or speak English. The system generates BSL glosses and renders animated skeleton-signing video.
Groq Llama 3.3 70B converts BSL glosses into fluent, grammatically correct English sentences.
Coqui XTTS v2 synthesises natural speech from translated text using voice cloning with speaker reference.
Trained on BSLDict with retrieval-based recognition achieving perfect Top-1 accuracy on dictionary signs.
Plain language, high contrast, keyboard navigation, text-first output. Designed for BSL users.
Translate between BSL glosses and English instantly. For the full system with video recognition, speech, and signing animation — explore the full demo.
Video recognition, camera input, speech output & signing animation
Benchmark results across recognition models and translation components.
| Model | Language | Top-1 | Top-5 |
|---|---|---|---|
| BSL Dict Retrieval | British | 100% | 100% |
| BSL-100 | British | 72.34% | 95.03% |
| BSL-500 | British | 59.26% | 89.04% |
| Pose Recognition | ASL | 44.44% | 81.62% |
| Multi-Lingual | ASL+LSF | 20.95% | 49.17% |
End-to-end pipeline unifying vision, language, speech, and animation.
| Component | Technology | Details |
|---|---|---|
| Sign Recognition | Video-SWIN-T | Retrieval on 5,203 pre-extracted 768-dim features |
| Speech Recognition | OpenAI Whisper | Base model, 16 kHz mono |
| Text-to-Speech | Coqui XTTS v2 | Voice cloning with speaker reference |
| Language Model | Groq Llama 3.3 70B | Gloss to natural English |
| Signing Animation | 2D Pose Animator | Skeleton signing with MP4 export |
| Vocabulary | 11,573+ glosses | BSL-1K + BSLDict datasets |
Cosine similarity on 768-dim SWIN features achieves perfect accuracy across 5,203 BSL dictionary signs with one sample per class.
Pre-computing features for all 5,203 videos takes ~1 hour on RTX 4060 (8 GB). Inference is near-instant after extraction.
Vision, language, speech, and animation integrated in one system with consistent API patterns and shared vocabulary.
System-wide BSL signing overlay for any app — VLC, Teams desktop, Zoom and more. Built on Electron.