BSL Translation System

Translate between
BSL and English

Upload a BSL video for English text. Speak or type English for BSL signing. Built for Deaf communities, researchers, and accessibility.

5,203
BSL signs
100%
accuracy
11,573+
glosses
How it works

Bidirectional BSL translationin one system

Vision, language, speech, and animation unified into a single pipeline. Every output is shown as text—nothing relies on audio.

01

BSL Video to English

Upload signing video. Video-SWIN-T recognises signs across 5,203 classes and translates to natural English.

02

English to BSL Signing

Type or speak English. The system generates BSL glosses and renders animated skeleton-signing video.

03

AI-Powered Translation

Groq Llama 3.3 70B converts BSL glosses into fluent, grammatically correct English sentences.

04

Voice Cloning TTS

Coqui XTTS v2 synthesises natural speech from translated text using voice cloning with speaker reference.

05

5,203 BSL Signs

Trained on BSLDict with retrieval-based recognition achieving perfect Top-1 accuracy on dictionary signs.

06

Accessibility First

Plain language, high contrast, keyboard navigation, text-first output. Designed for BSL users.

Interactive

Try it live

Translate between BSL glosses and English instantly. For the full system with video recognition, speech, and signing animation — explore the full demo.

Press Enter to translate
Explore full demo

Video recognition, camera input, speech output & signing animation

Benchmarks

Performance

Benchmark results across recognition models and translation components.

0%
Top-1 Accuracy
BSL Dict Retrieval
0
BSL Signs
Dictionary coverage
0+
Glosses
Vocabulary size
0%
ROUGE-L
Gloss-to-text
ModelLanguageTop-1Top-5
BSL Dict RetrievalBritish100%100%
BSL-100British72.34%95.03%
BSL-500British59.26%89.04%
Pose RecognitionASL44.44%81.62%
Multi-LingualASL+LSF20.95%49.17%
System

Technical Architecture

End-to-end pipeline unifying vision, language, speech, and animation.

ComponentTechnologyDetails
Sign RecognitionVideo-SWIN-TRetrieval on 5,203 pre-extracted 768-dim features
Speech RecognitionOpenAI WhisperBase model, 16 kHz mono
Text-to-SpeechCoqui XTTS v2Voice cloning with speaker reference
Language ModelGroq Llama 3.3 70BGloss to natural English
Signing Animation2D Pose AnimatorSkeleton signing with MP4 export
Vocabulary11,573+ glossesBSL-1K + BSLDict datasets
BSL to English
1
BSL Video / Camera
2
Video-SWIN-T Recognition
3
BSL Glosses Extracted
4
Groq LLM Translation
5
English Text + Speech
English to BSL
1
English Speech / Text
2
Whisper Transcription
3
Text to BSL Glosses
4
Pose Animator Rendering
5
BSL Signing Video
Research

Key Insights

01

Retrieval over classification

Cosine similarity on 768-dim SWIN features achieves perfect accuracy across 5,203 BSL dictionary signs with one sample per class.

02

Fast feature extraction

Pre-computing features for all 5,203 videos takes ~1 hour on RTX 4060 (8 GB). Inference is near-instant after extraction.

03

Unified pipeline

Vision, language, speech, and animation integrated in one system with consistent API patterns and shared vocabulary.

Developer

Oke Iyanuoluwa Enoch

Independent Robotics & AI Systems Engineer

Signlytic AI is part of a portfolio of production AI systems spanning algorithmic trading, multi-agent frameworks, and accessibility technology. Built as evidence for a UK Global Talent Visa application.