Echora
Literal speech recognition and evidence-aware assistive communication
- Qwen3-ASR model size
- 1.7B
- LoRA rank
- 8
- Decoding beams
- 5
- Composed test commands
- 480
Echora explores assistive communication for people whose speech is difficult for general-purpose recognizers to transcribe. The project includes literal ASR adaptation, evidence-aware message construction, listener-specific wording, and controllable speech delivery. Recognition alternatives remain separate from generated text, while explicit profiles and user actions govern personalization. The engineering focus is useful communication without treating fluent wording as proof of what was spoken.
Contents
Recognition
Dysarthric Speech Recognition
Qwen3-ASR 1.7B is adapted through selected upper acoustic layers, the multimodal projector, and decoder LoRA. Training combines TORGO recordings with controlled two- and three-word compositions built from the same speaker’s isolated words. Literal supervision retains incomplete phrases rather than completing them into plausible sentences, while normal-speech examples support retention of the foundation model’s broader recognition behavior.
Evaluation
Controlled Recognition Evaluation
The command test contains 480 composed utterances from held-out speaker M04. The comparison changes both the adapter and its literal prompt, with substantial source-prompt overlap across training and test. Results characterize this controlled task rather than natural conversation. Five-beam coverage measures whether the exact reference is available among alternatives, separately from final message accuracy.
| Measure | Earlier configuration | Command adaptation |
|---|---|---|
| Command WER | 72.58% | 51.58% |
| Exact reference in top five | 16.04% | 36.46% |
| Normal-speech WER | 5.23% | 5.23% |
The earlier configuration is the v1 adapter; command adaptation is command-v3. Command WER uses single-output decoding. Top-five coverage is a separate five-beam diagnostic with normalized reference matching. Normal-speech WER covers 262 utterances. WER is word error rate; lower is better.
Message construction
Evidence-Aware Message Construction
Literal hypotheses, candidate readings, and listener-facing messages have separate representations. Deterministic vocabulary and protected-term checks constrain generated readings to recognition evidence. Profiles supply declared names, wording preferences, and audience context, with personal detail expansions anchored to recognized words. Plain wording and alternative readings remain available for correction. Recognition evidence, contextual additions, and final phrasing can therefore be inspected separately.
| Representation | Content | Use |
|---|---|---|
| Literal hypotheses | Words returned by the recognizer | Selectable recognition evidence |
| Candidate readings | Filtered interpretations of that evidence | Alternatives for review and correction |
| Listener-facing message | Editable wording for the selected audience | Communication and controlled speech delivery |
Speech and memory
Revision-Bound Speech Delivery
A shared web/native lifecycle tracks message revisions, edits, cancellation, and speech delivery. Playback authorization binds the exact wording and pronunciation to the current revision; Stop, edits, and new requests invalidate earlier work. All distinct literal alternatives remain selectable. Temporary conversational references expire after ten minutes, while persistent wording requires an explicit Remember action and stays scoped to the selected profile.
- Current message
Editable wording and the selected literal alternative.
- Speech authorization
Exact words, pronunciation, and delivery bound to the current revision.
- Playback
Device or selected-provider speech; cancellation revokes pending playback.
Accessible interaction
Accessible Multilingual Interaction
Hindi/Hinglish processing separates display text from pronunciation preparation, protecting names and mixed-language spans during bounded transliteration. The web interface provides adjustable text and target sizes, dwell selection, scanning, contrast, and reduced motion. Delivery settings keep voice, tone, and pace distinct from editable message content, allowing wording and its spoken presentation to be controlled independently.
Technology
Model adaptation
- Qwen3-ASR
- PyTorch
- LoRA
Applications
- FastAPI
- React
- Expo
State
- SQLite