Echora

Literal speech recognition and evidence-aware assistive communication

Qwen3-ASR model size
1.7B
LoRA rank
8
Decoding beams
5
Composed test commands
480

Echora explores assistive communication for people whose speech is difficult for general-purpose recognizers to transcribe. The project includes literal ASR adaptation, evidence-aware message construction, listener-specific wording, and controllable speech delivery. Recognition alternatives remain separate from generated text, while explicit profiles and user actions govern personalization. The engineering focus is useful communication without treating fluent wording as proof of what was spoken.

Contents

Recognition

Dysarthric Speech Recognition

Qwen3-ASR 1.7B is adapted through selected upper acoustic layers, the multimodal projector, and decoder LoRA. Training combines TORGO recordings with controlled two- and three-word compositions built from the same speaker’s isolated words. Literal supervision retains incomplete phrases rather than completing them into plausible sentences, while normal-speech examples support retention of the foundation model’s broader recognition behavior.

Echora literal recognition architecture Recorded speech passes through Qwen3-ASR with adapted upper acoustic layers, its multimodal projector, and rank-eight decoder LoRA. Beam search retains up to five literal recognition hypotheses as separate alternatives. Recorded speech Selective adaptation Literal alternatives TORGO utterances Same-speaker compositions Qwen3-ASR 1.7B Upper acoustic layers 20–23 Multimodal projector Decoder LoRA Rank 8 Hypothesis 1 Hypothesis 2 Hypothesis 3 Hypothesis 4 Hypothesis 5 Up to five beam-search outputs
Echora literal recognition architecture Recorded speech passes through Qwen3-ASR with adapted upper acoustic layers, its multimodal projector, and rank-eight decoder LoRA. Beam search retains up to five literal recognition hypotheses as separate alternatives. Recorded speech TORGO utterances Same-speaker compositions Qwen3-ASR 1.7B Upper acoustic layers 20–23 Multimodal projector Decoder LoRA Rank 8 Selective adaptation Literal alternatives Hypothesis 1 Hypothesis 2 Hypothesis 3 Hypothesis 4 Hypothesis 5 Up to five beam-search outputs
Selective adaptation changes the recognizer while beam search preserves separate literal hypotheses. The diagram shows the recognition path, before message construction.

Evaluation

Controlled Recognition Evaluation

The command test contains 480 composed utterances from held-out speaker M04. The comparison changes both the adapter and its literal prompt, with substantial source-prompt overlap across training and test. Results characterize this controlled task rather than natural conversation. Five-beam coverage measures whether the exact reference is available among alternatives, separately from final message accuracy.

Recognition benchmark comparison
MeasureEarlier configurationCommand adaptation
Command WER72.58%51.58%
Exact reference in top five16.04%36.46%
Normal-speech WER5.23%5.23%

The earlier configuration is the v1 adapter; command adaptation is command-v3. Command WER uses single-output decoding. Top-five coverage is a separate five-beam diagnostic with normalized reference matching. Normal-speech WER covers 262 utterances. WER is word error rate; lower is better.

Message construction

Evidence-Aware Message Construction

Literal hypotheses, candidate readings, and listener-facing messages have separate representations. Deterministic vocabulary and protected-term checks constrain generated readings to recognition evidence. Profiles supply declared names, wording preferences, and audience context, with personal detail expansions anchored to recognized words. Plain wording and alternative readings remain available for correction. Recognition evidence, contextual additions, and final phrasing can therefore be inspected separately.

Evidence and message representations
RepresentationContentUse
Literal hypothesesWords returned by the recognizerSelectable recognition evidence
Candidate readingsFiltered interpretations of that evidenceAlternatives for review and correction
Listener-facing messageEditable wording for the selected audienceCommunication and controlled speech delivery

Speech and memory

Revision-Bound Speech Delivery

A shared web/native lifecycle tracks message revisions, edits, cancellation, and speech delivery. Playback authorization binds the exact wording and pronunciation to the current revision; Stop, edits, and new requests invalidate earlier work. All distinct literal alternatives remain selectable. Temporary conversational references expire after ten minutes, while persistent wording requires an explicit Remember action and stays scoped to the selected profile.

  1. Current message

    Editable wording and the selected literal alternative.

  2. Speech authorization

    Exact words, pronunciation, and delivery bound to the current revision.

  3. Playback

    Device or selected-provider speech; cancellation revokes pending playback.

Accessible interaction

Accessible Multilingual Interaction

Hindi/Hinglish processing separates display text from pronunciation preparation, protecting names and mixed-language spans during bounded transliteration. The web interface provides adjustable text and target sizes, dwell selection, scanning, contrast, and reduced motion. Delivery settings keep voice, tone, and pace distinct from editable message content, allowing wording and its spoken presentation to be controlled independently.

Technology

Model adaptation

  • Qwen3-ASR
  • PyTorch
  • LoRA

Applications

  • FastAPI
  • React
  • Expo

State

  • SQLite