RecallX
Vocabulary learning with source-aware ingestion and a shared review journal
RecallX brings vocabulary collected from reading into a shared, locally hosted review system. Source passages stay attached to the meaning being learned, while a central review journal keeps acknowledged schedules consistent across paired clients. Durable offline queues, recoverable document ingestion, and an experimental semantic assessment pipeline address the gaps between collecting a word, attempting recall, and recording reliable learning history.
Contents
Source context
Source-Aware Vocabulary Ingestion
Document-derived vocabulary stays linked to its original passage. Imported PDFs and images retain their bytes and hashes; page geometry and citations preserve where candidate meanings came from. Generated definitions and teaching material remain editable drafts, with human approval separating extraction from the canonical library.
Versioned meanings and rubrics preserve what a review actually tested. An answer given against an older meaning can remain in the historical record without updating the schedule for the revised content. This makes source context part of the review model.
Review history
Durable Review Synchronization
The backend owns FSRS scheduling and an append-only review journal. Each client saves an immutable review envelope before transmission, including its unique ID, sequence, and displayed content. Retrying after a lost response returns the original receipt instead of creating another review; an offline self-rating remains pending until the server acknowledges it.
Concurrent attempts are retained and replayed in a stable chronological order. Corrections append new records and trigger schedule replay, while resets advance the progress generation. Phone and browser caches provide continuity across interruptions without inventing independent due dates.
| Condition | Journal behavior |
|---|---|
| Lost acknowledgement | Retrying the same review ID returns its original receipt. |
| Concurrent reviews | Distinct attempts are retained and replayed in a stable chronological order. |
| Revised meaning | The original attempt stays in history without updating the schedule for revised content. |
| Correction | A new correction record triggers schedule replay. |
Assessment
Semantic Recall Assessment
An experimental assessment pipeline compares free-form answers with approved concepts using embeddings and natural language inference. It separates supporting, neutral, and contradictory evidence, then applies language-specific calibration and abstention rules. Ambiguous responses, missing rubrics, or unavailable models produce uncertainty and leave memory state unchanged.
Constructed English, Hindi, and Hinglish answers make noise sensitivity and language coverage part of evaluation. Self-rated flashcards remain available independently of the semantic assessment path, so review can continue when an automatic correctness decision is unavailable.
Local processing
Recoverable Document Processing
Document processing is backed by a persistent job outbox, worker leases, and page checkpoints. Native PDF text is used where suitable, with layout-aware OCR for scanned or complex pages. Checkpoints let retries skip completed stages, and lease tokens prevent cancelled or replaced workers from committing stale results.
Separate inference environments isolate the API, assessment models, and OCR runtimes. A shared lease coordinates heavy model residency on one machine, giving waiting assessment work priority between background pages. This trades cold-start latency for explicit control over local resource contention.
Technology
Clients
- Expo
- TypeScript
API and background jobs
- FastAPI
- Celery
Learning and storage
- FSRS
- SQLite
- PyTorch