RecallX

Vocabulary learning with source-aware ingestion and a shared review journal

RecallX brings vocabulary collected from reading into a shared, locally hosted review system. Source passages stay attached to the meaning being learned, while a central review journal keeps acknowledged schedules consistent across paired clients. Durable offline queues, recoverable document ingestion, and an experimental semantic assessment pipeline address the gaps between collecting a word, attempting recall, and recording reliable learning history.

Contents

Source context

Source-Aware Vocabulary Ingestion

Document-derived vocabulary stays linked to its original passage. Imported PDFs and images retain their bytes and hashes; page geometry and citations preserve where candidate meanings came from. Generated definitions and teaching material remain editable drafts, with human approval separating extraction from the canonical library.

Versioned meanings and rubrics preserve what a review actually tested. An answer given against an older meaning can remain in the historical record without updating the schedule for the revised content. This makes source context part of the review model.

Review history

Durable Review Synchronization

The backend owns FSRS scheduling and an append-only review journal. Each client saves an immutable review envelope before transmission, including its unique ID, sequence, and displayed content. Retrying after a lost response returns the original receipt instead of creating another review; an offline self-rating remains pending until the server acknowledges it.

Concurrent attempts are retained and replayed in a stable chronological order. Corrections append new records and trigger schedule replay, while resets advance the progress generation. Phone and browser caches provide continuity across interruptions without inventing independent due dates.

Review event semantics
ConditionJournal behavior
Lost acknowledgementRetrying the same review ID returns its original receipt.
Concurrent reviewsDistinct attempts are retained and replayed in a stable chronological order.
Revised meaningThe original attempt stays in history without updating the schedule for revised content.
CorrectionA new correction record triggers schedule replay.
Durable review synchronization A client stores an immutable review in its pending queue before sending it to the backend. The backend appends that review once to its journal and derives the FSRS schedule by chronological replay. Acknowledgements return to the client. Retries with the same review ID retain the original receipt. Client persistence Backend authority Immutable review Pending queue Review journal FSRS schedule ID and sequence Displayed content Persist before send Retry on reconnect Append once per ID Chronological replay Confirmed due date Save Sync Replay Acknowledgement and original receipt on retry A saved review stays pending until the backend acknowledges it.
Durable review synchronization The client persists an immutable review in a pending queue. The backend records the review once in an append-only journal, replays FSRS schedules, and returns an acknowledgement. An interrupted request can be retried with the same review ID and original receipt. Client persistence Immutable review Pending queue Review journal FSRS schedule ID and sequence Displayed content Persist before send Retry on reconnect Append once per ID Chronological replay Confirmed due date Backend Save Sync Replay Acknowledgement returns to the queue. Retries retain the original receipt.
An immutable review moves from durable client storage to the central journal. The backend owns the FSRS schedule and returns the same receipt when an interrupted request is retried.

Assessment

Semantic Recall Assessment

An experimental assessment pipeline compares free-form answers with approved concepts using embeddings and natural language inference. It separates supporting, neutral, and contradictory evidence, then applies language-specific calibration and abstention rules. Ambiguous responses, missing rubrics, or unavailable models produce uncertainty and leave memory state unchanged.

Constructed English, Hindi, and Hinglish answers make noise sensitivity and language coverage part of evaluation. Self-rated flashcards remain available independently of the semantic assessment path, so review can continue when an automatic correctness decision is unavailable.

Local processing

Recoverable Document Processing

Document processing is backed by a persistent job outbox, worker leases, and page checkpoints. Native PDF text is used where suitable, with layout-aware OCR for scanned or complex pages. Checkpoints let retries skip completed stages, and lease tokens prevent cancelled or replaced workers from committing stale results.

Separate inference environments isolate the API, assessment models, and OCR runtimes. A shared lease coordinates heavy model residency on one machine, giving waiting assessment work priority between background pages. This trades cold-start latency for explicit control over local resource contention.

Technology

Clients

  • Expo
  • TypeScript

API and background jobs

  • FastAPI
  • Celery

Learning and storage

  • FSRS
  • SQLite
  • PyTorch