# the01.dev — System Design

**Product:** the01.dev (brand "0.1% DEV — Become that developer") — an AI tutor built into deep CS video courses
**Role:** Co-founder; owned product and system architecture and implementation end to end; guided two engineers
**Stack:** React 19 + TypeScript + Vite + Tailwind + zustand · Firebase Auth · FastAPI + Pydantic 2 + asyncpg · LangChain 1.x (tutor agent) + LangGraph 1.2 (Lab) · OpenAI gpt-4.1-mini + text-embedding-3-small (+ whisper-1) · Neon Postgres (+ pgvector) · Razorpay · LangSmith + OpenTelemetry
**Hosting:** frontend on Netlify (the01.dev), backend on Railway

---

## 1. Quick Reference

| Dimension | Value |
|---|---|
| Courses | 10 deep CS courses (build a programming language, custom library, JS engine internals, real-time systems, ...) |
| Tutor persona | "Dev" — the same pipeline powers the Ask page and the in-video tutor |
| Product model | OBSERVE → MODEL → PERSONALIZE |
| Model provider | OpenAI only — gpt-4.1-mini (answers + reranker), text-embedding-3-small, whisper-1 |
| Retrieval | Hybrid vector + BM25, fused with RRF, reranked by gpt-4.1-mini |
| Tutor runtime | LangChain agent with 9 server-bound tools, streamed over SSE |
| Structured flows | LangGraph (Lab quiz flow, artifact generation) |
| Store | Neon Postgres; embeddings persisted in pgvector, search runs in application memory |
| Retrieval eval | Versioned set, ~900 player + ~110 Ask cases |
| Tests | 650+ backend tests |

---

## 2. Problem Statement

Deep CS material — compilers, engine internals, real-time systems — is hard to learn from video alone. A learner gets stuck on a concept mid-lecture and has nowhere to ask a question that is tied to the exact course and moment they are watching. Generic assistants answer in general terms, invent references, and know nothing about what this learner already understands.

The tutor therefore has to:

1. Look up the course material in code on every turn, not leave retrieval to the model's discretion.
2. Cite the exact lecture moment when it uses course content — and never invent a citation when the courses don't cover the question.
3. Know what *this* learner is weak at and adapt the answer, without letting the LLM write learner state directly.
4. Work inside the video player, answering "what does *this* mean?" from what is on screen.

---

## 3. Functional Requirements

- Browse and buy courses (Razorpay); access is checked server-side.
- Ask the tutor anything on the Ask page; answers stream with sources and deep links to `the01.dev/course/<id>?chapter=&sub=&t=`.
- In-video tutor answers "this" questions from the content at the current timestamp; jumps the video only on explicit "take me to..." requests, otherwise suggests up to 2 moments; key moments listed under the video.
- Quick "Quiz me": 3 questions on what was watched, graded server-side.
- Lab (LangGraph): plan objectives → approve → generate quiz → submit → loop per objective → summarize.
- Artifact generation: notes, flashcards, study plan.
- "Your Lab" topics marked solid / shaky / not tried (rule-based, no model call).
- Skill graph: per-learner concept scores, prerequisite gating, and an explanation of where a score came from.

---

## 4. Non-Functional Requirements

- **Identity from the token:** the student id comes from the verified Firebase ID token, never from the request body.
- **No false citations:** if retrieval finds nothing relevant, the tutor answers from general knowledge without citing a course.
- **Mastery is code-owned:** backend code updates mastery; the LLM can only propose observations through a validated tool.
- **Auditable:** an audit row is written for every AI turn (model, provider, risk level, prompt fingerprint).
- **Privacy:** PII is scrubbed before anything reaches the model.
- **Streaming UX:** tokens stream over SSE so the learner sees progress immediately.
- **Observable:** LangSmith + OpenTelemetry tracing across the chat pipeline.

---

## 5. High-Level Architecture

```
 Browser (React 19, Netlify)
   Ask page  ·  Video player tutor  ·  Lab  ·  Quiz me
        │  Firebase ID token
        ▼
 FastAPI (Railway)
   ├── Auth: verify Firebase token → student id
   ├── Chat pipeline ─────────────────────────────────────────────┐
   │     PII scrub → intent router → audit row                    │
   │     Course lookup (always, in code)                          │
   │        query understanding (if needed)                       │
   │        hybrid retrieval: vector + BM25 → RRF                 │
   │        gpt-4.1-mini reranker → coverage case                 │
   │     Context: grounding prompt + course facts                 │
   │              + weak concepts + learner memory                │
   │     LangChain agent (gpt-4.1-mini, 9 tools) ── SSE ──────────┘
   ├── Lab: LangGraph StateGraph (interrupt / Command(resume), MemorySaver)
   ├── Artifacts: LangGraph Send fan-out
   ├── Skill graph + learner memory (code-owned mastery updates)
   └── Payments: Razorpay
        │
        ▼
 Neon Postgres (+ pgvector for persisted embeddings)
 OpenAI: gpt-4.1-mini · text-embedding-3-small · whisper-1
 Tracing: LangSmith + OpenTelemetry
```

---

## 6. The Chat Pipeline

The Ask page and the in-video tutor share one pipeline, streamed over SSE.

### 6.1 Request guardrails
1. **Auth** — Firebase ID token verified; the student id is taken from the token.
2. **PII scrub** — personal data removed before any model call.
3. **Intent router** — rule-based. Quiz requests are handed off to the quiz flow; clearly out-of-scope requests get canned replies.
4. **Audit** — one row per AI turn: model, provider, risk level, prompt fingerprint.

### 6.2 Course lookup (always runs in code)
The tutor does not decide *whether* to search — code always does it:

1. **Query understanding** — only when the raw question needs rewriting.
2. **Hybrid retrieval** — vector search (text-embedding-3-small) + BM25, fused with Reciprocal Rank Fusion, scoped to the learner's owned courses (Ask) or the watched course (player).
3. **Reranker** — gpt-4.1-mini grades each candidate: `answers` / `partial` / `related` / `no`.
4. **Coverage case** — the result is classified as one of:

| Case | Meaning | Tutor behaviour |
|---|---|---|
| `ENROLLED_MATCH` | An owned course answers it | Answer from it, cite the moment |
| `ENROLLED_RELATED` | Owned course is related, not a direct answer | Answer, point to the related material |
| `NO_MATCH` | Nothing relevant | Answer from general knowledge, **no course citation** |
| `IN_COURSE` | Question about the watched course | Answer from that course |
| `IN_LECTURE` | Question about what's on screen now | Answer from the current timestamp's content |

### 6.3 Context assembly
Grounding prompt + course facts from retrieval + the learner's weak concepts (from the skill graph) + learner memory.

### 6.4 Agent
A LangChain agent on gpt-4.1-mini streams the answer. Its 9 tools are server-bound (they act only for the authenticated learner):

| Tool | Purpose |
|---|---|
| `search_course_content` | Further course search when needed |
| `get_enrolled_courses` | What the learner owns |
| `get_learning_strengths` | Strong / weak areas |
| `get_learning_metrics` | Progress metrics |
| `get_concept_confidence` | Score for a specific concept |
| `search_history` | Past conversation turns |
| `record_skill_observation` | Propose a validated observation (code decides the mastery update) |
| `explain_concept` | Show where a concept score came from |
| `find_moment` | Locate a lecture moment to cite or jump to |

### 6.5 Streaming and persistence
SSE events: `sources`, `token`, `navigate`, `suggest`, `done`. Citations deep-link to the exact lecture moment. The turn is saved to learner memory.

---

## 7. Personalization: OBSERVE → MODEL → PERSONALIZE

- **Observe** — evidence comes from quiz attempts and validated chat observations.
- **Model** — evidence updates a per-learner overlay on the course skill graph plus learner memory. Backend code computes mastery; the LLM never writes it directly.
- **Personalize** — the next answer reads the relevant slice (weak concepts, memory) into its context.

Skill-graph rules:
- A concept is **solid** at ≥ 80%.
- **Prerequisite gating:** no hard questions on a concept until its prerequisites are ≥ 70%.
- `explain_concept` shows the learner where a score came from.

---

## 8. Lab and Artifacts (LangGraph)

**Lab quiz flow** — a LangGraph `StateGraph`:

```
plan_objectives → [interrupt: learner approves] → generate_quiz
   → [interrupt: learner submits] → loop per objective → summarize
```

Suspension and resume use `interrupt()` / `Command(resume=...)` with a `MemorySaver` checkpointer. MemorySaver is in-memory, so a suspended flow does not survive a backend restart — a known limitation, not a durability claim.

**Artifact generation** — notes, flashcards and study plan are generated in parallel with LangGraph `Send` fan-out, then assembled.

**Quick "Quiz me"** — 3 questions on what was watched, graded server-side. **"Your Lab"** labels topics solid / shaky / not tried with rules, no model call.

---

## 9. Agent vs Workflow

| Use | When | Here |
|---|---|---|
| Agent (LangChain) | The model must choose its next step (which tool, whether to cite, whether to navigate) | Tutor chat |
| Workflow (LangGraph) | Every step is known in advance and a human gates progress | Lab quiz flow, artifact generation |

Retrieval sits *outside* the agent's discretion: the agent may call `search_course_content` again, but the first lookup always runs in code.

---

## 10. Data

| Data | Where |
|---|---|
| Users, purchases, course access | Neon Postgres |
| Course chunks + embeddings | Neon Postgres (pgvector column); loaded and searched in application memory |
| Skill-graph overlay, learner memory | Neon Postgres |
| Audit rows (one per AI turn) | Neon Postgres |
| Lab graph state | MemorySaver (in-memory) |
| Transcripts | Generated with whisper-1 |

Embedding cost for repeat queries is cut ~30% by caching.

---

## 11. Evaluation and Numbers

**Retrieval eval** (versioned, ~900 player + ~110 Ask cases), baseline → after Retrieval v2:

| Metric | Before | After |
|---|---|---|
| Coverage accuracy (player) | 0.81 | 0.99 |
| Coverage accuracy (Ask) | 0.75 | 0.99 |
| Recall@1 | 0.76 | 0.85 |
| MRR | 0.86 | 0.92 |
| Correct source in top-3 (Ask) | 0.64 | 0.94 |
| False citations | — | 0 |

Other verified numbers:
- Earlier 120-case answer eval: aggregate faithfulness 71% → 92%.
- ~30% lower repeat-query embedding cost via caching.
- 650+ backend tests.
- Latency: p50 ~2 s on single runs (not a load-tested figure).

---

## 12. Key Trade-offs

| Decision | Chosen | Rejected | Why |
|---|---|---|---|
| Who triggers retrieval | Code, every turn | Agent decides | Grounding should not depend on the model choosing to search |
| Retrieval | Hybrid vector + BM25 + RRF + LLM reranker | Vector only with a cosine cutoff | Coverage cases make the "no match" path explicit; measured gains in the retrieval eval (section 11) |
| No-match behaviour | Answer from general knowledge, no citation | Refuse | Learners still get help; trust is kept by never inventing a citation |
| Mastery writes | Backend code from validated evidence | LLM writes scores | Scores must be explainable and not drift with model output |
| Vector search | In application memory over pgvector-persisted embeddings | Dedicated vector DB / pgvector index in prod | No extra service to run at current corpus size |
| Model provider | OpenAI only | Self-hosted local models | One provider to operate |
| Structured flows | LangGraph with interrupts | Free-form agent | Steps are known; human gates need suspend/resume |

---

## 13. What Changed and Why (Oct 2026 pivot)

An earlier version ran on a different architecture. It was replaced:

| Earlier | Now | Reason |
|---|---|---|
| Hermes agent runtime | LangChain agent with server-bound tools | Tools run in-process for the authenticated learner; traced with LangSmith |
| Self-evolving prompt loop (GEPA) | Retired | Focus moved to retrieval quality, measured by the versioned eval |
| Local models via Ollama/vLLM | OpenAI only | One provider to operate |
| Single cosine-threshold gate (0.45) | Hybrid retrieval + reranker + coverage cases | Explicit coverage cases; Retrieval v2 lifted coverage accuracy to 0.99 |
| Generic Q&A tutor | OBSERVE → MODEL → PERSONALIZE with a skill graph | Answers adapt to what the learner actually knows |

RAGAS-based evaluation and the MCP server from the earlier version are dormant and not part of the current product.

---

## 14. Known Gaps

- MemorySaver keeps Lab state in memory — suspended flows are lost on restart.
- Vector search runs in application memory; a larger corpus would need an index (e.g. pgvector HNSW).
- Latency numbers are single-run p50, not load-tested.
