AI tutor for deep CS courses (UI brand "0.1% DEV — Become that developer"). Ten courses —
build a programming language, a custom library, JS engine internals, real-time systems.
The tutor, "Dev", follows one loop: observe → model → personalize. Quiz attempts and
validated chat observations update a per-learner skill graph and memory; the next answer reads
the relevant slice.
Separately, an earlier 120-case answer eval moved aggregate faithfulness from 71% to 92%. Caching cut repeat-query embedding cost by ~30%.
Key Design Decisions
Product Model
Observe → model → personalize
Evidence (quiz attempts, validated chat observations) updates a per-learner skill-graph overlay and learner memory. The next answer reads the relevant slice. Backend code updates mastery; the LLM never writes mastery directly.
skill graphlearner memorycode-owned mastery
Grounding
Course lookup always runs in code
The agent doesn't decide whether to search the courses. Every turn runs the lookup in code first: hybrid vector + BM25 retrieval, RRF fusion, then a gpt-4.1-mini reranker grades each candidate as answers / partial / related / no. The result is a coverage case the prompt is built from.
hybrid + RRFLLM rerankercoverage case
Honest Citations
NO_MATCH never invents a source
When nothing in the learner's courses covers the question, the tutor answers from general knowledge and does not attach a course citation. Retrieval is scoped to the learner's owned courses, or to the course being watched in the player.
NO_MATCH0 false citationsscoped search
In-Course Tutor
Tutor inside the video player
Answers "what is this?" questions from what's on screen at the current timestamp. Jumps the video only on an explicit "take me to…" request; otherwise suggests up to 2 moments. Key moments are listed under the video. Same pipeline as the Ask page.
IN_LECTUREfind_momentnavigate · suggest
Agent Tools
LangChain agent with 9 server-bound tools
search_course_content, get_enrolled_courses, get_learning_strengths, get_learning_metrics, get_concept_confidence, search_history, record_skill_observation, explain_concept, find_moment. All nine are server-bound.
LangChain 1.xgpt-4.1-miniSSE streaming
Lab — LangGraph
Human-in-the-loop quiz flow
Plan objectives → approve (interrupt) → generate quiz → submit (interrupt) → loop per objective → summarize, using interrupt() / Command(resume). State sits in MemorySaver, an in-memory checkpointer. Artifact generation (notes, flashcards, study plan) fans out in parallel with LangGraph Send.
LangGraph 1.2interruptMemorySaver (in-memory)Send
Skill Graph
Prerequisite-gated concept scores
Per-learner concept scores (solid ≥ 80%). No hard questions until prerequisites reach 70%. explain_concept shows where a score came from. Quick "Quiz me" asks 3 questions on what was watched, graded server-side; "Your Lab" marks topics solid / shaky / not tried by rule, with no model call.
solid ≥ 80%prereqs ≥ 70%rule-based Lab
Governance
Token-derived identity + audit log
The Firebase ID token is verified on every request and the student id comes from it, never from the request body. PII is scrubbed before routing. Every AI turn writes an audit row: model, provider, risk level and prompt fingerprint.
Firebase AuthPII scrubaudit log
Hosting & Ops
Netlify + Railway + Neon
Frontend on Netlify (the01.dev), FastAPI backend on Railway, Neon Postgres with embeddings persisted in pgvector (search runs in application memory). Tracing through LangSmith and OpenTelemetry. 650+ backend tests.
NetlifyRailwayNeonLangSmith
What Changed — October 2026 architecture pivot
Agent Runtime
Hermes + GEPA → LangChain agent + code-run lookup
The tutor replaced an earlier Hermes agent runtime and its GEPA self-evolving prompt loop with a LangChain agent, and moved course lookup out of the model's hands into code. Why: course lookup is too important to leave to the model.
retiredLangChain 1.x
Coverage Decision
Fixed cosine threshold → reranker-graded coverage
A single cosine-similarity cutoff used to decide whether the courses covered a question. It was replaced by an LLM reranker that grades each candidate and yields a named coverage case. Why: a fixed threshold misjudged coverage.
retiredLLM reranker
Model Serving
Local Ollama / vLLM stack → OpenAI only
The local-model serving stack was removed. OpenAI is now the only model provider: gpt-4.1-mini, text-embedding-3-small and whisper-1. Why: one provider is simpler to operate.