Core Architecture
Transactional outbox
One database transaction locks the user row, counts attempts, and inserts both the Submission and an outbox job row. A unique (student, question, attempt_number) constraint returns 409 on a clash. PostgreSQL is the authoritative record of every job.
PostgreSQLoutbox409 on clash
Queueing
Relay publishes references, not code
The relay polls Postgres every 0.5s with FOR UPDATE SKIP LOCKED (up to 50 jobs per tick) and publishes only {job_id, kind, HMAC token} to SQS, with a DLQ. A sweep requeues expired leases; a job is dead after 3 attempts.
AWS SQSDLQSKIP LOCKEDHMAC
Execution
Lambda claims, Django grades
Lambda claims the job from Django with the token and gets code, test inputs and a lease — never the expected outputs. It runs each test in a bounded child process and posts outputs back; Django grades and writes the result. JavaScript and Python are supported.
AWS LambdaleaseCloudFormation
Code Execution Safety
Bounded child process — an honest limit
Each test runs with rlimits, a 10s timeout, a 64KB output cap, an empty environment and an unprivileged user. A subprocess inside Lambda does not recreate the per-job isolation of Stage 1's Docker sandboxes. That is an accepted, documented risk.
rlimits10s timeout64KB cap
Result Delivery
Polling with backoff
The browser polls GET /api/submissions/<id>/status/ with backoff (1s ×5, then 2s, 3s … up to 10s) and picks up where it left off after a refresh. A stage tracker shows Submit → Queue → Execute → Grade → Feedback. SSE and Redis pub/sub were removed in Stage 3.
authenticated pollingbackoff
Async Non-Grading Work
Celery + Redis, narrowed
Celery and Redis now handle only plagiarism, AI tasks and scheduled jobs (Beat). Redis is a cache and broker only. Nothing on the grading path depends on it.
CeleryCelery BeatRedis
Plagiarism Detection
Behaviour first, similarity on suspects
Runs at exam close. Phase 1 scores paste ratio (0.40), speed vs difficulty baseline (0.30), tab switches (0.15) and submission surprise (0.15), and flags risk ≥ 0.20. Suspects (risk ≥ 0.35 or 3+ signals) are compared against the full cohort for that question: O(K×N), not O(N²). AST/identifier normalisation → TF-IDF → cosine ≥ 0.80. Connected components show cheating rings.
TF-IDFscikit-learnpre_save signal
Plagiarism Governance
The system never auto-accuses
Instructors confirm or dismiss every flag. Confirmed plagiarism withholds points. Each student gets one appeal per exam. Runs are audited and the policy is versioned. Checks run on exam close, on auto-close after the last submission, from a scheduled task, or on a manual re-run. Result: 19× more cheating cases detected.
human reviewappealsaudited
AI Feature Design
Single calls, off the grading path
Exam authoring, rubric scoring, Socratic hints, pair explanations, exam summaries and weekly recommendations are each one structured GPT-4o-mini call, not an agent. They are queued after commit, idempotent, and do nothing without an API key. Tests still decide correctness. AI authoring with instructor approval cut authoring effort by ~80%.
GPT-4o-miniidempotentafter commit
AutoEval
Cypress grading for web apps
Automated Cypress evaluation of 500+ student web apps per cohort. Turnaround went from about a week to about a day, saving roughly 500 instructor hours.
Cypress500+ apps / cohort
Stage 1 → Stage 3 — what changed on the execution path
From Redis/Celery workers and Docker sandboxes to an outbox → relay → SQS → Lambda pipeline.
| Dimension |
Stage 1 — Redis + Celery + Docker |
Stage 3 — Outbox + SQS + Lambda (current) |
| Job hand-off |
Redis/Celery workers |
Outbox row written in the same transaction as the Submission; relay publishes to SQS (+ DLQ) |
| Queue message |
Job dispatched to Celery workers via Redis |
Reference only: {job_id, kind, HMAC token}; Lambda claims the code from Django |
| Execution |
Disposable Docker sandbox per job |
Bounded child process per test inside Lambda |
| Isolation |
Per-job container isolation |
Weaker — subprocess does not recreate Docker's per-job isolation (accepted, documented) |
| Grading |
— |
Django grades; Lambda never sees expected outputs |
| Result delivery |
Server-Sent Events over Redis pub/sub |
Authenticated client polling with backoff |
Honest note
PostgreSQL holds the truth at every step. The queue carries only a signed reference and the runner never receives expected outputs. The cost is isolation: a child process inside Lambda is bounded (rlimits, timeout, output cap, empty env, unprivileged user) but is not a per-job container, and that risk is written down rather than glossed over.