← Portfolio CodeMas — Architecture
Architecture Evolution — three stages
Stage Architecture Result delivery
Stage 1 Django + Redis/Celery workers + disposable Docker sandboxes Server-Sent Events
Stage 2 Web and worker capacity split onto separate EC2 tiers Server-Sent Events
Stage 3 current PostgreSQL transactional outbox → relay → AWS SQS (+ DLQ) → AWS Lambda Authenticated client polling with backoff — SSE and Redis pub/sub removed
ACTORS API LAYER ASYNC EXECUTION 👨‍💻 Student Browser 👨‍🏫 Instructor Dashboard Stage tracker in the browser Submit → Queue → Execute → Grade → Feedback Django REST API DRF · Simple JWT · Gunicorn POST /api/submissions/create/ → one txn: lock user row · count attempts · INSERT Submission + outbox job unique (student, question, attempt_number) → 409 on clash · grades Lambda outputs · queues AI tasks after commit PostgreSQL Submission + outbox job authoritative state Outbox Relay polls every 0.5s FOR UPDATE SKIP LOCKED · ≤50/tick AWS SQS + DLQ small reference only {job_id, kind, HMAC token} AWS Lambda Runner claims job · holds a lease JavaScript + Python POST create GET status · backoff 1s→10s exams · flags one transaction claim publish invoke claim w/ token · post outputs code + test inputs + lease (never expected outputs) each test in a bounded child process: rlimits · 10s timeout · 64KB output cap empty env · unprivileged user sweep requeues expired leases job marked dead after 3 attempts
TRIGGERS AI FEATURES SERVICES Instructor: author exam context Submission graded tests already decided Failed submission tests didn't pass Pair flagged cosine ≥ 0.80 Exam closed is_active → False Weekly schedule Celery Beat Exam Authoring plan → approve → generate instructor approves first Rubric Scoring 4 dimensions · 0–2 each instructor-defined Socratic Hint a nudge, not the answer on failed submission Pair Explanation why two submissions match instructor decides Exam Summary cohort-level narrative after close Weekly Recommendation per-learner next steps scheduled Celery + Redis tasks queued after commit idempotent · no-op without API key GPT-4o-mini one structured call per task not an agent loop PostgreSQL drafts · scores · hints explanations · summaries write
Key Design Decisions
Core Architecture
Transactional outbox
One database transaction locks the user row, counts attempts, and inserts both the Submission and an outbox job row. A unique (student, question, attempt_number) constraint returns 409 on a clash. PostgreSQL is the authoritative record of every job.
PostgreSQLoutbox409 on clash
Queueing
Relay publishes references, not code
The relay polls Postgres every 0.5s with FOR UPDATE SKIP LOCKED (up to 50 jobs per tick) and publishes only {job_id, kind, HMAC token} to SQS, with a DLQ. A sweep requeues expired leases; a job is dead after 3 attempts.
AWS SQSDLQSKIP LOCKEDHMAC
Execution
Lambda claims, Django grades
Lambda claims the job from Django with the token and gets code, test inputs and a lease — never the expected outputs. It runs each test in a bounded child process and posts outputs back; Django grades and writes the result. JavaScript and Python are supported.
AWS LambdaleaseCloudFormation
Code Execution Safety
Bounded child process — an honest limit
Each test runs with rlimits, a 10s timeout, a 64KB output cap, an empty environment and an unprivileged user. A subprocess inside Lambda does not recreate the per-job isolation of Stage 1's Docker sandboxes. That is an accepted, documented risk.
rlimits10s timeout64KB cap
Result Delivery
Polling with backoff
The browser polls GET /api/submissions/<id>/status/ with backoff (1s ×5, then 2s, 3s … up to 10s) and picks up where it left off after a refresh. A stage tracker shows Submit → Queue → Execute → Grade → Feedback. SSE and Redis pub/sub were removed in Stage 3.
authenticated pollingbackoff
Async Non-Grading Work
Celery + Redis, narrowed
Celery and Redis now handle only plagiarism, AI tasks and scheduled jobs (Beat). Redis is a cache and broker only. Nothing on the grading path depends on it.
CeleryCelery BeatRedis
Plagiarism Detection
Behaviour first, similarity on suspects
Runs at exam close. Phase 1 scores paste ratio (0.40), speed vs difficulty baseline (0.30), tab switches (0.15) and submission surprise (0.15), and flags risk ≥ 0.20. Suspects (risk ≥ 0.35 or 3+ signals) are compared against the full cohort for that question: O(K×N), not O(N²). AST/identifier normalisation → TF-IDF → cosine ≥ 0.80. Connected components show cheating rings.
TF-IDFscikit-learnpre_save signal
Plagiarism Governance
The system never auto-accuses
Instructors confirm or dismiss every flag. Confirmed plagiarism withholds points. Each student gets one appeal per exam. Runs are audited and the policy is versioned. Checks run on exam close, on auto-close after the last submission, from a scheduled task, or on a manual re-run. Result: 19× more cheating cases detected.
human reviewappealsaudited
AI Feature Design
Single calls, off the grading path
Exam authoring, rubric scoring, Socratic hints, pair explanations, exam summaries and weekly recommendations are each one structured GPT-4o-mini call, not an agent. They are queued after commit, idempotent, and do nothing without an API key. Tests still decide correctness. AI authoring with instructor approval cut authoring effort by ~80%.
GPT-4o-miniidempotentafter commit
AutoEval
Cypress grading for web apps
Automated Cypress evaluation of 500+ student web apps per cohort. Turnaround went from about a week to about a day, saving roughly 500 instructor hours.
Cypress500+ apps / cohort
Architecture Trade-Offs — Quick Revision
Stage 1 → Stage 3 — what changed on the execution path
From Redis/Celery workers and Docker sandboxes to an outbox → relay → SQS → Lambda pipeline.
Dimension Stage 1 — Redis + Celery + Docker Stage 3 — Outbox + SQS + Lambda (current)
Job hand-off Redis/Celery workers Outbox row written in the same transaction as the Submission; relay publishes to SQS (+ DLQ)
Queue message Job dispatched to Celery workers via Redis Reference only: {job_id, kind, HMAC token}; Lambda claims the code from Django
Execution Disposable Docker sandbox per job Bounded child process per test inside Lambda
Isolation Per-job container isolation Weaker — subprocess does not recreate Docker's per-job isolation (accepted, documented)
Grading — Django grades; Lambda never sees expected outputs
Result delivery Server-Sent Events over Redis pub/sub Authenticated client polling with backoff
Honest note
PostgreSQL holds the truth at every step. The queue carries only a signed reference and the runner never receives expected outputs. The cost is isolation: a child process inside Lambda is bounded (rlimits, timeout, output cap, empty env, unprivileged user) but is not a per-job container, and that risk is written down rather than glossed over.
Result Delivery — SSE (Stage 1–2) vs Client Polling (Stage 3)
Why the last mile moved from push to polling.
Aspect SSE — earlier stages Client polling — current
Moving parts Redis pub/sub + a held-open stream per student One authenticated status endpoint
Server state Django holds a connection open until the result arrives Each request is independent
Page refresh Stream must be re-opened Polling resumes from the status endpoint
Cadence Server push Backoff: 1s ×5, then 2s, 3s … 10s
Progress UI Separate events per state Status response drives the stage tracker (Submit → Queue → Execute → Grade → Feedback)