# CodeMas — System Design

**Product:** CodeMas — AI coding-assessment platform (Masai School)
**Role:** Owned CodeMas end to end, worked directly with the CTO — submission architecture, AI features, plagiarism system, AutoEval
**Repo:** https://github.com/swanandapps/codemas (the public repo lags behind current work)

---

## Quick Reference

| Field | Detail |
|---|---|
| Platform | Proctored online coding assessments |
| Scale | Supports up to 10,000 simultaneously active learners |
| Result time | Typical results in under 10 seconds |
| Current pipeline (Stage 3) | Postgres transactional outbox → relay → AWS SQS → AWS Lambda; browser polls for status |
| Languages | JavaScript + Python |
| Plagiarism | Two-phase at exam close (behavioural → similarity); 19× more cheating cases detected |
| AI features | Single structured GPT-4o-mini calls, all off the grading path |
| AutoEval | Cypress grading of 500+ student web apps per cohort: ~1 week → ~1 day |
| Stack | Vue 3 + Vite + TS + Pinia · Django + DRF + Simple JWT · PostgreSQL · AWS SQS + Lambda · Celery + Redis · GPT-4o-mini · scikit-learn |

---

## Problem Statement

Instructors run coding exams for large cohorts. Every submission must run untrusted code in isolation, be graded against hidden tests, and return a result fast — while the platform detects cheating fairly and cuts the manual work of writing exams and grading projects.

Hard constraints:

1. **Correct attempt counting** — each attempt counted exactly once, even under double-clicks and retries.
2. **Never lose a submission** — a submission saved in the database must eventually be executed.
3. **Hidden tests stay hidden** — execution never sees expected outputs.
4. **Exam first** — plagiarism and AI work must never slow down or break grading.

---

## Functional Requirements

- Students submit JavaScript or Python code during an active exam and see progress: Submit → Queue → Execute → Grade → Feedback.
- Attempt limits enforced server-side.
- Plagiarism checks run automatically at exam close; instructors confirm or dismiss flags.
- Instructors author exams with AI assistance and approve before publishing.
- Rubric scoring, Socratic hints on failed submissions, plagiarism-pair explanations, exam summaries, weekly recommendations.
- AutoEval: Cypress-based grading of student web-app projects.

## Non-Functional Requirements

- **Scale:** up to 10,000 simultaneously active learners.
- **Latency:** typical results under 10 seconds.
- **Durability:** PostgreSQL is authoritative; queue messages carry only references.
- **Failure isolation:** AI tasks are idempotent, run after commit, and no-op without an API key.
- **Fairness:** the system never auto-accuses; humans decide.

---

## Architecture Evolution

| Stage | Architecture | What it fixed / what it cost |
|---|---|---|
| **1** | Django + Redis/Celery workers + disposable Docker sandboxes + Server-Sent Events | Strong per-job isolation; web and execution shared capacity |
| **2** | Web and worker capacity split onto separate EC2 tiers | Exam traffic and code execution scale independently |
| **3 (current)** | PostgreSQL transactional outbox → relay → AWS SQS → AWS Lambda; status by authenticated polling | Execution capacity elastic; SSE and Redis pub/sub removed; Postgres is the single source of truth |

---

## High-Level Architecture (Stage 3)

```
 Browser (Vue 3)
   POST /api/submissions/create/          GET /api/submissions/<id>/status/
        │                                   ▲  polls with backoff
        ▼                                   │  (1s ×5, 2s, 3s ... 10s)
 Django + DRF (Gunicorn) ───────────────────┘
   one transaction:
     lock user row → count attempts
     INSERT Submission + INSERT outbox job
        │
        ▼
 PostgreSQL (authoritative)
        │  relay polls every 0.5s
        │  FOR UPDATE SKIP LOCKED, ≤50 jobs/tick
        ▼
 Relay ── {job_id, kind, HMAC token} ──► AWS SQS (+ DLQ)
                                              │
                                              ▼
                                        AWS Lambda
   claims job from Django with token ◄────────┤
   receives code + test inputs + lease        │
   runs each test in a bounded child process  │
   posts outputs back ────────────────────────┘
        │
        ▼
 Django grades → writes results → queues AI tasks after commit

 Celery + Redis (Beat): plagiarism, AI tasks, scheduled jobs only
```

---

## Submission Pipeline (Stage 3)

### 1. Create — one database transaction
`POST /api/submissions/create/`:
- Lock the user row, count existing attempts.
- Insert the `Submission` and an outbox job row in the same transaction — the job exists if and only if the submission does.
- Unique constraint on `(student, question, attempt_number)` → a clash returns **409**.

### 2. Relay — outbox to queue
- Polls Postgres every **0.5 s** with `FOR UPDATE SKIP LOCKED`, at most **50 jobs per tick**, so multiple relays never grab the same job.
- Publishes only a small reference — `{job_id, kind, HMAC token}` — to SQS, with a dead-letter queue.
- A sweep requeues jobs whose lease expired; a job is marked dead after **3 attempts**.

### 3. Execute — Lambda
- Lambda claims the job from Django using the HMAC token.
- It receives the code, the **test inputs (never the expected outputs)** and a lease.
- Each test runs in a bounded child process: rlimits, **10 s timeout**, **64 KB output cap**, empty environment, unprivileged user.

### 4. Grade — Django
- Lambda posts raw outputs back; **Django** compares against expected outputs, grades and writes results.
- AI tasks (hints, rubric scoring) are queued after commit.

### 5. Status — polling
- The browser polls `GET /api/submissions/<id>/status/` with backoff: 1 s ×5, then 2 s, 3 s ... up to 10 s.
- Polling resumes after a page refresh. A stage tracker shows Submit → Queue → Execute → Grade → Feedback.

### Honest security note
A subprocess inside Lambda does **not** recreate Docker's per-job isolation from Stage 1. This is an accepted, documented risk, mitigated by rlimits, an empty environment, an unprivileged user and never sending expected outputs to the executor.

---

## Plagiarism Detection

Two phases, run at exam close.

**Triggers:** `Exam.is_active` True → False (pre_save signal), auto-close when the last student submits, a scheduled auto-manage task, or a manual "re-run checks".

### Phase 1 — Behavioural risk

| Signal | Weight |
|---|---|
| Paste ratio | 0.40 |
| Speed vs difficulty baseline | 0.30 |
| Tab switches | 0.15 |
| Submission surprise | 0.15 |

A student is flagged when risk **≥ 0.20**.

### Phase 2 — Code similarity
- **Suspects:** risk **≥ 0.35** OR high confidence (3+ signals).
- Each suspect is compared against the **full cohort** for that question: **O(K×N)**, not O(N²).
- Normalisation: Python AST / JavaScript identifier normalisation → TF-IDF → cosine similarity; pairs at **≥ 0.80** are flagged.

### Governance
- Plagiarism runs are audited; the policy is versioned.
- Cheating rings are detected as connected components of flagged pairs.
- Instructors confirm or dismiss — the system never auto-accuses. Confirmed plagiarism withholds points.
- One appeal per exam per student.

**Outcome:** 19× more cheating cases detected.

---

## AI Features

All are single structured GPT-4o-mini calls — **not agents** — off the grading path, idempotent, and no-ops without an API key.

| Feature | What it does |
|---|---|
| AI exam authoring | context → plan → instructor approves → generate → validate → publish (~80% less authoring effort) |
| Rubric scoring | 4 instructor-defined dimensions, 0–2 each; tests still decide correctness |
| Socratic hint | Guiding hint on a failed submission |
| Plagiarism-pair explanation | Explains why two submissions were flagged |
| Exam summary | Post-exam summary |
| Weekly recommendation | Scheduled |

---

## AutoEval

Cypress-based grading of 500+ student web apps per cohort. Turnaround dropped from ~1 week to ~1 day, saving ~500 instructor hours.

---

## Data and Infrastructure

| Component | Role |
|---|---|
| PostgreSQL | Authoritative store: users, exams, submissions, results, outbox jobs, plagiarism runs, policy versions |
| AWS SQS (+ DLQ) | Carries job references only |
| AWS Lambda | Code execution (defined in CloudFormation) |
| Celery + Redis | Plagiarism, AI tasks, scheduled jobs (Beat). Redis is cache + broker only |
| OpenAI GPT-4o-mini | AI features |
| scikit-learn | TF-IDF for similarity |

Not in the current stage: MongoDB, Nginx, Docker sandboxes, SSE, Redis pub/sub.

---

## Key Engineering Decisions

| Decision | Chosen | Alternative | Why |
|---|---|---|---|
| Enqueue | Transactional outbox | Send to queue inside the request | A crash between DB commit and send would lose or orphan jobs; the outbox commits both together |
| Queue payload | Reference + HMAC token | Full code + tests in the message | Postgres stays authoritative; expected outputs never leave Django |
| Who grades | Django | The executor | Executor sees inputs only; grading logic and answers stay server-side |
| Relay concurrency | `FOR UPDATE SKIP LOCKED` | Single relay / advisory locks | Multiple relays without double-publishing |
| Result delivery | Authenticated polling with backoff | SSE + Redis pub/sub | Survives refresh, no long-lived connections, no pub/sub layer |
| Similarity scope | Suspects vs full cohort, O(K×N) | All pairs, O(N²) | Behavioural phase narrows the search; full-cohort compare still catches the source |
| AI | Single structured calls off the grading path | Agents | Predictable, idempotent, can't block grading |
| Accusations | Human confirm / dismiss + appeal | Automatic penalties | Fairness |

---

## What Changed and Why

- **Stage 1 → 2:** web and execution were competing for the same machines; splitting tiers let each scale on its own.
- **Stage 2 → 3:** moved execution to SQS + Lambda behind a transactional outbox so execution capacity is elastic and no committed submission can be lost. SSE + Redis pub/sub were replaced by polling against Postgres, which is now the single source of truth. Trade-off: Lambda subprocess isolation is weaker than Docker per-job isolation (documented above).

---

## Numbers (verified)

- Supports up to **10,000 simultaneously active learners**.
- Typical results in **under 10 seconds**.
- **~80% less** exam-authoring effort.
- AutoEval: **~1 week → ~1 day** for 500+ web apps per cohort; **~500 instructor hours** saved.
- **19× more cheating cases detected.**
