Skip to content
← Back to selected work
Project Case Study

Code Grader

LLM-backed grading service focused on correctness, complexity, and style feedback.

Problem

Learners need feedback that is more explanatory than a pass/fail check, while submitted code still has to move through a backend grading flow in a controlled way.

Constraints

  • Keep grading orchestration in the backend instead of pushing evaluation logic into the interface.
  • Return feedback in categories that are understandable to a learner: correctness, complexity, and style.
  • Treat LLM output as assistance that needs structure, not as an unquestioned source of truth.

Architecture

A FastAPI backend receives submissions and coordinates LangChain/CodeLlama evaluation before returning structured feedback to the product surface.

01

Submission UI

02

FastAPI grading API

03

Submission normalization

04

LangChain orchestration

05

CodeLlama feedback pass

06

Structured feedback response

The UI submits code and problem context to the grading API

The backend normalizes inputs before model orchestration

LangChain coordinates the prompt and model interaction

Model output is shaped into correctness, complexity, and style feedback

The interface receives feedback as a product response, not raw model text

Tradeoffs

  • LLM-backed review can explain tradeoffs better than static checks alone, but it needs guardrails around vague or inconsistent output.
  • A backend grading boundary keeps the UI simpler, but it makes API contracts and error handling more important.
  • Category-based feedback is easier to consume, but it can hide uncertainty if the response format is too rigid.

Failure modes

  • Model output can be incomplete, contradictory, or too generic for the submitted code.
  • Large or malformed submissions can exceed practical prompt or processing limits.
  • A grading request can fail after submission, so the caller needs a clear retry or error state.

What I would improve next

  • Add deterministic pre-checks for syntax and obvious runtime constraints before invoking model feedback.
  • Store anonymized grading examples to compare feedback quality over time.
  • Make uncertainty visible when feedback cannot be confidently categorized.