Comparison
AI grading tools and autograders for programming assignments
Nine of these tools grade code by running tests on it, and one grades code written by hand on a paper exam. Four say they use AI in grading, and that is the part to try on your own assignments first.
The Graidable team
Published
Graidable makes one of the tools compared here. Every claim about another tool links to that vendor’s own page.
| Tool | Grading | AI feedback | Built for | Type | Free option |
|---|---|---|---|---|---|
| Autolab | Your own autograder | Not stated | Secondary and university | Open source | Yes |
| CodeGrade | Output, unit and code quality tests | Student assistant (add-on) | Computer science, data science, business school and AI courses | Hosted | Up to 50 students per course |
| CodeHS | Built-in test cases | Suggested grades (paid plans) | K-12 | Hosted | Free plan |
| Codio | Output and unit tests | Rubric grading by LLM | Colleges, bootcamps, training | Hosted | Free instructor account |
| Gradescope | Your own autograder, or by hand | Not stated | 2,600+ universities | Hosted | One-term trial |
| Graidable | Handwritten code on paper exams | Yes, from your key | Paper tests and exams | Hosted | Demo and sample class, no card |
| nbgrader | Assert tests in notebook cells | Not stated | Instructors who teach with Jupyter notebooks | Open source | Yes |
| Otter-Grader | Test files, in the style of pytest | Not stated | Python and R classes “at any scale” | Open source | Yes |
| Vocareum | Your own scripts | “AI-assisted grading” | Higher ed, bootcamps | Hosted | Not stated |
| zyLabs (zyBooks) | Output and unit tests | Hints; no AI grading | Academic institutions | Hosted | Evaluation copy |
Autolab
- Grading
- Your own autograder
- Type
- Open source
- Free option
- Yes
CodeGrade
- Grading
- Output, unit and code quality tests
- Type
- Hosted
- Free option
- Up to 50 students per course
CodeHS
- Grading
- Built-in test cases
- Type
- Hosted
- Free option
- Free plan
Codio
- Grading
- Output and unit tests
- Type
- Hosted
- Free option
- Free instructor account
Gradescope
- Grading
- Your own autograder, or by hand
- Type
- Hosted
- Free option
- One-term trial
Graidable
- Grading
- Handwritten code on paper exams
- Type
- Hosted
- Free option
- Demo and sample class, no card
nbgrader
- Grading
- Assert tests in notebook cells
- Type
- Open source
- Free option
- Yes
Otter-Grader
- Grading
- Test files, in the style of pytest
- Type
- Open source
- Free option
- Yes
Vocareum
- Grading
- Your own scripts
- Type
- Hosted
- Free option
- Not stated
zyLabs (zyBooks)
- Grading
- Output and unit tests
- Type
- Hosted
- Free option
- Evaluation copy
In this article
- The ten tools, in alphabetical order
- Pick by what you teach
- What is the difference between test-based autograding and AI feedback on code?
- Which tools use AI to grade coding assignments?
- Which is better for a large university course, and which for a high school class?
- How do these tools handle AI-written code?
- What happened to GitHub Classroom?
- Questions people ask
- How to test an autograder on your own assignment
This page compares ten tools that grade programming assignments: Autolab1, CodeGrade2, CodeHS3, Codio4, Gradescope5, Graidable6, nbgrader7, Otter-Grader8, Vocareum9 and zyLabs10. Nine run tests or grading scripts on code the student submits. The tenth, Graidable, is for the paper exam: it grades code that students write by hand. Four of them (CodeHS, Codio, Graidable and Vocareum) say they use AI in grading. Pick by who you teach, by whether you want to run the software yourself, and by whether tests cover everything you grade.
The ten tools, in alphabetical order
Each profile says what the vendor’s or project’s own pages stated when we checked. Six are hosted products and three are open source. Plans and prices change, so read the vendor’s page before you decide.
Autolab
An open-source autograding service from Carnegie Mellon.
- How it grades
- An autograder program you write runs on each submission in a container or virtual machine and prints the scores
- AI feedback on code
- Not stated
- Languages
- Any language
- Similarity check
- Compares submissions with each other and with past ones, using Stanford’s Moss
- LMS
- LTI course linking, with Canvas as the example in the docs; roster synchronization is what is supported today
- Who it’s for
- Programming and computer science classes at the secondary and university levels
- Hosted or open source
- Open source (Apache License 2.0); you install it on your own server
- Free option
- Yes: the software is open source
- Price
- No license fee
- Developed at Carnegie Mellon University. It has two parts: a web front end, and Tango, the server that runs grading jobs in Docker containers or AWS virtual machines.
- A scoreboard ranks students’ latest autograded scores under nicknames. You can also annotate code and grade by hand.
- The docs link to a demo site you can log in to before you install anything.
CodeGrade
Autograder, editor and AI assistant inside your LMS.
- How it grades
- Input/output, unit, code quality and structure tests built from drag-and-drop blocks, plus custom scripts
- AI feedback on code
- An AI assistant that answers students’ questions about their code under rules you set per assignment; a paid add-on
- Languages
- 175+ languages; anything you can install on Linux
- Similarity check
- Yes, on every plan, with JPlag: 9 languages, within a class, across semesters and against files you upload
- LMS
- Canvas, Blackboard, Moodle and Brightspace over LTI, on the Advanced plan
- Who it’s for
- Computer science, data science, business school and AI courses
- Hosted or open source
- Hosted; runs in the browser
- Free option
- Free plan for up to 50 students per course, with no card and no expiration
- Price
- Starter $24 and Advanced $39 per student per course; AI assistant add-on $15; institutional licensing on request
- You build tests in a block-based test builder. For advanced courses, each submission runs in a full Ubuntu virtual machine.
- Code quality tests can run tools such as Semgrep, ESLint, PyLint and cppcheck, and test results can fill in a rubric.
- You can read every conversation students have with the AI assistant.
- The pricing page puts LMS integration, the terminal, Jupyter notebooks and coding quizzes on the Advanced plan.
CodeHS
A coding and computer science platform for K-12 schools.
- How it grades
- Built-in test cases check a program’s output, its functions’ return values and its response to input; you can write your own autograders
- AI feedback on code
- AI Grading suggests a score and feedback against a rubric, and you approve it before the student sees it; on the paid plans
- Languages
- 10+ languages, including Python, Java, JavaScript and C++
- Similarity check
- A Plagiarism Report and a Similarity Matrix in the Academic Integrity App, on the paid plans
- LMS
- Google Classroom and Clever rostering on every plan; Google Classroom grade passback on the paid plans; Canvas, Blackboard, Schoology, D2L and Moodle over LTI on the School and District plans
- Who it’s for
- K-12 schools and districts, from grades K-5 to 9-12
- Hosted or open source
- Hosted; runs in the browser
- Free option
- Free plan with built-in autograders and your own autograders
- Price
- Starter, School and District plans by quote
- Autograders come with the exercises in its own courses. They run when a student checks or submits code.
- Its help page says output tests compare what a program prints with the expected result and “cannot check the student code”.
- AI Grading works on assignments that have solution code or an autograder. District and school admins can turn each AI tool on or off.
- Code History keeps snapshots of a student’s program and highlights pasted code in yellow.
Codio
Hands-on coding courses for colleges and training programs.
- How it grades
- Standard tests compare the output you expect with the student’s; advanced tests run unit tests, style checkers or your own scripts
- AI feedback on code
- “LLM-powered rubric grading for code quality and style”; Coach AI gives hints and explains error messages
- Languages
- Any language
- Similarity check
- Yes: compares the projects of all students in a class with the Dolos detection system
- LMS
- Canvas, Brightspace, Blackboard, Moodle and others over LTI, with single sign-on and grade passback; included in the standard offering
- Who it’s for
- Colleges and universities, bootcamps and workforce training; the site also mentions subsidized K-12 licenses
- Hosted or open source
- Hosted; runs in the browser
- Free option
- Free instructor accounts, and a free trial for new organizations
- Price
- Higher education: $48 per learner per semester or $90 per year. Bootcamps and workforce programs: $10 per learner per month. Volume discounts apply
- GitHub named Codio as a partner when it retired GitHub Classroom. Codio says you can move your grading tests over.
- Also grades multiple-choice, fill-in-the-blank and Parsons problems, and has rubric-based manual grading with comments in the code.
- Behavior Insights and Code Playback use logged keystrokes to show how a student’s code was written.
- Higher education institutions that sign an institution-pay agreement can get 50 free learner licenses for the first 12 months.
Gradescope
One place to grade paper, digital and code assignments.
- How it grades
- An autograder you write (a setup script and an autograder script) runs on Gradescope’s servers; you can also grade code by hand with a rubric and inline comments
- AI feedback on code
- Not stated
- Languages
- Any language: autograders run in Docker containers that you set up
- Similarity check
- Code Similarity covers 12 languages on the Institutional plan. It shows how similar two programs are and “does not automatically detect plagiarism”
- LMS
- Blackboard, Canvas, Moodle, Brightspace (D2L) and Sakai, on the Institutional plan
- Who it’s for
- Instructors in any subject; the home page counts 2,600+ universities
- Hosted or open source
- Hosted
- Free option
- An Institutional Trial for one term includes programming assignments; the free Basic plan that follows lists basic assignment types
- Price
- Institutional pricing on request
- Part of Turnitin. The same course can hold scanned paper exams, PDF homework and programming assignments.
- Students can submit from GitHub and Bitbucket, as many times as they want, and get results when the autograder finishes.
- One assignment can have autograded and manually graded questions.
- AI-assisted Grading, which groups similar answers for some question types, is part of the Institutional license.
GraidableOur tool
Built for paper exams, including code written by hand.
- How it grades
- Reads code that students write by hand on a scanned paper exam, and drafts points and feedback for each question from your answer key or rubric
- AI feedback on code
- Yes: a draft score and feedback for each question, from your answer key or rubric, for you to approve
- Languages
- Not stated
- Similarity check
- Not stated
- LMS
- Results come as feedback PDFs and a spreadsheet of each student’s points per question
- Who it’s for
- Teachers and instructors who set tests and exams on paper
- Hosted or open source
- Hosted
- Free option
- Demo and sample class, no card
- Price
- Student $9.99 and Professional $24.99 a month (10 and 60 students graded a month); the first plan starts with 10 free gradings; schools get a price on request
- For the written exam in a programming course: a class set of scanned papers, with code, math or written answers on the same pages.
- Each score shows where on the student’s page it comes from, and doubtful items get a “Check closely” flag with the reason.
- You review, change and approve every score; nothing reaches a student before you have.
nbgrader
Create and grade assignments in Jupyter notebooks.
- How it grades
- Test cells with assert statements: a cell that raises no error earns its points. Other cells can be graded by hand
- AI feedback on code
- Not stated
- Languages
- Python notebooks; other Jupyter kernels work but have not been extensively tested, the docs say
- Similarity check
- Not stated
- LMS
- Not stated
- Who it’s for
- Instructors who teach with Jupyter notebooks
- Hosted or open source
- Open source (BSD 3-Clause); installed by the instructor
- Free option
- Yes: the software is open source
- Price
- No license fee
- An assignment is an ordinary notebook. You write the solution and the tests, and nbgrader produces the student version with the solutions removed.
- Tests can be hidden from students and are put back when you grade.
- Students install nothing. An optional Validate extension lets them check their notebook against the visible tests.
- The docs cover running it with JupyterHub, and point to a Docker container for grading untrusted code.
Otter-Grader
An open-source autograder for Python and R from UC Berkeley.
- How it grades
- Test files: Python tests that raise an error when a case fails, in the style of pytest, or OK-format tests; R tests written like unit tests
- AI feedback on code
- Not stated
- Languages
- Python and R: scripts, Jupyter notebooks and R Markdown documents
- Similarity check
- Not stated
- LMS
- “Compatible with a few different LMSs, including Canvas and Gradescope”
- Who it’s for
- Python and R classes “at any scale”
- Hosted or open source
- Open source (BSD 3-Clause); you choose where it runs
- Free option
- Yes: the software is open source
- Price
- No license fee; “you provide the compute”
- Developed by the Data Science Education Program at UC Berkeley.
- Otter Assign builds the student version and the tests from one notebook. Otter Check lets students run the public tests while they work.
- Grades on your own machine, in parallel Docker containers, on a JupyterHub or as a Gradescope autograder.
- Otter Export turns notebooks into PDFs for the parts you grade by hand.
Vocareum
Cloud labs for computer science, AI and data courses.
- How it grades
- Auto-grading through custom Bash scripts and API endpoints; it can validate code output and check cloud configurations
- AI feedback on code
- “AI-assisted grading” is listed as one of its auto-grading methods
- Languages
- Not stated
- Similarity check
- Yes: plagiarism detection that compares the structure of submissions
- LMS
- Canvas, Blackboard, Moodle and D2L Brightspace over LTI 1.3, with roster sync and grade passback
- Who it’s for
- Higher education, bootcamps and certification programs, and online program providers
- Hosted or open source
- Hosted; cloud-based
- Free option
- Not stated
- Price
- $10 per monthly active user, with extra resource fees for AI, virtual machines and cloud labs
- A lab opens from the LMS as a Jupyter notebook, a VS Code IDE or a full Linux desktop.
- Grading scripts can check cloud tasks too, such as whether an AWS S3 bucket was created correctly.
- Its AI tutor, AI Compass, gives hints grounded in the instructor’s course materials and rubrics.
- Its higher education page says the platform is FERPA compliant.
zyLabs (zyBooks)
Coding labs inside zyBooks interactive textbooks.
- How it grades
- Deterministic test cases: output checks, unit tests with frameworks such as Python unittest and JUnit, and custom bash tests
- AI feedback on code
- AI Hints for students who are stuck; zyBooks says AI “plays no role in evaluating student work”
- Languages
- Over 50, including Python, Java, C, C++ and HTML
- Similarity check
- Yes: the Similarity Checker runs class submissions through MOSS after each deadline
- LMS
- Single sign-on and grade passback over LTI 1.3, with setup guides for Canvas, Blackboard, Brightspace and Moodle
- Who it’s for
- Courses at “over 1,800 academic institutions”; the catalog includes AP Computer Science
- Hosted or open source
- Hosted; a cloud IDE in the browser
- Free option
- Instructors can request an evaluation copy
- Price
- Students in a class typically pay $69 to $90 for a zyBook, depending on the content and add-ons
- zyBooks is a Wiley brand. Students do the labs in a cloud IDE that hosts VS Code, Jupyter Notebook and RStudio.
- You can generate a lab from a plain-language description, with test cases and a rubric, and edit it before it is published.
- Five integrity tools are built into the labs: Code Playback, the Similarity Checker, Paste Limiting, Student Behavior Insights and randomized assessments.
- The help center says LMS integration is optional: scores can also be downloaded as a CSV file.
Pick by what you teach
- If you teach a high school class and want the course material included, look at CodeHS. It says it is built for K-12, and its Free plan has autograders in the exercises.
- If you run a large university course with your own test suite, Autolab and Gradescope run the autograder you write, in any language. Gradescope’s programming assignments are on its Institutional plan.
- If you want tests without config files, CodeGrade has a block-based test builder, and Codio’s standard tests ask for the input and the output you expect.
- If your course is taught in Jupyter notebooks, nbgrader and Otter-Grader are open source. Otter-Grader also grades R.
- If you want the textbook and the labs in one product, zyLabs are part of zyBooks.
- If your labs need cloud accounts or a full Linux desktop, look at Vocareum.
- If the exam is on paper and students write code by hand, look at Graidable, the tool we make. It reads the scans and drafts points from your answer key.
If the same course has essays or paper exams, see AI essay graders compared and AI grading tools for handwritten math. The main comparison of AI grading tools covers every kind of grading.
What is the difference between test-based autograding and AI feedback on code?
A test-based autograder runs the program and checks what it does. AI feedback reads the code and comments on how it is written. They answer different questions, so most courses that use AI feedback still need the tests.
| Test-based autograding | AI feedback on code | |
|---|---|---|
| Looks at | What the program does when it runs | The code itself |
| Answers | Does it work? | Is it written well? |
| Good for | Correctness, fast results, resubmitting until the deadline, large classes | Style, structure, naming, comments, a first draft of rubric scores |
| Check in a trial | How long the first assignment takes to set up, and what a student sees when a test fails | Whether its scores match yours on work you have already graded, and whether you can change them before students see them |
Tests give every student the same check, in seconds, and the result doesn’t depend on who grades. Their limit is that they don’t read the code. CodeHS’s help page says it plainly: an output test that expects 10 can’t tell whether the student added 5 and 5, multiplied 5 by 2 or just printed 103.
Between the two sits a third kind of check that uses no AI: style checkers and linters. CodeGrade lists code quality tests that run tools such as PyLint and ESLint2, and Codio’s advanced tests can run style checkers4. If “code quality” on your rubric means consistent style and no unused variables, a linter may be all you need.
AI feedback is a draft. CodeHS describes its AI suggestions as “always a starting point, not the final grade”3. We think that is the right way to treat any of them: run the tool on submissions you have already graded, and read its comments before a student does.
Which tools use AI to grade coding assignments?
Four of the ten say they do: CodeHS, Codio, Graidable and Vocareum.
- CodeHS: AI Grading suggests a score and feedback from a rubric. You adjust them, then finalize the assignment or mark it as needing work, and “nothing reaches the student until you approve it”. It is part of the paid plans3.
- Codio: its page for former GitHub Classroom users lists “LLM-powered rubric grading for code quality and style”4.
- Graidable: the AI drafts points and feedback for code written by hand on a paper exam, from your answer key or rubric, and you approve each score6.
- Vocareum: its site lists “AI-assisted grading” next to Bash scripts and API endpoints as a way to auto-grade9. The pages we read don’t describe it further, so ask for a demo.
Two more put AI in front of students instead. CodeGrade’s assistant answers students’ questions under rules you write for each assignment2. zyLabs has AI Hints, and zyBooks says AI “plays no role in evaluating student work”10.
The pages we read for Autolab, Gradescope’s programming assignments, nbgrader and Otter-Grader describe grading by tests and by hand.
Which is better for a large university course, and which for a high school class?
For a large university course, start with the tools that run your own test suite and connect to your LMS. For a high school class, start with CodeHS, the one of the ten that describes itself as a platform for K-12 schools.
What the vendors say about large courses:
- Autolab puts grading jobs in a queue and hands each to an available container or virtual machine1.
- CodeGrade says results come back in seconds “even for classes of 500+ students”2.
- Gradescope says it runs “your autograder at scale” on its own servers5.
- Otter-Grader is designed for classes “at any scale” and grades in parallel Docker containers8.
- Vocareum says its engine can handle hundreds of submissions in minutes9.
- Codio says its auto-grading works “for classes of any size”4.
What they say about schools:
- CodeHS has courses for grades K-5, 6-8 and 9-12, and a Free plan with built-in autograders3.
- Autolab says its aim covers classes at the secondary level too. It is software you install and run yourself1.
- CodeGrade’s free plan covers up to 50 students per course2.
- Codio says it subsidizes K-12 student licenses4.
How do these tools handle AI-written code?
Mostly by showing you how the code was written. The vendors describe three approaches: recording the student’s work, limiting or flagging pasted code, and giving students an AI helper you can watch.
- Codio logs keystrokes. It says Behavior Insights lets teachers see “whether students have copied and pasted their work from tools like ChatGPT or elsewhere”, and Code Playback replays how a file was built4.
- zyBooks says it “cannot block external AI tools outright”. It offers Paste Limiting, Code Playback, a dashboard that flags outliers in time spent and pasted code, and lockdown browsers for exams10.
- CodeHS highlights pasted code in yellow in Code History. Its paid plans add a setting that turns off copy and paste in the editor3.
- CodeGrade lets you set what its AI assistant may do in each assignment, from guiding questions to full coding help, and shows you every conversation2.
A similarity check is a different thing. It compares submissions with each other, so it finds two students who handed in the same code, wherever the code came from. Gradescope says its Code Similarity tool “does not automatically detect plagiarism”5, and zyBooks says the judgment call stays with the instructor10.
What happened to GitHub Classroom?
GitHub retired it on 28 August 2026. Accounts, repositories and organizations created through Classroom were not affected11.
GitHub named two partners for the same workflow of handing out assignments, autograding and working in GitHub repositories. One is Codio, which is in the comparison above. The other is Classroom 50, which GitHub describes as a free, open-source alternative from the Fifty Foundation. It can autograde submissions with declarative tests or your own grading scripts, and it runs them with GitHub Actions11.
We also looked for Replit’s Teams for Education. Its page on replit.com returned “Page not found” when we checked, so it is not in the comparison12.
Questions people ask
Which autograders are free?
Three are open source, with no license fee: Autolab, nbgrader and Otter-Grader. You run them on your own machine or server.
Among the hosted tools:
- CodeGrade has a free plan for up to 50 students per course2.
- CodeHS has a Free plan that includes autograders3.
- Codio gives instructors free accounts and new organizations a free trial4.
- Gradescope offers an Institutional Trial for one term5.
What do the paid tools cost?
- CodeGrade: $24 or $39 per student per course, plus $15 for the AI assistant2.
- Codio: $48 per learner per semester or $90 per year in higher education4.
- Vocareum: $10 per monthly active user, plus resource fees for AI, virtual machines and cloud labs9.
- zyBooks: students in a class typically pay $69 to $90 for a zyBook10.
- CodeHS and Gradescope: by quote3,5.
Do they send grades to Canvas or another LMS?
Most of the hosted tools do. CodeGrade, Codio, Gradescope, Vocareum and zyBooks describe LTI integrations with grade passback, and CodeHS has Google Classroom grade passback on its paid plans. Check the plan: CodeGrade puts LMS integration on Advanced, Gradescope on Institutional, and CodeHS puts LTI on its School and District plans. Autolab’s docs describe roster sync, and Otter-Grader’s docs name Canvas and Gradescope.
Does the teacher stay in control of the grade?
With tests, yes: you write the tests and decide what each one is worth. CodeGrade, Codio and Gradescope also describe grading by hand next to the autograder, with rubrics and comments in the code.
With AI, check before you buy. CodeHS says nothing reaches a student until the teacher approves it3. For the others, ask to see the review step.
What happens to student code and data?
With the open-source tools, submissions are graded wherever you run the software. For hosted tools, read the vendor’s privacy pages and ask what any AI feature sends to which provider.
zyBooks is specific about this. It says no personally identifiable student information is sent to AI systems, and that its AI provider is contractually barred from using customer data for training10.
How to test an autograder on your own assignment
- Take an assignment you have already graded, with 15 to 30 submissions. Include one that doesn’t compile, one that loops forever, and one that prints the right answer in the wrong format.
- Build the tests for it in the tool, and time yourself. The first assignment is the slow one, so this is the cost that matters.
- Compare the tool’s score with yours, test by test. Totals hide errors that cancel out.
- Look at each disagreement. Is it an output match that is too strict, a case your tests missed, or a judgment call?
- Read what the student sees when a test fails. A message that names the failing case teaches something. A bare zero means an email to you.
If the tool has AI feedback, give it the same submissions and compare its rubric scores with yours, criterion by criterion.
Graidable is the tool we make. This part is about it.
Where Graidable fits
Graidable is the one tool here for the paper exam. Students write code, math or explanations by hand, you scan the class set, and we draft a score and feedback for each question from your answer key. You review, change and approve every one before students see anything. For the math and science side of the same course, see AI grading tools for handwritten math.
Sources
-
Autolab Project, home page, docs (including the guide for lab authors and LTI course integration) and GitHub repository. Checked 8 October 2026. ↑
-
CodeGrade, home page, pricing, autograder, code plagiarism checker and AI assistant. Checked 8 October 2026. ↑
-
CodeHS, home page, plans, and the help articles CodeHS Autograders, Using the AI Grading App and Guide to cheat prevention and detection. Checked 8 October 2026. ↑
-
Codio, auto-grading, pricing, plagiarism detection, LMS integration, home page and page for GitHub Classroom users. Checked 8 October 2026. ↑
-
Gradescope, home page, pricing, autograder documentation, and the guides Grading a Programming Assignment, Generating a Code Similarity Report and LMS workflows. Checked 8 October 2026. ↑
-
Graidable, Pricing and the product itself; the tool is ours. Checked 8 October 2026. ↑
-
nbgrader, documentation (highlights, creating and grading assignments, FAQ) and GitHub repository. Checked 8 October 2026. ↑
-
Otter-Grader, documentation (including test files) and GitHub repository. Checked 8 October 2026. ↑
-
Vocareum, home page, pricing and higher education. Checked 8 October 2026. ↑
-
zyBooks, zyLabs, AI tools and content, academic integrity, home page, and the help articles Payment and Getting started with LTI. Checked 8 October 2026. ↑
-
GitHub, retirement announcement for GitHub Classroom (27 August 2026) and Export or migrate GitHub Classroom data. Checked 8 October 2026. ↑
-
Replit, replit.com/teams-for-education. Checked 8 October 2026. ↑