RevisionDojo is a strong option for rapid, criterion-level IB coursework feedback. However, no verified public study currently establishes its average error against matched final IB-moderated marks, so it cannot responsibly be described as the most accurate AI grader for IB based on predictive evidence alone.
The RevisionDojo IB Coursework Grader can identify rubric-related weaknesses, annotate relevant passages, estimate criterion marks, and suggest revision priorities. These functions may be useful even when the estimated total does not exactly predict the mark eventually recorded by the IB.
A reliable accuracy claim requires frozen AI predictions, identical submitted files, authenticated official results, and transparent subject-level calculations. This article explains what the existing evidence shows, what remains unknown, and how students should interpret AI-generated marks.
The direct verdict on RevisionDojo AI grader accuracy
As of 27 September 2026, the defensible conclusions are:
- RevisionDojo provides IB-specific, rubric-aware feedback for supported IAs, EEs, TOK work, and other listed coursework.
- Reports can include estimated marks, criterion breakdowns, annotations, strengths, weaknesses, and revision priorities.
- RevisionDojo has not published a sufficiently detailed dataset matching frozen AI predictions with final moderated or externally assessed IB marks.
- There is no verified public percentage showing how often estimates fall within 1 mark, 2 marks, or one grade boundary of the official result.
- Available evidence does not justify naming RevisionDojo, Clastify, or another platform as the most accurate AI grader for IB.
This does not establish that the grader is inaccurate. It means its predictive accuracy remains publicly unquantified. Students should treat an estimated mark as a provisional second opinion and use the detailed comments to find weaknesses they can verify against the rubric.
What would count as a valid moderated-mark study?
A credible study needs matched pairs. For every submission, researchers must preserve the AI prediction before the official result becomes known and compare it with the final mark for the same submitted version.
| Study element | Required disclosure |
|---|---|
| Sample | Total matched predictions and official marks |
| Coverage | Sessions, subjects, levels, and coursework types |
| Prediction timing | Confirmation that results were unknown when predictions were frozen |
| Version control | Evidence that the graded and submitted files were identical |
| Benchmark | Final moderated component mark, examiner mark, or another verified result |
| Verification | Method used to authenticate marks and candidate records |
| Exclusions | Rules for incomplete files, incorrect rubrics, and missing results |
These details prevent misleading comparisons. If a student substantially revises an IA after receiving an estimate, comparing the earlier draft with the final result measures both grader error and the effects of revision. Predictions generated after results are available also create a risk of data leakage.
Moderated marks require careful interpretation
Under the IB assessment process, internal assessment is generally marked by teachers and externally moderated using sampled work. Moderation evaluates how accurately and consistently a school applied the criteria, and an adjustment may affect candidates whose work was not individually reviewed by a moderator.
A final moderated component mark is therefore not always equivalent to an examiner independently re-marking that exact script. A rigorous study should distinguish between:
- Work directly included in the moderation sample
- Work affected by a school-level moderation adjustment
- Externally assessed work marked directly by IB examiners
- EE and TOK components with assessment arrangements different from conventional subject IAs
Reporting these categories separately would be more informative than publishing one pooled accuracy percentage.
How the error margin should be calculated
For coursework item , signed error is:
Here, is the AI-predicted raw mark and is the verified final mark. A positive value means the prediction was too high; a negative value means it was too low.
The principal metric should be mean absolute error, or MAE:
MAE answers the practical question, “How many raw marks away was the grader on average?” A complete analysis should also report mean signed error, median absolute error, exact agreement, agreement within 1 mark, agreement within 2 marks, and the largest observed error.
IB coursework components have different maximum marks, so raw errors should not be pooled carelessly. Subject-level results can use raw marks, while an overall analysis should use normalized percentage-point errors or another clearly explained standardization method.
How subject-level agreement should be reported
No matched RevisionDojo dataset containing the necessary results has been publicly verified. A future report should use a format such as this:
| Subject | Matched scripts | Mean absolute error | Mean signed error | Exact match | Within 2 marks |
|---|---|---|---|---|---|
| Subject and component | Verified count | Raw-mark MAE | Directional bias | Percentage | Percentage |
| Overall result | Total count | Standardized MAE | Overall bias | Percentage | Percentage |
The current headline must therefore remain: Average error margin: data not publicly established.
Confidence intervals are essential because an MAE based on 15 submissions is much less stable than one based on several hundred. Subject-level accuracy should not be promoted when the sample is too small or unrepresentative.
What RevisionDojo moderation data actually shows
RevisionDojo has published a separate analysis of IB coursework moderation. It reports approximately 4,000 collected coursework samples, including 2,459 verified samples with both a teacher-awarded mark and final moderated mark across 33 subjects.
This evidence shows why a teacher's initial mark should not automatically be treated as the final benchmark. It does not validate Jojo AI's predictions because the analysis compares teacher marks with moderated marks, not AI estimates with moderated marks.
Teacher moderation statistics cannot be converted into AI grader accuracy figures, even if both concern IB coursework. Doing so would answer a different research question and produce an unsupported accuracy claim.
Where an AI coursework grader may be weakest
No subject can be identified as RevisionDojo's statistically weakest area without matched subject-level results. Greater caution is nevertheless appropriate in several situations:
- Holistic judgments: TOK and essay-based criteria may depend on the coherence, significance, and development of an argument as a whole.
- Visual or technical evidence: Mathematical notation, handwritten work, graphs, diagrams, appendices, audio, or poor scans may be interpreted incompletely.
- Specialist context: A plausible method or interpretation can contain a subtle disciplinary error that requires subject expertise.
- Borderline markbands: Small differences in best-fit judgment may move a response between adjacent marks.
- Incorrect setup: Choosing the wrong subject, level, task, or rubric invalidates the estimate.
- Incomplete drafts: Missing tables, citations, reflections, or appendices prevent the grader from assessing evidence that is not present.
RevisionDojo's practical advantage is that students can take a disputed comment into Jojo AI, review the relevant concept, and return to the rubric-linked report. This improves usability, but integration is not evidence of predictive accuracy.
RevisionDojo compared with Clastify
Students often ask which IB coursework grader is most accurate. A fair answer requires both systems to grade the same unseen submissions under identical conditions before official results are disclosed.
| Comparison point | RevisionDojo | Clastify |
|---|---|---|
| Coursework feedback | Rubric-aware reports and annotations | AI coursework grader described as a beta tool |
| Public matched validation | No sufficiently detailed dataset verified | No comparable dataset verified |
| Human support | Coursework support and tutors are available | Separate examiner review is advertised |
| Wider workflow | Integrates with Jojo AI and RevisionDojo study tools | Connects with exemplar and coursework resources |
| Appropriate interpretation | Estimate does not replace teacher judgment | AI feedback does not replace human review |
Clastify presents its AI grader as a draft-stage tool and recommends human examiner review for a more authoritative evaluation. RevisionDojo may suit students who want feedback integrated with an IB-focused study system, but claiming that either service is more accurate would require a controlled comparative study.
How students and teachers should use the grader
The most reliable use is diagnostic rather than predictive:
- Upload the complete draft you want assessed.
- Select the exact subject, level, task, and current rubric.
- Read the criterion breakdown before focusing on the total.
- Check every comment against evidence in the draft.
- Prioritize two or three revisions with the greatest rubric impact.
- Discuss disputed or consequential judgments with a teacher or supervisor.
- Rewrite independently and preserve version history.
RevisionDojo's guide to AI-powered IB coursework grading explains how to interpret annotations. Its IA grader accuracy review reaches the same central conclusion: criterion feedback may be useful even though a precise moderated-mark accuracy rate has not been established.
Students must also follow school AI policies. AI-generated text is not automatically the student's own work, so feedback should not become undisclosed authorship. RevisionDojo's IB AI ethics guide explains how to use assistance transparently.
Limitations of the available evidence
Current evidence has several important limitations:
- No public matched prediction dataset is available for calculating MAE.
- Anonymized data and reproducible analysis have not been released for independent review.
- The graded draft may not always match the final submitted version.
- Model updates and syllabus revisions can change performance over time.
- Uneven subject representation could conceal weaknesses in smaller categories.
- Moderated IA marks may reflect school-level adjustment rather than individual re-marking.
- Voluntary submissions may not represent all users, schools, or achievement levels.
A future report should publish subject-level counts, predefined metrics, confidence intervals, model versions, exclusion rules, and mark-verification procedures. Until then, a precise accuracy percentage would be unsupported.
Conclusion
RevisionDojo offers a practical IB coursework grader for identifying rubric-linked weaknesses, receiving annotations, and planning revisions. Its main strength is the connection between the grader, Jojo AI, and a wider IB study environment.
The evidence supports a limited conclusion: no verified public matched study currently provides RevisionDojo's average error against final moderated marks. Treat the score as provisional, verify comments against the current rubric, and retain human judgment for consequential decisions.
For a structured second reading of an IA, EE, or TOK draft, use the RevisionDojo Coursework Grader and Jojo AI to clarify feedback rather than generate assessed writing.

