BlogFor teachers
How English teachers can use AI marking responsibly
A balanced guide for English teachers: where AI marking helps, where it doesn't, and how to integrate it into your department workflow without losing pedagogy.
AI marking divides English departments. In some schools the head of department is enthusiastic, the second in department is cautious, and at least one experienced colleague is openly hostile to the whole idea. We have sat in those staff rooms. We have heard every version of the argument. The strongest opinions, on both sides, often come from teachers who have not actually tried using the tool yet — which is fair enough, because nobody has the spare time at half-term to test a new platform on the off chance it might be useful.
This piece is for the teachers and HoDs in the middle. It is not "AI replaces you" — it does not. It is not "AI is dangerous" — that framing flattens a useful distinction between bad and good uses. It is a balanced framework: where AI marking helps, where it falls short, and a slow integration path that does not require you to gamble pedagogy on a vendor's marketing copy.
Where AI marking genuinely helps
The honest case for AI marking is narrower than the marketing claims and broader than the sceptics admit. The places it most reliably helps are the ones where the work is high-volume, repetitive, and pattern-driven — the kind of marking that, if you are honest, you do less well by Sunday evening than you did on Friday.
First-pass marking on volume essays. A mock cohort of 180 students sitting Paper 1 generates 180 essays that all need a band, an AO breakdown, and a few targeted comments. AI marking handles that first pass quickly, so the teacher arrives at the pile already knowing which scripts need close attention and which are sitting comfortably in their target band. The same is true of weekly homework essays — a class of 30 timed paragraphs is a tractable thing to skim when an AI has already flagged the AO gaps.
Catching feature-spotting and AO-coverage gaps early. Students who write three paragraphs of language analysis and forget structure altogether tend to do it consistently. AI marking spots the missing AO5 or the thin AO3 in the first essay rather than the fifth, which means the intervention can happen in week two rather than week eight. Earlier diagnosis is the single most underrated benefit.
Consistent feedback, no end-of-stack fatigue. Every honest marker knows the twenty-eighth essay gets less attention than the third. AI does not get tired, does not lose patience with weaker essays, and does not start writing shorter comments at 10pm because the kettle is calling. The consistency is not better than a fresh teacher's best work. It is better than a tired teacher's average work.
Per-student analytics over a half-term. When essays go through the same system you get a per-student picture that is hard to construct from a stack of paper: this child consistently low on AO2; that one plateaued on AO1 since November. Patterns that took a year to surface now surface in three weeks.
Frees teacher time for high-leverage interventions. This is the point most often missed in the AI debate. The aim is not to remove the teacher from the loop. The aim is to move the teacher from the routine middle of the loop to the high-impact ends of it: the 1:1 conversation with a student about why their argument is not landing, the live feedback on a coursework draft, the moderation discussion that needs human judgement. AI takes the easy bit so the teacher can do the hard bit better.
Where AI marking falls short
The honest list of limits is just as important as the honest list of strengths. A department that adopts AI marking without a clear-eyed view of where it fails will end up with a worse classroom than one that did not adopt it at all.
Coursework with departmental house-style. Most departments mark coursework against an internal interpretation of the spec — a particular emphasis on sustained argument, a preferred way of weighting AO2 against AO5, a tradition of expecting two contextual references rather than one. AI defaults to the public mark scheme. It does not know your department's house-style unless you tell it, and even then the calibration is imperfect. For coursework, AI is a draft tool. The final mark is a human one.
Final pieces requiring human judgement. Spoken language assessment is the obvious example. The whole exercise is about presence, register, and the interaction between speaker and audience. There is no AI substitute for a teacher in the room. The same logic applies, more loosely, to any piece where the assessment is fundamentally about the person rather than the page.
Moderation conversations. AI can flag a script that sits awkwardly between two bands. It cannot replace the hour-long discussion in which two examiners argue about whether the structure mark should be a 6 or a 7. The flagging is genuinely useful — it surfaces the cases that need debate. The debate itself is human.
Pastoral context. A struggling student's essay needs a teacher who knows the student is struggling. Knows the bereavement, the move, the EHCP, the fact that they are working three years above their literacy assessment because they have an older sibling who reads to them. The AI sees the words on the page. The teacher sees the child.
A 3-tier integration model
Rather than asking "should we use AI marking?" — which is a yes/no question that hides most of the interesting design decisions — it is more useful to ask "at what tier should we use it?". We find three tiers covers most departmental use cases.
Tier 1: AI as first-pass for routine homework. Students submit, AI marks, teacher reviews flagged cases only. This is the highest-leverage tier and the easiest to start with. The teacher's contribution shifts from marking everything to spot-checking everything and deeply marking the 20% that actually needed eyes.
Tier 2: AI as student self-checker. Students mark their draft against AI feedback, redraft, and submit the redraft to the teacher for the real mark. The pedagogical bonus here is that students start internalising the AO criteria — they have to read the feedback, decide whether they agree, and act on it. The teacher marks the better second draft, which is more interesting work and a clearer signal of the child's actual ceiling.
Tier 3: AI as moderation aid. Teacher marks a paper. AI marks the same paper. Teacher and AI compare. Disagreements become discussion points — sometimes the AI is wrong, sometimes the teacher reads something they had missed, sometimes the disagreement reveals a genuine ambiguity in the script. This is the tier most useful for new teachers calibrating to a board, and for departments running internal moderation.
You do not have to commit to all three. Most departments start at Tier 1 with one teacher and one class, and only widen the use once the workflow is settled.
What the literature says (broadly)
Research on automated essay scoring is genuinely mixed, and we want to represent that fairly rather than cherry-pick. The strongest evidence base supports AI in the formative role — feedback to help students improve a draft — rather than in the summative role of awarding the final grade. That maps neatly onto the practical experience of departments using these tools: AI is a much better practice partner than it is a final examiner.
The wider regulatory position, taken broadly across bodies like the OECD and Ofqual, is that human-in-the-loop is non-negotiable for high-stakes assessment. Final grades for qualifications stay with humans. AI is a tool that supports those humans. We agree with that position and design accordingly.
We are deliberately not citing specific studies in this piece, because the AI-and-assessment field moves quickly and we would rather you trust our framing than have us link to a paper from 2023 that has already been superseded. If you want a deeper read, your subject association is a better starting point than any vendor blog.
Setting student expectations
How you frame AI marking to students matters more than the tool itself. The line we recommend is short: "AI marks for AO coverage; I mark for the argument you actually made." That distinction does most of the pedagogical work. It tells students the AI is good at the structural, mark-scheme-aligned features and that the human is reading them as a writer.
Two further habits help. First, teach students to disagree with AI feedback. A child who can read an AI comment and explain why it is wrong is doing exactly the kind of critical reading the qualification rewards. The skill of pushing back on an automated judgement is becoming a core literacy in its own right. Second, make the workflow visible: AI feedback, then student response, then teacher review. When students can see the pipeline they understand they are not being graded by a machine — they are being supported by one.
Department-level adoption — the slow path
The fastest way to lose a department on AI marking is to mandate it from above. The slow path works better.
Term 1. One teacher trials it with one class. Ideally a Year 10 group on routine homework — low stakes, high volume. The teacher keeps notes on what helps and what gets in the way.
Term 2. Two or three teachers run parallel trials. The findings get a slot in a department meeting. The conversation is honest: where did it work, where did it not, what would you change. Anyone who hated it gets to say so.
Term 3. Cohort-wide use, but only for the workflows the trials proved out. Shared protocols on how AI feedback is framed for students, when teachers override the AI, and what gets escalated for moderation.
Year 2. AI integrated into the mocks workflow. By this point the department has its own informal expertise, and the tool is shaped to your house-style rather than the other way round.
This is slower than the rollout your senior leadership team probably wants. It is also the only version that actually lands.
What we would love teachers to push back on
We are a small team and the calibration of our marker improves when teachers tell us what is wrong with it. Specifically:
- When the AI is wrong — overrating or underrating a script in a way you can pin down.
- When it over-credits fluency at the expense of argument or evidence.
- When students misuse it as an answer-machine rather than a feedback partner.
- When the tone of the feedback is unhelpful for a particular cohort.
We update calibration based on this feedback. The platform gets better when teachers push back, not when they say nothing.
Where to go next
If this framing matches how you think about marking, the related pages are:
- /teachers — the teacher tools overview, including class dashboards and per-student analytics.
- /schools — department licence options.
- /pricing — current pricing.
- /blog/how-ai-marking-actually-works — the companion piece on what the marker actually does under the hood.
The aim of all of this is the same. Keep the teacher at the centre. Take the routine work off their plate. Make the high-leverage minutes more frequent and more focused. That is what responsible AI marking looks like in an English department, and it is the only version we are interested in building.
Keep revising on The English Hub
Put this into practice with our teaching resources and practice area. Short, regular sessions on the areas you find hardest are what move your grade.