Grading essays has always been one of the most demanding tasks in education – time-consuming, prone to inconsistency, and dependent on the individual judgment of each instructor. With classrooms growing and digital learning scaling rapidly, the question is no longer just theoretical: can computers do this job? Automated Essay Scoring (AES) has evolved from a fringe idea in the 1960s into a field backed by serious research and widely deployed systems. Here’s a clear look at how AI-based subjective assessment actually works, which tools lead the space, what techniques they rely on, and where they still fall short.
Table of Contents
- The evolution of essay evaluation: from human grading to AI-powered tools
- Automated essay scoring models: IEA and E-rater
- Intelligent Essay Assessor (IEA)
- E-rater
- Techniques used in AI grading
- Latent Semantic Analysis (LSA)
- Syntactic analysis
- Rhetorical structure analysis
- Challenges in AI-based essay grading
- Handling creativity and originality
- Nuanced arguments and deep reasoning
- The subjectivity problem
- Where does this leave educators?
The evolution of essay evaluation: from human grading to AI-powered tools
Before technology entered the picture, high-stakes essays were typically scored by two trained human raters. If their scores differed significantly, a more experienced third rater would resolve the disagreement. This system worked, but it was expensive, slow, and vulnerable to fatigue, bias, and inconsistency – especially when grading thousands of scripts across large institutions.
The first attempt at automating this process came in 1966, when educator Ellis Page developed Project Essay Grader (PEG). Page’s insight was that a computer could be trained to recognize surface-level features of writing – spelling, sentence length, punctuation – that tend to correlate with human judgments of quality. Progress stalled for decades, but the 1990s changed everything: better computing power and advances in natural language processing (NLP) unlocked far more capable systems.
By the late 1990s and early 2000s, multiple competing AES platforms had emerged. These systems were developed to assist teachers in low-stakes classroom assessment as well as testing companies in large-scale high-stakes assessment, helping address persistent issues of time, cost, reliability, and generalizability in writing evaluation. Today, AES tools are used in everything from classroom formative feedback to standardized exams like the GRE and TOEFL.
Automated essay scoring models: IEA and E-rater
Among the many AES systems developed over the past three decades, two stand out for their academic influence and real-world deployment: the Intelligent Essay Assessor (IEA) and E-rater, both associated with Educational Testing Service (ETS).
Intelligent Essay Assessor (IEA)
IEA was developed by Peter Foltz and Thomas Landauer and first used to score essays in 1997 for undergraduate courses. It later became a product of Pearson Educational Technologies and has been used in both commercial applications and state and national exams.
What makes IEA distinctive is its approach to training. IEA uses three sources to analyze an essay: pre-scored essays from other students, expert model essays and knowledge source materials, and an internal comparison of an unscored set of essays. This allows it to compare each submission against a broad base of domain-relevant content, rather than just checking for surface errors.
In terms of accuracy, IEA has shown strong results. Landauer (2003) used IEA to score more than 800 students’ answers in middle school, with results showing a 0.90 correlation value between IEA and human raters – a level of agreement comparable to two expert humans rating the same essays. IEA also has an advantage in scale: unlike human raters who cannot individually compare hundreds of essays to each other, IEA can systematically cross-reference every submission in a dataset.
Another notable feature is plagiarism detection. IEA identifies extremely similar essays regardless of paraphrasing, synonym substitution, or sentence rearrangement – making it harder for students to disguise copied work.
E-rater
The e-rater engine is an AI system developed by ETS that uses Natural Language Processing to evaluate writing proficiency by providing automatic scoring and feedback on grammar, mechanics, word use and complexity, style, organization, and more. It was first used commercially in February 1999, under the leadership of researcher Jill Burstein.
The E-rater system is upgraded annually; the current version uses 11 features divided into two areas: writing quality (grammar, usage, mechanics, style, organization, development, word choice) and other linguistic indicators. Unlike IEA, which focuses primarily on semantic content, E-rater evaluates both style and content, making it a more comprehensive tool for writing assessment.
E-rater powers ETS’s Criterion Online Writing Evaluation Service, which gives students diagnostic feedback on their writing drafts. It is also used in high-stakes testing contexts – but critically, in high-stakes settings, the E-rater engine is always used in conjunction with human ratings, not as a standalone judge. This hybrid approach reflects both the system’s capabilities and its known limitations.
Techniques used in AI grading
Behind these systems are several core technical methods. Understanding these helps clarify both what AI can assess well and where it runs into difficulty.
Latent Semantic Analysis (LSA)
Latent Semantic Analysis is a statistical model of word usage that permits comparisons of semantic similarity between pieces of textual information. It is the primary technique powering IEA’s content evaluation.
In simple terms, LSA works by analyzing large collections of text to map relationships between words based on how often they appear together in similar contexts. This allows the system to understand that “global warming” and “carbon footprint” are semantically related to “climate change” – even if they don’t appear in the same sentence. When a student’s essay is submitted, LSA places it within a semantic space built from a large body of domain texts, then compares it against expert-scored essays to assess content quality.
LSA has proven highly effective for evaluating content relevance and coherence. Research shows that scoring systems incorporating semantic analysis outperform surface-only systems by 15-20% in accuracy. However, LSA has a well-documented limitation: it is a “bag of words” model that ignores the structure of sentences, meaning it can fail to distinguish between sentences containing similar words but opposite meanings.
Syntactic analysis
Syntactic analysis is how AES systems evaluate the structural correctness of writing. NLP analyzes essays by examining vocabulary, grammar, and sentence structure, going beyond simple error detection by aiming to understand the content and context of the writing. Systems like E-rater examine features such as verb usage, sentence variety, part-of-speech patterns, and grammatical error rates to build a picture of a student’s structural competence.
This dimension of grading is where AES systems are most reliable. Detecting a run-on sentence, a misplaced modifier, or incorrect subject-verb agreement is a well-defined task that NLP handles consistently – and often more reliably than a fatigued human grader working through a large batch of papers.
Rhetorical structure analysis
Beyond grammar, effective writing requires logical organization and persuasive structure. Rhetorical analysis in AES addresses this by evaluating how well a student constructs and supports an argument. E-rater uses natural language processing and information retrieval to develop modules that capture features such as syntactic variety, topic content, and organization of ideas or rhetorical structures from training essays pre-scored by expert raters.
In practice, this means the system checks whether a thesis is clearly stated, whether body paragraphs contain relevant supporting evidence, and whether ideas transition logically between sections. Some AES systems rely on Rhetorical Structure Theory (RST) as a framework for understanding how the parts of an essay relate to one another – for instance, identifying whether a paragraph is functioning as an elaboration, a contrast, or a conclusion relative to the main argument.
Challenges in AI-based essay grading
Despite impressive benchmark results, AI-based essay scoring faces real limitations – especially when it comes to the qualities that make writing genuinely insightful or original.
Handling creativity and originality
Research has shown that while AI can effectively assess objective criteria like grammar, structure, and style, it struggles with subjective elements such as creativity, critical thinking, and originality. An essay that presents a counterintuitive argument or deliberately subverts conventional structure may score poorly on an AES system even if it demonstrates sophisticated thinking – simply because it doesn’t match the patterns the system was trained to reward.
AES tools often struggle to capture and evaluate the creativity and originality of student writing, a limitation that has been a consistent point of criticism in the educational assessment community. This is especially problematic in humanities disciplines where unconventional argumentation is not just acceptable but often valued.
Nuanced arguments and deep reasoning
AI systems are trained on patterns from previously scored essays. When a student’s argument is genuinely novel or requires deep contextual knowledge to evaluate, the system has no reliable framework to assess it. Research by Flodรฉn (2025) found that while AI grading of essay exams yielded somewhat comparable results to human grading, teachers still expressed concern over AI’s limitations in assessing creativity or nuance.
There is also a documented scoring bias issue. Studies show that AI often grades more leniently on low-performing essays and more harshly on high-performing ones, suggesting the system’s calibration breaks down at the extremes of the quality spectrum – precisely where accurate assessment matters most.
The subjectivity problem
Essay grading is inherently interpretive. Two experienced human teachers may legitimately disagree on the quality of an ambitious but flawed piece of writing. AES systems, trained to match human scores on average, end up encoding a kind of median judgment – which may not reflect the full range of valid assessments. The complexity and quality of rubrics directly influence AI performance, especially when rubrics contain subjective expressions and evaluative language that doesn’t map neatly onto measurable linguistic features.
There is also the question of gaming. Students who understand how AES systems work can potentially inflate their scores by using longer sentences, more sophisticated vocabulary, and clear structural markers – without necessarily improving the quality of their thinking. AES machines appear to be less reliable than human readers for any kind of complex writing test, which is why, in practice, high-stakes assessments are always scored by at least one human rater alongside the AI system.
Where does this leave educators?
The consensus emerging from research is that AI technologies, particularly machine learning and NLP, demonstrate significant potential in automating assessment processes and delivering personalized feedback, but successful implementation requires careful integration with human expertise. AES tools are most valuable as a first-pass filter, a source of instant formative feedback, and a consistency check – not as a replacement for the nuanced judgment that skilled educators bring to evaluating complex written work.
The future likely lies in hybrid models: AI handles volume, speed, and structural consistency, while human graders focus their attention on the dimensions that machines genuinely cannot replicate – insight, originality, and the quality of reasoning that makes an essay more than the sum of its grammatical parts.
What do you think? As AI systems become more sophisticated, should they be allowed to serve as the primary grader for high-stakes essay assessments in higher education – or should human oversight always remain a requirement? And if AI can already match human-rater agreement on standardized writing tasks, what does that tell us about what we’re actually measuring in those tasks?
References
- https://en.wikipedia.org/wiki/Automated_essay_scoring
- https://www.essaygrader.ai/blog/automated-essay-scoring
- https://pmc.ncbi.nlm.nih.gov/articles/PMC7924549/
- https://files.eric.ed.gov/fulltext/EJ843855.pdf
- https://www.ets.org/erater/about.html
- https://link.springer.com/article/10.3758/BF03204765
- https://citejournal.org/volume-8/issue-4-08/english-language-arts/automated-essay-scoring-versus-human-scoring-a-correlational-study/
- https://www.navgood.com/en/article-details/nlp-essay-grading-article-784eb
- https://www.researchgate.net/publication/224223186_Automated_essay_scoring_using_Generalized_Latent_Semantic_Analysis
- https://www.researchgate.net/publication/363663827_Automated_essay_scoring_AES_systems_Opportunities_and_challenges_for_open_and_distance_education
- https://link.springer.com/article/10.1007/s44217-024-00320-6
- https://ascode.osu.edu/news/ai-and-auto-grading-higher-education-capabilities-ethics-and-evolving-role-educators
- https://link.springer.com/article/10.1007/s44163-026-01002-y
- https://link.springer.com/article/10.1007/s44163-025-00517-0
Leave a Reply