Grading essays has always been one of the most demanding tasks in education – time-consuming, prone to inconsistency, and dependent on the individual judgment of each instructor. With classrooms growing and digital learning scaling rapidly, the question is no longer just theoretical: can computers do this job? Automated Essay Scoring (AES) has evolved from a fringe idea in the 1960s into a field backed by serious research and widely deployed systems. Here’s a clear look at how AI-based subjective assessment actually works, which tools lead the space, what techniques they rely on, and where they still fall short.

Table of Contents

The evolution of essay evaluation: from human grading to AI-powered tools

Before technology entered the picture, high-stakes essays were typically scored by two trained human raters. If their scores differed significantly, a more experienced third rater would resolve the disagreement. This system worked, but it was expensive, slow, and vulnerable to fatigue, bias, and inconsistency – especially when grading thousands of scripts across large institutions.

The first attempt at automating this process came in 1966, when educator Ellis Page developed Project Essay Grader (PEG). Page’s insight was that a computer could be trained to recognize surface-level features of writing – spelling, sentence length, punctuation – that tend to correlate with human judgments of quality. Progress stalled for decades, but the 1990s changed everything: better computing power and advances in natural language processing (NLP) unlocked far more capable systems.

By the late 1990s and early 2000s, multiple competing AES platforms had emerged. These systems were developed to assist teachers in low-stakes classroom assessment as well as testing companies in large-scale high-stakes assessment, helping address persistent issues of time, cost, reliability, and generalizability in writing evaluation. Today, AES tools are used in everything from classroom formative feedback to standardized exams like the GRE and TOEFL.

Automated essay scoring models: IEA and E-rater

Among the many AES systems developed over the past three decades, two stand out for their academic influence and real-world deployment: the Intelligent Essay Assessor (IEA) and E-rater, both associated with Educational Testing Service (ETS).

Intelligent Essay Assessor (IEA)

IEA was developed by Peter Foltz and Thomas Landauer and first used to score essays in 1997 for undergraduate courses. It later became a product of Pearson Educational Technologies and has been used in both commercial applications and state and national exams.

What makes IEA distinctive is its approach to training. IEA uses three sources to analyze an essay: pre-scored essays from other students, expert model essays and knowledge source materials, and an internal comparison of an unscored set of essays. This allows it to compare each submission against a broad base of domain-relevant content, rather than just checking for surface errors.

In terms of accuracy, IEA has shown strong results. Landauer (2003) used IEA to score more than 800 students’ answers in middle school, with results showing a 0.90 correlation value between IEA and human raters – a level of agreement comparable to two expert humans rating the same essays. IEA also has an advantage in scale: unlike human raters who cannot individually compare hundreds of essays to each other, IEA can systematically cross-reference every submission in a dataset.

Another notable feature is plagiarism detection. IEA identifies extremely similar essays regardless of paraphrasing, synonym substitution, or sentence rearrangement – making it harder for students to disguise copied work.

E-rater

The e-rater engine is an AI system developed by ETS that uses Natural Language Processing to evaluate writing proficiency by providing automatic scoring and feedback on grammar, mechanics, word use and complexity, style, organization, and more. It was first used commercially in February 1999, under the leadership of researcher Jill Burstein.

The E-rater system is upgraded annually; the current version uses 11 features divided into two areas: writing quality (grammar, usage, mechanics, style, organization, development, word choice) and other linguistic indicators. Unlike IEA, which focuses primarily on semantic content, E-rater evaluates both style and content, making it a more comprehensive tool for writing assessment.

E-rater powers ETS’s Criterion Online Writing Evaluation Service, which gives students diagnostic feedback on their writing drafts. It is also used in high-stakes testing contexts – but critically, in high-stakes settings, the E-rater engine is always used in conjunction with human ratings, not as a standalone judge. This hybrid approach reflects both the system’s capabilities and its known limitations.

Techniques used in AI grading

Behind these systems are several core technical methods. Understanding these helps clarify both what AI can assess well and where it runs into difficulty.

Latent Semantic Analysis (LSA)

Latent Semantic Analysis is a statistical model of word usage that permits comparisons of semantic similarity between pieces of textual information. It is the primary technique powering IEA’s content evaluation.

In simple terms, LSA works by analyzing large collections of text to map relationships between words based on how often they appear together in similar contexts. This allows the system to understand that “global warming” and “carbon footprint” are semantically related to “climate change” – even if they don’t appear in the same sentence. When a student’s essay is submitted, LSA places it within a semantic space built from a large body of domain texts, then compares it against expert-scored essays to assess content quality.

LSA has proven highly effective for evaluating content relevance and coherence. Research shows that scoring systems incorporating semantic analysis outperform surface-only systems by 15-20% in accuracy. However, LSA has a well-documented limitation: it is a “bag of words” model that ignores the structure of sentences, meaning it can fail to distinguish between sentences containing similar words but opposite meanings.

Syntactic analysis

Syntactic analysis is how AES systems evaluate the structural correctness of writing. NLP analyzes essays by examining vocabulary, grammar, and sentence structure, going beyond simple error detection by aiming to understand the content and context of the writing. Systems like E-rater examine features such as verb usage, sentence variety, part-of-speech patterns, and grammatical error rates to build a picture of a student’s structural competence.

This dimension of grading is where AES systems are most reliable. Detecting a run-on sentence, a misplaced modifier, or incorrect subject-verb agreement is a well-defined task that NLP handles consistently – and often more reliably than a fatigued human grader working through a large batch of papers.

Rhetorical structure analysis

Beyond grammar, effective writing requires logical organization and persuasive structure. Rhetorical analysis in AES addresses this by evaluating how well a student constructs and supports an argument. E-rater uses natural language processing and information retrieval to develop modules that capture features such as syntactic variety, topic content, and organization of ideas or rhetorical structures from training essays pre-scored by expert raters.

In practice, this means the system checks whether a thesis is clearly stated, whether body paragraphs contain relevant supporting evidence, and whether ideas transition logically between sections. Some AES systems rely on Rhetorical Structure Theory (RST) as a framework for understanding how the parts of an essay relate to one another – for instance, identifying whether a paragraph is functioning as an elaboration, a contrast, or a conclusion relative to the main argument.

Challenges in AI-based essay grading

Despite impressive benchmark results, AI-based essay scoring faces real limitations – especially when it comes to the qualities that make writing genuinely insightful or original.

Handling creativity and originality

Research has shown that while AI can effectively assess objective criteria like grammar, structure, and style, it struggles with subjective elements such as creativity, critical thinking, and originality. An essay that presents a counterintuitive argument or deliberately subverts conventional structure may score poorly on an AES system even if it demonstrates sophisticated thinking – simply because it doesn’t match the patterns the system was trained to reward.

AES tools often struggle to capture and evaluate the creativity and originality of student writing, a limitation that has been a consistent point of criticism in the educational assessment community. This is especially problematic in humanities disciplines where unconventional argumentation is not just acceptable but often valued.

Nuanced arguments and deep reasoning

AI systems are trained on patterns from previously scored essays. When a student’s argument is genuinely novel or requires deep contextual knowledge to evaluate, the system has no reliable framework to assess it. Research by Flodรฉn (2025) found that while AI grading of essay exams yielded somewhat comparable results to human grading, teachers still expressed concern over AI’s limitations in assessing creativity or nuance.

There is also a documented scoring bias issue. Studies show that AI often grades more leniently on low-performing essays and more harshly on high-performing ones, suggesting the system’s calibration breaks down at the extremes of the quality spectrum – precisely where accurate assessment matters most.

The subjectivity problem

Essay grading is inherently interpretive. Two experienced human teachers may legitimately disagree on the quality of an ambitious but flawed piece of writing. AES systems, trained to match human scores on average, end up encoding a kind of median judgment – which may not reflect the full range of valid assessments. The complexity and quality of rubrics directly influence AI performance, especially when rubrics contain subjective expressions and evaluative language that doesn’t map neatly onto measurable linguistic features.

There is also the question of gaming. Students who understand how AES systems work can potentially inflate their scores by using longer sentences, more sophisticated vocabulary, and clear structural markers – without necessarily improving the quality of their thinking. AES machines appear to be less reliable than human readers for any kind of complex writing test, which is why, in practice, high-stakes assessments are always scored by at least one human rater alongside the AI system.

Where does this leave educators?

The consensus emerging from research is that AI technologies, particularly machine learning and NLP, demonstrate significant potential in automating assessment processes and delivering personalized feedback, but successful implementation requires careful integration with human expertise. AES tools are most valuable as a first-pass filter, a source of instant formative feedback, and a consistency check – not as a replacement for the nuanced judgment that skilled educators bring to evaluating complex written work.

The future likely lies in hybrid models: AI handles volume, speed, and structural consistency, while human graders focus their attention on the dimensions that machines genuinely cannot replicate – insight, originality, and the quality of reasoning that makes an essay more than the sum of its grammatical parts.

What do you think? As AI systems become more sophisticated, should they be allowed to serve as the primary grader for high-stakes essay assessments in higher education – or should human oversight always remain a requirement? And if AI can already match human-rater agreement on standardized writing tasks, what does that tell us about what we’re actually measuring in those tasks?

How useful was this post?

Click on a star to rate it!

Average rating 5 / 5. Vote count: 1

No votes so far! Be the first to rate this post.

We are sorry that this post was not useful for you!

Let us improve this post!

Tell us how we can improve this post?

References
  1. https://en.wikipedia.org/wiki/Automated_essay_scoring
  2. https://www.essaygrader.ai/blog/automated-essay-scoring
  3. https://pmc.ncbi.nlm.nih.gov/articles/PMC7924549/
  4. https://files.eric.ed.gov/fulltext/EJ843855.pdf
  5. https://www.ets.org/erater/about.html
  6. https://link.springer.com/article/10.3758/BF03204765
  7. https://citejournal.org/volume-8/issue-4-08/english-language-arts/automated-essay-scoring-versus-human-scoring-a-correlational-study/
  8. https://www.navgood.com/en/article-details/nlp-essay-grading-article-784eb
  9. https://www.researchgate.net/publication/224223186_Automated_essay_scoring_using_Generalized_Latent_Semantic_Analysis
  10. https://www.researchgate.net/publication/363663827_Automated_essay_scoring_AES_systems_Opportunities_and_challenges_for_open_and_distance_education
  11. https://link.springer.com/article/10.1007/s44217-024-00320-6
  12. https://ascode.osu.edu/news/ai-and-auto-grading-higher-education-capabilities-ethics-and-evolving-role-educators
  13. https://link.springer.com/article/10.1007/s44163-026-01002-y
  14. https://link.springer.com/article/10.1007/s44163-025-00517-0

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *

Instruction in Higher Education

1 Instructional System

  1. Learning and Instruction
  2. Concept of System
  3. Instructional System
  4. Systems Approach to Instruction
  5. Selection of Instructional Inputs
  6. Effectiveness and Efficiency
  7. Role of the Teacher in the Instructional System

2 Input Alternatives – Teacher Controlled

  1. What is a Lecture?
  2. Steps in a Lecture
  3. Different Approaches to Content Treatment and Information Processing
  4. Lecture in Combination with Other Methods and Media
  5. Versatility of Lecture
  6. Demonstration
  7. Team Teaching

3 Input Alternatives – Learner Controlled

  1. Input Alternatives – Learner Controlled: The Concept
  2. Self-Learning
  3. Forms of Self-Learning
  4. Programmed Instruction/Learning
  5. Personalised System of Instruction
  6. Computer-Assisted Instruction
  7. Project Work
  8. Group-Controlled Learning Experiences
  9. Co-operative Learning Method
  10. Group Investigation

4 Evolving Instructional Strategies

  1. What is an instructional strategy?
  2. Bloom’s Taxonomy of Educational Objectives: Cognitive Domain
  3. Affective Domain of the Taxonomy of Educational Objectives
  4. Psychomotor Domain of the Taxonomy of Educational Objectives
  5. Specifying the Objectives in Behavioral Terms
  6. Difference Between Instructional Objectives, Goals of Education, Terminal Behaviors, and Learning Outcomes
  7. Evolving Instructional Strategy
  8. Dale’s Cone of Experience
  9. Evolving Instructional Strategies – Some Parameters

5 Unit and Topic Planning

  1. Unit Plan
  2. Planning the Daily Topic/Lesson
  3. Statement of General and Specific Objectives
  4. Introduction or Opener
  5. Presentation or Development Section
  6. Recapitulation or Closing Section
  7. Example of a Lesson Plan

6 Teacher Competence in Higher Education

  1. The Concept of Teacher Competence
  2. Teacher Competencies at the Tertiary Level
  3. Classification of Teacher Competencies
  4. Repertoire of Teaching Competencies
  5. How to Improve Classroom Practice
  6. Teacherโ€™s Self-Improvement

7 Skills Associated with a Good Lecture

  1. Content Organisation
  2. Preparing Lecturing Notes
  3. Activities During the Introductory Phase of a Lecture
  4. Activities During the Development Phase
  5. Activities During the Consolidation Phase
  6. Skills Associated with the Delivery of a Lecture
  7. Questioning Skills
  8. Pitfalls Associated with Lecturing

8 Skills Associated with the Conduct of Interaction Sessions

  1. Nature and Importance of an Interaction Session
  2. Tasks Undertaken in an Interaction Session
  3. Types of Discussion
  4. Formats for Group Discussion
  5. Arranging an Interaction Session
  6. Conducting an Interaction Session
  7. Follow-up of an Interaction Session
  8. Seating Plan for an Interaction Session
  9. Norms During an Interaction Session

9 Skills of Using Communication Aids

  1. Classroom Instruction and Communication Aids
  2. Classification of Communication Aids
  3. Skills of Using Some Non-Projected Aids
  4. Skills of Using Some Projected Aids
  5. Computer and Computer-Assisted Instruction Learning
  6. Integration of Communication Aids with Interaction Techniques
  7. Improvisation of Teaching Aids

10 Emerging Communication and Information Technologies

  1. Future Trends: Emerging Technologies in Education
  2. Audio-Video Technology
  3. Computer Technology
  4. Telecommunications and Networks
  5. Internet and Intranet

11 Status of Evaluation in Higher Education-I

  1. Historical background of examinations and examination reform
  2. The introduction of standardized tests
  3. The testing movement
  4. The reform movement in India
  5. Educational evaluation in the teaching-learning process
  6. Basic concepts in educational evaluation
  7. Role of objectives and evaluation in the teaching-learning process
  8. Tests and Examinations
  9. Examination as the stumbling block for qualitative assessment
  10. Defects in present-day examinations
  11. Examinations dominate teaching

12 Status of Evaluation in Higher Education-II

  1. Examination reforms – Significant aspects
  2. Reformulation of syllabus
  3. Nature of examinations and question papers
  4. Question banks
  5. Internal assessment
  6. Grading
  7. National testing service

13 Evaluation Situations in Higher Education-I

  1. Norm-referenced testing and criterion-referenced testing
  2. Formative and summative tests
  3. Cognitive and non-cognitive assessment of learning outcomes
  4. Tools and techniques for assessment of cognitive and non-cognitive outcomes

14 Evaluation Situations in Higher Education-II

  1. Evaluation of Laboratory Work
  2. Evaluation of Students’ Performance in Seminars or Similar Group-Controlled Learning Situations
  3. Evaluation of Project Work and Dissertation
  4. Internal Assessment Versus External Examination
  5. Various Types of Evaluation

15 Mechanics of Evaluation- I

  1. Framing-test items and question papers
  2. Outlining the subject matter content
  3. Identifying and stating the desired learning outcomes
  4. Different forms of test items or questions
  5. Essay type items/questions
  6. Short-answer type questions
  7. Very short answer type questions
  8. Selection type or fixed response type items or questions
  9. Essay type and objective type items compared
  10. Preparing a good question paper
  11. Preparing a Table of Specifications (Blueprint)

16 Mechanics of Evaluation-II

  1. Essential characteristics of an effective tool of evaluation
  2. Parameters concerning an evaluation item
  3. Item analysis
  4. Question banks
  5. Examination reform and question banks

17 Processing Evaluation Data

  1. Marking and grading systems
  2. The Marking system
  3. The standard error of measurement
  4. The Grading system
  5. Merits and limitations of grading system
  6. University Grants Commission recommendations on the grading system
  7. Upgraded data
  8. Test norms
  9. Computation of test norms

18 Alternative Evaluation Procedures

  1. Alternative Techniques of Evaluation
  2. Observational Technique
  3. Observation Schedule
  4. Anecdotal Records
  5. Rating Scales
  6. Checklists
  7. Score Cards
  8. Self-Reporting Techniques
  9. Interview
  10. Portfolio
  11. Questionnaires
  12. Inventories
  13. Peer Appraisal
  14. Processing Qualitative Evaluation Data
  15. Reporting the Results of Evaluation

19 Online/Web-Based Student Assessment

  1. Computers in Student Evaluation
  2. Electronic Delivery of Objective Tests
  3. Possibilities in Subjective Tests
  4. Methodologies of Essay Evaluators
  5. Other Tests Suitable for Online/Web-Based Assessment
  6. Advantages of Online/Web-Based Student Assessment
  7. Offline Use of Computers in Student Assessment