Evaluation is one of the most critical-yet often underappreciated-components of educational program design. Whether you are an instructor refining a course mid-semester or an administrator assessing the success of a curriculum, evaluation gives you the evidence needed to make informed decisions. It is not simply a tool for assigning grades; it is a systematic process of collecting and interpreting data to determine how well an educational program is working and where it needs to go next. Understanding the types of evaluation available-and the sources from which meaningful feedback can come-can transform how instructors design, deliver, and improve their courses.
Table of Contents
- What is educational evaluation?
- Formative evaluation: improving learning in progress
- Key characteristics of formative evaluation
- Summative evaluation: measuring overall achievement
- Key characteristics of summative evaluation
- Using formative and summative evaluation together
- Diagnostic evaluation: knowing where students start
- Sources of evaluation: who provides the feedback?
- Self-assessment
- Peer evaluation
- Teacher evaluation
- Why combining multiple evaluation sources matters
- Norm-referenced vs. criterion-referenced evaluation
- Evaluation in the context of instructional improvement
What is educational evaluation?
Educational evaluation refers to the systematic collection, analysis, and interpretation of data related to educational programs, teaching methods, and student learning outcomes. As researchers Palomba and Banta (1999) define it, assessment is “the systematic collection, review, and use of information about educational programs undertaken to improve learning and development.” Evaluation goes one step further – it draws on judgment to determine the overall value of an outcome based on that data. In essence, assessment is process-oriented while evaluation is product-oriented: assessment gathers the information, and evaluation interprets it to guide decisions about instruction and program design.
There are several types of evaluation used in educational settings, classified primarily by when they occur and what purpose they serve. The most fundamental distinction is between formative and summative evaluation.
Formative evaluation: improving learning in progress
Formative evaluation is used while learning is ongoing to monitor student progress, provide feedback, and adjust instruction as needed. Its primary goal is not to judge final outcomes but to identify strengths, challenges, and misconceptions so that both instructors and learners can course-correct before it’s too late. It is diagnostic in nature – it tells you where students are struggling so that teaching can be adapted in real time.
Key characteristics of formative evaluation
Formative evaluation is generally low-stakes, meaning it carries little or no grade weight, which reduces student anxiety and promotes honest participation. It happens continuously throughout a course rather than at a single fixed point. Common formative techniques include:
- In-class observations – noticing students’ non-verbal cues or responses during a lecture
- Reflection journals – periodic written reflections that help students track their own thinking
- Question-and-answer sessions – both planned and spontaneous discussions to gauge understanding
- Short quizzes and homework exercises – low-stakes checks that reveal gaps before major assessments
- Student feedback forms – structured responses on how students perceive instruction quality
The value of formative evaluation lies in the critical, timely information it provides. As instructional researchers note, small misunderstandings can be addressed right away before they become larger, harder-to-fix learning gaps. When teachers use formative data well, they can slow down where needed, offer additional examples, or regroup students for targeted support – all while instruction is still happening.
Summative evaluation: measuring overall achievement
Summative evaluation takes place at the end of an instructional period – after the learning has been completed – to evaluate how well students have achieved the stated learning objectives. Unlike formative evaluation, it focuses on final results rather than ongoing progress. It is used to assign grades, certify mastery, determine program effectiveness, and make decisions about course continuation or improvement.
Key characteristics of summative evaluation
Summative evaluations are typically high-stakes and formally graded. They capture a cumulative picture of what students have learned over a defined period. Common summative techniques include:
- Final examinations – comprehensive tests covering an entire unit or course
- End-of-term projects or presentations – tasks that require students to apply skills in integrated ways
- Portfolio reviews – collections of work over time that demonstrate growth and mastery
- Standardised tests – external assessments against defined benchmarks
Well-designed summative assessments go beyond testing memorization – they challenge students to apply skills in authentic contexts. When crafted thoughtfully, they promote learning rather than simply measure it. Importantly, summative results are not just useful for students: at the program level, they help administrators and curriculum designers identify trends, evaluate instructional effectiveness, and allocate resources more strategically.
Using formative and summative evaluation together
These two forms of evaluation are most powerful when used in tandem. Educational program specialists point out that relying only on formative evaluation means you never gain a comprehensive look at program outcomes, while relying solely on summative evaluation means leaving opportunities for improvement untapped. Think of formative assessment as the rehearsal and summative assessment as the performance – students need both guided practice with feedback and clear opportunities to demonstrate mastery.
Diagnostic evaluation: knowing where students start
Before either formative or summative evaluation can be effective, it helps to understand where students are starting from. Diagnostic evaluation occurs before instruction begins. Its purpose is to identify students’ existing knowledge, strengths, and weaknesses so that instruction can be designed accordingly. A diagnostic quiz at the start of a course, for instance, might reveal which foundational concepts students have already mastered – allowing an instructor to skip ahead, or alternatively, to build in remediation for gaps before teaching new content.
Diagnostic evaluation is not about grading – it is about gathering baseline data. When used thoughtfully, it ensures that instruction meets learners where they actually are rather than where the curriculum assumes they are.
Sources of evaluation: who provides the feedback?
Evaluation does not come from a single source. To build a comprehensive picture of instructional effectiveness, educators draw on multiple perspectives. Three key sources are self-assessment, peer evaluation, and teacher evaluation – each offering a distinct vantage point that, when combined, produces richer, more actionable feedback.
Self-assessment
Self-assessment asks learners – and instructors – to reflect on their own performance. For students, it means critically examining their own work against defined criteria, identifying what they understand well and where they need more effort. Self-assessment can be used at all stages of learning: before a lesson to activate prior knowledge, during learning to monitor progress, and after a task to evaluate quality of output. When students engage in this kind of structured self-reflection, they develop metacognition – the ability to think about their own thinking – and become more independent, self-directed learners.
For instructors, self-assessment involves reflecting on teaching goals, challenges, and accomplishments. Formats vary and may include reflective statements, activity reports, goal-setting exercises, or structured rubrics. The Danielson Framework for teaching, for example, identifies critical reflection as the mechanism through which teachers assess the effectiveness of their work and take steps to improve it. The limitation of self-assessment, whether for students or teachers, is the risk of bias – people may overestimate their effectiveness or avoid acknowledging weaknesses. This is why self-assessment works best when combined with external feedback sources.
Peer evaluation
Peer evaluation is a collaborative strategy in which students critically examine the work of their classmates and provide structured, constructive feedback. It is related to self-assessment but adds an external layer – students must apply course criteria to someone else’s work, which in turn deepens their understanding of what quality work actually looks like. The process builds collaborative skills, encourages honest dialogue, and helps students develop the evaluative thinking they will use throughout academic and professional life.
For peer evaluation to be effective, it must be structured. Students need clear rubrics, explicit instructions, and practice before they can evaluate peers meaningfully. Research from the New South Wales Department of Education recommends involving students in defining success criteria, using exemplar work to make quality visible, and providing sentence starters to help students frame feedback constructively and respectfully.
Among teachers, peer evaluation takes the form of peer observation and review. Peer review of teaching is primarily formative – its main purpose is to provide faculty with meaningful feedback to help them set goals and improve their instructional abilities. Faculty are uniquely positioned to evaluate aspects of teaching that students cannot: the accuracy of content, the currency of course materials, the appropriateness of rigor, and contributions to curriculum development. When implemented well, peer review fosters a culture of collegial dialogue and shared professional growth.
Teacher evaluation
Teacher evaluation is the process of reviewing how effectively educators support student learning and development. It encompasses classroom observations (both formal and informal), reviews of teaching materials, student feedback surveys, performance-based assessments, and analysis of student outcomes. The primary goal of teacher evaluation is to promote accountability, enhance instructional quality, and support ongoing professional development – not simply to judge performance.
In the context of educational program evaluation, teacher evaluation contributes data that self-assessment and peer feedback alone cannot provide. Formal observations, for instance, allow evaluators to assess specific teaching behaviors against established standards or frameworks. Teaching evaluation frameworks, such as those guided by principles of excellence in teaching, use structured rubrics across multiple performance levels – from mastery to areas needing immediate improvement – to give teachers actionable, evidence-based feedback.
Why combining multiple evaluation sources matters
No single source of evaluation gives a complete picture. Student self-assessment surfaces metacognitive awareness but may lack objectivity. Peer evaluation builds collaborative skills but requires careful scaffolding to be rigorous. Teacher evaluation provides expert judgment but may miss dimensions of the student experience. Research on evaluation techniques consistently finds that combining multiple methods and sources produces the most valid and actionable picture of educational effectiveness. When all three sources – self, peer, and teacher – are integrated alongside both formative and summative data, evaluation moves from being a bureaucratic requirement to a genuine engine of instructional improvement.
In practical terms, this might look like a course where students complete self-reflection forms after each module (formative, self-source), peer-review each other’s major assignment drafts (formative, peer source), sit a comprehensive final project (summative), and where the instructor collects mid-semester student feedback alongside their own structured self-evaluation. Each piece of data informs the next cycle of teaching and course design.
Norm-referenced vs. criterion-referenced evaluation
Beyond timing and source, evaluations also differ in how they interpret performance. Criterion-referenced evaluation measures a student’s performance against a fixed standard or set of learning objectives – it describes what a learner can and cannot do, independent of how others perform. Norm-referenced evaluation, by contrast, compares a student’s performance relative to peers who took the same assessment, emphasising rank and percentile standing. In educational program design, criterion-referenced approaches are generally more aligned with learning outcome-focused evaluation, since the goal is mastery of specific skills rather than rank among a cohort.
Evaluation in the context of instructional improvement
Ultimately, the purpose of evaluation in educational programs is not to label or sort students – it is to generate the feedback loop that makes teaching and learning continuously better. As instructional researchers Hanna and Dettmer (2004) argued, educators should develop a range of assessment strategies that match all aspects of their instructional plans from the very beginning of a course – not as an afterthought. When evaluation is built into program design intentionally, and when data from multiple types and sources is used to inform decisions, educational programs become more responsive, more equitable, and more effective at achieving their intended goals.
What do you think? If you were designing a course from scratch, how would you decide which combination of evaluation types and sources would give you the most accurate picture of student learning? And to what extent do you think self-assessment alone is reliable enough to drive instructional decisions, without external input from peers or instructors?
References
- https://pmc.ncbi.nlm.nih.gov/articles/PMC9468254/
- https://poorvucenter.yale.edu/teaching/teaching-resource-library/formative-summative-assessments
- https://www.formative.com/read/formative-vs-summative
- https://www.k-state.edu/assessment/toolkit/basics/formativesummative.html
- https://distancelearning.institute/instructional-design/types-of-evaluation-in-education/
- https://ceop.ku.edu/taking-our-programs-end-zone-formative-v-summative-evaluation
- https://www.21kschool.com/us/blog/types-of-evaluation-in-education/
- https://www.niu.edu/citl/resources/guides/instructional-guide/peer-and-self-assessment.shtml
- https://citt.ufl.edu/resources/assessing-student-learning/designing-effective-peer-and-self-assessment/
- https://teaching.pitt.edu/resources/assessment-of-teaching-self-assessment/
- https://www.uwlax.edu/catl/guides/teaching-improvement-guide/how-can-i-improve/peer-evaluation/
- https://education.nsw.gov.au/teaching-and-learning/professional-learning/teacher-quality-and-accreditation/strong-start-great-teachers/refining-practice/peer-and-self-assessment-for-students
- https://pmc.ncbi.nlm.nih.gov/articles/PMC2384183/
- https://www.educationadvanced.com/blog/7-examples-of-teacher-evaluation-methods
- https://teaching.utk.edu/assessment/teaching-evaluation-frameworks/
- https://www.researchgate.net/publication/333633265_Formative_and_Summative_Evaluation_Techniques_for_Improvement_of_Learning_Process
- https://humanperitus.in/types-of-evaluation/
- https://www.niu.edu/citl/resources/guides/instructional-guide/formative-and-summative-assessment.shtml
Leave a Reply