How do educators know whether a curriculum is truly working? Not just whether students are passing tests, but whether the entire educational program – its goals, design, delivery, and outcomes – is doing what it is supposed to do? This is the central question that curriculum evaluation tries to answer. Over the decades, researchers and educators have developed a range of structured models to guide this process, each offering a different lens through which to examine a curriculum’s effectiveness. Understanding these models is essential for anyone involved in curriculum design, teaching, or educational administration.
Table of Contents
- What curriculum evaluation actually means
- The Tyler model: evaluation through objectives
- How the Tyler model works in practice
- Stake’s countenance model: a broader view of curriculum reality
- Congruence and contingency: the model’s two analytical lenses
- The CIPP model: evaluation for decision-making and improvement
- Breaking down the four components
- CIPP in formative and summative contexts
- Scriven’s goal-free evaluation: letting outcomes speak for themselves
- Provus’s discrepancy evaluation model: pinpointing the gap
- Eisner’s connoisseurship model: evaluating what numbers can’t capture
- Choosing the right model: a practical framework
What curriculum evaluation actually means
Curriculum evaluation is far more than grading students or measuring test scores. According to the Indian Journal of Continuing Nursing Education, it is a systematic process of gathering and applying descriptive and judgmental information about a program’s merit, worth, and significance. In simpler terms, it asks: Is this curriculum achieving what it set out to achieve? Are the goals relevant? Is the delivery effective? And crucially, what needs to change?
The answers to those questions depend heavily on the model used to guide the evaluation. Different models prioritize different things – some focus on whether stated objectives were met, others examine the broader context and process, and a few deliberately ignore stated goals altogether to catch outcomes that might otherwise be missed. A well-chosen model provides the structure, language, and procedures needed to turn raw observations into actionable insight.
The Tyler model: evaluation through objectives
Developed by Ralph Tyler in 1949, the Tyler Model is the oldest and most widely cited framework for curriculum evaluation. Its logic is straightforward: a curriculum should begin with clearly defined educational objectives, and evaluation should determine whether those objectives have been achieved. Tyler proposed four guiding questions that structure the entire process – What educational purposes should the program seek? What learning experiences can achieve these purposes? How should these experiences be organized? And how can we determine whether the purposes are being attained?
How the Tyler model works in practice
In practice, evaluators using the Tyler Model start by identifying measurable behavioral objectives – specific, observable outcomes that students should be able to demonstrate after instruction. Learning experiences are then selected and organized to support those objectives. Finally, assessment data is compared against the objectives to judge success. Tyler’s objective-based approach is particularly effective for programs with well-defined, measurable outcomes such as basic skills training or professional competency programs.
The model’s biggest strength is its clarity. The step-by-step structure makes it easy to implement and communicate, which is why it remains widely used decades after its introduction. However, it has well-documented limitations. Critics point out that it tends to favor easily measurable outcomes while potentially overlooking harder-to-assess results like critical thinking, creativity, or social development. It also provides little guidance on how to handle unexpected outcomes or individual learner differences. In short, what the Tyler Model gains in simplicity, it sometimes sacrifices in depth.
Stake’s countenance model: a broader view of curriculum reality
Robert Stake introduced the Countenance Model in 1967 as a response to the perceived narrowness of objectives-based evaluation. Where Tyler focuses on outcomes alone, Stake’s model examines three interconnected categories of data – antecedents, transactions, and outcomes.
Antecedents refer to the conditions that exist before instruction begins: student prior knowledge, teacher backgrounds, available resources, and institutional context. Transactions capture the actual teaching and learning interactions that take place during curriculum delivery – the dynamic exchanges between students, teachers, and materials. Outcomes encompass all consequences of the educational process, both intended and unintended, immediate and long-term.
Congruence and contingency: the model’s two analytical lenses
What makes the Countenance Model distinctive is its dual analytical framework. Congruence asks whether what was intended actually happened – do the observed antecedents, transactions, and outcomes align with what was planned? Contingency explores the relationships among the three data categories – how do prior conditions shape instructional transactions, and how do those transactions drive outcomes?
Stake also emphasized that evaluation involves both description (documenting what is actually happening) and judgment (assessing the merit and worth of what is happening from multiple stakeholder perspectives). This dual emphasis represents a significant departure from Tyler’s more technical, goal-focused assessment. Stake’s model is especially valued in diverse educational settings where context matters enormously and where curriculum effectiveness cannot be reduced to a single metric. Its main drawback is that it is time-intensive and requires experienced evaluators to apply effectively.
The CIPP model: evaluation for decision-making and improvement
The CIPP Model, developed by Daniel Stufflebeam in the late 1960s, is arguably the most comprehensive curriculum evaluation framework in wide use today. The acronym stands for Context, Input, Process, and Product, and each component addresses a different dimension of the curriculum and a different set of decisions facing educators. It emerged as a direct alternative to objectives-driven approaches that dominated evaluation thinking at the time.
What sets the CIPP Model apart from Tyler’s framework is its fundamental orientation. As Stufflebeam himself stated, the purpose of evaluation in this model is “not to prove but to improve.” This philosophy makes CIPP inherently forward-looking rather than simply retrospective.
Breaking down the four components
Context evaluation is the starting point. It examines the environment in which the curriculum operates – the needs of learners, existing resources, institutional goals, community expectations, and any problems or opportunities present in the setting. The question it answers is: What needs to be done? This stage helps evaluators identify the specific challenges and contextual factors that could influence a program’s success.
Input evaluation follows by examining the resources and strategies available to achieve those identified needs. It looks at curriculum design, staffing, materials, technology, and budget. The central question here is: How should we accomplish our goals? This phase ensures that the plan for curriculum delivery is both viable and well-resourced before implementation begins.
Process evaluation monitors the actual delivery of the curriculum once it is underway. It tracks whether plans are being implemented as designed, identifies problems in real time, and provides feedback that allows for mid-course corrections. This component helps decision-makers determine whether the curriculum is being implemented as designed, and it plays a vital formative role throughout the program’s lifecycle.
Product evaluation assesses the outcomes of the curriculum – both the intended results and any unintended ones, over both the short and long term. It asks whether the curriculum succeeded, and whether the program should continue, be modified, or discontinued. Together, the four components create a continuous feedback loop that supports ongoing curriculum refinement.
CIPP in formative and summative contexts
The CIPP model operates in both formative and summative modes. In its formative mode, it guides evaluators by asking: What needs to be done? How should it be done? Is it being done? Is it succeeding? In its summative mode, it revisits those same questions retrospectively to render an overall judgment on the program. This dual capacity makes CIPP one of the most versatile tools available to curriculum evaluators. Its primary limitation is complexity – implementing all four components thoroughly requires substantial time, expertise, and resources, which can be challenging for smaller institutions.
Scriven’s goal-free evaluation: letting outcomes speak for themselves
Michael Scriven’s Goal-Free Evaluation (GFE) takes a counterintuitive approach: the evaluator deliberately avoids learning about the stated goals of the curriculum before or during the evaluation. According to Scriven, the purpose is to find out what a program is actually doing without being cued by what it is trying to do. If the program is achieving its stated goals, those achievements will still appear in the data. If it is producing significant unintended consequences – positive or negative – those will surface too, without being filtered out by a goal-focused lens.
This approach serves as a powerful corrective to evaluator bias. When evaluators know what a program is supposed to achieve, they may unconsciously focus their data collection on those targets and overlook important side effects. The goal-free model broadens the evaluator’s attention to a wider range of program outcomes, making it a useful complement to goal-based approaches rather than a complete replacement. Scriven himself acknowledged that GFE is not appropriate as a standalone strategy in every context, particularly when a client is specifically interested in goal attainment. It works best alongside more structured models like Tyler’s or CIPP.
Provus’s discrepancy evaluation model: pinpointing the gap
Malcolm Provus introduced the Discrepancy Evaluation Model (DEM) in 1971 with a clear conceptual focus: evaluation is the process of identifying the gap between what a program was designed to do and what it is actually doing. Provus defined evaluation as the process of agreeing upon program standards, determining whether a discrepancy exists between the program and those standards, and using that discrepancy information to identify weaknesses.
The DEM proceeds through five stages: program definition (establishing the design and standards), installation (verifying that the program has been set up as designed), process (assessing the delivery and its relationship to intended changes), product (evaluating whether the program met its goals), and cost (assessing efficiency relative to comparable programs). At each stage, a comparison is made between the actual state of the program and the standard set in stage one. Any gap becomes the focus of corrective action.
The DEM is particularly effective during curriculum development and early implementation phases, when systematic problem identification is more valuable than final judgment. This model excels at identifying and correcting problems during the course of program development, making it a practical tool for iterative curriculum improvement.
Eisner’s connoisseurship model: evaluating what numbers can’t capture
Elliot Eisner proposed the Connoisseurship Model as a response to the overemphasis on measurable outcomes in curriculum evaluation. Drawing on the metaphor of an art critic – someone who possesses deep, cultivated knowledge that allows them to perceive and articulate quality – Eisner argued that expert educational judgment has a legitimate and necessary role in evaluation. The model involves two interconnected processes: connoisseurship, the private act of appreciating and perceiving the qualities of a curriculum, and criticism, the public act of articulating that appreciation in ways that help others understand what is happening in the classroom and why it matters.
This model is especially well-suited to evaluating the experiential and qualitative dimensions of learning – the texture of classroom interactions, the richness of learning materials, the intellectual climate of a school – that conventional outcome measures often miss. Its limitation is the reliance on expert judgment, which raises questions about objectivity and consistency, and it may carry less weight with stakeholders who prefer quantitative data.
Choosing the right model: a practical framework
No single evaluation model is universally superior. The most appropriate choice depends on the specific purpose of the evaluation, the nature of the curriculum, the questions being asked, and the resources available. As a practical guide:
The Tyler Model works best when objectives are clear, measurable, and stable – for instance, in professional training programs or basic skills curricula where outcomes can be defined precisely. The CIPP Model is the best choice for comprehensive program reviews, complex curricula, or situations where ongoing improvement decisions need to be informed by data at every stage. Stake’s Countenance Model is most valuable when contextual factors and stakeholder perspectives are central to understanding curriculum effectiveness. Scriven’s Goal-Free Evaluation is most useful as a supplementary lens to catch unintended outcomes alongside other approaches. The Discrepancy Model suits early-stage programs where identifying and closing implementation gaps is the priority. And Eisner’s Connoisseurship Model fills the gap when qualitative depth and experiential insight are essential to a complete evaluation.
In practice, combining principles from different models often produces the strongest evaluations. For example, using a goal-free lens within a CIPP framework can reveal unintended outcomes that a purely component-focused evaluation might miss. The key is always to match the model – or combination of models – to the specific evaluative questions at hand.
What do you think? If you were evaluating a curriculum at your school or institution, which model would you choose first – and what factors would drive that decision? And do you think modern curricula, with their emphasis on competencies and holistic learning, can be adequately captured by objective-based models like Tyler’s, or do they demand more context-sensitive approaches like CIPP or Stake’s?
References
- https://journals.lww.com/ijcn/fulltext/2017/18020/curriculum_evaluation__using_the_context,_input,.1.aspx
- https://zoneofeducation.org/curriculum-evaluation-models/
- https://distancelearning.institute/curriculum-development/effective-curriculum-evaluation-models/
- https://iiardjournals.org/get/IJSSMR/VOL.%2011%20NO.%207%202025/Strenghts%20and%20Weaknesses%20of%20Evaluation%20196-205.pdf
- https://link.springer.com/chapter/10.1007/978-94-009-6669-7_7
- https://ijrehc.com/uploads2024/ijrehc05_142.pdf
- https://en.wikipedia.org/wiki/Goal-free_evaluation
- https://www.researchgate.net/publication/324891193_EVALUATION_MODELS_IN_EDUCATIONAL_PROGRAM_STRENGTHS_AND_WEAKNESSES
- https://healthnet.org.np/training/msoffice/powerpoint/ww196.htm
Leave a Reply