Criterion-referenced tests (CRTs) are one of the most purposeful tools an educator can use – and yet, many teachers construct them without a clear process, which undermines their effectiveness. Unlike norm-referenced tests that rank students against one another, a CRT is designed to measure whether each student has achieved a specific, pre-defined set of learning outcomes. The question it answers is simple: Has this student mastered what was taught? Constructing a CRT well requires a structured, step-by-step approach – one that begins long before you write a single test question.
Table of Contents
- What makes a criterion-referenced test different
- Step 1: Identify the subject area
- Step 2: Delineate the testing domain
- Step 3: Specify the learning objectives
- Step 4: Construct the test items
- Step 5: Ensure validity
- Step 6: Ensure reliability
- Step 7: Set the performance standard (cut score)
- Step 8: Review, pilot, and revise
- CRTs as a tool for teaching, not just testing
What makes a criterion-referenced test different
Before diving into construction, it helps to be clear about what a CRT actually is. First introduced by educational psychologist Robert Glaser in 1963, criterion-referenced testing evaluates student performance against a fixed domain of knowledge or skills – not against the performance of other test-takers. A student who scores 85% on a well-constructed CRT has demonstrated mastery of 85% of the defined content domain. That result is meaningful on its own, without any comparison to classmates.
This is fundamentally different from norm-referenced testing. CRTs focus solely on whether each student meets or exceeds established standards, providing detailed insights into individual performance – including where they are strong and where they still need support. This makes them particularly valuable for formative assessment and instructional planning.
Common real-world examples include driving tests, citizenship exams, and professional licensure assessments. The goal is not to find out who is best – it’s to find out who is ready.
Step 1: Identify the subject area
Every CRT begins by identifying what area of knowledge or skill is being tested. This is not simply choosing a broad topic – it requires aligning the subject area with the curriculum and the outcomes you actually intend students to achieve.
Subject areas can range from broad themes to highly focused skills. For example, “photosynthesis in plants” is too broad if you only taught one aspect of it. It is not enough to say that a criterion-referenced test is going to cover fractions – it cannot be determined if the test will cover only adding and subtracting, or multiplying and dividing, or all four concepts. Precision at this stage prevents ambiguity throughout the rest of the construction process.
Ask yourself: What did students actually learn in this instructional period? What content areas were taught, and to what depth? The answers define the boundaries of your test’s subject area.
Step 2: Delineate the testing domain
Once the subject area is established, the next step is to delineate the testing domain – the specific, well-defined body of knowledge, skills, or behaviors that the test will cover.
The foundation of criterion-referenced testing is a clearly defined content domain. This requires explicit specification of learning objectives, competencies, or performance standards. The domain must be narrow enough to be well-defined, yet comprehensive enough to meaningfully represent the skill or knowledge area.
Domain delineation is critical because it determines which test items are appropriate. Items in a CRT are selected based on whether they accurately reflect mastery of the content domain – not based on how well they differentiate between high and low performers. CRTs select items that accurately reflect mastery of the content domain, correctly answered by those who know the material and missed by those who don’t, regardless of how many test-takers fall into each category.
Think of the domain as a map: your test items must represent all the important territories within it. Leaving out significant sections creates a test that is technically complete but practically incomplete.
Step 3: Specify the learning objectives
Learning objectives are the heart of any criterion-referenced test. They define exactly what students should know or be able to do, and every test item must trace back directly to one of them.
Effective learning objectives are SMART: Specific, Measurable, Achievable, Relevant, and Time-bound. A well-structured objective typically includes an action verb, a description of the skill or knowledge to be acquired, and a criterion for success. For example: “By the end of this unit, students will be able to identify three causes of soil erosion and explain how each contributes to land degradation.”
The choice of action verb matters enormously. Bloom’s Taxonomy offers a widely-used framework for writing objectives at the right cognitive level. The taxonomy moves through six levels – Remembering, Understanding, Applying, Analyzing, Evaluating, and Creating – and each level signals both the depth of learning expected and the type of assessment item needed. A lower-level objective might use the verb “identify” or “list,” while a higher-level objective might call for “analyze” or “evaluate.”
When teaching lower-division or introductory courses, many objectives will target lower-order skills; more advanced courses should assess students at higher levels of the taxonomy. Matching objectives to the correct cognitive level ensures the test is appropriate for the learners being assessed.
Avoid vague verbs like “understand” or “appreciate” – these cannot be directly measured. As a practical guideline, aim for 4-6 objectives per unit, written so they complete the phrase: “Upon completion of this course, students will be able to…”
Step 4: Construct the test items
With a clearly delineated domain and specific objectives in hand, you can now write the actual test items. Each item must directly assess one of the learning objectives – this alignment is non-negotiable in a criterion-referenced test.
A well-constructed CRT should have an adequate number of items to cover each competency. Ideally, there should be at least five questions addressing each concept, particularly when the test carries significant consequences such as grade promotion.
Item format should match what the objective requires. Tests can be constructed with open-ended questions and tasks that require students to use higher-level cognitive skills such as critical thinking, problem solving, and analysis, while multiple-choice and true-false formats are better suited to recall and recognition. Use the format that best allows students to demonstrate the targeted learning outcome – not the format that is easiest to grade.
When writing items, keep language clear and unambiguous. Avoid trick questions. The purpose is to find out whether students have learned the material, not to catch them off guard. Each item should have one defensibly correct answer or, for open-ended items, a clearly defined scoring rubric that specifies what a successful response looks like.
Step 5: Ensure validity
A test is only useful if it actually measures what it claims to measure. This is what validity means in assessment, and it is one of the two most important technical qualities a CRT must have.
There are three main types of validity to examine during construction:
Content validity asks whether the test covers all relevant content areas defined by the learning objectives. If students studied five causes of the French Revolution but the test only addresses two, content validity is compromised. For CRTs, the content validity approach involves systematically analyzing the degree to which test items measure what the teacher claims to test – often by placing items side-by-side with the course objectives and teaching materials.
Construct validity asks whether the test truly measures the intended learning outcome. A test claiming to assess critical thinking should not rely primarily on fact memorization – that would measure a different construct entirely.
Criterion-related validity asks whether performance on this test correlates appropriately with performance on other assessments of the same outcomes. Students who do well on this test should generally do well on other valid measures of the same content.
To safeguard validity, review each item against its corresponding objective. Have a colleague examine the test. Ask: Is there any item here that does not belong? Is there any objective that has no corresponding item?
Step 6: Ensure reliability
Reliability refers to the consistency of test results. A reliable test produces similar results when administered under similar conditions – the scores reflect student learning, not chance variation.
For any assessment test to be reliable, it is important to have a representation of the entire content as well as adequate sampling. A test based on only one or two narrow questions cannot reliably represent a student’s mastery of an entire domain. The more items you include per objective, the more reliable your assessment becomes.
Research has shown that criterion-referenced scaling generally results in higher inter-rater reliability than norm-referenced approaches – but this advantage only holds when the criteria are explicitly and precisely defined. Vague criteria invite inconsistent scoring, which undermines reliability.
Practical steps to improve reliability include: using unambiguous language in both questions and rubrics, piloting the test with a small group before full administration, and having more than one reviewer check the scoring guide for open-ended items.
Step 7: Set the performance standard (cut score)
A criterion-referenced test often requires a performance standard – the minimum level of achievement that indicates mastery. This is sometimes called a cut score.
It is important to understand that the criterion is not the cut score itself – the criterion is the domain of subject matter the test is designed to assess. The cut score is simply the decision point established for that context. For a classroom unit test, a teacher might set mastery at 75% correct. For a high-stakes professional licensure exam, a panel of subject-matter experts determines the cut score through a structured process.
The cut score should be set based on what level of performance genuinely indicates mastery of the objectives – not on a desire to pass or fail a certain percentage of students. This is one of the key distinctions between criterion-referenced and norm-referenced thinking.
Step 8: Review, pilot, and revise
No test is complete after the first draft. Before administering a CRT, it should go through a structured review process. Check each item for clarity, alignment with objectives, and freedom from bias. Seek feedback from a colleague familiar with the content area.
Where possible, conduct a pilot test with a small group of students. Analyze which items most students got wrong – not to make the test easier, but to determine whether the difficulty is due to unclear wording or genuinely unlearned material. Items that are confusing or poorly worded skew results and lower both validity and reliability.
Revision is not a sign of weakness in test construction – it is a sign of rigor. A properly constructed CRT allows you to classify the people who take it into clear groups: those who have mastered the content and those who have not. That classification is only trustworthy when the test itself has been carefully reviewed and refined.
CRTs as a tool for teaching, not just testing
A well-constructed criterion-referenced test does more than measure student achievement – it informs instruction. When most students fail a particular item, the test is telling the teacher something important: that objective was not yet fully learned, and instruction needs to revisit it. When results show widespread mastery, the teacher knows it is safe to move forward.
By identifying learning gaps, weaknesses, and areas of improvement through criterion-referenced tests, teachers can provide a supportive environment and design instruction for individual learning needs. This transforms assessment from a one-time event into an ongoing dialogue between teaching and learning.
CRTs are also closely linked to formative assessment practices. Because CRTs are usually used to measure short-term objectives, they tend to be formative rather than summative in nature – meaning they are most powerful when used during learning, not just at the end of it. Used this way, they become a navigational tool, guiding both students and teachers toward genuine mastery.
What do you think? When you design a test for your students, how deliberately do you trace each question back to a specific learning objective? And if your students consistently struggle with certain items, how does that change your approach to the next unit of instruction?
References
- https://www.edglossary.org/criterion-referenced-test/
- https://en.wikipedia.org/wiki/Criterion-referenced_test
- https://www.ebsco.com/research-starters/social-sciences-and-humanities/criterion-referenced-testing
- https://www.cogn-iq.org/learn/theory/criterion-referenced-test/
- https://www.teachfloor.com/blog/how-to-write-learning-objectives-using-blooms-taxonomy
- https://tips.uark.edu/using-blooms-taxonomy/
- https://tigerlearn.fhsu.edu/the-revised-blooms-taxonomy-as-a-framework-for-writing-learning-objectives/
- https://senate.ucsf.edu/course-actions/blooms-taxonomy
- https://teval.jalt.org/sites/default/files/18-1-29%20Brown%20Statistics%20Corner.pdf
- https://pmc.ncbi.nlm.nih.gov/articles/PMC10666833/
- https://pmc.ncbi.nlm.nih.gov/articles/PMC10498947/
- https://eric.ed.gov/?id=ED189168
- https://www.21kschool.com/us/blog/criterion-referenced-test/
- https://methods.sagepub.com/ency/edvol/encyclopedia-of-measurement-and-statistics/chpt/criterionreferenced-tests
Leave a Reply