When a student receives a test result, two very different questions can be asked about it: “How did this student perform compared to everyone else?” or “Did this student master what was taught?” These two questions sit at the heart of two major approaches to educational evaluation – norm-referenced evaluation and criterion-referenced evaluation. Both are widely used in schools and institutions around the world, but they serve distinct purposes, produce different kinds of information, and are best suited for different contexts. Understanding how they differ – and when to use each – is one of the most practical things an educator can know.
Table of Contents
- What is norm-referenced evaluation?
- Key characteristics of norm-referenced evaluation
- When norm-referenced evaluation is useful
- Limitations of norm-referenced evaluation
- What is criterion-referenced evaluation?
- Key characteristics of criterion-referenced evaluation
- When criterion-referenced evaluation is useful
- Limitations of criterion-referenced evaluation
- A direct comparison: norm-referenced vs. criterion-referenced evaluation
- Can both approaches be used together?
- Choosing the right approach for the right purpose
What is norm-referenced evaluation?
Norm-referenced evaluation measures a student’s performance by comparing it to the performance of a defined group – called the norm group – who have taken the same assessment. The goal is not to determine what a student knows in absolute terms, but to determine where they stand relative to their peers.
According to EBSCO Research, norm-referenced tests rank students based on their scores, allowing educators to understand how an individual’s performance compares to others who have taken the same assessment. Results are typically reported as percentile ranks, stanines, or normal curve equivalents – figures that only have meaning in relation to how the group performed overall.
Well-known examples include the SAT, the ACT, and most IQ tests. These assessments are built around the assumption that scores will naturally fall along a bell curve – with most students clustering in the middle and fewer students at the high and low ends. By design, it is statistically impossible for all students to score above the 50th percentile on a norm-referenced measure.
Key characteristics of norm-referenced evaluation
Three defining features shape how norm-referenced evaluation works in practice:
Relative comparison: A student’s score is meaningful only in relation to the group. If a student answers 70% of questions correctly but most peers answer 85%, that student’s score falls below average – regardless of how much they actually know.
Rank ordering: A study published in PubMed Central notes that norm-referenced scales require raters to determine where individuals reside relative to others without specific regard to a competency continuum. The result is a ranking, not a diagnosis of what a student knows or can do.
Normal distribution assumption: These evaluations assume that abilities naturally spread across a population. There will always be top, middle, and bottom performers – regardless of how well instruction has worked.
When norm-referenced evaluation is useful
Norm-referenced evaluation works well when the primary goal is selection or comparison. Research from Classtime identifies its advantages clearly: it is particularly useful for identifying high and low performers within a larger group, making it beneficial for college admissions, scholarship allocations, and other competitive scenarios where ranking is essential. It can also help identify exceptional students who might benefit from advanced programs or additional support. National-level standardized tests used for university entrance decisions are a classic application.
Limitations of norm-referenced evaluation
Despite its usefulness for ranking, norm-referenced evaluation has significant drawbacks in instructional contexts. A report in PubMed notes that the norm-referenced approach is often insensitive to instruction – it provides information about relative strengths and weaknesses compared to peers but does not estimate the absolute level of performance achieved. In short, a teacher cannot use a percentile rank to determine what specific skills a student needs to develop next. It tells you where a student stands, not what they know.
What is criterion-referenced evaluation?
Criterion-referenced evaluation takes an entirely different approach. Instead of comparing students to each other, it measures each student’s performance against a predetermined standard or set of learning objectives. The question is not “How did this student compare to the group?” but “Has this student mastered what was expected?”
EBSCO Research describes criterion-referenced testing as an assessment approach designed to evaluate what students know and can do based on specific educational outcomes, focusing solely on whether each student meets or exceeds established standards. A familiar real-world example is a driving test – you either demonstrate safe driving competencies and pass, or you don’t. It doesn’t matter how many others passed or failed on the same day.
In classroom settings, criterion-referenced evaluation might mean a student needs to demonstrate they can solve a set of algebra problems, correctly identify parts of speech, or describe the stages of the water cycle – and their result reflects how well they’ve met those specific targets, not how they compare to classmates.
Key characteristics of criterion-referenced evaluation
Fixed standards: The Center for Applied Linguistics explains that with criterion-referenced scores, performance is compared to a preset standard. The performance of other students is entirely irrelevant – it is possible for all students to achieve the expected level, and that would be a positive outcome, not a statistical anomaly.
Mastery focus: According to research published in the Journal of Intelligence, criterion-referenced approaches allow educators to know what a student can do in terms of achieved levels of mastery, rather than merely in relative terms to what others can do. This is critical for skills-based learning where competency must be demonstrated before progressing.
Diagnostic value: Because results are tied to specific objectives, Extramarks notes that criterion-referenced assessments show exactly what skills a student has mastered and what still needs work – providing clear, actionable insights rather than a relative ranking.
When criterion-referenced evaluation is useful
Criterion-referenced evaluation is the preferred choice when the goal is ensuring that all students achieve essential competencies. The Distance Learning Institute points out that professional licensing exams, end-of-unit tests, and certification programs typically use criterion-referenced scoring because the goal is determining competency, not ranking candidates. Medical licensing, bar exams, and teacher certification tests all follow this logic – every candidate must meet the standard, regardless of how others perform.
In classroom instruction, it supports formative assessment cycles. Teachers can identify which students have mastered a concept and which need further support, then adjust instruction accordingly. Classtime’s assessment research affirms that for teachers, criterion-referenced assessments provide a roadmap for instruction – helping identify what students already know and what they still need to learn, allowing teachers to tailor their approach to meet individual needs.
Limitations of criterion-referenced evaluation
Criterion-referenced evaluation is not without challenges. Setting appropriate cut scores – the threshold between “mastered” and “not yet mastered” – is genuinely difficult. EBSCO Research highlights that since each state or institution can choose its own standard, there can be significant disparity in what “proficiency” actually means across different regions. There is also the concern that if standards are set too low, criterion-referenced assessments can lead to grade inflation – in theory, all students could pass, which some critics argue reduces the meaningfulness of the achievement.
A direct comparison: norm-referenced vs. criterion-referenced evaluation
To put the differences in concrete terms, consider these contrasting dimensions:
Purpose: Norm-referenced evaluation is designed to rank and differentiate students. Criterion-referenced evaluation is designed to measure mastery against defined goals.
Score interpretation: In norm-referenced evaluation, a score of “70th percentile” means the student outperformed 70% of the comparison group. In criterion-referenced evaluation, a score of “70%” means the student correctly demonstrated 70% of the expected skills – two very different pieces of information.
Possibility of universal success: In norm-referenced systems, it is mathematically impossible for all students to perform above average. In criterion-referenced systems, it is entirely possible – and desirable – for every student to meet the standard.
Instructional usefulness: NWEA’s assessment research makes the point clearly: what ultimately matters with any assessment is whether it can effectively guide instruction. Criterion-referenced data is generally far more actionable for classroom teachers, while norm-referenced data is more useful for system-level comparisons and selection decisions.
Can both approaches be used together?
Yes – and in practice, many assessment systems do exactly that. Renaissance Learning explains that many universal screeners report both types of scores simultaneously: risk categories based on criterion-referenced proficiency levels, alongside national percentile scores for norm-referenced comparison. The combination makes these tools suitable for multiple purposes – from identifying students who need additional support to evaluating program effectiveness.
The Distance Learning Institute describes a practical model: progress monitoring tools in Response to Intervention (RTI) programs use norm-referenced benchmarks to identify students who are at risk, then use criterion-referenced measures to track whether specific intervention goals are being met. Teachers can simultaneously monitor whether a student is closing gaps relative to peers and whether they are mastering targeted skills.
A high school might use criterion-referenced assessments for most day-to-day classroom tests and unit exams – where the goal is confirming mastery – while also administering norm-referenced standardized tests for college admissions and scholarship eligibility decisions. Neither approach is universally superior; the right choice depends on the question being asked.
Choosing the right approach for the right purpose
The guiding principle is straightforward: match the evaluation method to the goal. If the purpose is to select a fixed number of candidates from a large pool – scholarship recipients, university admissions, competitive program placement – norm-referenced evaluation provides the ranking data needed to make those decisions. If the purpose is to confirm that students have achieved specific learning targets, to plan instruction, or to certify professional competence, criterion-referenced evaluation is more appropriate.
Vaia’s educational research notes that criterion-referenced assessments align well with mastery-based learning and standards-based grading – both of which prioritize depth of understanding over relative performance. As educational systems increasingly move toward competency-based models, criterion-referenced evaluation is gaining ground as the primary tool for measuring student growth and instructional effectiveness.
What matters most, regardless of the method, is that evaluation results are used thoughtfully – to inform instruction, support students, and improve learning – rather than simply to sort and label.
What do you think? If you had to design an assessment system for your school or classroom from scratch, which approach would form the foundation – and why? And in contexts where both norm-referenced and criterion-referenced data are available, how should teachers decide which information to act on first?
References
- https://www.ebsco.com/research-starters/social-sciences-and-humanities/norm-referenced-testing
- https://pmc.ncbi.nlm.nih.gov/articles/PMC10498947/
- https://www.classtime.com/en/norm-referenced-vs-criterion-referenced-assessment
- https://pubmed.ncbi.nlm.nih.gov/2586296/
- https://www.ebsco.com/research-starters/social-sciences-and-humanities/criterion-referenced-testing
- https://www.cal.org/twi/EvalToolkit/5when2usetests.htm
- https://pmc.ncbi.nlm.nih.gov/articles/PMC9397071/
- https://www.extramarks.com/blogs/schools/criterion-referenced-assessments/
- https://distancelearning.institute/instructional-design/evaluation-measures-in-education/
- https://www.classtime.com/en/criterion-referenced-assessment
- https://www.nwea.org/blog/2024/norm-vs-criterion-referenced-in-assessment-what-you-need-to-know/
- https://www.renaissance.com/2018/07/11/blog-criterion-referenced-tests-norm-referenced-tests/
- https://www.vaia.com/en-us/explanations/english/tesol-english/criterion-referenced/
Leave a Reply