Assessment is at the heart of every higher education system. How a university chooses to test its students – and how it interprets those scores – has real consequences for learning, grading, and academic decisions. Two frameworks sit at the center of this conversation: norm-referenced testing (NRT) and criterion-referenced testing (CRT). While both aim to evaluate student performance, they do so from fundamentally different angles. Understanding the distinction is not just an academic exercise – it shapes how educators design assessments, how students are graded, and how institutions make decisions about placement, progression, and certification.
Table of Contents
- What is norm-referenced testing?
- Common examples of NRT in higher education
- What is criterion-referenced testing?
- Common examples of CRT in higher education
- Key differences between NRT and CRT
- Purpose
- Assessment criteria
- Score interpretation
- Score outcomes
- Pros and cons of norm-referenced testing
- Advantages
- Limitations
- Pros and cons of criterion-referenced testing
- Advantages
- Limitations
- Which approach is right for higher education?
What is norm-referenced testing?
Norm-referenced testing is a method of assessment that evaluates a student’s performance by comparing it to the performance of a defined group – typically called the norm group. The goal is not to measure what an individual knows in absolute terms, but to determine where that individual stands relative to others who took the same test.
Before an NRT is released for public use, it is administered to a representative sample of students. The scores from this sample establish the “norm.” When other students take the test later, their scores are interpreted against that baseline. Results are typically reported as percentile ranks, stanines, or normal curve equivalents – all of which communicate relative standing rather than absolute achievement.
A defining statistical feature of NRT is that, by design, it is impossible for all students to score above the 50th percentile. Norms are structured so that 25% of students will always fall in the bottom quartile, regardless of how much the group actually knows.
Common examples of NRT in higher education
The most recognizable examples of norm-referenced assessments in the context of higher education admission and placement include the SAT, ACT, GRE, and GMAT. These tests are designed to rank applicants relative to one another, helping institutions make comparative decisions during selective admissions. University entrance examinations that use a competitive cutoff based on class-wide performance follow the same logic. Even some internal university assessments – such as bell-curve grading systems – operate on norm-referenced principles.
What is criterion-referenced testing?
Criterion-referenced testing takes a different approach entirely. Rather than comparing students to each other, CRT measures what test takers know and can do relative to a predetermined performance level on a specified set of educational goals and outcomes. In other words, the benchmark is a fixed standard – not the performance of the peer group.
A student either meets the criterion or does not, and this determination is entirely independent of how others perform. On criterion-referenced assessments, it is entirely possible for all students to achieve the minimum achievement expectation – there is no built-in requirement for a portion of students to fall below a threshold.
Scores in CRT are typically reported as performance categories such as proficient, basic, below basic, or as pass/fail outcomes tied to a defined cut score. The focus is on what the student has mastered, not where they rank.
Common examples of CRT in higher education
Criterion-referenced assessments are widespread in higher education, especially in professional and vocational programs. Licensing and certification exams – such as bar exams for lawyers, medical licensing exams (like the USMLE for doctors or NCLEX for nurses), and engineering board exams – are among the clearest examples. The NCLEX exam for nurses in the United States is a well-known example of a criterion-referenced assessment, with a fixed cut score designed to confirm that candidates meet minimum competency standards for safe practice. End-of-semester subject exams that require a student to score 50% or 60% to pass – regardless of the class average – also operate on criterion-referenced logic.
Key differences between NRT and CRT
While both testing approaches use the same raw scores as their starting point, the critical difference lies in the context in which a student’s score is interpreted. Here is a structured comparison of the two frameworks across the dimensions that matter most in higher education:
Purpose
NRT is designed to differentiate among students and produce a reliable rank order. Its primary use is in selection, placement, and comparison – for instance, identifying which students are eligible for a scholarship, a competitive program, or remedial support. CRT, by contrast, is designed to determine whether students have achieved specific learning objectives. Its purpose is mastery verification and instructional accountability.
Assessment criteria
In NRT, test content is selected based on how well it ranks students – items that nearly everyone gets right or wrong are often removed because they don’t help discriminate between high and low performers. In CRT, content is determined by how well it aligns with the intended learning outcomes. Items are chosen based on educational relevance, not their ability to spread scores across a distribution.
Score interpretation
Norm-referenced scores tell you how a student performed compared to others, but offer little specificity about the student’s strengths or weaknesses in terms of content. Criterion-referenced assessments, on the other hand, give more explicit information about a student’s level of achievement on specific content – but do not communicate how they performed relative to peers.
Score outcomes
NRT scores are relative – a student’s score means something only in relation to the group. CRT scores are absolute – a student’s score indicates whether they met a defined standard, regardless of what anyone else did. If a student scores 90% and the cut score is 80%, the criterion-referenced interpretation is that they passed. If the class average is 95%, the norm-referenced interpretation is that the same student performed below average – the same score, two entirely different meanings.
Pros and cons of norm-referenced testing
Advantages
Effective for ranking and selection: NRT is particularly useful when the goal is to identify top performers from a large group – such as national university entrance examinations or competitive scholarship programs. The primary strength of norm-referenced assessment is its ability to produce a rank order, making it very useful for selecting relatively high and low achievers among students.
Handles test difficulty variability: Because scores are compared within the group, if a test turns out to be too easy or too hard for a class, the norm-referenced comparison can still reflect levels of student achievement, since all students faced the same test under the same conditions.
Useful for program evaluation: At the institutional level, NRT data can help universities benchmark their students against national or international cohorts, informing decisions about curriculum quality and program standards.
Limitations
Tells you little about what students actually know: A significant disadvantage of norm-referenced assessment is that it gives little information about what a test-taker actually knows or can do, and it cannot measure students’ progress or learning outcomes.
Potential for cultural and demographic bias: Norm-referenced tests are structured around traditional Western values such as individual achievement, competitiveness, and emphasis on objectivity. If the norm group does not adequately represent the test-takers’ demographic background, the results can be misleading and unfair.
Competitive pressure: When grades are assigned on a curve, students are competing against each other rather than working toward defined learning goals. This can undermine collaborative learning environments and create unnecessary anxiety.
Pros and cons of criterion-referenced testing
Advantages
Directly supports learning objectives: Criterion-referenced assessments excel in instructional planning and allow for individualized learning paths. By focusing on specific objectives, these assessments provide a clear picture of what a student has mastered and what areas need improvement, making it easier for educators to tailor instruction.
Fairer and more transparent: Criterion-referenced testing can be considered fairer because a person’s score is not dependent on the performance of others. Students know in advance what is expected of them, and achievement of the standard is accessible to anyone who puts in the work.
Higher inter-rater reliability in structured programs: Research published in PubMed Central comparing evaluation approaches in graduate medical education found that criterion-referenced scaling produced consistently higher inter-rater reliability than norm-referenced scaling across all competencies studied – suggesting CRT may yield more consistent and valid evaluation data in structured professional programs.
Limitations
Cannot differentiate among high performers: When many students meet the criterion, CRT offers no way to distinguish between those who barely passed and those who excelled. This limits its usefulness in contexts where selection or ranking is necessary.
Risk of grade inflation: Criterion-referenced approaches may increase grade inflation and passing rates, and may not effectively identify the worst performers – particularly if the performance standard is set too low.
Standard-setting is complex: Defining what “mastery” looks like requires careful, expert-driven processes. If the cut score is set arbitrarily or without rigorous validation, the entire assessment loses credibility.
Which approach is right for higher education?
The honest answer is that most well-designed assessment systems in higher education do not rely exclusively on one framework. Many assessments couple norm-referenced scores with criterion-referenced performance categories, serving multiple purposes simultaneously. A university might use criterion-referenced grading for course assessments – ensuring students meet defined learning outcomes – while using norm-referenced entrance examinations to manage competitive admissions.
The choice ultimately depends on the purpose. When the goal is to verify mastery, support instruction, or certify competence, criterion-referenced testing is the more appropriate tool. When the goal is to rank, select, or compare students across a population, norm-referenced measurement serves better. As assessment experts at NWEA note, the more important questions for any institution are not whether a test is norm- or criterion-referenced, but whether the assessment is trustworthy and whether it can effectively guide instruction. Both questions deserve serious attention before any evaluation system is put in place.
What do you think? If you were designing an assessment system for a professional degree program – such as medicine, law, or engineering – which approach would you prioritize, and why? And do you think universities today rely too heavily on ranking students against each other, at the expense of measuring what they have actually learned?
References
- https://www.ebsco.com/research-starters/social-sciences-and-humanities/norm-referenced-testing
- https://www.cal.org/twi/EvalToolkit/5when2usetests.htm
- https://eric.ed.gov/?id=ED410316
- https://assess.com/norm-referenced-vs-criterion-referenced-testing/
- https://www.michiganassessmentconsortium.org/wp-content/uploads/LP_NORM-CRITERION.pdf
- https://www.classtime.com/en/norm-referenced-vs-criterion-referenced-assessment
- https://study.com/learn/lesson/norm-referenced-test-vs-criterion-referenced-test-what-is-a-norm-referenced-test.html
- https://dpi.wi.gov/sites/default/files/imce/sped/pdf/sl-lim-normref.pdf
- https://pmc.ncbi.nlm.nih.gov/articles/PMC10498947/
- https://www.renaissance.com/2018/07/11/blog-criterion-referenced-tests-norm-referenced-tests/
- https://www.nwea.org/blog/2024/norm-vs-criterion-referenced-in-assessment-what-you-need-to-know/
Leave a Reply