When we talk about assessing student learning in higher education, the traditional examination system is almost always the default. It has been used for generations, and its familiarity has earned it a kind of institutional permanence. But familiarity should not be confused with effectiveness. A comprehensive review published in Higher Education (Springer, 2023) found that support for high-stakes exams is “largely rooted in opinion and pragmatism, rather than being justified by scientific evidence or pedagogical merit” – while evidence of their drawbacks is substantial. So, what exactly is broken about the present-day examination system? Let’s break it down.
Table of Contents
Over-reliance on final exams: discouraging continuous learning
One of the most deep-rooted problems with the current system is how much it leans on a single, high-stakes final examination to determine a student’s grade – often for an entire semester or year of work. This structure sends an unintended message: what you learn throughout the term matters less than how you perform on one day.
Research on student motivation consistently shows that high-stakes terminal examinations reduce intrinsic motivation and negatively impact self-regulated learning. When students know that a single exam decides everything, many default to last-minute cramming rather than building genuine understanding over time. They memorize patterns from past papers, rehearse model answers, and then forget most of it shortly after. This is not learning – it is performance under pressure.
Students themselves recognize this. Studies on student preferences consistently show that learners favor a broader range of assessment tasks spread across a semester, provided they are not over-assessed. High-stakes examinations have also been criticized for their poor consequential validity – research shows they narrow what is actually taught in classrooms, as faculty focus instruction on what the exam tests rather than on the full scope of meaningful learning. The curriculum, in effect, bends to serve the exam rather than the other way around.
The solution is not to abolish assessment but to distribute it. Continuous assessment models – involving assignments, quizzes, projects, and presentations spread across the term – keep students engaged throughout the learning journey and provide a far more accurate picture of their academic growth.
Lack of critical thinking assessment: memorization over analysis
Traditional university exams are overwhelmingly designed around factual recall. Multiple-choice questions, short answers, and even many essay prompts ask students to reproduce what they were taught – not to question it, extend it, or apply it to new situations.
This is a significant mismatch with the demands of modern professional life. Employers across sectors increasingly report that graduates lack the analytical and problem-solving capabilities needed at work. Yet university assessments continue to reward students who can accurately recall content rather than those who can think through complex problems. Research on exam-oriented education has confirmed that students in such systems show a reduced ability to demonstrate and apply practical knowledge – a direct consequence of prioritizing test-taking skills over genuine comprehension.
Consider a common scenario in a social sciences exam: a student might be asked to list the factors that contributed to a particular policy failure, but rarely asked to evaluate those factors, compare competing interpretations, or propose what could have been done differently. The former tests memory; the latter tests intellect. Higher-order cognitive skills – analysis, synthesis, evaluation – are exactly what Bloom’s Taxonomy identifies as the goals of advanced education, yet they remain underrepresented in most examination formats.
Reforming this requires a shift in how questions are designed. Assessments that present unfamiliar scenarios, require students to argue a position with evidence, or ask for the evaluation of competing theories can test critical thinking without abandoning academic rigor. The goal is to move from “what do you know?” to “what can you do with what you know?”
Subjectivity in essay-type questions: the problem of inconsistent grading
Essay-based exams are commonly used in humanities, social sciences, law, and education because they allow students to demonstrate depth of understanding. But they come with a well-documented problem: the same essay can receive significantly different marks depending on who grades it and when.
A meta-analysis by Malouff and Thorsteinsson (2016), drawing on 20 studies and nearly 2,000 graders, found statistically significant bias in the subjective grading of student work, including essays. Biasing factors included a student’s prior academic performance, race and ethnicity, perceived attractiveness, and even their name. A 2020 study by USC’s David Quinn found that nearly identical essays received different marks based solely on whether a name mentioned in the text sounded racially familiar – a striking illustration of how unconscious bias can distort grading.
Research published in PMC further highlights that examiners differ substantially in how they interpret marking rubrics, how they define correctness, and how comfortable they are with the inherent subjectivity of essay assessment. Those with a scientific or clinical background tend to view answers dichotomously – right or wrong – while those with education qualifications are more comfortable with nuanced, partial credit judgements. This inconsistency means that a student’s grade can vary not just between different examiners, but depending on the same examiner’s mood or workload on a given day.
The halo effect adds another layer of concern. A study published in Cogent Psychology found that faculty who had previously seen a student give a strong oral presentation subsequently graded that student’s written work more favorably – even when the written work itself was identical to work submitted by other students. Prior impressions were shaping outcomes in ways that had nothing to do with the written assessment itself.
Addressing this requires structural safeguards: blind marking (removing student identifiers), detailed rubrics with clearly articulated performance criteria, and moderation processes where a second examiner reviews a sample of marked scripts. Research confirms that when teachers use specific grading rubrics, racial and other forms of implicit bias in assessment are significantly reduced.
Language barriers in university exams: an uneven playing field
In most countries, university examinations are conducted in a dominant language – usually English, or the national language of instruction. For students who did not grow up speaking that language, this creates a structural disadvantage that has little to do with their actual subject knowledge.
A 2023 MIT Graduate Communication Survey found that non-native English-speaking (NNES) students reported significantly greater difficulty in reading and writing than their native-speaking peers. Over 30% of NNES students said that anxiety about their oral academic skills had significantly impacted their performance. These are not students who lack intellectual ability – they are students who are operating in an additional language while simultaneously being assessed on subject content.
Research by conservation scientist Tatsuya Amano at the University of Queensland, based on a survey of over 900 academic researchers, found that non-native English speakers from countries with moderate proficiency spent around 47% more time reading academic material than native speakers – and those from low-proficiency countries spent up to 91% more time. In timed exam conditions, this kind of disadvantage is unaccommodated and invisible.
The problem extends beyond reading speed. Research on international students in UK universities has documented that many students struggle to express their knowledge in exams not because they lack understanding, but because they cannot articulate it fluently in English under time pressure. A student may fully grasp an economic theory or a legal principle but lose marks because their written expression does not match the register expected by examiners – register being a feature of language proficiency, not subject mastery.
This issue demands institutional attention. Possible remedies include allowing additional exam time for verified non-native speakers, permitting the use of bilingual dictionaries in non-language subjects, designing assessment tasks that allow for varied modes of expression (such as diagrams, structured frameworks, or presentations), and training examiners to assess content separately from language fluency. Times Higher Education notes that the success of non-native speaking students depends heavily on whether institutions genuinely appreciate linguistic diversity and take practical steps to meet their needs – rather than treating language proficiency as a proxy for academic capability.
The case for reforming how we assess
Each of the four problems discussed here – the over-reliance on final exams, the neglect of critical thinking, biased essay grading, and language disadvantage – points to the same underlying issue: the traditional examination system was not designed to be fair, inclusive, or comprehensive. It was designed to be efficient and scalable.
That may have made sense in a different era. It does not serve today’s learners well. Scholars in pharmacy education at the University of Toronto have called on institutions to deliberately redesign grading and assessment practices so they are “more evidence-based and less reliant on historical tradition.” The same call applies to higher education broadly.
Reform does not mean making assessment easier. It means making it more accurate. A mix of continuous assessment, performance-based tasks, peer evaluation, structured rubrics, and inclusive exam accommodations can together produce a far more honest picture of what a student knows and can do – without sacrificing academic standards. The goal of evaluation, after all, is to support and measure learning. When the tool gets in the way of that goal, the tool needs to change.
What do you think? Is a single final exam ever a fair way to assess a full semester of learning – or does it inevitably reward performance over understanding? And how should universities better support students who are being assessed in a language that is not their mother tongue?
References
- https://link.springer.com/article/10.1007/s10734-023-01148-z
- https://adelaidebooks.org/discussions-on-the-disadvantages-of-exam-oriented-education
- https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10159463/
- https://journals.sagepub.com/doi/abs/10.1177/0004944116664618
- https://www.edutopia.org/article/the-evidence-backed-grader/
- https://pmc.ncbi.nlm.nih.gov/articles/PMC11069304/
- https://www.tandfonline.com/doi/full/10.1080/23311908.2014.988937
- https://fnl.mit.edu/may-june-2024/non-native-english-speaking-graduate-students-still-face-significant-disadvantages/
- https://www.insidehighered.com/news/faculty-issues/2023/07/21/profound-disadvantage-nonnative-english-speakers
- https://www.jmu.edu/global/isss/resources/global-campus-toolkit/files/barriers.pdf
- https://www.timeshighereducation.com/campus/breaking-language-barriers-supporting-nonnative-englishspeaking-students
- https://pmc.ncbi.nlm.nih.gov/articles/PMC10159463/
Leave a Reply