FPSC Part-III (Pedagogy) — Topic 3 of 4. Covers the concepts, types, and qualities of educational testing, measurement, assessment, and evaluation.
Key Definitions (in increasing order of scope)
| Term |
Meaning |
Example |
| Test |
A specific tool/instrument (a set of questions or tasks) used to sample behavior or knowledge. |
A 10-question fractions quiz. |
| Measurement |
The process of assigning numbers/scores to that performance (quantitative — "how much"). |
Hamza scored 4 out of 10 on the quiz. |
| Assessment |
The broader process of gathering information about student learning using various tools (tests, observations, projects). |
Combining the quiz score with homework completion and class participation notes. |
| Evaluation |
The broadest term — making a value judgment/decision based on measurement and assessment data (e.g., pass/fail, grade, program effectiveness). |
Deciding Hamza needs extra help and will get a "C" grade this term. |
Exam angle: Ordering from narrowest to broadest — Test → Measurement → Assessment → Evaluation — is a recurring MCQ.
Formative vs Summative Evaluation
- Formative Evaluation: ongoing, during instruction; purpose is to improve learning/teaching (e.g., quizzes, class questioning, drafts). Low/no stakes.
- Summative Evaluation: at the end of instruction; purpose is to judge/certify achievement (e.g., final exams, board exams). High stakes.
- Analogy often quoted: "When the cook tastes the soup, that's formative; when the guest tastes the soup, that's summative." — (Robert Stake)
- Example (Formative): Teacher checks students' rough drafts and gives feedback before the final essay is due.
- Example (Summative): The final board exam at the end of Grade 10 that decides promotion.
Norm-Referenced vs Criterion-Referenced Tests
- Norm-referenced: compares a student's performance to that of other students (a norm group); reports rank/percentile (e.g., most IQ tests, many entrance tests). Standard item difficulty is usually set around 50% so scores spread out and rank well. Example: "You scored in the 90th percentile" — meaning better than 90% of test-takers.
- Criterion-referenced: compares performance to a fixed standard/criterion regardless of others' scores (e.g., a driving test, mastery tests). Example: "You correctly solved 8/10 problems, which meets the 80% mastery standard" — everyone who reaches 80% passes, regardless of others.
Classification of Tests
Two main frameworks: by Method (how it's built) and by Purpose (why it's given), plus a few special-purpose test labels.
1. Classification by Method
A. Objective Type — one correct answer, high reliability (no grader bias)
- Selection Type (student picks the answer)
- MCQs — most popular objective item
- Parts: Stem (question — must be positive form, avoid "NOT"), Answer (correct choice), Distractors (wrong choices)
- Rules: avoid qualifiers (often, sometimes, generally) & absolutes (always, never) — they hint the answer; keep all choices similar length
- True/False — binary choice (True/False, Yes/No, Agree/Disagree)
- Pros: objective, no grader bias
- Cons: 50% guessing chance
- Matching — links two related columns
- Column A = Premises; Column B = Responses
- Supply Type (student produces the answer)
- Short Answer — brief answer (word/phrase/number). Good for simple recall; not for problem-solving/creativity
- Completion (Fill in the Blanks) — insert missing word. No options given — guessing eliminated
B. Essay (Subjective) Type — for complex outcomes: critical thinking, creativity, synthesis