Assessment Rubrics for Pitch Skills: Five Formative Tools for Teachers

Five formative rubrics for judging a learner's pitch listening with Hearkle: observation-led, low-stakes, and honest about their own limits.

Short answer: this is a set of five classroom rubrics for noticing how a learner's pitch skills are developing while they play Hearkle. Each rubric covers one dimension, pitch matching, higher-or-lower discrimination, interval recognition, pitch memory, and self-reflection, and describes observable behaviour at four levels rather than handing out a score. They are teaching tools, not tests, and they are not a hearing assessment of any kind.

Why formative, and why zero data

Pitch listening is learnable, and like most learnable skills it improves fastest when assessment feeds back into teaching instead of sitting in a grade book. That is the core idea of formative assessment, or assessment for learning: assessment used to shape the next step rather than to sum up the last one, a practice classroom research associates with some of the larger gains available in ordinary teaching (Black and Wiliam, 1998). The rubrics below are written in that spirit: they help you notice what a learner is doing as they play, describe it plainly, and decide what to teach next.

A word on data. Hearkle keeps no account, no login, and no server-side record of anyone's play; scores live only on the device for the length of a session and vanish when the tab closes. That "zero data" design shapes how these rubrics work: there is no stored dashboard of results to mine, so the assessment evidence is simply what you and the learner observe together in the moment. That keeps the stakes low, keeps a child's listening data private, and keeps the focus where formative assessment wants it: on the conversation, not the record.

Read against the learner's own starting point, and not as a hearing test

Pitch perception varies naturally from person to person and from day to day, and much of that variation reflects listening conditions, the device, and practice rather than any fixed limit. So these rubrics describe growth relative to a learner's own earlier work, not against an absolute standard or an age norm. A learner is "Extending" when they have moved well beyond where they themselves began, not because they have beaten a benchmark. One boundary matters more here than in any other subject: none of this is a clinical or diagnostic instrument. Hearkle is a practice game, not an audiogram, and a rubric level is a snapshot of behaviour in one setting, never a verdict on a learner's hearing. If you ever have a genuine concern about a learner's hearing, refer to a qualified audiologist. For how the game frames pitch, see the methodology note.

Rubric 1: Pitch matching accuracy

Accuracy looks at how close a learner's produced pitch lands to the target when they match or reproduce a tone: where the pitch sits, not how they feel about it.

LevelWhat it looks like
EmergingMatches land clearly high or low, with large errors in both directions, and the learner may not yet hear that a match missed.
DevelopingMatches cluster nearer the target and bigger misses get noticed; the learner can usually say whether an attempt was too high or too low.
SecureMost matches sit close to the target at comfortable tones; the learner reliably lands near the pitch and holds it through a short task.
ExtendingAccuracy holds up under harder conditions, wider ranges and quicker tasks, and the centre stays stable even when a single attempt slips.

Rubric 2: Higher-or-lower discrimination

Discrimination looks at how reliably a learner tells higher from lower, and how small a difference they can still call, the just-noticeable difference (see how small a pitch difference we can hear). A learner can be confident on wide gaps yet unsure on close ones, which is a discrimination limit, not a matching problem.

LevelWhat it looks like
EmergingReliable only on wide, obvious gaps; close tones are guessed, and direction calls swing from round to round.
DevelopingCalls clear differences well and begins to catch narrower ones, though close pairs are still hit and miss, and best at comfortable ranges.
SecureReliably calls direction across a range of gaps, including fairly small ones, and knows when a pair is too close to be sure.
ExtendingDetects very small differences consistently, stays accurate when tones are quick or subtle, and reports uncertainty honestly.

Rubric 3: Interval recognition

Interval recognition looks at whether a learner hears and reproduces the distance between two tones, the heart of relative pitch, rather than fixating on single notes (see musical intervals explained).

LevelWhat it looks like
EmergingHears "two notes" but not the gap between them; reproductions copy a note rather than the interval, and change size when the start moves.
DevelopingReproduces familiar intervals from the same start, but the size drifts when asked to begin on a new pitch.
SecureReproduces an interval accurately from a new starting pitch, showing that the relationship, not the note, is being heard.
ExtendingNames and reproduces a range of intervals, transfers them to the voice or an instrument, and hears them inside short melodies.

Rubric 4: Pitch memory

Pitch memory looks at how well a learner holds a pitch across a silent gap and reproduces it, the skill that tone memory training builds and that holding a note depends on.

LevelWhat it looks like
EmergingThe pitch is lost almost as soon as the tone stops; reproductions after a pause land far from the original.
DevelopingHolds a pitch across a short gap but drifts as the pause grows, and does better when quietly humming through the silence.
SecureReproduces a remembered pitch accurately after a normal pause, keeping an inner reference through the gap.
ExtendingHolds a pitch through longer gaps and mild interference, and recovers the reference even after a distracting sound.

Rubric 5: Self-reflection

Self-reflection looks at how well a learner can judge and talk about their own pitch listening. In formative assessment it is the most valuable dimension of all: a learner who can hear their own errors no longer depends on you to spot them.

LevelWhat it looks like
EmergingCannot yet say whether an attempt was too high, too low, or on target, and relies on the game's feedback to know.
DevelopingDescribes a result in general terms ("that one felt sharp") and begins to connect it to something they did.
SecureJudges their own pitch accurately before seeing any feedback, names what went wrong, and suggests a fix for the next try.
ExtendingReflects across sessions, spots recurring patterns in their own errors, and plans practice around them.

How to use these rubrics

These tools reward light hands. A few cautions keep them fair and stop them from claiming more than they can.

Watch out for the device and the room. Pitch games depend heavily on the sound path: a tinny laptop speaker, background noise, and audio output lag all change what a learner can hear, and studies of web-based tasks find that browsers and devices add real, machine-dependent variation to timing and playback (Anwyl-Irvine and colleagues, 2021). Headphones in a quiet room give the fairest reading. Compare like device with like device where you can, and treat any absolute figure a learner sees as rough, not exact.

A single measurement is not an ability level, and never a diagnosis. One round is a noisy sample. Fatigue, mood, an unfamiliar device, or a lucky streak can move a result far more than a real change in skill, and the just-noticeable difference itself varies with training and conditions (Moore, 2013). Measurement guidance stresses gathering enough evidence before drawing a conclusion (AERA, APA and NCME, 2014). Place a learner on a level only after watching across several sessions, hold that placement loosely, and never read a hearing condition from it.

Read for growth, keep it a conversation. The point is not to rank a class but to find each learner's next useful step and talk it through. Share the rubric language so learners can locate themselves, invite them to set the next target, and revisit it together. Used that way the rubrics stay what they are meant to be: a support for learning, not a label.

Rubrics describe what you see; the game gives you something to see. Sit a learner down with a short round of Hearkle, watch one dimension at a time, and let the plain-language levels do the rest. Free, and it stores nothing.

Play Hearkle →

Frequently asked questions

Are these rubrics a hearing test or a diagnostic instrument?

No. They are formative teaching tools for describing observable behaviour during a practice game. They are not designed or validated to assess hearing or identify any condition, and they should never be used to diagnose one. If you have a concern about a learner's hearing, refer to a qualified audiologist rather than reading it from a rubric.

Why measure against a learner's own baseline instead of a class standard?

Because pitch perception varies naturally between people and with practice and conditions, an absolute cut-off would penalise ordinary variation and reward a strong starting point rather than genuine learning. Comparing a learner with their own earlier work keeps the assessment about growth.

How many sessions should I watch before placing a learner on a level?

More than one, and ideally several across different days. A single round is a noisy sample that a poor speaker, a noisy room, tiredness, or luck can distort. Look for a pattern that repeats before you settle on a level, and revise it freely as you see more.

Does the device or room affect the results?

Yes, a great deal. Speakers, headphones, background noise, and audio lag each change what a learner can hear, so the exact figures partly reflect the setup rather than the ear. Use headphones in a quiet room, compare results on the same kind of device, and treat absolute numbers as approximate.

Does Hearkle store any of this assessment data?

No. Hearkle keeps no account and no server-side record; scores stay on the device for the session only and vanish when the tab closes. The assessment evidence is what you and the learner observe together, not a stored profile.

Keep reading

Sources: Black, P. and Wiliam, D. (1998), "Assessment and Classroom Learning," Assessment in Education, 5(1), 7-74, doi:10.1080/0969595980050102; Moore, B. C. J. (2013), An Introduction to the Psychology of Hearing (6th ed., Brill); Anwyl-Irvine, A., Dalmaijer, E. S., Hodges, N. and Evershed, J. K. (2021), "Realistic precision and accuracy of online experiment platforms, web browsers, and devices," Behavior Research Methods, 53, 1407-1425, doi:10.3758/s13428-020-01501-5; American Educational Research Association, American Psychological Association and National Council on Measurement in Education (2014), Standards for Educational and Psychological Testing.

Understand it, now train it. The daily games stay free forever. A Personal Pass adds progress tracking (trends + export); Teacher / School tools let you assign it and follow a whole class. For teachers & self-learners →

Updated · By Lajos Toldi — educator and researcher in adaptive intelligent tutoring systems at AI24EduLabs (part of AI24Labs). Scientific claims are checked against authoritative sources; see our editorial policy.