A teacher evaluation is one of the few moments in a school year when a leader sits with a single teacher's practice and says, out loud and in writing, what is working and what should change next. Done well, it builds trust and sharpens instruction. Done as a compliance exercise, it produces a folder no one reads and a rating no one believes. This guide is about doing it well, and doing it in a way that still holds up when a district or an appeal asks to see your evidence.
It is written to be framework-neutral. Whether your district uses a Danielson-style framework, a state model, or a locally built rubric, the underlying craft, the way you prepare, gather evidence, score, and confer, is the same. Where a step depends on your specific instrument, we say so.
1. What an evaluation is actually for
Every evaluation system is asked to do two jobs at once: verify that a teacher meets the standard, and help that teacher get better. Those jobs pull in different directions. Accountability wants a defensible rating. Growth wants an honest, low-stakes conversation. When systems collapse the two, teachers learn to perform for the rating and stop surfacing the real problems in their practice.
The most useful mental model is to keep the two purposes visible and sequenced. Run formative work all year, the informal visits, the feedback, the coaching, so that growth happens in a space that feels safe. Reserve the summative rating for the end of the cycle, where it becomes a summary of evidence the teacher has already seen, not a verdict delivered cold. If you want to go deeper on why the developmental and the evaluative should not live in the same conversation, we wrote a separate piece on separating coaching from evaluation.
This matters for reasons beyond morale. In 2009, TNTP's report The Widget Effect documented that in districts using binary satisfactory/unsatisfactory ratings, more than 99 percent of teachers were rated satisfactory. An evaluation that cannot tell teachers apart cannot target support, defend a personnel decision, or reward excellence. The craft below is what turns an evaluation from a rubber stamp into a real signal.
2. Before the observation
Most of the quality of an evaluation is decided before you ever walk into the room.
Know your instrument cold. You cannot score against a rubric you half-remember. Re-read the criteria you will be rating and, for each one, be clear on what distinguishes an effective rating from a developing one. If your district lets evaluators bring their own instrument or configure a state model, settle the form and the scoring rules before the cycle opens, not after. Our library of observation form templates includes ready-to-use starting points for a full-lesson formal observation and a short walkthrough that you can adapt.
Hold a pre-conference when the observation is announced. A short pre-conference tells you what you are about to watch: the lesson's objective, where it sits in the unit, what the teacher has planned for students who are ahead or behind, and what the teacher wants feedback on. It turns a cold observation into an informed one, and it signals that you are there to understand the lesson, not to catch it.
Align on the model you are using. If your teachers are evaluated under a specific state framework, make sure both you and they are working from the same version of it. Our state-by-state evaluation guides lay out how the common state models are structured, which is a useful shared reference before a cycle begins.
Sort out logistics in advance. Confirm timing, whether the visit is announced or unannounced, and how you will capture evidence in the moment. Deciding all of this on the way to the classroom is how observations end up thin.
3. During the observation: gathering evidence
The single most important discipline of an observation is this: collect evidence, not judgments.
Script low-inference notes. Write down what you actually see and hear, and only that. Quote the teacher's questions and the students' responses. Note how many students had their hands up, what was on the board, how a transition ran, what the task actually asked students to do. Time-stamp the moments that matter. A low-inference note reads "9:14, teacher asks 'why do you think the author repeated that word?' and calls on three students, each of whom cites a line of text." It does not read "good questioning."
Why low-inference notes matter: judgments cannot be discussed, but evidence can. When you tell a teacher "your questioning was weak," there is nothing to work with. When you show them that seven of your scripted questions could be answered yes or no, the next step writes itself. Evidence is also what makes a rating defensible if it is ever challenged.
Watch the students, not just the teacher. The clearest evidence of instruction is what students are actually doing: the work they produce, the questions they ask, how they respond when they are stuck. A charismatic teacher can hold a room while students do very little; a quiet one can have every student thinking hard. Evidence keeps you honest about which is which.
Keep the load light enough to sustain across many teachers. A five-minute walkthrough and a full-lesson formal observation call for different instruments, and trying to run a formal instrument in a five-minute window produces bad evidence. Match the form to the visit. When it comes time to write feedback, a well-stocked bank of specific, evidence-based language helps; we keep a set of sample observation comments you can adapt rather than starting from a blank page.
4. Scoring against the rubric
Scoring is where good evidence pays off and where bias creeps in. The move is to sort your evidence against the rubric criteria after the observation, letting the evidence drive the rating rather than deciding the rating first and hunting for evidence to justify it.
Match evidence to criteria, one at a time. Take each criterion you are rating and ask which of your scripted notes speak to it. Place the criterion where the preponderance of evidence puts it. If you cannot point to evidence for a rating, you cannot defend it, and you should note that the evidence was insufficient rather than guess.
Guard against the predictable biases. Recency bias weights the last five minutes over the first forty. Halo bias lets one strong criterion pull up the rest. Similarity bias rewards teachers who teach the way you taught. Familiarity with the evidence, not with the teacher, is the corrective.
Do not hang a year on one lesson. The Measures of Effective Teaching (MET) project, funded by the Bill and Melinda Gates Foundation and summarized in its 2013 culminating brief, reported that reliable observation ratings depend on more than one lesson and, ideally, more than one trained observer, since adding a second observer improves reliability. That helps explain why many systems require a mix of announced and unannounced visits across the year. If your system allows more than one evaluator, calibrate them: have observers score the same lesson and reconcile where they diverge, so a teacher's rating reflects the practice and not the luck of which evaluator walked in.
5. The post-observation conference
The conference is where an evaluation either changes practice or gets filed. A few habits separate the two.
- Start with the teacher's reflection. Ask how they thought the lesson went and what they would change. Teachers who name their own next step own it in a way they never own yours.
- Ground every point in evidence. Trade "your pacing dragged" for "students finished the independent task by 9:30 and then waited nine minutes for the next one." Specific, observable, undeniable.
- Pick one or two high-leverage next steps. A list of twelve fixes changes nothing. One well-chosen move that a teacher can practice this week changes the next lesson.
- Attach support to the ask. Name what help comes with the next step: a coaching cycle, a colleague's room to visit, a specific piece of professional development. Feedback without support is just a grade.
Write the summary while the lesson is fresh, tie it to the evidence you scripted, and make sure the teacher has seen the substance of it before any rating becomes final. A summative rating should be the last line of a conversation that has been going all year, not the opening line of a new one.
6. Connecting the rating to growth
A rating that goes nowhere teaches teachers that the whole exercise is theater. The evaluation should hand off cleanly to the systems that actually develop people.
Connect next steps to professional goals and SLOs so the growth areas you named show up in the goals the teacher revisits all year. Route developmental support into instructional coaching, kept separate from the evaluative rating so the coaching stays safe. And feed identified needs into professional development so the district is building learning around what observations actually surface, not around a guess. When evaluation, coaching, goals, and PD share the same evidence, a rating stops being an endpoint and becomes the start of a plan.
7. Coverage, timelines, and documentation
The best-run individual observation still fails the district if it is one of the forty that got done and the other three hundred did not. Two things keep an evaluation program whole.
Coverage. Know, at any point in the year, which teachers have been observed, how many times, by whom, and against which required components. Mid-year is when a coverage gap is still fixable; at the end of the year it is a compliance problem. This visibility is exactly what a good evaluation system is supposed to give a leader.
Documentation. Keep the evidence, the scored rubric, the teacher's acknowledgment, and the signatures together and time-stamped. If a rating is ever questioned, the record either speaks for itself or it does not. Audit-ready is not bureaucracy for its own sake; it is what protects a fair rating and a fair process for everyone involved.
8. Common pitfalls
- Rating first, evidence second. If you have decided the score before you sort your notes, the notes will always agree with you. Let the evidence lead.
- Adjectives instead of evidence. "Engaging," "rigorous," and "student-centered" are conclusions. A teacher cannot act on a conclusion, only on the observation behind it.
- One visit, one verdict. A single lesson is a snapshot. Judgments that carry weight rest on more than one.
- Feedback with no support. Naming a weakness without offering help reads as a gotcha and rarely changes anything.
- Ratings that no longer distinguish anyone. If nearly everyone lands on the same rating, the instrument has stopped telling teachers apart. Calibrate your evaluators and trust your evidence.
The short version
Know your rubric before you walk in. Script what you see, not what you conclude. Sort the evidence against the criteria afterward, and rest the year's rating on more than one lesson. Lead the conference with the teacher's reflection, ground it in evidence, choose one or two next steps, and attach real support to them. Then connect all of it to the goals, coaching, and PD that do the actual developing, and keep the record clean enough to stand on its own.
Sources
- TNTP, The Widget Effect: Our National Failure to Acknowledge and Act on Differences in Teacher Effectiveness (2009).
- Bill and Melinda Gates Foundation, MET Project, Ensuring Fair and Reliable Measures of Effective Teaching: Culminating Findings from the MET Project's Three-Year Study (2013).