Classroom Observation: A Guide for School Leaders
What classroom observation is, which type answers which question, and how many observers a rating needs before it means anything. With a coverage plan.
In short
Decide three things before the school year starts: which question each visit answers, who observes, and how many minutes each teacher gets. The third decision is the one most schools skip, and it is the one that determines whether a rating reflects the teacher or the observer.
Classroom observation is a structured visit that produces evidence about teaching. It only works when three things are decided in advance: which question the visit answers, who observes, and how much observation time each teacher gets over a year.
Most guides to classroom observation stop at the first question. They list what to look for, describe the types of visit, and explain that feedback should be timely and specific. All of that is true and none of it is the hard part. The hard part is that one person watching one lesson produces a score that says almost as much about the observer as it does about the teaching — and the fix is a scheduling decision you make in August, not a technique you apply in the room.
What is classroom observation?
A planned visit to a classroom during instruction, in which an observer records what happens against an agreed structure, and then uses that record for either a formal judgement or a coaching conversation. The structure is what separates an observation from dropping in.
Two things make an observation useful rather than decorative. First, the observer records what happened rather than what they concluded — times, counts, and quotes, not adjectives. Second, the record is attached to a question that was agreed before the visit. An observation with no question behind it produces notes nobody opens again.
Which type of observation answers which question?
Formal observations answer whether a teacher meets a standard. Coaching observations answer whether a specific practice is improving. Walkthroughs answer what is happening across a building. Interaction observations answer how good the exchanges in a room are. Running the wrong type for your question is the most common way observation data ends up unused.
| The question you need answered | Observation type | Typical length | What it produces |
|---|---|---|---|
| Is this teacher meeting the standard? | Formal observation | Full lesson | A scored rubric with evidence, on the record |
| Is the thing this teacher is working on moving? | Coaching observation | 15–25 minutes | Low-inference notes and one next step, off the record |
| What is happening across the building this month? | Walkthrough | 5–15 minutes | A few tallied look-fors, aggregated, not attributed to individuals |
| How good are the interactions in this room? | Interaction observation | Timed cycles | Interaction quality scores inside a fixed protocol |
The four rows are not interchangeable, and the failure mode is specific: a walkthrough score used as if it were a formal rating. Walkthrough data is aggregate data. It tells you that questioning is thin across the sixth grade. It does not tell you that one named teacher's questioning is thin, and using it that way is how staff stop trusting the whole system.
How many observers does an observation rating need?
More than one. In the MET project's 2013 analysis, a rating built from one full lesson scored by the teacher's own administrator reached a reliability of 0.51. Adding a second lesson scored by that same administrator moved it to 0.58. Having the second lesson scored by a different administrator moved it to 0.67 — more than twice the gain, for the same amount of observation time.
Reliability here runs from 0 to 1 and describes how much of a score reflects consistent aspects of the teacher's practice rather than the particular observer or the particular day. A score of 0.51 is not a scandal. It is just a number that should stop you writing a paragraph of career-affecting prose on the strength of one visit.
| Who observes | Reliability |
|---|---|
| One full lesson, the teacher's own administrator | 0.51 |
| Two full lessons, both scored by that same administrator | 0.58 |
| Two full lessons, the second scored by a different administrator from the same school | 0.67 |
| One full lesson by the administrator, plus three 15-minute visits by three different observers | 0.67 |
How long should each observation be?
Long enough to see the components you are scoring. The MET project found that observations based on the first 15 minutes of a lesson were about 60 percent as reliable as full-lesson observations while taking one-third of the observer time — efficient, but not a replacement. The same analysis found that several components of the Danielson framework were frequently not visible in the first 15 minutes.
So short visits are a good way to add perspectives and a bad way to replace the formal one. Keep at least one or two full-length observations per teacher per year, and spend everything above that on more observers rather than longer visits. If your rubric contains components about lesson closure or assessment of learning, a 15-minute visit at the start of a period will systematically fail to see them, and teachers will notice that pattern before you do.
What does a year of observation actually cost in observer time?
For a 40-teacher school, roughly 60 hours of observation time a year — and that figure barely moves between a weak plan and a strong one. What changes is how the 60 hours are distributed. The same budget, split across more observers, produces a materially better rating.
| Plan | Per teacher | Total, 40 teachers | Load per observer | Reliability |
|---|---|---|---|---|
| A — the principal does everything | Two 45-minute lessons, one observer | 60 hours | 60 hours on one person | 0.58 |
| B — split the two lessons | Two 45-minute lessons, two observers | 60 hours | 30 hours each, across two | 0.67 |
| C — one formal plus short visits | One 45-minute lesson plus three 15-minute visits | 60 hours | About 15 hours each, across four | 0.67 |
Plan A is what most schools run by default, and it is the worst row in the table on every measure that matters. It concentrates 60 hours on the person with the least free calendar, and it produces the least reliable rating. Nothing about moving to Plan B costs money. It costs a conversation about who is allowed to observe, which is a harder thing to schedule than the observations themselves.
One thing worth checking before you plan around more observers: whether your observation software charges per observer. Voxento's schools plan does not — the pricing page lists unlimited users with no per-observer or per-teacher fees — but several tools in this category price by seat, which quietly makes the reliable configuration the expensive one. It is worth adding to the list of questions to ask before buying observation software.
How do you record evidence instead of judgment?
Write what a camera would have caught: times, counts, and direct quotes. A conclusion like 'students were engaged' gives the teacher nothing to work with, because the only available responses are to agree or to disagree with you. Evidence gives you both something to reason about together.
| What observers usually write | What is actually usable |
|---|---|
| Good classroom management | 9:05 — transition from carpet to desks took 40 seconds, no verbal redirection needed |
| Students were engaged | 9:14 — 6 of 24 students raised hands; teacher took 3 responses, all from the front two rows |
| Questioning could be stronger | 9:22 — 11 questions in 8 minutes; 9 had one correct answer, 2 asked students to explain reasoning |
| The learners seemed tense | 9:31 — three students asked a neighbour what to do before starting the task |
This is also the discipline that makes multiple observers workable. Two people who write conclusions will disagree and have no way to resolve it. Two people who write timestamps and counts can compare records and find out which of them was watching the wrong thing.
What has to happen after the visit?
A conversation, on a date fixed before the observation, in which the teacher does most of the talking and leaves with one thing to change. Follow-up is where observation either becomes coaching or becomes filing.
This is the part with the strongest evidence behind it. A 2018 meta-analysis of 60 causal studies by Kraft, Blazar and Hogan found pooled effects of teacher coaching of 0.49 standard deviations on instructional practice and 0.18 on student achievement. Read the caveat too: the same paper reports that average effects from larger-scale programmes were only a fraction of those from small ones. Plan for the direction of that finding, not the size of it.
- Fix the debrief date when you schedule the visit, not afterwards. A date, not 'within a week'.
- Open by asking the teacher what they noticed. If you open with your notes, you get a meeting instead of a conversation.
- Bring two or three pieces of evidence, not the whole record. The full record is available if they want it.
- Leave with one change, named specifically enough that both of you would recognise it in the next visit.
- Write the next visit into the calendar before either of you stands up.
Which framework should you score against?
Whichever one your district has already adopted. Framework choice matters far less than observer count and follow-up speed, and switching frameworks costs a year of retraining that buys very little reliability.
If you are genuinely choosing, whether for a new district or because nobody uses the framework on paper, the practical differences between the main options are in observation length, observer training load, and what you can defend in an evaluation conference. That comparison is its own piece of work: see the comparison of Danielson, Marzano and CLASS. If you are choosing the software rather than the framework, start with the buying guide for teacher observation software instead.
The coverage plan: build yours in one sitting
This is the artifact. It takes about an hour and it replaces the annual argument about why observations did not happen. Do it before the year starts, on paper, and keep it visible.
- List every teacher who must be observed this year, and mark the ones for whom a formal observation is required by your contract or state rules. That count is your floor. Everything else is discretionary.
- Name your observers, and not only the principal: assistant principals, instructional coaches, department leads, and peers from another building all count. Write down how many minutes a week each one can realistically give, and be pessimistic about it.
- Assign every teacher a plan: A, B or C from the table above. Non-tenured teachers and anyone on a support plan get B or C. Everyone else can sit on a lighter cycle without anything being lost.
- Multiply it out. Teachers times minutes per plan gives your annual observation minutes. Divide by your observers' weekly minutes across 36 weeks. If it does not fit, cut coverage rather than reliability — observe fewer teachers properly rather than all of them badly.
- Block the time on a named calendar now, week by week. Observation time that is not on a calendar in August does not happen in February, and no software fixes that.
- Add one final column to every row: the date feedback is due. A date, not a duration. This is the column that will be wrong first, which is exactly why it needs to be written down.
Once the plan exists, the tooling question is only whether one observer can run a whole observation — form, evidence, score, next steps, sent — without switching apps, because every handoff is where the debrief date slips. That is the argument for running it in one workspace, and the four starting forms in the observation templates library map to the four observation types in the table above.
Frequently asked questions
- How many classroom observations should a teacher get per year?
- Start from your contractual minimum, then add short visits rather than long ones. A defensible pattern is one or two full-length formal observations plus three or four short visits from different observers. The MET project's analysis suggests that spreading a fixed amount of observation time across more observers raises reliability more than adding lessons with the same observer does.
- Should classroom observations be announced or unannounced?
- It depends on which question the visit answers. Formal observations that feed an evaluation are usually announced, because the teacher is entitled to prepare and because a scheduled lesson is a fairer test of planning. Coaching visits and walkthroughs are more useful unannounced, because you are looking at the ordinary day rather than the prepared one. Mixing the two — unannounced visits that quietly feed an evaluation — is the version that destroys trust.
- Can one person do all the observations in a school?
- They can, and the rating will be weaker for it. In the MET project's figures, two lessons scored by the same administrator reached a reliability of 0.58, while two lessons split between two administrators reached 0.67 for the same observation time. If only one person is available, spend the time on fewer teachers rather than shorter visits on all of them.
- Do short walkthroughs count as classroom observations?
- They count as evidence, not as ratings. MET found 15-minute observations were about 60 percent as reliable as full-lesson ones, and that several rubric components were often not visible in the first 15 minutes of a lesson. Use walkthroughs to spot building-wide patterns and to add observer perspectives, and keep at least one full-length observation per teacher for anything that goes on the record.
Sources
- MET Project, Ensuring Fair and Reliable Measures of Effective Teaching: Culminating Findings from the MET Project's Three-Year Study (2013)
- Ho, A. D. & Kane, T. J., The Reliability of Classroom Observations by School Personnel, MET Project (2013)
- Kraft, M. A., Blazar, D. & Hogan, D., The Effect of Teacher Coaching on Instruction and Achievement: A Meta-Analysis of the Causal Evidence, Review of Educational Research 88(4), 547–588 (2018)
Written by
Muhammad Amin — Co-founder, Voxento
I co-founded Voxento and build the platform. I work directly with the schools and training teams running observations and AI roleplay on it, which is where most of what I write here comes from.
Related reading
- What Is Observation Studio? A Simple Guide for School Leaders
The workspace where an observation runs end to end: pick a structure, capture evidence, score, and share next steps — without handoffs between tools.
- Danielson vs Marzano vs CLASS: Which Framework Fits Your School?
Danielson has 22 components, Marzano Focused has 23 elements, CLASS scores interaction quality in timed cycles. How to pick the right one.
- How to Choose Teacher Observation Software in 2026
Score tools on workflow fit, rubric fidelity, follow-up speed and procurement readiness — in that order. Includes a weighted scorecard for demos.
Put these ideas into practice with Voxento.
See how observations, evaluations, coaching, and walkthroughs work in one workspace.