Key Questions

1. To what extent are administrator observation ratings and student survey reports affected by the characteristics of students in the classroom?

2. How would adjusting observation ratings for student characteristics change the distribution of teacher rankings, and which teachers would be most affected?

Overview

School districts rely heavily on teacher performance ratings—especially classroom observations—to make high-stakes decisions about tenure, promotion, and dismissal. These ratings are intended to reflect teaching quality. But what if they also reflect the students teachers are assigned? Understanding whether evaluations measure teacher effectiveness or classroom context is critical for designing fair and effective accountability systems.

Exploiting naturally occurring year-to-year variation in classroom composition within teachers, this article examines whether teacher performance ratings assigned by evaluators and students are influenced by classroom context. We find that teachers with higher-achieving and less disruptive students, holding constant the teacher and school, receive systematically higher performance ratings. These effects are robust across model specifications, placebo tests, and multiple dimensions of teaching practice. By contrast, classroom demographics show no consistent association with performance ratings. A policy that adjusts evaluator scores for classroom characteristics, analogous to value-added models, increases the relative ranking of Black teachers by 8 percentage points, highlighting equity impacts of considering classroom context.

Key Findings

  • Classroom composition significantly affects teacher performance ratings 
    • Teachers received higher ratings when assigned students who had higher prior achievement, attended more regularly, and exhibited fewer behavioral challenges. 
    • A one standard deviation (SD) increase in classroom quality index (based on baseline academic and behavioral measures) raised classroom observation scores by 0.07 SD and student survey ratings by 0.13 SD.
       
  • Academic and behavioral factors—not demographics—drive the effect 
    • Prior test scores, GPA, and attendance were the main drivers, while most demographic characteristics (race, income, special education) did not consistently affect evaluator ratings. Student surveys were somewhat more sensitive to demographics. 
    • Furthermore, our exploration of mechanisms indicated that these effects were not reflected in outcome-based measures of teacher productivity (i.e., student growth measures), pointing to evaluator bias rather than genuine changes in teacher effectiveness as the likely driver.
       
  • Adjusting ratings changes who is identified as “effective” 
    • Simulating a policy that adjusts ratings for classroom context shows: 
      • Black teachers’ rankings improved by ~8 percentile points. 
      • Teachers serving more disadvantaged students benefited the most, although more research needs to be done to understand this effect.

Share

FacebookTwitterLinkedinEmail