Assessing Professor Effectiveness

What 89,893 RateMyProfessor records say about gender bias, teaching tags, and what drives a rating

A reproducible statistical analysis of the RateMyProfessor dataset: is there a pro‑male rating bias, which teaching “tags” are gendered, and how well can we predict a professor’s rating — or whether they get a “pepper”? Every figure below is generated by the ape package; click Code to see how.

89,893
Professors analyzed
11
Questions answered
42
Passing tests
0.81
Classifier AUC

Key findings

Question Verdict Statistic Detail
Q1 Pro-male rating bias? Yes, small MWU p=3.7e-04 d=0.086
Q2 Difference in spread? Yes Levene p=0.0024 var ratio=0.91
Q4 Gendered tags? 17/20 tags FDR @ α=0.005 top: hilarious, caring
Q5 Difference in difficulty? No MWU p=0.79 —
Q7 Rating ~ numeric Strong R²=0.81 top: prop_retake
Q8 Rating ~ tags Good R²=0.74 top: tough_grader
Q9 Difficulty ~ tags Moderate R²=0.60 top: tough_grader
Q10 Predict ‘pepper’ Good AUC=0.81 balanced + Youden J
Q11 NY vs NJ ratings No difference MWU p=0.06 —

At a glance

Code
_, figs = questions.q1()
show(figs, "q1_rating_by_gender")
Figure 1: Average rating by gender (professors with ≥10 ratings). Men sit slightly higher — a small but statistically significant gap.
Code
_, figs = questions.q10()
show(figs, "q10_roc")
Figure 2: ROC for the ‘pepper’ classifier (AUC ≈ 0.81). Standardization is fit inside the pipeline; the threshold is chosen by Youden’s J.

How to read this site

  • Data — where the dataset comes from, what each of the 89,893 records contains, the 20 teaching tags, a preview, and why we require ≥ 10 ratings.
  • Methods — cleaning choices, the α = 0.005 threshold, and our statistical approach (FDR control across many tests, honest out-of-sample model evaluation, directional tests where the hypothesis calls for one).
  • Gender bias — Q1–Q3, Q5–Q6: bias in the mean, the spread, and effect sizes, plus difficulty.
  • Tags — Q4, Q8, Q9: which tags are gendered, and tag-driven rating/difficulty models.
  • Models — Q7 and Q10: predicting rating and “pepper”, with from-scratch regression validated against scikit-learn.
  • Extensions — the NY-vs-NJ bonus and how we’d push the analysis further.

NYU DS-GA 1001 capstone, group CAP 85 (Deepanshu Mody, Evan Beck, Samarth Agarwal). The analysis is packaged as tested, importable code (ape) with a reproducible pipeline. Regression and classification (Q7–Q10) and this engineering write-up are Deepanshu’s contributions.