The data

Where it comes from, what each record contains, and why we clean it the way we do.

Dataset

The data come from RateMyProfessor.com, where students rate instructors on quality and difficulty and tag their teaching style. The professor scraped, anonymized, and collated the site into three aligned files — one row per professor, 89,893 professors in total, in the same order across all three files.

numeric file : 89,893 professors × 8 columns
tags file    : 89,893 professors × 20 columns
qual file    : 89,893 professors × 3 columns

What each professor record contains

rmpCapstoneNum.csv — the quantitative facts about each professor:

Column Meaning
avg_rating Mean of all quality ratings (1–5)
avg_difficulty Mean of all difficulty ratings (1–5)
num_ratings How many ratings the averages are based on
received_pepper Whether students judged the professor “hot” (a 🌶️ on RMP) — boolean
prop_retake Proportion of students who said they’d take the class again
num_online How many ratings came from online classes
high_conf_male Flagged male with high confidence — boolean
high_conf_female Flagged female with high confidence — boolean

rmpCapstoneQual.csv — three qualitative fields: major / field, university, and US state.

The 20 teaching tags

rmpCapstoneTags.csv holds the raw count of each tag a professor received (a student can award up to three tags per rating). Because counts grow with popularity, we normalize each tag by the professor’s number of ratings before analysis.

Tag Tag
0 Tough grader So many papers
1 Good feedback Clear grading
2 Respected Hilarious
3 Lots to read Test heavy
4 Participation matters Graded by few things
5 Don't skip class Amazing lectures
6 Lots of homework Caring
7 Inspirational Extra credit
8 Pop quizzes! Group projects
9 Accessible Lecture heavy

A look at the raw numbers

A few real rows (numeric file):

The single most important cleaning choice is a minimum-ratings threshold. An average built on one or two ratings is unstable — it piles up at the integer values 1, 2, …, 5. Requiring ≥ 10 ratings trades sample size for trustworthy averages:

(a) Average rating: all professors (with the tell-tale spikes at whole numbers from tiny samples) vs. those with ≥ 10 ratings (smooth and reliable).
(b)
Figure 1
(a) Distribution of the number of ratings per professor (clipped at 60). The dashed line marks our ≥ 10 threshold; most professors fall below it.
(b)
Figure 2

Missing data

Some fields are missing — professors with no ratings, and the would-retake proportion (often absent). We handle this per question: dropping incomplete rows where a model needs the field, and otherwise working with what’s available.

(a) Share of missing values by column in the numeric file.
(b)
Figure 3

From raw to analysis-ready

all professors                     : 89,893
  with >= 10 ratings               :  9,841
    and a high-confidence gender    :  7,105  (3987 M / 3118 F)

The exact cleaning, the α = 0.005 significance threshold, the RNG seed, and our statistical choices are detailed on the Methods page.