Data, cleaning choices, and statistical approach

For the dataset itself — the source, every variable, the 20 tags, and a preview — see the Data page. This page covers how we prepare it and the statistics we run.

Preprocessing decisions

Code
from ape import data

num = data.load_numeric()
ge10 = data.filter_min_ratings(num)        # average is meaningful only with enough ratings
gendered = data.gender_subset(ge10)        # confidently male XOR female
print(f"all professors:            {len(num):>6}")
print(f"with >= 10 ratings:        {len(ge10):>6}")
print(f"...and confident gender:   {len(gendered):>6}  ({int((gendered['high_conf_male']==1).sum())} M / {int((gendered['high_conf_female']==1).sum())} F)")
all professors:             89893
with >= 10 ratings:          9841
...and confident gender:     7105  (3987 M / 3118 F)
  • Reliability threshold. A mean built on one rating can be 1 or 5 by chance, so we keep only professors with ≥ 10 ratings (9,841 of them). This is the single most important cleaning choice; we trade sample size for trustworthy averages.
  • Gender. We use only rows flagged male or female with high confidence (7,105), dropping ambiguous/both/neither.
  • Tags. A student can award up to 3 tags per rating, so raw counts scale with popularity. We normalize each tag by the professor’s number of ratings — one consistent scheme across every tag question.
  • Significance. Following the brief (Benjamin et al., 2018), we use α = 0.005 everywhere.
  • Reproducibility. All randomness is seeded (RANDOM_SEED = 10676128).

Statistical approach

A few deliberate choices keep the results defensible:

NoteRank-based, non-parametric tests

Ratings are bounded and skewed, so group comparisons use Mann–Whitney U and Kolmogorov–Smirnov rather than a t-test, and Levene’s test for differences in spread. Where a hypothesis is directional (e.g. pro-male rating bias), we use a one-sided test and estimate the effect on the full sample rather than searching subgroups.

NoteMultiple-comparison control

When testing all 20 tags at once, we control the false-discovery rate with Benjamini–Hochberg at α = 0.005 — appropriate when many effects are expected, and less conservative than Bonferroni.

NoteHonest model evaluation

For every model, all preprocessing (standardization) is fit inside cross-validation on the training folds only, via scikit-learn Pipelines — so reported R²/RMSE/AUC are genuine out-of-sample estimates. The pepper classifier handles class imbalance with balanced class weights and selects its decision threshold by Youden’s J. Effect sizes come with bootstrap 95% confidence intervals.

How we’d push it further

  • Confounders done right. Model rating with gender and covariates (difficulty, experience, school) jointly — ideally a mixed-effects model with a random effect per university, since ratings cluster by school.
  • Power. Several gender effects are tiny (d ≈ 0.09); a power analysis would tell us which “non-results” are genuine nulls versus underpowered.
  • Calibration. For the pepper classifier, calibrated probabilities (Platt/isotonic) would make the predicted scores trustworthy, not just the ranking (AUC).

The engineering

Everything here is importable, tested code: pytest validates the from-scratch OLS/ridge/lasso against scikit-learn, the data-cleaning invariants, and the effect-size math. See the repository.