For the dataset itself — the source, every variable, the 20 tags, and a preview — see the Data page. This page covers how we prepare it and the statistics we run.
Preprocessing decisions
Code
from ape import datanum = data.load_numeric()ge10 = data.filter_min_ratings(num) # average is meaningful only with enough ratingsgendered = data.gender_subset(ge10) # confidently male XOR femaleprint(f"all professors: {len(num):>6}")print(f"with >= 10 ratings: {len(ge10):>6}")print(f"...and confident gender: {len(gendered):>6} ({int((gendered['high_conf_male']==1).sum())} M / {int((gendered['high_conf_female']==1).sum())} F)")
all professors: 89893
with >= 10 ratings: 9841
...and confident gender: 7105 (3987 M / 3118 F)
Reliability threshold. A mean built on one rating can be 1 or 5 by chance, so we keep only professors with ≥ 10 ratings (9,841 of them). This is the single most important cleaning choice; we trade sample size for trustworthy averages.
Gender. We use only rows flagged male or female with high confidence (7,105), dropping ambiguous/both/neither.
Tags. A student can award up to 3 tags per rating, so raw counts scale with popularity. We normalize each tag by the professor’s number of ratings — one consistent scheme across every tag question.
Significance. Following the brief (Benjamin et al., 2018), we use α = 0.005 everywhere.
Reproducibility. All randomness is seeded (RANDOM_SEED = 10676128).
Statistical approach
A few deliberate choices keep the results defensible:
NoteRank-based, non-parametric tests
Ratings are bounded and skewed, so group comparisons use Mann–Whitney U and Kolmogorov–Smirnov rather than a t-test, and Levene’s test for differences in spread. Where a hypothesis is directional (e.g. pro-male rating bias), we use a one-sided test and estimate the effect on the full sample rather than searching subgroups.
NoteMultiple-comparison control
When testing all 20 tags at once, we control the false-discovery rate with Benjamini–Hochberg at α = 0.005 — appropriate when many effects are expected, and less conservative than Bonferroni.
NoteHonest model evaluation
For every model, all preprocessing (standardization) is fit inside cross-validation on the training folds only, via scikit-learn Pipelines — so reported R²/RMSE/AUC are genuine out-of-sample estimates. The pepper classifier handles class imbalance with balanced class weights and selects its decision threshold by Youden’s J. Effect sizes come with bootstrap 95% confidence intervals.
How we’d push it further
Confounders done right. Model rating with gender and covariates (difficulty, experience, school) jointly — ideally a mixed-effects model with a random effect per university, since ratings cluster by school.
Power. Several gender effects are tiny (d ≈ 0.09); a power analysis would tell us which “non-results” are genuine nulls versus underpowered.
Calibration. For the pepper classifier, calibrated probabilities (Platt/isotonic) would make the predicted scores trustworthy, not just the ranking (AUC).
The engineering
Everything here is importable, tested code: pytest validates the from-scratch OLS/ridge/lasso against scikit-learn, the data-cleaning invariants, and the effect-size math. See the repository.
Source Code
---title: "Methods"subtitle: "Data, cleaning choices, and statistical approach"---For the dataset itself — the source, every variable, the 20 tags, and a preview — see the**[Data](data.qmd)** page. This page covers how we prepare it and the statistics we run.## Preprocessing decisions```{python}#| label: cleaningfrom ape import datanum = data.load_numeric()ge10 = data.filter_min_ratings(num) # average is meaningful only with enough ratingsgendered = data.gender_subset(ge10) # confidently male XOR femaleprint(f"all professors: {len(num):>6}")print(f"with >= 10 ratings: {len(ge10):>6}")print(f"...and confident gender: {len(gendered):>6} ({int((gendered['high_conf_male']==1).sum())} M / {int((gendered['high_conf_female']==1).sum())} F)")```- **Reliability threshold.** A mean built on one rating can be 1 or 5 by chance, so we keep only professors with **≥ 10 ratings** (9,841 of them). This is the single most important cleaning choice; we trade sample size for trustworthy averages.- **Gender.** We use only rows flagged male *or* female with high confidence (7,105), dropping ambiguous/both/neither.- **Tags.** A student can award up to 3 tags per rating, so raw counts scale with popularity. We normalize **each tag by the professor's number of ratings** — one consistent scheme across every tag question.- **Significance.** Following the brief (Benjamin et al., 2018), we use **α = 0.005** everywhere.- **Reproducibility.** All randomness is seeded (`RANDOM_SEED = 10676128`).## Statistical approachA few deliberate choices keep the results defensible:::: {.callout-note}## Rank-based, non-parametric testsRatings are bounded and skewed, so group comparisons use **Mann–Whitney U** and**Kolmogorov–Smirnov** rather than a t-test, and **Levene's test** for differences in spread.Where a hypothesis is directional (e.g. *pro-male* rating bias), we use a **one-sided** test andestimate the effect on the full sample rather than searching subgroups.:::::: {.callout-note}## Multiple-comparison controlWhen testing all 20 tags at once, we control the **false-discovery rate** with**Benjamini–Hochberg** at α = 0.005 — appropriate when many effects are expected, and lessconservative than Bonferroni.:::::: {.callout-note}## Honest model evaluationFor every model, all preprocessing (standardization) is fit **inside cross-validation on thetraining folds only**, via scikit-learn `Pipeline`s — so reported R²/RMSE/AUC are genuineout-of-sample estimates. The pepper classifier handles class imbalance with balanced class weightsand selects its decision threshold by **Youden's J**. Effect sizes come with **bootstrap 95%confidence intervals**.:::## How we'd push it further- **Confounders done right.** Model rating with gender *and* covariates (difficulty, experience, school) jointly — ideally a **mixed-effects model** with a random effect per university, since ratings cluster by school.- **Power.** Several gender effects are tiny (d ≈ 0.09); a power analysis would tell us which "non-results" are genuine nulls versus underpowered.- **Calibration.** For the pepper classifier, calibrated probabilities (Platt/isotonic) would make the predicted scores trustworthy, not just the ranking (AUC).## The engineeringEverything here is importable, tested code: `pytest` validates the from-scratch OLS/ridge/lassoagainst scikit-learn, the data-cleaning invariants, and the effect-size math. See the[repository](https://github.com/deepanshumody/Analysis_RMP_Ratings).