Lab 3 — Instructor answer key

Companion file: solution.R (in this folder), the lab chunk code with both TODOs completed, in chunk order. Run it from the lab root (Rscript solutions/solution.R) and it prints every checkpoint in the handout, including the two AIC values Q2 quotes.

Contents: TODO code answers · model estimands · expected checkpoints · model answers Q1–Q7 · understanding notes

TODO code answers

TODO Answer
1a glm(hospital_death ~ 1, family = binomial(link = "logit"), data = coh). This is the intercept-only (“null”) model. plogis(coef(fit0)) returns exactly the crude proportion, 0.107: the simplest regression is a descriptive summary, which is the lecture’s “machinery is always descriptive” point in one line. Note nobs(fit0) = 1,274 — this model dropped nobody, unlike Step 0’s.
2a glm(hospital_death ~ age + sex + baseline_map + baseline_hr, family = binomial, data = dat_model). This is the Step-0 model minus the exposure, fit on dat_model. Fitting all the Step-2 models on dat_model guarantees every model is fit and judged (not ideal) on the same 1,271 rows. For this particular formula, fitting on coh happens to be harmless (it still contains baseline_hr, so glm drops the same 3 rows in the same order). But for fit_only (~ early_vasopressor, provided) it would not be: fit on coh, that model uses all 1,274 rows and predict() returns 1,274 values, which the AUC call would misalign against 1,271 outcomes (R recycles the logical index). Same-rows-by-construction beats same-rows-by-luck.

Model estimands (the two writing templates)

Descriptimand (Step 1). “Among adult ICU patients with new-onset hypotension (first valid MAP < 65 mmHg during a first ICU stay, no prior vasopressor) who were alive and in the ICU 60 minutes later (person, anchored at the landmark), in US hospitals participating in the eICU tele-ICU program (place) during 2014–2015 (time): the proportion who die before hospital discharge (measure of occurrence: incidence proportion over the hospitalization; outcome: in-hospital death), reported overall and stratified by sex (prespecified stratification, which is the Lesko paper element 4).”

The anchor (landmark survivors) is an eligibility criterion of the descriptive question too, inherited from Lab 1’s protocol. Narrower target populations (e.g., “patients in eICU-participating units”) are valid if the sample-to-target discussion matches; the point is that a population is named, so the gap to the analytical sample can be clearly articulated (not ignored).

Predictimand (Step 2). “For adult ICU patients meeting the Lab 1 eligibility criteria in the unit(s) where the tool will run, in the period after deployment (deployment population, which is not the training sample), at the 60-minute landmark (moment of prediction): predict in-hospital death (outcome) from age, sex, baseline MAP, baseline heart rate, and early-initiation status, all of which are fully knowable at the landmark (inputs available at the moment of prediction). Judge performance by out-of-sample discrimination and calibration in deployment-like data; or better, a”use-driven”” loss such as the sensitivity achievable within the unit’s monitoring capacity (a flag budget) (loss/performance measure).”

The failure not to commit here is the circular reasoning such as “judged by AUC because that’s the norm”. An availability audit (which data are available at tool use) should catch that followup_days (and anything post-landmark, e.g., a later first_vasopressor_offset) is the future written as a column; early_vasopressor is admissible exactly at the landmark because its definition window closes there.

Expected checkpoints

Verified 2026-09-27 against the canonical Lab 1 dataset (1,292 rows → 1,274 complete-case). No seeds, no randomness; sex enters with Female as the reference level (alphabetical).

setup:   nrow(lab1) = 1292; nrow(coh) = 1274 (18 blank outcomes dropped)
Step 0:  glm(hospital_death ~ early_vasopressor + age + sex +
             baseline_map + baseline_hr), family binomial
         term            est      SE      p      OR
         (Intercept)   -6.877   1.083  <.001     —
         early_vaso     0.429   0.509   .400   1.54   <- largest |coef|
         age            0.046   0.007  <.001   1.047
         sexMale        0.290   0.195   .137   1.34
         baseline_map  -0.012   0.013   .339   0.99
         baseline_hr    0.024   0.004  <.001   1.024
         nobs(fit) = 1271  (3 stays missing baseline_hr dropped silently;
           all 3 are unexposed survivors, so Step 1's counts are unaffected)
         AIC(fit) = 793.62  (null deviance 864.79 on 1270 df) <- Q2's
           "adjusted model" AIC
Step 1:  proportion 136/1274 = 0.107; plogis(coef(fit0)) = 0.107,
           nobs(fit0) = 1274
         AIC(fit0) = 867.47  <- Q2's "null model" AIC
         intercept-only model on the SAME 1,271 rows as fit: AIC 866.79
           (= fit's null deviance + 2) — the like-for-like comparison
           Q2's collaborator did not make
         by sex: Male 86/754 = 0.114, Female 50/520 = 0.096;
           crude sex OR = (86/668)/(50/470) = 1.21 vs model sexMale OR 1.34
Step 2:  mean predicted = observed = 0.107 (in-sample MLE; cannot fail)
         AUC full model 0.7246 | without exposure 0.7225 (diff 0.0021)
           | exposure alone 0.5128
Step 3:  exposure row: logOR 0.429, SE 0.509, OR 1.54, 95% CI 0.57–4.16
         2x2 (coh): exposed 21 alive / 6 dead; unexposed 1,117 / 130
           crude OR (6/21)/(130/1117) = 2.45 (exact 2.4549); Lab 2's crude RR = 2.13
         age row: OR 1.047/year ≈ 1.58/decade, p < .001 (no longer
           discussed in the handout; still useful for Q7(d))

If a student’s Step-0 table differs, suspects can be: an uncompleted Lab 1 pipeline variant of the dataset, filtering baseline_hr NAs before Step 0 (nobs becomes 1,271 everywhere and the Step-1 “dropped nobody” beat disappears), or recoding sex (reference level changes the sign/label of the sexMale row).

Model answers, Q1–Q7

Q1. 0.107 vs. Lab 2’s 30-day KM 0.335 — which describes the world? The 0.107 does: it is the proportion of these patients who actually died in hospital, in a world where live discharge is a competing event that ends the possibility of in-hospital death. The KM figure treats discharge as censoring (that is, it estimates the risk in a hypothetical world where discharged patients remain hospitalized and at risk indefinitely). In Lesko et al.’s terms, censoring a competing event “implies an intervention on the data”: legitimate as an explicitly counterfactual (“conditional”) risk, but not a description of the world these patients inhabit. For the descriptimand as written, 0.107 (the cumulative incidence by end of stay) is the estimate.

Q2. The adjusted sex OR (1.34) as “the” sex difference, because its model’s AIC (793.62) beats the null model’s (867.47)? No, the AIC argument answers a question that was never asked. AIC is a penalized in-sample likelihood: it ranks models by how well they fit these data, not by which question their coefficients answer. A lower AIC for the adjusted model says that age, MAP, HR, and exposure carry information about death (nobody should dispute this) and says nothing about whether a sex contrast conditional on those variables is the contrast the descriptimand named. “It’s the better model” is the “true-model myth” (C&MB): pick the best-fitting model, then read every conclusion off its coefficients (another way to think about it: better with respect to what?). What the 1.34 actually answers is: “how do the odds of death differ between men and women of the same age, baseline MAP, heart rate, and exposure status”, which is a conditional contrast within strata the descriptive question never mentioned, whose meaning also leans on the model’s constraints (linearity in age, no interactions). It is not wrong (though in the marginal vs conditional adjustment lecture, we’ll discuss why it might be wrong); it is the answer to a different, and here unasked, question. Two smaller issues: (i) the two AICs are not even comparable — fit0 used 1,274 rows and the Step-0 model 1,271, and AIC only compares models fit to identical data; on the same 1,271 rows the intercept-only AIC is 866.79 (the Step-0 printout’s null deviance 864.79 + 2). This changes nothing about the argument but shows the comparison was made carelessly. (ii) The handout calls 1.34 “age-and-MAP-adjusted”; the model also conditions on heart rate and exposure. A legitimate descriptive reason to move past the crude 1.21 is a declared standardization target: if the descriptimand is a sex comparison in a named target population whose age distribution differs from the sample’s (or comparisons across populations are intended), directly standardizing the sex-specific risks to a stated standard population is a different but prespecified descriptimand — Lesko’s element 4, with the auxiliary variable’s role (nuisance, standardized over) declared in advance. The distinction is purpose stated up front vs. “the better model.”

Q3. Two assumptions the descriptive claim needs; one it does not. Needs (any two): (i) sampling/selection: the analytical sample stands in for the named target population (the demo is a non-random subsample of a proprietary database; hospitals self-selected into tele-ICU; the landmark restricts to 60-minute survivors); (ii) missingness: “Expired” in the discharge field means in-hospital death, and the 18 blank statuses are missing at random; (iii) target population: if standardizing, that the standard population is the intended one; (iv) measurement: devices were used to obtain measurements that we converted to data. Are we describing measurement properties or actual underlying biological / physiological mechanims?. Does NOT need: conditional exchangeability (no-confounding). A descriptive estimand attributes nothing to any cause (there is no counterfactual contrast for a confounder to distort) so confounding has no referent. (Lesko warns that “confounding” concerns are often smuggled into descriptive work as reflexive adjustment.)

Q4. Ranking predictors by p-value. The wrong procedure because a p-value measures evidence against a zero coefficient at this sample size and this predictor prevalence, not contribution to out-of-sample performance under the stated loss: the exposure has the largest coefficient and p = 0.40 (rare predictor → huge SE) while contributing ~0.002 to AUC. The two rankings measure different things and only performance under the loss is the predictive arbiter (C&MB 5.2.2/5.2.4; their gas-enema example did selection by p and never assessed prediction at all). Add overfitting: in-sample p-screening plus in-sample performance is doubly optimistic; a real roadmap validates out of sample.

Q5. Flags change treatment; what happens to the algorithm’s performance at next year’s recalibration? If clinicians treat flagged patients early and the treatment works, flagged patients die less often than the model predicted: the model appears to over-predict among the flagged, calibration degrades, and a naive refit learns a weaker (or inverted) risk signal. The model partially “unlearns” the pattern that made the flags useful. The training data cannot warn you because they were generated under the old treatment policy; deployment changes the data-generating process, and the change is caused by the model itself. That is Vansteelandt & Steen’s point about predictions used for treatment decisions: the moment forecasts feed interventions, “what would happen if we act” is a causal quantity, and a purely associational model has no answer to it.

Q6. 2014–2015 training data vs. the deployment population. Any two concrete shifts, e.g.: (i) practice drift such as post-2016 sepsis guidance (SEP-1 bundles, earlier norepinephrine) changes both the treatment mix and outcome rates can lead to systematic miscalibration even if discrimination survives; (ii) case mix: non-tele-ICU hospitals, different age structure or admission thresholds can lead to predictor distributions shift, performance must be re-established; (iii) measurement: different monitor mix (arterial line vs. cuff MAP), different EHR discharge-status coding, leading to the “same” predictors and outcome being measured differently. Each is a “transportability” failure the training data has no inforamtion on; the fix is evaluation (and possibly recalibration) in deployment-like data.

Q7. Sorting the six assumptions. Identification layer: (a) exchangeability given the four covariates, (c) positivity, (e) consistency / well-defined strategies. Model layer: (b) constant conditional effect (no product terms), (d) functional form (linearity in age — and MAP and HR — on the log-odds scale). Neither: (f) — the missing-at-random assumption for the 18 blank outcomes belongs to a different conversation (missing data / analytical-sample formation) and is inherited by all three of the lab’s questions, not just the causal one.