In this section we’ll use an actual study and our CNODES Simulated Data Sets to walk through the process of applying propensity scores. The study we’re using is a past CNODES project that has been adapted to work with the simulated data, which are less complex than actual administrative data. If you have not done so already, the study protocol can be downloaded here:
Research question
Does high potency statin use increase the risk of incident diabetes?
Data sources
If you haven’t already downloaded the CNODES Simulated Data Sets, you’ll want to go do that now. You can return to the Getting Started section of this module for further instructions. These data include a cohort data set, hospitalization data set, dispensations data set and medical claims (physician visits) data set. The cohort data set includes demographics, index date, exposure, outcome and covariates.
Exposure
Higher potency statins (HPS): rosuvastatin, atorvastatin, and simvastatin (Exposure = 1)
vs
Lower potency statins (LPS): all other statins (Exposure = 0)
Outcome measure
Incident diabetes (i-diab): first occurrence of a hospitalization with any diagnosis for diabetes or physician visit with diabetes diagnosis (ICD-9 250; ICD-10 codes E10, E11, E12, E13, E14) or a prescription for insulin or an oral antidiabetic medication (ATC A10A, A10B, A10X).
Index date is the date of the outcome (i.e. diagnosis date, death date or drug dispensation date).
Predefined covariates at baseline
Cohort entry year, health care utilization: [> 4 distinct drugs (generic chemical name; yes=1 or no=0), > 4 diagnoses (yes=1 or no=0)], hypertensive disease (ICD9 401.x – 405.x, , ICD10 I10.x-I15.x), hypercholesterolemia (ICD9 272.x, ICD10 E78.x), peripheral vascular disease (ICD9 443.x, ICD10 I73.x), congestive heart failure (ICD9 428.x, ICD10 I50.x), loop diuretics (ATC C03CA).