Selection bias
Selection bias occurs when systematic error in the recruitment or retention of study participants (i.e., sampling from a source population) leads to distortion in the measure of association between exposure and outcome. This bias can result from the sampling approach used by researchers or from factors related to the study variables that influence study participation and/or retention.4 The result is that the observed measure of association differs from the “true” measure of association.
Selection bias in case-control studies
Selection bias in case-control studies arises due to differential selection of cases or controls.
To illustrate, imagine that all cases and controls in a source population of 10,000 individuals were included in a case-control study that measured the association between an exposure and outcome (Table 4).1 We’ll consider this the hypothetical “gold standard” in that the odds ratio represents the true association if we were able to study the entire source population. To focus on selection bias, we will assume there is no confounding or information bias.
Table 4. Hypothetical case-control study including the entire source population

Now imagine that we sample from this population with an overall sampling probability of 50% for cases and 10% for controls (Table 5). Here, we achieve the same sampling probability for both exposed and unexposed within each of the case and control groups. Individuals within the cases and within the controls have the same probabilities of being selected regardless of exposure status. The odds ratio is unbiased.
Table 5. Unbiased case-control study

Finally, imagine that we use the same overall sampling probabilities for cases and controls as just described. However, unknown to the researcher, the sampling probabilities are different for the exposed (60%) and unexposed (40%) within the cases (Table 6).1 Although the overall sampling probability for cases is still 50%, the probability of selection depends on exposure status. Here, differential probabilities of selection resulted in an erroneous overestimated measure of association compared with the true measure of association (OR = 6.0 vs OR = 4.0).
Table 6. Biased case-control study

Other types of selection bias that can arise in case-control studies include control selection, Berkson’s bias/paradox, and differential surveillance or diagnosis.
Control selection
Consider a case-control study assessing the association between coffee and pancreatic cancer. In this study, cases and controls were selected from a group of patients who were hospitalized by the same physicians who had diagnosed the cases’ disease.1 The idea of this approach was to make the selection of cases and controls similar. It turned out that individuals in the control group were often hospitalized for gastrointestinal conditions, who were also less likely to drink coffee due to their condition.25 As a result, the control group had a low prevalence of exposure, which overestimated the odds ratio for the association between coffee drinking and pancreatic cancer. In contrast, when population-based controls were used (coffee drinking is representative of the source population), the study found no strong association between coffee and pancreatic cancer.26
Berkson’s bias
Berkson’s bias is a form of selection bias in case-control studies that arises when cases and controls are selected from a hospital.1 Individuals who are admitted to hospital are more likely to be exposed to risk factors or disease than the general population. Therefore, controls selected from a hospital tend to overrepresent the prevalence of risk factors and disease and thus bias the associations toward the null.
Selection bias in cohort studies
Selection bias in cohort studies can occur due to systematic error in recruitment or retention. Selection bias can occur in retrospective cohort studies (where the outcome is already known) if the selection of exposed and unexposed participants is somehow related to the outcome. Selection bias in prospective study designs is different from that in retrospective study designs because participants are selected before the outcome has occurred, but it may occur for reasons such as differential losses to follow-up.
Loss to follow-up
In cohort studies, selection bias may occur due to differential retention of participants related to exposure and outcome. This is referred to as differential losses to follow-up, where individuals who are lost to follow-up during the study are different from those who remain under observation until the outcome or the end of the follow-up period.1 Individuals who are lost to follow-up (i.e., due to death, non-response, non-compliance, migration) can have different probabilities of the outcome than those who remain, resulting in biased measures of association.
Healthy user bias
Healthy user bias occurs when treated individuals are healthier than untreated individuals due to other factors, such as tendency for health-seeking behaviours and higher socioeconomic status. Let’s return to the example we discussed in the first section. Past observational studies suggested that hormone replacement therapy (HRT) could reduce coronary heart disease (CHD) in women. However, randomized controlled trials showed that HRT might actually increase the risk of CHD in women. The spurious association found in the observational studies were due to the following issues of health user bias:
- Women on HRT were more health conscious, had higher socioeconomic status, and better access to health care than women who were not on HRT.
- Self-selection of women who tend to be healthier into the HRT user group.
- Compliance bias due to individuals who adhere to medication tend to be healthier than those who do not.
Indication bias
Indication bias (or confounding by indication) occurs when the reasons for a particular treatment (e.g., disease severity, prognosis) is related to the outcome.27 When an individual’s disease severity is high, they’re at an increased risk of an undesired outcome, and thus may be more likely to undergo treatment.
For example, let’s say that you’re conducting a cohort study among individuals with hypertension, comparing the risk of stroke in antihypertensive medication users versus non-users. Consider that individuals who have a greater number of risk factors for cardiovascular disease (e.g., family history, symptoms, lifestyle factors) are more likely to undergo antihypertensive therapy compared with those with fewer risk factors. Therefore, the risk of stroke would be higher among those who receive antihypertensive treatment because of their worse health status. Indication bias is an important consideration in pharmacoepidemiology.