Information bias
Information bias can occur due to systematic error in the measurement of exposure or outcome.1 Misclassification can result from inaccuracies in methods used to collect data, and can be described in two forms: differential and nondifferential.10
Differential misclassification
Differential misclassification occurs when measurement error depends on the values of other variables, or when the rate of measurement error differs for different study groups.10 For example, differential misclassification of exposure is when individuals with the disease are less or more often misclassified as having the exposure than the individuals without the disease. Differential misclassification of disease is when misclassification of disease if different for the unexposed group and exposed groups.4 Differential misclassification can lead to spurious associations in either direction — an apparent association when one does not truly exist, or an apparent lack of association when one does truly exist.10
Let’s illustrate differential misclassification using an example.4
Consider a hypothetical cohort study of smoking and emphysema. Emphysema is a condition that can go undetected without medical attention. Therefore, if smokers are more likely to seek medical attention than non-smokers (e.g., due to negative health effects of smoking), then emphysema could be detected more frequently in smokers than in non-smokers. This is an example of differential misclassification of the disease dependent on exposure status. Underdiagnosis of emphysema in non-smokers would spuriously strengthen the association between smoking and emphysema.
Recall bias is a type of differential misclassification bias, which can occur when one study group reports true exposures differently than another study group, or when one study group recalls exposures that did not actually occur while another study group does not. Recall bias is possible in case-control studies where participants’ memory of past exposures is required. For example, cases may recall certain exposures more frequently, especially if they attribute the exposure to the disease. Consider a case-control study of congenital malformations where information is required from the mother. Mothers of infants with malformations (i.e., cases) may more frequently recall past exposures that might have contributed to the unfortunate outcome, compared with mothers of healthy infants (controls).4
Nondifferential misclassification
In contrast, nondifferential misclassification occurs when measurement error does not depend on the values of other variables, or when the rate of measurement error is the same for all study groups. Nondifferential misclassification is not related to exposure status in a cohort study, or to case/control status in a case-control study. Instead, it results from issues inherent to the data collection methods that are used.10 The effect of nondifferential misclassification is more predictable in nature versus that of differential misclassification — it tends to shift the measure of association towards the null (i.e., we’re less likely to detect an association even if one really exists).10
Let’s illustrate nondifferential misclassification using an example.10
Consider a hypothetical cohort study of alcohol drinking and incidence of laryngeal cancer. Assume that the true incidence rate for drinkers is 50 per 10,000 persons, and for non-drinkers is 10 per 10,000 persons (rate ratio = 5.0). Let’s say that half of true drinkers fail to report as drinkers and are misclassified as non-drinkers (but this misclassification is unrelated to whether they develop laryngeal cancer), resulting in an incidence rate of 30 per 10,000 persons rather than 10 per 10,000 persons. We see that the effect of nondifferential misclassification biased the association towards the null (rate ratio = 1.7).