• No results found

4.3 Model building

4.3.2 Selection of covariates

The selection of medical, lifestyle, and socio-demographic covariates for this research was based on the literature review discussed in Chapter 2. This means that all covariates were previously shown to be risk factors that contributed in explaining survival variations in AMI patients or primary cardiovascular risk groups. Interaction effects within and between all groups of risk factors were not studied before, except for limited interactions with the main exposure, sex, and age. For this research,

95

all second-order interactions were tested to examine survival variations in greater detail. To obtain the leanest model possible where the prediction error is minimised, a selection procedure was carried out to select the most important interaction effects. The prediction error is the difference in observed and predicted outcome values (Harrell, 2001; Hosmer et al., 2008). The prediction error includes the bias between observed and predicted outcome values and the variance of the predicted outcome values. The minimal prediction error is observed with the true underlying model of the data. A model that includes less covariates than the true model ‘underfits’ the data by not capturing the underlying trend of the data. In other words, this model excludes important covariates in predicting the outcome and therefore has high bias and low variance of the predicted outcome values. A model that includes more covariates than the true model ‘overfits’ the data by not only capturing the underlying trend but also the noise of the data. In other words, this model includes important and unimportant covariates in predicting the outcome, and therefore has low bias and high variance of the predicted outcome values. To find the true model, several stepwise methods were proposed.

Stepwise methods select the most important covariates to include in a model solely based on their significance that is specified by some mathematical criterion (Harrell, 2001; Hosmer et al., 2008). Forward selection starts with an empty model and adds one covariate at the time where the covariate that is added first explains the greatest percentage of (the remaining unexplained) variation in the outcome. This process is continued until no covariate significantly contributes to the model in explaining the variation in the outcome. Backward elimination starts with a full model that includes all covariates and removes one covariate at the time where the covariate that is removed first contributes the least to the model. This process is continued un- til all the covariates in the model contribute significantly in explaining variations in

96

the outcome. Bidirectional elimination starts with an arbitrary model and considers adding, removing, or swapping a covariate with each step until the model does not change. This stepwise method is thus a combination of forward selection and back- ward elimination. All methods have the risk of excluding important and/or including unimportant covariates (Harrell, 2001; Hosmer et al., 2008). Backward elimination is preferred over the other stepwise methods because it is the least likely to exclude im- portant covariates that are only significant in relation with other covariates (Harrell, 2001; Hosmer et al., 2008).

Stepwise methods are based on some mathematical information criterion. This information criterion is an estimate of the prediction error. With the model selection process, the model with the lowest information criterion would be selected. There are numerous information criteria of which Akaike information criterion (AIC) and Bayesian information criterion (BIC) are most commonly used, and defined as (Hastie et al., 2009):

AIC = −2LL + 2k, (4.48)

BIC = −2LL + k ln(m), (4.49)

where LL is the partial log likelihood of the Cox’s model, k is the number of coefficients estimated by the model, and m is the limiting sample size, which is the number of events in survival analysis (Harrell, 2001; Hastie et al., 2009). As the sample size increases to infinity, model selection based on AIC has a non-zero and model selection based on BIC has a zero probability of overfitting the data, i.e. selecting unimportant covariates thereby making the model too complex (Hastie et al., 2009). However, with a finite limiting sample size, model selection based on BIC could select too few covariates, making the model too simple. As the stepwise method based on

97

some information criterion is an automated process, it is important to check whether the resulting survival model is biologically plausible (Hosmer et al., 2008).

For this research, backward elimination of second-order interaction effects was carried out. For the automated process, BIC was used as information criterion. It was assumed that only interaction effects would be removed from the model and not main effects as their significance in explaining survival prospects were established in previous studies. The automated backward elimination process where BIC was minimised, resulted per age cohort in a survival model with 20 to 25 interaction effects with p-values ranging from <.001 to .500. Backward elimination was continued until all interaction effects were significant at 1% level in the decomposition of Cox’s model by ANOVA (analysis of variance). Due to large sample sizes, the significance level was set at 1% to obtain only the interaction effects that contribute the most to the model in explaining survival variations. The low significance level comes with the cost that interaction effects between covariates that have relatively rare categories will most likely be excluded. This manual backward elimination process resulted per age cohort in a survival model with three to eight interaction effects. A common set of effects in the survival models was chosen to have the same interpretation of the effects at each age cohort. If an interaction effect was found in the majority of the models (≥50%), it was included in all models. For the analysis that makes use of multiple imputation, the variable selection had a two-step approach. At the first step, interaction effects found in the majority of the models on the imputed datasets of an age cohort, were included in the age-specific model. At the second step, interactions found in the majority of the age-specific models, were included in the final overall model. It was checked whether the interaction effects were biologically plausible by comparing them with previous studies and clinical guidelines.

98

Related documents