• No results found

The published literature on lung cancer prediction models indicates that the focus has been in publishing new models rather than validating existing prediction model. As a result, there was a varied standard of validations for models with some being extensively validated while others have only been considered in the original article. The epidemiological models have been the most extensively reviewed, which is expected since they were published earlier. Since then there has been a shift to consider incorporating genetic markers in prediction models. These published clinical models are relatively new and therefore have rarely been evaluated outside the original article. Finally, the TSCE models were seldom used and have never been considered beyond the original study. These targeted high risk, niche populations so were not practical in new environments.

When validating the epidemiological models the results are often presented for different models using distinct datasets so direct comparisons between models has been difficult (Table 3.5). The calibration was commonly reported and in most instances the model demonstrated a good calibration based on the

Hosmer-Lemeshow p-values. This suggests a model could be used to accurately provide predictions for an individual’s risk. The AUC was also commonly reported for the epidemiological models. The AUC results ranged from a weak 0.57 (Spitz Model) to a very strong 0.86 (PLCO Model). The leading reported

AUC results were for the PLCO, PLCOM2012, PLCOM2014 and Hoggart models. The results would suggest

these models have the highest potential to be a successful selective screening tool. To assess the ability of a model as a selective screening tool the prediction rules need to be validated, unfortunately, only for

the PLCOM2012 Model have these been extensively evaluated. This model at the optimal risk threshold

(1.34%) where it improved upon the previously implemented NLST criteria to identify a target group for screening.

At this stage, the optimal risk threshold for the remaining epidemiological models needs to be identified and the prediction rules externally validated. Janssen et al. (2008) recommends that all models should be externally validated to quantify their predictive performance through calibration, discrimination and classification predictors [45]. Additionally, comparing all the models in the same dataset is recommended. This should address limitations with the current reporting, where direct comparisons between the models is difficult. This may also identify a different leading model for distinct populations, based on ethnicity or location, which may allow different optimal screening programmes to be conducted worldwide.

The review found that the models which consider clinical factors (such as blood testing and scans) do not offer a universal improvement over epidemiological models. The discrimination was commonly reported with an AUC between 0.639 − 0.875 and the majority of results between 0.7 − 0.79. The lack of improvement for clinical models was highlighted by comparisons between the COSMOS (Bach), Extended Spitz and Extended LLP Models and their original models where any improvement (AUC/calibration) was minimal. The slightly disappointing results may be a result of not identifying the correct genetic factors associated with lung cancer risk. Further work is being conducted to identify the key markers which could lead to new prediction models that improve upon the existing literature. However, it is important that the new clinical markers assist in developing a robust prediction model rather than publishing a poor prediction model, as has been observed with some published clinical models. The performance of a clinical model is critical; and before a clinical model would be considered as a selective screening tool the costs to apply the model will need to be justified, which can be demonstrated by a high performing model to identify a high risk target population.

Clinical models are still at a very early stage of development and have usually been published since 2010. As a direct consequence they have not been properly tested at this stage and commonly only validated in the original article in an internal validation. A conclusion of the systematic review was that it would be beneficial if every clinical model was external validated to develop stronger conclusions about the models’ performances.

In conclusion the review found some flaws in the current validations of the models that will be addressed in our study. Firstly, models have inconsistently been compared with each other in the same dataset making it difficult to draw direct comparisons between models. This has been observed for models developed for many different diseases, and the current reporting of validations across different studies has been poor [46]. Secondly, prediction rules have been poorly reported in validation studies and for some models these have never been reported (Table 3.5). It could be argued that the lack of clear guidelines, through assessing the prediction rules, on how to optimally apply models has prevented these being incorporated in selective screening programmes. Therefore, this study will address the poor validation reporting for lung cancer prediction models. The models will also be compared to the current screening programmes to evaluate if there is conclusive evidence that a leading prediction model or criteria has been identified. This will be conducted for the epidemiological lung cancer prediction models using a series of datasets received from the International Lung Cancer Consortium (ILCCO). These models will be considered because the variables will be commonly collected in the ILCCO datasets. Additionally, these models demonstrated the potential to improve upon current screening guidelines whether the UKLS or NLST criteria. Therefore, a leading model could be identified and implemented.

3.9

Summary

The systematic review identified 29 different models that predict lung cancer risk in individuals. These were classified as epidemiological, clinical assessment and TSCE models. There were ten epidemiological lung cancer prediction models identified; the Bach, LLP, Spitz, African-American, Hoggart, Pittsburgh, two

PLCO versions, PLCOM2012 and PLCOM2014 models. The epidemiological models showed potential to be

utilised as a clinical utility with some good performances for calibration, discrimination and prediction rules in validations. However, the models had often been validated independently in distinct datasets so direct comparisons between the models was difficult. The current reporting of validations was inconsistent with some validation studies not considering prediction rules. Additionally, when a model had been evaluated in multiple studies the prediction rules were often reviewed at different risk thresholds. This resulted in no clear indication into a leading model or being able to provide recommendations into the risk threshold that will allow the model to perform optimally. The standard of validations did not allow a confident recommendation into a leading model and how this may be utilised as a selective screening tool.

Based on the systematic review the next stage of research should be to address the lack of consistent reporting for models in the validations. An external validation, comparing all the models in the same dataset, is required which should provide better understanding into the different model performances.

To perform an external validation individual patient level datasets (IPD) will be collected and prepared, which is detailed in the next chapter. The collected datasets will then be analysed in the subsequent chapter to identify any limitations with the datasets that could influence the model results in the validation. Then the validation will be conducted and the results presented and analysed.

CHAPTER

4

Dataset Collection, Preparation and Imputation

4.1

Introduction

To conduct external validations on the identified lung cancer prediction models datasets were made available through the International Lung Cancer Consortium (ILCCO). Prior to performing the validation, there needs to be confidence that the datasets would allow the models to be accurately evaluated. Therefore, reported information on the variables in the datasets were reviewed and modified when required. This included harmonising the reporting of the variables so they were applicable to the models’ specifications, removing participants with unreliable information and imputing missing information. The modifications made to any of the variables in the datasets were recorded and reported in this chapter.

4.2

Objectives

This chapter aims to detail how the datasets were obtained and prepared to perform an external validation of lung cancer prediction models. The chapter will:

1. Detail the ILCCO application process to acquire datasets for the project. 2. Introduce the datasets that were made available by ILCCO.

3. Present which models were applicable to which dataset based upon having complete information for each variable required by the model.

4. Present how the datasets required modifications to be compatible for the models.

• Report any modifications to the variable information to allow it to be in the correct form required by the models.

• Detail any participants removed because of unreliable information that could negatively affect the model validation.

5. Impute missing information in the datasets where possible. • Introduce the imputation objectives.

• Present different imputation methods and identify an appropriate method. • Conduct and report the imputation.