3.4 Discussion
5.2.5 Retrieval of gene sequences from C jejuni ST-474 gene predictions
The BLAST+ application was downloaded from The National Center for Biotechnology
Information (NCBI)6 to perform stand-alone BLAST searches on the ST-474 gene pre-
dictions in order to find the gene/nucleotide sequences from theC. jejuniNCTC 11168
genome. For a given metabolic housekeeping gene from the CJJ11168 genome, the BLAST hits with maximum identity, alignment length and highest bitscore were retained and the respective open reading frames (ORFs) were retrieved from the BLAST database in FASTA format. The best hits for all of the 25 genes that corresponded with the above selection criteria are consolidated and presented as a blast table in Appendix C. For the
5.2 Experimental procedures 127
reference genomes C. jejuni doylei269.97 and CJJCG8486, the orthologs for the sdhA
gene could not be obtained from their genome sequences. As a result thesdhA gene se-
quence from these genomes could not be included in the comparative analysis along with the other seventeen genomes.
While analysing the seven housekeeping genes, it was observed that the glmM and the
pgmgenes were two different genes involved in two different functions in the genome.
Thepgm gene is a phosphoglyceromutase involved in carbohydrate metabolism (in the
interconversion of 2-phosphoglycerate) with a locus tag and gene ID of Cj0434 and
904759, respectively.7 Whereas, theglmM gene is a bacterial phosphoglucosamine mu-
tase (PNGM) involved in the interconversion of glucosamine-6-phosphate and glucosamine-
1-phosphate in the biosynthetic pathway. The locus tag and gene ID of theglmM gene are
Cj0360 and 904683, respectively. The primers used for the amplification and sequencing
of thepgmgene in the MLST scheme ofC. jejuni were retrieved and assembled against
both pgm and glmM full length genes using Geneious software v5.3.4.8) There was a
perfect assembly with glmM gene in contrast to pgm. Further details are provided in
Figure C.1 in Appendix C. Therefore,glmM was chosen for analysis in place ofpgmin
the comparison of MLST housekeeping genes. There could have been a possibility that
pgm was mispelt in the place of pngm in the MLST scheme for C. jejuniwhen it was
originally established. In a recent report by Sheppard et al. (2011), the authors have ac-
knowledged the renaming of the allelespgmanduncA to glmM andatpA respectively in
later genome annotations and their names (pgm anduncA) being retained in the MLST
scheme to maintain consistency.
5.2.6
Analysis of Guanine-Cytosine content, codon usage, selection
pressure and recombination
The overall GC and GC3 contents of individual housekeeping genes were compared us- ing DnaSP v5 (Librado & Rozas 2009). The number of genes that shared identical GC and GC3 contents between ST-474 and the reference genomes were examined to predict the closer ancestor of ST-474 with respect to the GC contents of the investigated genes. Codon based maximum likelihood analyses were used to investigate selection pressures
7URL: (http://www.ncbi.nlm.nih.gov/protein/218562018)
on individual codons using the Muse & Gaut (1994) and Tamura & Nei (1993) methods
implemented with the HyPhy software package within MEGA v5. A test statisticdN -
dS was used for detecting codons that have undergone positive selection, wheredS and
dN denote the number of synonymous substitutions per site (s/S) and the number of non-
synonymous substitutions per site (n/N) within a codon, respectively (Tamura & Dudley
2007). A positive value for the test statistic is an indication of an over abundance of non- synonymous substitutions. In addition, Tajima’s D test was conducted to test the selection pressure on individual genes as well as individual sites using DnaSP v5 (Tajima 1989, Hudson & Kaplan 1985). CBI and the scaled chi square codon bias indices for individual
genes from allC. jejunigenomes were estimated using DnaSP v5. The mean relative evo-
lutionary rates of nucleotide sites and codons were estimated using MEGA v5 (Tamura & Dudley 2007) for all of the 25 genes using gene alignments from all the 19 genomes. The model used in MEGA analyses the first, second, third and the non-coding positions for their substitution rates within a codon, and the probability of all the four nucleotides (A, T, G and C) getting substituted within a site and/or a codon in order to predict the evolutionary rates of sites and codons (Jukes & Cantor 1969, Tamura & Dudley 2007). Inferences on recombination within each gene under investigation were drawn using Dual- Brothers within Geneious v.5.3.4 (Minin et al. 2005). This function uses a double change- point model that detects spatial variation in the phylogenetic tree topology and spatial variation of the nucleotide substitution process (Minin et al. 2005). The gene sequences were aligned using Geneious v5.3.4 and the recombination detection functionality was applied on each gene alignment to infer changes in the topologies and substitution pro- cesses. In addition, the aligned sequences were examined using DnaSP v5 to identify the sites involved in recombination. The number of recombination events, referred to as
Rm was estimated using DnaSP v5. The relationship between the GC variance and the
recombination events was analysed using a linear model by having Rm as a dependent variable and log GC variance and length as independent variables.