• No results found

3.4 Discussion

5.2.5 Retrieval of gene sequences from C jejuni ST-474 gene predictions

The BLAST+ application was downloaded from The National Center for Biotechnology

Information (NCBI)6 to perform stand-alone BLAST searches on the ST-474 gene pre-

dictions in order to find the gene/nucleotide sequences from theC. jejuniNCTC 11168

genome. For a given metabolic housekeeping gene from the CJJ11168 genome, the BLAST hits with maximum identity, alignment length and highest bitscore were retained and the respective open reading frames (ORFs) were retrieved from the BLAST database in FASTA format. The best hits for all of the 25 genes that corresponded with the above selection criteria are consolidated and presented as a blast table in Appendix C. For the

5.2 Experimental procedures 127

reference genomes C. jejuni doylei269.97 and CJJCG8486, the orthologs for the sdhA

gene could not be obtained from their genome sequences. As a result thesdhA gene se-

quence from these genomes could not be included in the comparative analysis along with the other seventeen genomes.

While analysing the seven housekeeping genes, it was observed that the glmM and the

pgmgenes were two different genes involved in two different functions in the genome.

Thepgm gene is a phosphoglyceromutase involved in carbohydrate metabolism (in the

interconversion of 2-phosphoglycerate) with a locus tag and gene ID of Cj0434 and

904759, respectively.7 Whereas, theglmM gene is a bacterial phosphoglucosamine mu-

tase (PNGM) involved in the interconversion of glucosamine-6-phosphate and glucosamine-

1-phosphate in the biosynthetic pathway. The locus tag and gene ID of theglmM gene are

Cj0360 and 904683, respectively. The primers used for the amplification and sequencing

of thepgmgene in the MLST scheme ofC. jejuni were retrieved and assembled against

both pgm and glmM full length genes using Geneious software v5.3.4.8) There was a

perfect assembly with glmM gene in contrast to pgm. Further details are provided in

Figure C.1 in Appendix C. Therefore,glmM was chosen for analysis in place ofpgmin

the comparison of MLST housekeeping genes. There could have been a possibility that

pgm was mispelt in the place of pngm in the MLST scheme for C. jejuniwhen it was

originally established. In a recent report by Sheppard et al. (2011), the authors have ac-

knowledged the renaming of the allelespgmanduncA to glmM andatpA respectively in

later genome annotations and their names (pgm anduncA) being retained in the MLST

scheme to maintain consistency.

5.2.6

Analysis of Guanine-Cytosine content, codon usage, selection

pressure and recombination

The overall GC and GC3 contents of individual housekeeping genes were compared us- ing DnaSP v5 (Librado & Rozas 2009). The number of genes that shared identical GC and GC3 contents between ST-474 and the reference genomes were examined to predict the closer ancestor of ST-474 with respect to the GC contents of the investigated genes. Codon based maximum likelihood analyses were used to investigate selection pressures

7URL: (http://www.ncbi.nlm.nih.gov/protein/218562018)

on individual codons using the Muse & Gaut (1994) and Tamura & Nei (1993) methods

implemented with the HyPhy software package within MEGA v5. A test statisticdN -

dS was used for detecting codons that have undergone positive selection, wheredS and

dN denote the number of synonymous substitutions per site (s/S) and the number of non-

synonymous substitutions per site (n/N) within a codon, respectively (Tamura & Dudley

2007). A positive value for the test statistic is an indication of an over abundance of non- synonymous substitutions. In addition, Tajima’s D test was conducted to test the selection pressure on individual genes as well as individual sites using DnaSP v5 (Tajima 1989, Hudson & Kaplan 1985). CBI and the scaled chi square codon bias indices for individual

genes from allC. jejunigenomes were estimated using DnaSP v5. The mean relative evo-

lutionary rates of nucleotide sites and codons were estimated using MEGA v5 (Tamura & Dudley 2007) for all of the 25 genes using gene alignments from all the 19 genomes. The model used in MEGA analyses the first, second, third and the non-coding positions for their substitution rates within a codon, and the probability of all the four nucleotides (A, T, G and C) getting substituted within a site and/or a codon in order to predict the evolutionary rates of sites and codons (Jukes & Cantor 1969, Tamura & Dudley 2007). Inferences on recombination within each gene under investigation were drawn using Dual- Brothers within Geneious v.5.3.4 (Minin et al. 2005). This function uses a double change- point model that detects spatial variation in the phylogenetic tree topology and spatial variation of the nucleotide substitution process (Minin et al. 2005). The gene sequences were aligned using Geneious v5.3.4 and the recombination detection functionality was applied on each gene alignment to infer changes in the topologies and substitution pro- cesses. In addition, the aligned sequences were examined using DnaSP v5 to identify the sites involved in recombination. The number of recombination events, referred to as

Rm was estimated using DnaSP v5. The relationship between the GC variance and the

recombination events was analysed using a linear model by having Rm as a dependent variable and log GC variance and length as independent variables.