• No results found

We present a slim and interactive browser application capable of visualizing Hi-C contact maps alongside complementary data tracks. Besides Hi-C contact maps genome-wide data, such as ChIP-seq and RNA-seq, can be included in the layout. Bekvaem can be utilized for the visualization of any genomes including mammalian genomes.

6.4. Conclusion 77

Figure 6.4: Screenshot of the web interface of Bekvaem. The user can easily switch between different genomes of interest by loading the respective data files in the selection windows of the browser interface. A Hi-C contact map of wild type E. coli cells is depicted. The resolution of the balanced Hi-C contact map is 10 kbp and the linear color scale ranges from a contact probability of 0 up to 0.035. The complementary data tracks show the ChIP-seq profile of Fis (top) as well as the RNA-seq profile (below) in early exponential phase. The read counts of both experiments are depicted on a linear scale. Data from Lioy et al. [47] and Kahramanoglou et al. [126].

Figure 6.5: Screenshot of the web interface of Bekvaem. A heat map of the intra- and inter- chromosomal Hi-C interactions among chromosomes 2 to 11 of the mouse genome is depicted. The resolution of the balanced Hi-C contact map is 20 kbp and the linear color scale ranges from a contact probability of 0 up to 0.02. The complementary data tracks show the ChIP-seq profile of CTCF (top) as well as the RNA-seq profile (below). The read counts of both experiments were processed using a sliding window size of 1 kbp and are depicted on a linear scale. Data from Nora et al. [115].

6.4. Conclusion 79 0.000 0.008 0.004 0.007 0.006 0.005 0.003 0.002 0 10 20 ChIP-seq counts chr8:0Mb chr8:50Mb chr8:100Mb chr9:20Mb chr9:70Mb chr9:120Mb genome position chr8:0Mb 0.001 chr8:100Mb chr8:50Mb chr9:20Mb chr9:70Mb chr9:120Mb 0.000 0.008 0 10 20 ChIP-seq counts 0.004 0.006 0.002 chr8:0Mb chr8:50Mb chr8:100Mb chr9:20Mb chr9:70Mb chr9:120Mb genome position

A

B

Figure 6.6: PDF export of Bekvaem. The Hi-C contact map of chromosomes 8 and 9 of the mouse genome are depicted in A. square and B. triangular shape alongside the CTCF ChIP-seq profile. The resolution of the balanced Hi-C contact map is 20 kbp and the linear color scale ranges from a contact probability of 0 up to 0.008. The ChIP-seq read counts were processed using a sliding window size of 1 kbp and are depicted on a linear scale. Data from Nora et al. [115].

Chapter 7

Domain Boundary Detection in Hi-C

Maps

A Probabilistic Graphical Model Approach

References

The results presented in this chapter are adapted from

• A. Hofmann, F.Z. Rashid, F. Crémazy, R.T. Dame and D.W. Heermann (2019),

Domain Boundary Detection in Hi-C Maps: A Probabilistic Graphical Model Approach, in preparation, to be submitted to PLoS ONE.

AH and DWH developed the method. AH performed the analysis of the Hi-C maps.

Chapter Summary

To understand the nature of a cell, one needs to understand the structure of its genome. For this purpose, experimental techniques such as Hi-C detecting chromosomal contacts are used to probe the three-dimensional genomic structure. These experiments yield topo- logical information, consistently showing a hierarchical subdivision of the genome into self-interacting domains across many organisms. Current methods for detecting these do- mains using the Hi-C matrix, i.e. a doubly-stochastic matrix, are mostly based on the assumption that the domains are distinct, thus non-overlapping. For overcoming this sim- plification and for being able to unravel a possible nested domain structure, we developed a probabilistic graphical model that makes no a priori assumptions on the domain struc- ture. Within this approach, the Hi-C matrix is analyzed using an Ising like probabilistic graphical model whose coupling constant is proportional to each lattice point (entry in the contact matrix). The results show clear boundaries between identified domains and the background. These domain boundaries are dependent on the coupling constant, so that one matrix yields several clusters of different sizes, which show the self-interaction of the genome on different scales.

7.1 Introduction

Early work using optical microscopy with fluorescent markers established that chromo- somes are not randomly organized in the nucleus [127]. Exactly how the chromosomes are organized could not be further revealed by this method, even though multi-color ex- periments pushed the experimental boundary [128]. At this stage several models have been proposed how the genome is physically organized in space [82, 97, 129–133]. With the 3C technology [134] new data on 3D genome organization became available. Whereas the information coming from the microscopy experiments gives a physical relationship be- tween between points in space, i.e., Euclidean distances on single cell data, the 3C data (and later the Hi-C data [3]) yields topological information loosing the embedding into Euclidean space, i.e., only neighborhood relationships are revealed attached with a certain probability. Furthermore, the information represents an average over many cells. In a way this is very much information one would classify as of mean-field type. Thus, the chal- lenge is to develop a model that is consistent with the mean-field result in the sense that it succeeds to re-embed the topological information into Euclidean space, i.e., geometrical information and topological information need to be reconciled.

A crucial part of this process is to identify the structures and substructures that appear in Hi-C contact maps. Most prominently are the TADs (topologically associated domains). Their defining characteristic is that the interaction frequency within domains is much higher as opposed to that across domains, i.e. the contact matrix resembles a block-diagonal matrix.

There are various different methodological approaches identifying the domain structure in Hi-C contact maps. A first attempt was presented in Dixon et al. [7] and is based on a two-step strategy. Firstly, the 2D contact information is condensed to the directionality index, a 1D measure encoding both downstream and upstream chromatin interactions. In