1.2 The Simulation of Biological Systems
1.2.3 Coarse-grained Discrete Methods(Resolution:Nano/Mesoscale)
The scaling of the MD method means it often cannot feasibly be applied to the afore- mentioned ‘large’ systems at atomistic resolution due to the runtime required. Coarse- grained (CG) methods have been developed to lessen the computational load by re- ducing the number of degrees of freedom in the system but conserving the general MD method shown in Figure 1.2.
Chapter 1. Introduction 12
Within large molecules, various intermediate structures are formed as we move up in length-scale: atoms form residues, residues form secondary structure, secondary to ter- tiary and so on. Each of these intermediaries can be considered as the fundamental unit in a coarse-grained simulation technique using the general MD method. Provided that effective force-fields that describe the interactions between these units can be defined, the MD algorithm is able to produce dynamical trajectories that generate the dynam- ics of these intermediate structures at a much lower computational cost than all-atom MD. This discrete coarse-graining approach can group arbitrary numbers of atoms and functional groups together, and so naturally there are multiple different CG methods available to be applied to different length-scales. Izvekov et al. developed a method which generalises the creation of a CG force-field from higher-resolution structural data [50], but perhaps a more well-known example is the MARTINI model [51]. The MARTINI model uses a potential similar to that of MD, but where the fundamen- tal units are groups of atoms. On average, every 4 heavy atoms are represented as a single CG particle, with specific charged, polar, non-polar and apolar particle ‘types’, each with additional hydrogen bonding capabilities. The bonded force parameters are determined from the underlying atomic structure, and non-bonded interactions are still treated using a Lennard-Jones potential for Van der Waals and Coulomb electrostat- ics for charged particles [52]. Once parametrised the CG simulation progresses in a manner identical to MD, as per Figure1.2.
By reducing the number of degrees of freedom in this fashion we lose dynamical in- formation at smaller length-scales, which often corresponds to the fastest motions in the system. However, as we saw in Section 1.2.2, our computational limitations are due both to the number of particles and the smallest time-scale in the system, and so CG simulations using the MARTINI model provide a substantial speed increase in comparison to all-atom MD. Spatial resolution, then, forms somewhat of a trade-off with simulation speed.
This acceleration has allowed CG MD simulations of much larger systems than all- atom MD. Applications include, but aren’t limited to, self-assembly processes, protein conformational change and lipid membrane formation and interactions [52]. A recent study probed the interaction and permeation of particles through a lipid bilayer mem- brane [53]. The lengths of these simulations enabled comparison with experimental results, which in turn allowed conclusions to be drawn about the underlying mecha- nisms responsible for experimental observations. For example, Oroskar et al. were able to show that a gold nano-particle functionalised with hydrophobic ligands, designed
Chapter 1. Introduction 13
to transport drug molecules, weakened the membrane following permeation. The hy- drophobicity of the nano-particle displaced lipid molecules from the membrane, with larger nano-particles causing permanent dissociation and potentially irreparable dam- age. Hydrophilic nano-particles on the other hand were still able to permeate through the membrane, yet caused only minor disruption to the membrane with almost imme- diate recovery of membrane structure once the nano-particle exited. These simulations give evidence that CG MD, specifically those implementing the MARTINI model in this case, capture a significant proportion of the relevant dynamics one would obtain from all-atom MD, but with a much smaller runtime.
1.2.3.1 Gaussian Network Models
Proteins in general can be viewed as a network of ‘backbone’ carbon atoms with func- tional groups attached to them. The backbone carbon network, or Cα atoms, are (mostly) responsible for the overall flexibility of the molecules, so the next level of coarse-graining we may consider is including only those Cα atoms in our calculations, and removing electrostatic interactions altogether. A model in which molecular struc- tures are represented as a series of Cα atoms connected by Hookean springs is known as a Gaussian Network Model (GNM) [54]. The simplest CG potential to describe a GNM containing N Cα atoms is significantly less complex than the MARTINI model,
VGN M = 1 2 N X i=0 N X j=i γij(Rij − R0ij)2, (1.5)
where Rij = Ri− Rj is the distance between each pair of atoms in the structure, R0ij
being the equilibrium distance, and γij is the associated stiffness. Here, there is no
‘type’ defined for each particle, as was the case for MARTINI. Instead, each particle is simply described by its position in space with interactions modelled as pure elastic connections with neighbouring atoms. Equation1.5can be rewritten in matrix form,
VGN M =
1 2∆ ~R
TK∆ ~R, (1.6)
where K is the stiffness matrix of the system, populated with the values γij and applied
to the vector of deviations from equilibrium ~R. A cut-off distance rc is also specified to simplify the calculation further, such that for R0ij > rc, Kij = 0. In other words,
elastic connections are neglected in the calculations for atoms that are too far away from one another.
Chapter 1. Introduction 14
The GNM potential is simple enough that we do not need to run a simulation at all to study the resulting dynamics. It can be shown (see Chapter 3) that all of the possible dynamical information of a GNM is contained within the stiffness matrix, K. Diagonalisation of K gives a series of elastic normal modes. The eigenvectors describe the relative motions of each atom in the mode, and the eigenvalues describe the relative stiffness of each mode.
Although the stiffness matrix is able to incorporate inhomogeneous spring constants γij
between each pair of Cα atoms (known specifically as an Anisotropic Network Model (ANM)), it is still based on a linearisation of the force-field between atoms. Thus, the eigendecomposition gives only first-order dynamical information regardless of the spring constants chosen. Hence, any long-range or non-linear effects are not included in a GNM. However, GNMs have been suprisingly successful in their analysis of the motion of large proteins. A GNM of the entire ribosome, a ∼ 3MDa assembly of 55 protein subunits and 3 rRNA segments, was able to show strong positive and negative correlations between the motions of different regions within the superstructure [55]. Although they verified previous experimental observations, these correlations alone imply a ratchet-like motion that gives an immediate insight into the mechanism by which a segment of RNA may be translated through the ribosome.
1.2.3.2 Dissipative Particle Dynamics
Coarse-grained MD and network models effectively cover the entire range of possible methodologies for coarse-graining large molecules via structural averaging. However, an additional spatially discrete method considers only the overall shape of the system, rather than the underlying atomistic structure.
Dissipative Particle Dynamics (DPD) takes the overall volumetric structure of a bi- ological object and populates that volume with close-packed spherical particles [56]. These spherical particles have no relation to the underlying, higher resolution struc- ture, they are used purely to fill space within the defined volume. On each of these particles we apply a stochastic force, representing the effect of temperature, and a vis- cous force, representing the internal and external frictional forces. We can also include any conservative force we wish, giving us a general force on any particle i, ~Fi, of the
form, ~ Fi= X j6=i ~ Fijc + ~Fijd+ ~Fijt, (1.7)
Chapter 1. Introduction 15
The thermal and dissipative forces are mathematically coupled through the fluctuation- dissipation theorem, such that the energy within the system conforms to equipartion at equilibrium [21]. These terms are both given generalised forms such that the frictional interaction range and strength can be varied a priori without the need for re-derivation of the governing equations. The fluctuation-dissipation theorem will discussed in more detail in Chapter2.
In addition to this coupling we have the conservative forces, ~Fijc. The form of this force used in DPD is often arbitrary, with the only ‘constraint’ other than it being conservative is that it remains as a soft potential that inhibits particle overlap. DPD has a number of advantages compared to single particle models such as simple Brownian or Langevin dynamical models. Unlike single particle models, the total fric- tional force within DPD is determined as a superposition of interactions between all pairs of connected particles. This ensures that Newton’s 3rd law is obeyed, such that
~
Fij = − ~Fji ∀ i, j, and so momentum is explicitly conserved in addition to preseving
equilibrium thermodynamics. As a consequence, if we ensure all additional forces are conservative, a simulation of unbound DPD particles converges to the dynamics pre- dicted by the Navier-Stokes equation. DPD, therefore, is a fluctuating fluid dynamical model [57].
While Brownian dynamics have been used to model large, overdamped systems where the relative local viscosity is so large that momentum is dynamically unimportant [58], if the inertial forces are large and the convergence to hydrodynamical behaviour is important, DPD is the appropriate choice. Recalling our earlier example of the red blood cell at the mesoscale, a 2012 study by Li et al. [59] used a conservative DPD force to model molecular chirality, and were thus able to observe the self-assembly of sickle hemoglobin into fibres. Following this, they modelled the RBC itself as a DPD system with appropriate conservative forces matching experimental observations, and inserted the fibres inside the RBC. They saw that extended growth of the fibres within the RBC can cause it to form the abnormal half-moon shape associated with sickle cell disease, showing the emergence of a macroscopic observable as a result of underlying fluctuating fluid dynamics.