• No results found

Gene conversion, linkage, and the evolution of repeated genes dispersed among multiple chromosomes.

N/A
N/A
Protected

Academic year: 2020

Share "Gene conversion, linkage, and the evolution of repeated genes dispersed among multiple chromosomes."

Copied!
16
0
0

Loading.... (view fulltext now)

Full text

(1)

Copyright 0 I990 by the Genetics Society of America

Gene Conversion, Linkage, and the Evolution

of Repeated Genes Dispersed

Among Multiple Chromosomes

Thomas Nagylaki

Department of Ecology and Evolution, The University of Chicago, Chicago, Illinois 60637 Manuscript received April 3, 1990

Accepted for publication June 4, 1990

ABSTRACT

The evolution of the probabilities of genetic identity within and between the loci of a multigene family dispersed among multiple chromosomes is investigated. Unbiased gene conversion, equal crossing over, random genetic drift, and mutation to new alleles are incorporated. Generations are discrete and nonoverlapping; the diploid, monoecious population mates at random. The linkage map is arbitrary, but the same for every chromosome; the dependence of the probabilities of identity on the location on each chromosome is formulated exactly. The greatest of the rates of gene conversion, random drift, and mutation is E << 1. Under the assumption of loose linkage (i.e., all the crossover

rates greatly exceed e, though they may still be much less than Yz), explicit approximations are obtained for the equilibrium values of the probabilities of identity and of the linkage of disequilibria. The probabilities of identity are of order one [i.e., O(l)] and do not depend on location; the linkage disequilibria are of O ( E ) and, within each chromosome, depend on location through the crossover rates. It is demonstrated also that the ultimate rate and pattern of convergence to equilibrium are close to that of a much simpler, location-independent model. If intrachromosomal conversion is absent, the above results hold even without the assumption of loose linkage. In all cases, the relative errors are of O ( t ) . Even if the conversion rate between genes on nonhomologous chromosomes is considerably less than between genes on the same chromosome or homologous chromosomes, the probabilities of identity between the former genes are still almost as high as those between the latter, and the rate of convergence is still not much less than with equal conversion rates. If the crossover rates are much less than Yz, then most of the linkage disequilibrium is due to intrachromosomal conversion. If linkage is loose, the reduction of the linkage disequilibria to O ( e ) requires only 0(-h E )

generations.

I

n a recent paper (NAGYLAKI 1988), the evolution of the probabilities of genetic identity within and between the loci of a multigene family was investi- gated analytically. Unbiased gene conversion, equal crossing over, random genetic drift, and mutation to new alleles were incorporated. In this model, the greatest of the rates of gene conversion, random drift, and mutation is E

<<

1 . Under the assumption of loose linkage ( i e . , all the crossover rates greatly exceed E ,

though they may still be much less than %), explicit approximations were obtained for the equilibrium values of the probabilities of identity and of the link- age disequilibria. T h e former are of order one [ie.,

O( l)] and do not depend on location; the latter are of

O ( c ) and depend on location through the crossover

rates. It was demonstrated also that the ultimate rate and pattern of convergence to equilibrium are close to that of a much simpler, location-independent

model. In the absence of mutation, the ultimate rate of convergence gives the characteristic time for the decay of genetic variability. If intrachromosomal con- version is absent, the above results hold even without the assumption of loose linkage. In all cases, the rela- tive errors are of O ( E ) . If the crossover rates are much

Genetics 126: 261-276 (September, 1990)

less than 1/2, then most of the linkage disequilibrium is due to intrachromosomal conversion. If linkage is loose, the reduction ofthe linkage disequilibria to O ( c ) requires only 0(-ln E ) generations.

In NAGYLAKI (1988), the linkage map is arbitrary, and the location dependence of the probabilities of identity was formulated exactly. It was posited also that the pairwise conversion rates between genes at nonhomologous loci are location independent. For

repeated genes dispersed among multiple chromo- somes, however, this assumption may often be un- realistic because pairwise conversion rates between genes on nonhomologous chromosomes should be appreciably smaller than the corresponding rates be- tween genes on homologous chromosomes. T h e rather scarce data generally support this expectation, though the former rates are surprisingly high, appar- ently roughly comparable to or within about an order of magnitude of the latter UINKS-ROBERTSON and

PETES 1985, 1986; LICHTEN, BORTS and HABER

1987). T h e purpose of this paper is to treat this important case by extending the model and analyses of NAGYLAKI (1 988) to multiple chromosomes.

(2)

spectively, review the pertinent empirical and retical literature. Consult WATTERSON (1 989a,b) for the most recent mathematical investigations. SAWYER

( 1 989) has devised statistical tests for detecting gene conversion.

OHTA and DOVER (1 983) were the first to study the evolution of multigene families dispersed among mul- tiple chromosomes. They included unbiased gene con- version, equal crossing over, random genetic drift, and mutation to new alleles, and they supposed that each of these evolutionary forces is weak. They ne- glected (1) the position dependence of the probabili- ties of identity, (2) the linkage disequilibrium between loci on nonhomologous chromosomes, (3) interactions between loci on homologous chromosomes, and

(4)

symmetric heteroduplexes,

WEIR, OHTA and TACHIDA (1985) incorporated interchromosomal linkage disequilibrium into their recursion relations (which they did not analyze), but otherwise made the same assumptions as OHTA and DOVER (1 983). Neglecting interchromosomal linkage disequilibrium does not, however, reduce the recur- sion relations of WEIR, OHTA and TACHIDA (1985) to those of OHTA and DOVER (1 983), which have some terms missing.

Here, we shall posit only that gene conversion, random drift, and mutation (but not necessarily cross- ing over) are weak. Then approximations 1 and 2 above receive considerable support from the work of NACYLAKI and BARTON (1986) and the analyses in this paper, respectively, and assumption 4 is probably reasonable in many cases ($ ORR-WEAVER and SZOS-

TAK 1985). It is, however, important to remove as- sumption 3 because, as discussed above, pairwise con- version rates between genes on homologous chromo- somes may often exceed the corresponding rates between genes on nonhomologous chromosomes, which are an essential feature of this problem.

It is necessary to compare the crossover rates with the strength of the other evolutionary forces. In this paper, we posit that there are n repeats on each of k chromosomes, each of which has the same linkage map. Let rrz denote the crossover rate between distinct loci y (= 1,

2,

. . .

, n ) and z (Zy) on the same chro- mosome. We neglect intralocus crossing over. Set

r,,,i,, = min r y z , T,,,;,~ = max ryz. (1)

We assume that the greatest of the rates of gene conversion (defined more precisely in the next sec- tion), random drift, and mutation is E

<<

1. For some

of our analyses, we shall posit loose linkage: E

<<

rmin,

which may, of course, hold even if rmin

<<

$4.

Unlike OHTA and DOVER (1 983) and WEIR, OHTA and TACH- IDA (1985), we shall not impose the restriction rmax

<<

v2.

y . 2 Y ' L

lfl

fl

FIGURE ].-The respective pairwise conversion rates a=,, as, a,,,

ah, and cy, for intrachromatid, sister-chromatid, classical, semi- classical, and ectopic interactions.

T h e nomenclature of conversional interactions in multigene families has not yet become standardized. Our terminology, summarized in Figure 1, differs from that of PETES and HILL (1988), which would be inconvenient here. In this paper, the adjectives (i)

intrachromatid, (ii) sister-chromatid, (iii) classical, (iv)

semiclassical, and (v) ectopic refer to interactions be-

tween genes (i) on the same chromatid, (ii) at nonho- mologous loci on sister chromatids, (iii) at homologous loci on homologous chromosomes, (iv) at nonhomol- ogous loci on homologous chromosomes, and (v) on nonhomologous chromosomes. (In our model, gene conversion occurs immediately after chromosome du- plication, so interactions between genes at homolo- gous loci on sister chromatids have no effect.) Collec- tively, intrachromatid and sister-chromatid interac- tions are intrachromosomal, whereas classical, semiclassical, and ectopic interactions are interchro- mosomal.

In the next two sections, we shall formulate our exact, general difference equations and establish some preliminary results. In the following section, we shall investigate the equilibrium values of the probabilities of identity and of the linkage disequilibria. Then w e shall examine the ultimate rate and pattern of conver- gence to equilibrium. In the final section, we shall summarize and discuss our results.

FORMULATION

Generations are discrete and nonoverlapping; the diploid, monoecious population mates at random. T h e life cycle starts with infinitely many gametes. We use probabilities of identity to summarize the genetic structure of the population; these provide much im- portant biological information, but do not fully specify the state of the population. T h e term "identity" must be interpreted in accordance with the type of data available: at the most detailed level, it refers to identity of the DNA sequences of two genes; if less information

is available, it can signify coincidence of restriction sites or the ability to hybridize.

(3)

Evolution of Repeated Genes 263

Let f , ( t ) denote the probability that two genes on distinct randomly chosen gametes, at locus y on ho- mologous chromosomes, are identical. Thenf, repre- sents the expected homozygosity at locus y immedi- ately after fertilization; hy = 1

-

f,, the expected heterozygosity, is a measure of genetic variability in the population at locus

y.

Let gl,yz(t) denote the prob- ability that two genes on a randomly chosen gamete, one at locus y and the other at a different locus z (z # y) on the same chromosome, are identical. Clearly,

gl,yz(t) is an index of the amount and pattern of homology between repeats within a chromosome. Let l l , y z ( t ) denote the probability that two genes on dis- tinct randomly chosen gametes, one at locus y and the other at a different locus z (z # y) on a homologous chromosome, are identical. Thus, ll,yz incorporates variation both between loci and between homologous chromosomes. Let g2,zf(t) denote the probability that two genes on a randomly chosen gamete, one at locus

z and the other at locus { (we allow { = z ) on a nonhomologous chromosome, are identical. Finally,

let Z2Jt) denote the probability that two genes on distinct randomly chosen gametes, one at locus z and the other at locus {on a nonhomologous chromosome, are identical.

Thus, we have five sets of probabilities of identity: one within chromosomes (gl,yz), two between homol- ogous chromosomes ( J and Z1,yz), and two between

nonhomologous chromosomes (g2,=f and Z2,zr). Since we are considering n repeats on each of

k chromo-

somes with the same linkage map, and since we shall posit a chromosome-independent conversion pattern, we have therefore assumed thatf,, gl,yz, and are the same for every chromosome and that g2,rf and 12,z(

are the same for every pair of chromosomes. If this simplifying assumption holds initially, then it holds for every subsequent generation. Since our recursion relations will be linear, their equilibrium is unique and therefore unaffected by our assumption. Further- more, Theorem 2.4 and Remark 2.9 of BOUCHER and

NAGYLAKI ( 1 988) imply that this assumption does not affect the ultimate rate and pattern of convergence to equilibrium. So, all our results will be independent of this simplification.

We posit the cycle shown below; x designates the vector of the probabilities of identity, and the prime signifies the next generation. T h e population number is infinite, except immediately after population regu- lation, when it is N .

Gametes

-

Zygotes

-

00, x fertilization w, x mutation

Adults Adults

w, x * chromosome duplication w, x *

-

Adults

-

conversion w, x** regulation

Adults

-

Gametes

N , x

* *

gametogenesis QJ, x!

Since gametes fuse wholly at random, a proportion

1/N of zygotes are produced by self-fertilization and the corresponding probabilities of identity within and between zygotes are equal.

We suppose that every allele mutates to new alleles at rate u (0 u

<

1). This model of infinitely many alleles was proposed by MAL~COT ( 1 946, 1948, 195 1 )

for identity by descent and by WRIGHT (1948) and

KIMURA and CROW (1964) for identity in state. After

mutation, we have

(f?'

g&z, - T y z 9 g h , k ? z f )

= 4 f , , g1,yzt h , y z , g2,zf, h , Z f ) ? (2)

where v = ( 1

-

u)'.

T o incorporate gene conversion, we posit the fol- lowing: (i) An interaction between two alleles cannot produce a third allele. (ii) Each interaction involves the formation of heteroduplexes between two re- peated genes or double-strand-break repair (SZOSTAK

et al. 1983). T h e heteroduplexes may be either sym- metric (HOLLIDAY 1964) or asymmetric (MESELSON

and RADDING 1975). (iii) Interactions occur immedi-

ately after chromosome duplication. (iv) A given pair

of genes participates in at most one interaction per generation ( i e . , neither gene interacts, they interact with each other, or one of them interacts with a third gene and the other does not interact). If one of the five types of interaction depicted in Figure 1 occurs, the two genes that participate in it are chosen at random. (v) All mismatches are repaired. (vi) Parity obtains in the initiation of asymmetric heteroduplex formation, the repair of mismatches, and the occur- rence of double-strand breaks. (vii) If symmetric het- eroduplexes are formed, the direction of correction of one heteroduplex is independent of that of the other. (viii) Crossing over is not associated with gene conversion. Consult NAGYLAKI and PETES (1982) and

NAGYLAKI (1983, 1984a) for discussion of these as- sumptions.

(4)

Y

w

2

5

c

9

Y

2

5

c

9

FIGURL.: 2.-The probabilities of identity before (singly starred variables) and after (doubly starred variables) gene conversion. Three pairs of chronlosomes are shown for each individual. The singly subscripted letters adjacent to the chromatids represent genes

a t the distinct loci y. w , z, {, [, and 7.

interaction probabilities

As a check, note that the sum of the denominators in (3) is

(yz)4nk(4nk

-

2 ) .

( 4 )

Since self-interactions do not occur and we have ex- cluded interactions between genes at homologous loci on sister chromatids, therefore (4) is precisely the number of gene pairs that can interact, as it must be. Solely for the purpose of our derivation, we shall say that a gene is converted if its DNA is replaced by DNA from another gene or by DNA synthesized from that of another gene. We adhere to the convention that each strand of a homoduplex formed by two identical genes is “corrected” with probability ‘/2. Let

a + b and a

$,

b denote the events that a is converted

to b and that it is not, respectively. If y, u, and 6 ( y

+

u

+

6 = 1) represent the respective probabilities of

asymmetric heteroduplexes, symmetric heterodu- plexes, and double-strand-break repair, then (NAGY-

LAKI 1984a)

p

= P ( a + b I I&) =

% ( 2

-

y), ( 5 4

q = P ( a + b , b $, a

I

I a b ) = % ( I i- 6). (5b)

We set

From (6) we see easily that 1 C p S 2 . Furthermore, p = 1 if and only if there are no symmetric hetero- duplexes (a = 0), and p = 2 if and only if both asymmetric heteroduplexes and double-strand-break repair are absent ( u = 1).

T h e definitions

a, =

”’

j = w , s, 6 , nk(n

-

1) ’

a ( ) = - , PPO nk

enable us to reduce (3) to

P ( L , , , ) =

aw/(2p),

W I c J = a s / ( 2 p ) , (84

P(zalm,) = a0/(4p), P ( z a l c : , ) = ab/(4p), (8b)

P(Zaldl) = a e / ( S P ) . (84

Then (5), (6), and (8) yield the unconditional proba- bilities

P(a1 + c1) = !haw,

P ( a l + c l , c1

$, a l )

= L/27aw, (9a)

P(a1 + c y ) = !has,

P ( u ~ + C Y , c:!

$,

a ] ) = ?‘ZTLY,, (9b) P(a1 + a:*) = ?Lao,

P(a1 3 u3, a3

+

a ] ) = 1/47a0, ( 9 4

P(a1 4 c:3) = % a b ,

P(al + c 3 , c3

$,

a l ) = % T a b , (9d) P(a* + d l ) = % a e ,

P ( a l + d l , d l

$,

a ] ) = %ra,. (9e)

Many of our results simplify considerably if the pair- wise classical and semiclassical conversion rates are equal (ao = a b ) . It will be convenient to introduce the total pairwise intrachromosomal conversion rate ( a ] ) and the sum of the pairwise intrachromosomal and semiclassical conversion rates (cry):

=

+

ax, =

+

01,

+

a b . (10)

(5)

Evolution of Repeated Genes 265

interchromosomal homologies within individuals

(f;"

*

, 1:; , and I $ $ ) will differ from the correspond- ing homologies between individuals (F,**, Lf,T, and L;$). We define the averages

analogous definitions hold for Zi,=, and the

starred variables. T o derive our equations, we use the nomenclature of genes in Figure 2 and apply the method in NACYLAKI and BARTON (1986) and NACYLAKI (1988). Long but straightforward calcula- tions lead to

+

[ l

-

a0

-

7aw

-

as

+

Yz(n

-

l)ae(Ttz

+

Tt<)

+

[ l - nap

-

n(k

-

l)ae]l&c

+

1/2n[a2

+

(k

-

2)ae](T:*z

+

T&)#

In the gametes of the next generation, we have (WRIGHT 193 1 ; M A L ~ C O T 1946,1948; KIMURA 1963)

f;

= 8 ( l

+f?*)

+

( 1

-

28)Fy**, ( 1 3 4

g ; , y z = ( 1

-

ryz)g?$

+

ry2lt;T7 ( 1 3b)

= O(g;lr;

+

1;;)

+

(1 - 28)Lt$, ( 1 3 ~ )

g L r =

G(g$$

+ l $ Z ) ,

(1 3 4

1 4 ~ = 8 ( g $ f

+

I $ $ )

+

( 1

-

28)L;$, (13e)

where 8 = 1 / ( 2 N ) . We have neglected in ( 1 3) the second-order terms that arise because sister chroma- tids may differ after conversion.

Substituting (12) into ( 1 3) and then (2) into the result leads to the recursion relations for our model. This system depends on the order of the evolutionary forces in the life cycle. However, our assumptions concerning gene conversion are plausible only if it has a low probability per gene, and we lose no biological generality by positing weak mutation and random drift. I f

t = max(u, 8, a0

+

nap

+

nka,)

<<

1 , (14)

but ryz is arbitrary for every y and z, we obtain (y # z)

f j = 8 + [ 1 - 2 u - O

-

(n

-

1)a2

-

n(k - l ) a e ] f r ( 1 5 4

+

(n

-

q a p

TI

,y

+

n(k - I ) ~ ,

ip,y,

g;,yr = ( 1 - r y z ) ( T ( Y w

+

a s )

+

%( 1

+

T)ry,(Yb

+

%[( 1

-

ryz)ab

+

'?y.a]]

* [ A

+.L+

(n

-

l ) ( t ; , y

+

t;.J]

(6)

+{(1 -rYz)[l - 2 ~ - a o + a b + ( l - ~ ) a , (15b)

-

-

n(k

-

1 ) a e ]

+

r y z ( a 0

-

ab)lgl,yz

+

{( 1

-

ryz)(aO

-

a b )

+

ryz[ 1

-

2u

-

a0

+

I/4n(k

-

1 )ae(g2,J

+

g 2 , z

+

L Y

+

T*,J,

+

[ l

-

224

-

0

-

nap

-

n(k

-

l)ae]zl,yz

+

%(n

-

1)a2(

+

T1.J

+

%n(k

-

1)ae(T2,y

+

T 2 J ,

+

'/4(n- l ) a e g I , z + g l , < + T I , , +

TIii,<)

+!/2(1-2u-na2

+

1/2(3

-

-

nap

-

n(k - l)a,])~!l,~,

K~~

= y2a2(fy

+ f ~ )

+

egl,yi

(1 5 4

g6,zr= '/4(1

+

?)ae '/4%(& ff)

( 1 5 4

-

[n(k

-

1)

-

' / 4 ( 1

-

~ ) ] a e l ( g 2 , z <

+

h , z < )

+

%n[ap

+

(k

-

2)ae].

* (&,z

+

g2,f

+

L Z

+

T*,f),

+

%(a

-

l)ae(Tl,z

+

TI,,)

+

eg2,zr

+

[ 1

-

2u

-

0

-

nap

-

n(k - l)ae]12,1r

+

!hn[a*

+

(k

-

2)a,](T2,z

+

&J).

1 6 , z f = 1 / 2 a e ( f r

+

fr>

(1 5 4

We have simplified writing by not indicating that (1 5) has been linearized in the rates of mutation, conver- sion, and random drift. As a result of this linearization, these evolutionary forces are additive, and (15) is independent of their order.

In general, the system (1 5) involves 2n2

+

n inde- pendent homologies, and in the next section we shall see that some position dependence and linkage dis- equilibrium are present even at equilibrium. Despite this high dimension and complexity, we shall derive accurate approximations that provide a rather com- plete analytic understanding of (1 5).

PRELIMINARY RESULTS

In this section, we discuss some general properties of (1 5) that motivate, illuminate, and aid the more detailed analyses in the next two sections.

Linkage disequilibrium: Crucial to our analyses is the investigation of the deviations from linkage equi- librium, a subject of intrinsic biological interest. We utilize

as condensed measures of average intrachromosomal and interchromosomal linkage disequilibrium. In an infinite population (6 = 0 ) , it is easy to express them

in terms of gametic and allelic frequencies (NAGYLAKI

1988). Subtracting (15c) and (1 5e) from (15b) and (1 5d), respectively, and appealing to (1 6) lead to (y #

2 )

D I'+ = P I . J h , Y z

+

q l , y z , (1 7 4

D i z r = P 2 . @ 2 , z <

+

q z . z r , (1 7b)

in which

= (1 - ryJ[ 1

-

2u

-

a0

+

a b

-

nayP

-

n(k - 1)ae] ( 1 8 4

+

ryz(a0

-

a b )

-

0,

q 2 ~ f = ' / 4 a e [ l

+

7

-

fL

-fr

+

(1 - T)&,,r

+

(n

-

l)(&z

+

Dl,,)] ( 1 9b)

+

'/4n[(~g

+

(k

-

2)c~~](Dp,~

+

&,<),

Dl.? = gl,?

-

TI,^,

Dn.2 = g2.Z

-

r 2 . z . (20)

-

Observe that

P l , y = 1

-

r y z

+

O(E), P 2 , L f =

'/2

+

O(€) (21)

as E 3 0.

Position dependence: Since the inhomogeneous terms ( q l , y z and q 2 , z f ) in ( 1 7 ) are generally nonzero,

therefore (15) neither preserves nor converges to linkage equilibrium. Furthermore, crossing over pro- duces and maintains position dependence through the linkage map { r y z ) . To prove these assertions, we ex- hibit directly how (15) generates both linkage dis- equilibrium and position dependence even if the link- age map initially depends only on separation and if initial linkage equilibrium and exchangeability are assumed. Put x = y

-

z and posit ruz = rx and the initial condition

(fy,

g 1 , y z , l l , y z , gP,Zf, 1 2 . 2 0 =

(f,

1 1 , 1 1 , 1 2 , 1 2 ) (22)

for every y, z (y # z), and

{.

We iterate (15), use (1 1) and (17), and display the position dependence of our variables in an obvious notation: ( J ' , g;,=, 1

I'

, g2/, l i ) ,

(f",

gCrz, Ifx, g&, 14'), and

(f;,

g & , 11,, g2,=t,

12%).

Since genes on nonhomologous chromosomes as- sort independently, we might naively have expected position independence of the corresponding probabil-

(7)

Evolution of Repeated Genes 267

ities (g2,,f and 1 2 , 2 f ) . We have just demonstrated, how-

ever, that position dependence within chromosomes induces position dependence between chromosomes. Only in the special case of two repeats per chromo- some (n = 2) does symmetry lead to the preservation of exchangeability, though linkage disequilibrium is still generated.

Parameter dependence: Scrutiny of (15) reveals that (i) Equations 15a, c, and e are independent of the molecular parameter 7 , the linkage map ( r Y I L ] , and the

classical conversion rate ao, and they depend on a,, a,, and ab only through their sum, a2 ; (ii) Equation 15d involves T , but otherwise satisfies (i), and (iii)

Equation 15b shares none of these invariance prop- erties. Since our approximations for the equilibrium homologies and the asymptotic rate and pattern of convergence will ultimately be derived from (15a, c, e) in linkage equilibrium, they will also possess the invariance properties (i). Only the invariance proper- ties (ii) apply to our approximate equilibrium value of

Df,,f ( 6 2 , 2 c ) , however, because its calculation requires (1 5d) as well. None of these invariance properties is expected for f i l , y 2 , because its evaluation_requires also

(15b); nevertheless, for loose linkage, will turn out to be approximately independent of ao.

For classical gene conversion, symmetric heterodu- plexes occur with high probability at some loci in some organisms, but their probability is often low (FOGEL, MORTIMER and LUSNAK 198 1; NAGYLAKI and PETES 1982; ORR-WEAVER and SZOSTAK 1985). If these results generalize to interactions between re- peated genes, then T = 1 should often fit the data,

and in this case (1 5 ) depends on a, and a, only through their sum, a ] .

The asymptotic rate of convergence: From the row sums of the nonnegative matrix of coefficients in (1 5), we find that its (real) maximal eigenvalue, Xo,

satisfies (GANTMACHER 1959, pp. 63, 68)

1

-

2u

-

max[8, s, '/4(1

+

7)a,]

( 2 3 4

I X0 I 1 - 2u, where

s = (1 - rrn)(7a,

+

a,)

+

!h(1

+

~ ) r , a b , (23b)

rrn =

{

rlnax, !h( 1

+

7)ab 2 7%

+

as,

rnmin, !h( 1

+

7)ab

<

T a w

+

as. ( 2 3 4 We conclude that (15) converges to a unique equilib- rium at the asymptotic rate Xo.

Special cases: Several special cases are of interest. Equations 14 and 19 inform us that

41.Y' = O(t), q 2 , r c = O ( € ) (24)

as + 0. From (1

7)

and (21) we see that after a sufficiently long time

1

D ~ . = c ( t )

I

<<

1 for every z and

5; whereas

I

D l , y 2 ( t )

I

<<

1 holds generically after suf-

ficient time only if t

<<

ry2. Consequently, for our

general analyses, we shall suppose that linkage is loose

(t

<<

rlni,,), or more precisely, that t -+ 0 with ( r ]

fixed. If, however, intrachromosomal conversion y: 1s

absent (a1 = 0), then all terms in q1,y2 except the last, negligible one are proportional to r,,, and this obser- vation will enable us to derive approximations without positing loose linkage, i.e., these results will hold uni- formly in the linkage map (ry2] as E 0.

If mutation and random drift are absent (u = 8 = 0), the allelic frequencies in the entire multigene family are conserved because conversion is unbiased. Therefore, even without mutation, in an infinite pop- ulation genetic variability is preserved. Let ai denote

the frequency of the allele A, in gametes. Then the probability that two genes chosen at random from distinct gametes are identical is

- f + T

nk 1 - n - l - ' I + n ( k - nk

')

T2

= a?, (25)

i

where

j = 1, 2. Since a, is conserved for every i, so is (25), as is easily verified from (1 5a, c, e).

In the special cases (i) a, = 0 or k = 1, (ii) n = 1, and (iii) a p = a t , the system (1 5) can be related to the system (68) in NAGYLAKI (1988).

If ap = 0 or k = 1, then (1 5, a, b, c) simplify to the single-chromosome Equation 68 of NAGYLAKI (1988), and (17a), (lsa), and (19a) agree with Equations 69 and 70a of NAGYLAKI (1988).

Next, set n = 1 in (15a, d, e). With a single gene on each chromosome, there can be, of course, no position dependence. For comparison, take (unrealistically) ryz = !h in Equation 68 of NAGYLAKI (1988), thereby eliminating position dependence from that system. Then obvious notational changes show at once that (1 5a) and (1 5e) are respectively identical to (68a) and (68c) of NAGYLAKI (1988). We conclude that the approximate equilibrium homologies and the asymp- totic rate and pattern of convergence are the same as those in NAGYLAKI (1984b, 1988) with n and a re- placed by

k and

ae, respectively. Furthermore, if T =

1, then (1 5d) is the same as (68b) of NAGYLAKI (1 988), and therefore the idenification of the two models becomes exact. In this case, a multigene family with one gene on each of k chromosomes is equivalent to a (fictitious) family of

k

unlinked genes on a single chromosome.

Finally, suppose a2 = a e , replace z!: in (1 5e) by yz (y # z), and insert gl,Yz = and l l , y r = 12.yz into the right

(8)

a=, the approximately equilibrium homologies and the asymptotic rate and pattern of convergence are iden- tical to those in NAGYLAKI (1984b, 1988) with nk

instead of n. This simplification is due to approximate linkage equilibrium (and hence independence of the linkage map) and the conversional symmetry of the special case a2 = a e .

The exchangeable approximation: Neglecting po-

sition dependence in (15) yields the exchangeable approximation:

f'

= 0

+

[ 1

-

2u - 6

-

(n

-

1)a2

-

n(k

-

l)ae3f

( 2 7 4

+

( n

-

l)anl1

+

n(k - l)aelp, g; = (1 - r)(T(Y,

+

(us)

+

%( 1

+

7)rab

+

[( 1

-

r)ab

+

ral]f

+ { ( l - r ) [ l - 2 u - a o - 7 a , - a ,

-

( n

-

1)ab

-

n(k - 1)a)e]

+

r[aO

+

(n

-

2)abI)gl (27b)

+

(( 1

-

r)[aO

+

(n

-

2)ab]

+ r [ l - 2 u - a o - a 1

-

%(2n

-

3

+

T)(Yb

-

n(k

-

l)ae])Z]

+

Yzn(k

-

l)ae(g2

+

l z ) ,

~ ; = ~ ~ f + e g ~ + [ 1 - 2 ~ - ~ - ~ ~

g; = 1/4( 1

+

7)a,

+

1/2aef

(274

-

n(k - l)ae]Z1

+

n(k

-

l ) a e Z p ,

+

Yz(n

-

l)ae(gl

+

ZI)

+

Yz(

1

-

2u ( 2 7 4

-

[n

-

1

-

7)]ae)(g2

+

lz),

16 = a,f+ (n

-

l)a,l1

+

eg,

+

(1

-

2u

-

e

-

na&

( 2 7 4

in which r signifies the mean crossover rate

(6

NAGYLAKI and BARTON 1986).

If classical and semiclassical conversion are absent

(ao = a b = 0), there are no symmetric heteroduplexes (T = I), and crossing over is slow ( r

<<

%), then (27)

reduces to the model of WEIR, OHTA and TACHIDA (1 985). If we impose the approximation gz = 1 2 , then (27a, b, c) reduce to the corresponding equations of OHTA and DOVER (1983), but (27d, e) do not lead to their equation for the homology between chromo- somes. Refer to the introduction for further discus- sion.

We shall derive all our analytic results from the exact model (15) and use the exchangeable model (27) only for numerical comparisons. T h e computa- tions of NAGYLAKI and BARTON (1 986) indicate that the exchangeable approximation for intrachromoso- mal conversion in a tandem array is always qualita- tively correct and that it is quantitatively accurate if

crossing over is either much faster (E

<<

rmin) or much slower ( t

>>

rlllax) than the other evolutionary forces. For loose linkage, the linkage disequilibria are small and position dependence is weak. For tight linkage, the linkage disequilibria may be substantial, but posi- tion dependence is weak because it is a consequence

of small terms proportional to ruz. T h e system (27) is exact if n = 2.

EQUILIBRIUM

Here, we derive approximations for the equilibrium values of the probabilities of identity and of the link- age disequilibria. First, we deduce general formulas under the assumption of loose linkage. Next, we re- move this restriction for pure interchromosomal con- version; only the expres:ion for the intrachromosomal linkage disequilibrium D l , y r requires modification. Fi-

nally, we present some numerical results.

The general model: We fix ( r Y L ) and let E -+ 0. At

equilibrium, (1

7),

(2 l), and (24) inform us that

6LYZ

= O(€), ( 2 8 4

= W E ) . (28b)

We write (1 5a, c, e) at equilibrium, invoking (16) and (28) to eliminate i l , y r and

&r:

j j = e + [ 1 - 2 u - e - ( n - l ) a p - n ( k - l ) a e ] j ;

1

+

Mn(k - 1)ae(t-2,y

+

T 2 . J

+

O(E2),

i 2 , r f = Y Z a e ( f z

+ j ~

+

~ z ( n

-

l)ae(tl,z

+

TI,{)

+

[ 1

-

2u

-

na2

-

n(k - l)ae]lp,r{ (294

+

%n[az

+

(K

-

2)ae](t;,,

+

t-2,f)

+

O(t').

From (29) we infer that the unique equilibrium of (15) has the form

j ;

=f+

O(€),

k l + = i l

+

O(E), ( 3 0 4

i l . y r = i l

+

O(€),

=

f,

+

O(t), i2,zl =

i p

+

O(t) (30b) as t + 0 with ( r y z ) fixed. Substituting (30) into (29)

and letting t -+!l,Awepbtain three simultaneous linear

equations for

(f,

Z I

, Z Z ) , which have the solution

(9)

Evolution of Repeated Genes 269

A = 2u(2u

+

nka,)[2u

+

8

+

nay:!

+

n(k

-

1)ae]

+

8(2ua2

+

na,[az

+

(k

-

l)a,]). (3 Id)

On account of the general discussion of parameter dependence presented in the previous section, the approximation (3 1) is independent of 7 , ( r y L ) , and ( Y O ,

and it depends on a, , a,, and a b only through a2. In agreement with the study of special cases presented below (26), the result (3 1) reduces to the simpler one in NAGYLAKI (1984b) if (i) a, = 0 or k = 1, (ii) n = 1, or (iii) a2 = a,.

Some limiting cases of (3 1) are instructive and prq- vide checks. As u + 0, all variability disappears:

f,

il

,

i2

+ 1. As N + 00 (0 + 0), all homology disappears:

f,

i l , i2

+ 0. As a 2 , a, + 0, all interlocus homology disappears and the intralocus homology converges to the classical value for the balance between mutation and random drift (MAL~COT 1946, 1948, 1951; KI-

MURA and CROW 1964):

i,

,

f2

*

0 and

j +

e/(2u

+

e)

= 1/(1

+

4 ~s f o . ~ ) (32) Some simple inequalities illuminate the rather com- plicated solution (3 1). Su:cessive use of (32), (31b), (3 la), and (3 Id) leads to f

<fo.

Replacing B by O i l in

the numerator of (3 lb) immediately informs us that

f

>

il. Therefore,

O < i ] < f < f O < l . ( 3 3 4

Deleting 2u from the denominator of (31c) demon- strates that the probability of identity between genes on nonhomologous chromosomes is less than the weighted average of the probabilities of identity be- tween genes on homologous chromosomes:

l , < - f + -

-

1 - n - 1 - 11

n n

Since

i,

<

f,

replacing

f1

by

f

in the numerator of (3 IC) yields another upper bound on

iz,

whereas re- placing? by

i,

there yields a lower bound:

T h e last two numerical examples in Table 1 show that both

il

>

i,

a?d

ilA<

i,

are possible, b,ut

iy

all the examples with l 1

<

1 2 , it was found that

1 1

z 1 2 .

Thus, as in previous models (NAGYLAKI 1984a, b, 1988; NAGYLAKI and BARTON 1986), interaction with other loci lowers the intralocus homology

(f

<

fo).

T h e homology between loci is less than that within loci

(i,,

i2

<

f).

T o estimate the linkage disequilibria, we insert (2 I ) ,

(28), and (30) into (1

7)

and (1 9) at equilibrium. We find (y # z)

bI,,,

= r ~ I ( 1

-

r y z ) ~

+

% f f b C I

+

0 ( 2 ) , (34a)

&,<

= ! / z a , C 2

+

O ( 2 ) (34b)

B = a l ( l

-f)

-

(1 - T)a,(l

-

il), (35a)

as t + 0 with ( r Y L ) fixed, in which the constants

C, = 2(1

-f)

-

(1 - 7)(1

-

G,),

(35b)

j = 1, 2, can be evaluated from (3 1).

In accordance with the general discussion of param- eter dependence preseFted in the previous section, our approximation for D2,,, is independent of {ry,) and a. and depends on a,, a,, and a b only through a2

,

whereas

bl,y,

shares none of these invariance proper- ties except approximate independence of 0 0 . If semi- classical conversion is absent (ab = 0), then

I

f i l , y z

I

decreases as r,. increases. This generalizes the corre- sponding result for a single chromosome (NAGYLAKI

1988). In the absence of symmetric heteroduplexes (7

= l), from (35) we see at once that B , C1, C 2

>

0, so

(34) implies that

61,yz,

b2,,<

>

0, and

bI,Yz

decreases with r,,. Negative linkage disequilibria do occur if 7

<

1 (NAGYLAKI 1984a, 1988; NAGYLAKI and BARTON 1986). I f rY,

<

1/2, then the intrachromosoma!-conver-

sion term generally dominates in (34a),

I

D l , Y r

I

de-

creases with r,,, and

1

b2,r<

1

<<

I

fil,,,

I

.

Had we omitted

the terms of order try, in (1 5b), then the term in (34a) due to semiclassical conversion would have been miss- ing.

Pure interchromosomal conversion: In this sub- section, we demonstrate that in the absence of intra- chromosomal conversion (a1 = 0), the results (28),

(30), (31), and (34b) hold for arbitrary linkage, ie.,

uniformly in (rYz) as t + 0, and we derive a uniform

replacement for the loose-linkage formula (34a). From (1 7b), (1 9b), and

(2

1) we see that (28b) holds uniformly in {r,,] as t + 0. Since we have already proved (28a) for fixed r,,, we may assume here for convenience that ryz

<

Vi.

Then the assumption

CY,

= 0 and (14) and (18a) enable us to show

1

-

pl,,.

>

ryz

+

Y4t. (36)

Application of (14) and (28b) to (1 9a) yields

I

i l . y r

I

cryrab

+

dta, (37)

for suitable positive constants c and d. With the aid of (36) and (37), we deduce easily from (1 7a)

I f i ~ , ~ ,

I

c a b

+

4dae = O(t), (38)

which establishes the uniformity of (28a).

By the proof in the previous subsection, we now conclude that (30) and (31) with a2 = a b hold uni-

formly in ( r y z J as t + 0.

(10)

TO evalute 6,,yz, note first from (1 8a) that

1 - P I , Y Z = byz[l

+

O(t)], ( 3 9 4

where

byz = rrz

+

2u

+

0

+

a.

(39b)

+

( n

-

I)ab

+

n(k - l)ae.

Substituting (19a), (30), (28a), (34b), (35b), and (39a) into (1 7a) leads to

6,,YL

= b;’{V2ryrab[~1

+

~ ( t ) ]

(40)

+

’/4n(k - 1)a4[c2

+

0(t)])

uniformly in { r y z ] as t + 0. T h e constants C1 and

C2

can be calculated from (35b) and (31).

This is our only result that depends on the classical conversion rate ao. For loose linkage (t << rmin), the second term in (40) is negligible, and we recover the special case of (34a) with a , = 0. If 7 = 1, then

61,yz

>

0 and Dl,yz increases asAry. increases. In the absence of ectopic conversion,

I

Dl,yr

I

increases with ryz for all 7 (NAGYLAKI 1988). It is important to recognize that the intrachromosomal disequilibrium cannot be cor- rectly approximated in this delicate case if the terms of order try, are neglected in (1 5b), for this would lead to the absence of the first term in (40), which would then be accurate only if the unrealistic condition r,,,

<<

t held. If ryz

<<

%, then (40) is generally much

smaller than (34a).

Numerical results: Previous numerical investiga- tions (NACYLAKI 1984a; NAGYLAKI and BARTON

1986) and the above analytic results indicate that 7

seldom has a strong effect on the behavior of this model; furthermore, as remarked in the previous sec- tion, 7 = 1 often fits the data. Therefore, we set T =

1 in all our computations. Then the exact model (1 5 )

depends on a, and a, only through a , . Since all our approximations except (40) are independent of (YO,

and the dependence in (40) is usually weak, therefore we make the natural simplifying assumption a0 = ab

throughout.

T o parametrize crossing over, we posit that the repeats on each chromosome are equally spaced on the linkage map. This should be most accurate for tandem repeats. Let signify the probability of equal crossing over between consecutive repeats. Suppose that at most one crossover occurs per generation among the repeats on each chromosome, as is reason- able if there is complete positive interference or, more likely, if (n

-

2 ) p

<<

1. Then

r = %(n

+

1)p (41)

represents the probability of crossing over between two randomly chosen loci on the same chromosome

(OHTA 1983; NAGYLAKI 1984a). We use r not only in the exchangeable approximation (27), but also in (34a) and (40).

Even with the above simplifications, the equilibria depend on eight parameters: n , k,

N

(or

e),

u, P, a , ,

a b , and a e . We vary these one at a time, as displayed

on the left in Table 1. Unless otherwise specified, the parameters values in Table 1 are n = 2, k = 3, N = 5

For each parameter value on the left, the first line is our perturbation approximation, calculated from (31), (34), and if a1 = 0, (40). We obtain the second

line by evaluating numerically the equilibrium of the exchangeable model (27) and then employing the exchangeable form of (1 6). T h e relative error of the perturbation approximation,

X io4 , = IO-’,

p

=

IO-^

, and a , = a b = a, = lo-‘.

ax

= @ ( P )

-

X ( e ) ) / X ( e ) (42)

in each column, appears in the third line. Recall that the exchangeable approximation is exact if n = 2, which holds throughout Table 1 except in its first section. Computations much more extensive than those exhibited in Table 1 support the conclusions summarized below.

In Table 2, we display the observed qualitative behaviors of the exchangeable approximation. In- crease (i), decrease (d), maximum (M), and minimum (m) refer to behavior of each variable as the parameter on the left increases. T h e notation d,M means that both decrease and a single maximum can occur (de- pending on the values of the fixed parameters), and Mm signifies that a maximum is followed by a mini- mum.

T h e perturbation approximation applies if either a 1

>

0 and E

<<

rmin =

p,

or a , = 0 and t

<<

1. In

agreement with our analytic results for the perturba- tion approximation with 7 = 1, the linkage disequili-

bria in the exchangeble approximation are positive and I s 1 decreases (increases) as p increases if a1

> 0

T h e perturbation approximation underestimatfrs the homologies; the perturbation values of 61 and D2 may be too low or too high. If a ,

>

0, the app-oxi- mations (31) for the homologies and (34b) for 02are much more accurate than (34a) is for 61, which has anAerror roughly of order t/rmin =

t/P.

If a1 = 0, then

A D , , is roughly of order t, and the other errors are

generally even smaller, so the accuracy is very high. When the conditions for their validity are satisfied, our analytic approximations (31), (34), and (40) are sufficiently accurate for quantitative comparisons.

Our most important conclusion is that ectopic con- version is surprisingly effective in producing homol-

ogy. Even if the ectopic conversion rate (a,) is consid- erably less than the intrachromosomal ( a ] ) and semi- classical ( a b ) conversion rates, the probabilities of identity betwee? genes on nonhomologous chromo- somes (& and 1 2 ) are still almost as high as tho2e between genes on the same

(gL)

or homologous ( 1 1 )

chromosomes. In particular, as long as there is sub-

(11)

Evolution of Repeated Genes 27 1

TABLE 1

The probabilities of identity and the linkage disequilibria at equilibrium

Parameter Value

i

0.62844 10 0.63361 -8.16 X IO-'' n

0.14469 100 0.14823

-2.39 X

i l i, i 2

0.61916 0.61916 0.61885 0.62540 0.62448 0.62415 -9.98 X lo-'' -8.52 X IO-:' -8.49 X

0.14255 0.14255 0.14254 0.14633 0.14610 0.14610 -2.59 X -2.43 X IO-' -2.44 X IO-'

i,

0.61885 0.62415 -8.49 X lo-'

0.14254 0.14609 -2.43 X 10"

61 6,

9.19 X 3.75 X

1.01 x IO-' 3.72 x IO-"

1.02 X 10" -8.17 X IO-"

2.54 X 10"' 8.55 X lo-" 2.33 X 8.74 X 10"' 8.91 X lo-? -2.14 X lo-' 0.89419 0.88100 0.88100 0.87881 0.87881 1.06 X IO-' 1.06 X IO-" 3 0.89505 0.88306 0.88209 0.87977 0.87977 9.72 X 1.06 X lo-"

-9.61 x -2.33 X IO-" -1.23 x -1.10 x IO-:' -1.10 x IO-" 8.89 x IO-? -1.02 x lo-:' 0.56020 0.54974 0.54974 0.54947 0.54947 4.40 X lo-' 4.40 X IO-''

k

20 0.56199 0.55473 0.55164 0.55131 0.55130 3.09 X lo-$ 4.41 X

-3.18 X IO-" -9.00 X -3.45 X lo-" -3.33 X -3.32 X IO-' 4.25 X 10-I -2.90 X lo-'' 0.97688 0.96247 0.96247 0.96007 0.96007 2.31 X 2.31 X IO-' 1 0.97708 0.96312 0.96292 0.96040 0.96039 2.05 X 2.31 X 10-7

-2.02 X 10"' -6.78 X -4.66 X -3.34 X -3.34 X 1.30 X 10" -1.74 X 10"' 0.29704 0.29266 0.29266 0.29193 0.29193 7.03 X IO"' 7.03 X IO-" 100 0.29896 0.30118 0.29463 0.29387 0.29386 6.55 X lo-' 7.08 X lo-"

-6.44 X lo-" -2.83 X IO-' -6.71 X lo-'' -6.60 X -6.57 X lo-" 7.35 X -6.51 X lo-''

1 0 - 4 ~

0.98816 0.98668 0.98668 0.98643 0.98643 1.18 X 1.18 X 10-7 1 0.98827 0.98692 0.98681 0.98655 0.98655 1.09 X 10-4 1.18 X 1 0 - 7

-1.09 X 10-4 -2.47 x 10-4 -1.37 X 10-4 -1.23 X 10-4 -1.23 X 10-4 8.98 X 10-2 -4.27 X 1 0 - 5

lO*U

0.48667 0.42407 0.42407 0.41397 0.41397 5.13 X IO-" 5.13 X

100 0.48868 0.43163 0.42691 0.41618 0.41618 4.73 X IO-' 5.16 X lo-" -4.12 X IO-' -1.75 X 10" -6.64 X lo-'' -5.31 X -5.30 X lo-'' 8.62 X 10" -5.22 X 10"

0.89419 0.88100 5 0.89578 0.88480 -1.77 X lo-'' -4.30 X 10"' I 04p

0.89419 0.88100 50 0.89438 0.881 44

-2.09 X -5.04 X

~

0.89442 0.87689 0 0.89443 0.87689

~~

-5.13 x -7.44 x

1 Oha,

0.89439 0.87741 1 0.89448 0.87764

0.88100 0.88301 -2.27 X IO-'

0.8788 1

0.88059 -2.02 x 10"

0.87881 0.88059 -2.02 x

2.12 x lo-:' 1 . s o x lo-'' 1.78 X 10"

~~

1.06 X lo-" 1.06 X lo-''

- 1.94 X IO-' 0.88100

0.88123 -2.68 X 10"'

0.87689 0.87689 -6.29 X

0.87741 0.87754

0.87881 0.87902 -2.40 X 10-4

0.87689 0.87689 -7.50 X IO-"

0.87713 0.87724

0.8788 1 0.87902 -2.39 X

2.12 X 10-4

2.08 x 1 0 - ~ 1.77 x lo-' 0.87689

0.87689 -6.30 X

0.877 13 0.87724

1.01 x

1.01 x lo-" -2.16 X

1.07 X 10-4

9.94 x

1.06 X lo-" 1.06 X lo-" 7.41 x

-1.02 X -2.52 X -1.38 X -1.22 X -1.20 X 7.18 X IO-' -4.01 X

0.89439 0.87741 0.87741 0.87713 0.877 13 1.06 x lo-'' 1.06 x IO-" 1 0.89526 0.87957 0.87858 0.87814 0.87814 9.85 X 1.06 X 10"'

-9.70 X -2.45 X 10" -1.33 X IO-" -1.15 X lo-" -1.15 X 7.09 X IO-' -1.07 X IO-" 1 060h

0.89371 0.88962 0.88962 0.88284 0.88284 1.07 X 1.06 X lo-" 100 0.89446 0.89128 0.89044 0.88362 0.88361 8.45 X IOW4 1.06 X

-8.45 X -1.87 X lo-' -9.18 X -8.83 X -8.82 X 2.69 X 10" -8.02 X

0.89984 0.87718 0.877 18 0.80774 0.80774 1.00 X lo-:' 1.00 x

1 0.90068 0.87937 0.87842 0.80868 0.80868 9.51 X 10-4 1.00 x 10-7 -9.27 X -2.49 X lo-' -1.41 X IO-" -1.17 X IO-' -1.17 X lo-' 5.29 X IO-' -1.14 X lo-'

1

ofiff<

0.89305 0.89062 0.89062 0.89095 0.89095 1.07 X lo-' 1.07 X

100 0.89375 0.89209 0.89135 0.89167 0.89166 7.39 X 10-4 1.07 X 10-5

-7.80 x 1 0 - ~ -1.65 x -8.18 x -8.11 x 1 0 - 4 -7.99 x 4.46 x 10-1 -1.80 X 10-4

For each parameter value, the first, second, and third lines represent the perturbation approximation, the exchangeable approximation, and the error of the former relative to the latter, respectively. T = 1 and a0 = a b throughout. The default values o f the other parameters are

(12)

TABLE 2

The qualitative behavior of the probabilities of identity and the linkage disequilibria at equilibrium

Parameter ,? i , i, i2 1:, 6, by n d d d d , M d , M i , d , M i

k d d d d d i, M i

N d d d d d 1 1

U d d d d d i, M i

P i , d i , d i , d i , d i , d i , d i , d

i, 111 i I 1 I 1 1

a, 111 m m i, Mm i, Mm i, Mm i

ah d , m i , M i , M i , M i , d i , M i , M

T = 1 and 010 = DL,, throughout. T h e symbols, i, d , M, and rn

denote increase, decrease, maxinlum, and minimum, respectively.

stantial homology, if a, = l / ~ ~ a l = %oab, then

i2

and

f 2 are still within about 10% of

i1

and

i,

.

Sometimes

a, can be even smaller without substantial relative reduction in i 2 and i2.

CONVERGENCE

In this section, we derive approximations for the rate and pattern of convergence to equilibrium. First, we deduce general formulas under the assumption of loose linkage. Next, we prove for pure interchromo- somal conversion that these formulas hold without this restriction. Finally, we present some numerical results.

The general model: We fix ( r y z ] and let E + 0.

First, we need the dynamical generalization of (28). We iterate (1 7a), starting in generation to :

D I + ( ~ ) = DI.~~(~o)P\T::'

1-111- 1 (43)

+

c

P : , y r q l , r r ( t

-

1 - j ) . j = O

Taking absolute values of each term in (43) and then extending the sum to infinity, we find ( t 2 t o )

I

Dl&)

I

5

I

DI,YL(tO)

I

P::;!

(44)

+

(1 - P l J l sup

I

q l . y & )

I.

t b l O

If to = 0 and tyz represents the shortest time such that

then

Setting

tl

In t

In( 1 - ryz) ' t y z

-

and (24) give

D l , , ( t ) = O(t), t 2 ty.. (47)

we obtain

D I . Y L ( t ) = O(E), t 2 t l , (49)

for every y and z , as E 3 0 with ( r ] fixed. Since the

crossover rates ryz are fixed, the time tl = O(-ln E ) , which will often be only 5 or 10 generations. If r,,;,,

<<

Y 2 , however, then

:"

t l z --, In

E

rlrlin

which can be much larger.

For D P , r t , the above argument leads to

for every z and {, where

Clearly, t2

<

t l , so both (49) and (51) hold for t 2 tl

.

Notice that the typical time t l for the reduction of

linkage disequilibrium to O ( E ) is much shorter than the characteristic time for convergence to equilibrium because, by (23), the latter is of O ( ~ / E ) .

We use ( 1 6), ( 4 9 , and (51) to eliminate gl,yz and g 2 , z g

from (15a, c, e) for t 2 t l :

f:=e+[1-2u-e-(n-1)a2-72(k-1)a,1fy

+

(n - 1)a2t;.y

+

n(k

-

l)a,T*,y, ( 5 3 4 1 L Y Z = %aYP(fy

+f.)

+

Y2n[a2

+

( k - 2)ae](T2,.

+

T 2 , { )

+

O(2) as t --$ 0 with (r,.) fixed.

For the moment, let us ignore the error terms in (53b, c). In that case, (53) preserves position inde- pendence: if

f,

= J

LI,, = 1 1 , 1 2 J f = 12

f; = f ' ,

I;,,,

= 1 ; , l & = 1;

for every y, z ( y # z), and {, then

for every y, z (y # z), and S; where f ' =

e

+

[ 1

-

2u

-

e

-

(n

-

-

n(k - l)ae]f

+

(n

-

1)a'LlI

+

n(k - l ) a , / z ,

z;

= aYpf

+

[ l - 2u

-

a2

(13)

Evolution of Repeated Genes 273

1; = ae f

+

(n

-

l)a,ll

+

(1

-

2u

-

na,)Zs. (56c) Furthermore, it is easy to see that the nonnegative matrix of coefficients in (53), S, is irreducible. Con- sequently, Theorem 2.4 of BOUCHER and NAGYLAKI (1 988) implies that its maximal eigenvalue is identical to that of the 3 X 3 matrix, Q, in (56), which we denote by K ~ . Since the diagonal elements of S are

positive and S is irreducible, it is aperiodic (FELLER 1968, p. 426). Therefore, from Remark 2.9 of BOUCHER and NAGYLAKI (1988), we conclude that (53) and (56) have the same asymptotic rate and pattern of convergence. Thus, the decay of position dependence is faster than the asymptotic rate of con- vergence to equilibrium.

We write Q and its eigenvalues as

Q = (1 - 2u)I

-

R , K = 1 - 2U

-

f , (57)

where I designates the 3 X 3 identity matrix. Then the eigenvectors v of Q satisfy Rv = f v . A glance at (56) reveals that the elements of the 3 X 3 matrix R

are of O ( E ) , so f = O ( E ) , whereas if v is normalized to unit length, its components will be of O(1). It follows that the error terms in (53b, c) lead to errors of O ( E * )

in the maximal eigenvalue of X. in (53) and of O ( E ) in the corresponding eigenvector wo. Thus,

X" = K O

+

O(E2) (5 8 4

as E + 0 with {rYL} fixed. Writing the components of

w' as w: , where i = 1, 2, and 3 refer respectively to

f,

I , ,

and 12 in (53), and j refers respectively to the

position subscripts y, yz, and zS; from Theorem 2.4 of BOUCHER and NAGYLAKI (1 988) we have

w ; = vp

+

O(E) (58b)

as E + 0 with ( r Y L ) fixed.

tion for [:

Equations 56 and 57 yield the characteristic equa-

f 3

-

a l f 2

+

a 2 f

-

as = 0, ( 5 9 4 where

a l = 8

+

na2

+

n(2k - l)a,, (59W

a2 = 8(a,

+

nka,)

+

n2ka,[a2

+

(k

-

I)a,], (59c)

a3 = n8ae[a:!

+

(k

-

l ) a e ] . ( 5 9 4

Let f o represent the smallest positive root of (59a). In accordance with our general discussion of param- eter dependence, the approximation (59) is independ- ent of 7 , { rYz

1,

and a o , and it depends on a,, as, and

a b only through 0 2 . In agreement with the study of

special cases below (26), the result (59) reduces to the simpler one in NAGYLAKI (1984b) if (i) ac = 0 or

k

=

1, (ii) n = 1, or (iii) a2 = a,. (A spurious root must be deleted in each of these cases.) If a2 = a, = 0, the roots of (59) are 0, 0, and 8, as they must be (WRIGHT

193 1 ; MALBCOT 1946, 1948; KIMURA 1963). If 8 = 0, the roots are

f = 0, nka,, n[a2

+

(k

-

l)a,]. (60)

We define the scaled characteristic time for the decay of genetic variability in the absence of mutation as

T = 2 / ( N f o ) . (61)

If mutation is present, (57) gives the scaled character- istic convergence time

T/(1

+

N U T ) . (62)

Some asymptotic results are instructive and provide

By (60), f o + 0 as 8 + 0 with the other parameters a check on numerical calculations.

fixed. From (59) we deduce

to""--

a3 8

-

1

a2 nk 2nkN'

so

T

-

4nk (63b)

as 8 + 0. Since random drift is much slower than conversion in this limit, it is the former that deter- mines the ultimate rate of convergence, which in (63a) is simply that for 2nkN genes at a single locus. [Our derivation is validated by the observation that insert- ing (63a) into (59a) produces 03, 02, 8, and 8 for the orders of the successive terms.]

Since zero is a root of (59a) if a, = 0, we must have

40

+ 0 as a, + 0 with the other parameters fixed. T h e above method now informs us

T

-

2/(nNa,) (64b)

as a, + 0, which is again determined by the slowest process, ectopic conversion.

Finally, if a2 .--, 0 and a, + 0 with a2/a, fixed and positive, then only

t 3

is negligible in (59a), and in the limit f~ satisfies the quadratic

f 2

-

(a2

+

nka,)[

+

na,[a2

+

(k

-

l)a,] = 0, whence

f o

-

%(a2

+

nka,

-

[(Cy2

+

nka,)'

(65)

-

4na,(a2

+

(k

-

l)ae)]"2).

Pure interchromosomal conversion: Here, we demonstrate that in the absence of intrachromosomal conversion (a1 = 0), the main results (56) to (59), and therefore (60) to (65), of the previous subsection hold for arbitrary linkage, i e . , uniformly in ( r y L ) as E + 0.

Figure

FIGURE ].-The and for  intrachromatid,  sister-chromatid,
TABLE 1 The  probabilities of identity  and  the  linkage  disequilibria at equilibrium
TABLE 2
TABLE 3

References

Related documents

Malaysia is a multi-racial country with the main ethnic groups Malays, Chinese and Indians and hence it would be interesting to study the reactions, fears and

Reduction with or without appendicectotomy performed in 62 cases (59.6 percent) carried the lowest mortality risk (6.7 percent and 8.5 percent respectively, Table V) Resection

(A) Type I IFN production by SV40T MEFs from matched wild-type and Sting -deficient mice stimulated with 0.1 ␮ M CPT for 48 h, measured by bioassay on LL171 cells (data shown as

This unique quality of HRV/HRC analyses has led to interest in its application in pre-hospital triage and emergency clinical decision making; interest which has been bol- stered by

PTPN1 knock-down, cell proliferation and tyrosine phosphorylation analyses, and RT-qPCR mRNA expression was assessed on SH-SY5Y, SMS-KCNR, and IMR-32 human NB cell lines..

• Severe acne- refer early for oral isotretinoin if large nodulocystic lesions, scarring or no rapid response to treatment. • Moderately severe acne which has not responded to 2 x

Cuando el malware tiene la forma de un archivo ejecutable en algún lugar del disco duro, normalmente se inicia sin la intervención del usuario, invocado de forma automática, como

Aangezien 26,5% van de respondenten onvoorwaardelijk vindt dat er één gezamenlijke identiteit geformuleerd kan worden en ook nog eens 60% van de respondenten vindt dat