• No results found

Nucleotide sequence of a cloned duck hepatitis B virus genome: comparison with woodchuck and human hepatitis B virus sequences.

N/A
N/A
Protected

Academic year: 2019

Share "Nucleotide sequence of a cloned duck hepatitis B virus genome: comparison with woodchuck and human hepatitis B virus sequences."

Copied!
11
0
0

Loading.... (view fulltext now)

Full text

(1)

Vol.49, No. 3 JOURNALOFVIROLOGY, Mar. 1984, p.782-792

0022-538X/84/030782-11$02.00/0

Copyright © 1984, American Society for Microbiology

Nucleotide Sequence of

a

Cloned

Duck Hepatitis

B

Virus Genome:

Comparison with Woodchuck and Human Hepatitis

B

Virus

Sequences

ELISABETH MANDART, ALAN KAY, AND FRANCIS GALIBERT*

Laboratoire d'Hematologie Experimentale, Centre Hayem, HopitalSaint-Louis, 75475 ParisCe'dex 10, France Received 16 June 1983/Accepted 6 October 1983

The nucleotide sequenceofanEcoRI duckhepatitis B virus (DHBV) clone waselucidated by usingthe

Maxam and Gilbert method. This sequence, which is 3,021 nucleotides long, was comparedwiththe two previously analyzed hepatitis B-like viruses (human andwoodchuck). From this comparison, itwas shown

that DHBV is derived from an ancestor common to the two others but has a slightly different genomic organization. There wasnointergenic region between genes5 and 8, whichwerefused intoa singleopen

reading frame in DHBV. Genes for the surface andcoreproteinswereassignedtoopenreading frames 7 and

5/8. Amino acid comparisons showed some structural relationship between gene 6 product and avian reverse transcriptase, suggesting either evolution from a common ancestor or convergence to some

particular structure tofulfill aspecific function. This should be correlated with the synthesis ofan RNA

intermediate duringDNAreplication. This is also takenas anargumentinfavor of thehypothesis thatgene

6 codesfor the DNApolymerase that isfound within the virion. DNA sequencecomparison also showed that thetwomammalian hepatitis B viruses are more homologous to each other thanthey areto DHBV, indicating thatDHBVstarts toevolveonitsownearlier than the twootherviruses,asdo birds compared

with mammals. Fromthis it is proposed that the virusesevolved inafashion parallel to the speciesthey infect.

Duck hepatitis B virus (DHBV), which was recently

isolated

(18), is the fourth member ofa new viral family called Hepadnaviridae, whose prototypeishuman hepatitis

Bvirus(HBV) (24). The fourmembersofthisgrowing family

have several characteristics in common, such as ultrastruc-tire, antigenic makeup, DNA size, and structure.

Similar-ities inthe

pathological

field have been describedaswell (11,

17, 33, 34).

In the past few years, by molecular cloning (4, 5) and

nucleotide sequence analysis (3, 8-10, 35, 36), our

knowl-edge of the biology of these viruses has

increased.

Two genes, one coding forthecoat protein and the other

coding

for thecore

protein,

havebeen identifiedonthe genomes of

HBV andwoodchuck hepatitis B virus (WHV). Twoother

possible coding regions have been

mapped,

but their func-tion and their products have not, asyet, been determined.

In spite ofourgrowing knowledge aboutthe structure of

various viral components, the

biological study

of these

viruses is still complicated by the absence of cell cultures

susceptible toinfection.

Some ofthese

complications

have been overcome

by

the

discovery ofDHBV (18), which is not only able to infect ducks, an animaleasily kept in colony, but can alsoinfect andmultiplyinembryonic eggs. These

findings

caused usto undertake the elucidation of the nucleotide sequence of the DHBV genome. Duringthecourseofthisstudy, the useful-nessof the duck-DHBV modelwasprovenbythe

finding

of Summers and Mason (31), who, using an in vitro system, found evidenceimplicating an RNA intermediate in DHBV DNA replication.

Inthis paper,wereport thecompletenucleotide sequence of the genome of DHBV and compare its

primary

structure with thatof the HBV and WHVgenomes.

*Correspondingauthor.

782

MATERIALS AND METHODS

Enzymes and chemicals. Restriction endonucleases came from NewEngland Biolabs and were used as recommended bythe manufacturer. DNA polymerase I canie from Boeh-ringerMannheimBiochemicals, andbacterialalkaline phos-phatase and polynucleotide kinase were from P. L. Biochemicals.

Chemicals used for nucleotide sequence analysiswere as described, previously (14). [-y-32P]ATP (specific activity,

>2,500 Ci/mmol) and

ot32P-labeled

nucleotide triphosphate (specificactivity, >3,000Ci/mmol)werefrom New England Nuclear Corp.

Preparation ofEcoandXho DHBV DNAs. X-DHBV recom-binants wereconstructed andgiven to us by W. S. Mason et al. The cloned DNAs were referred to as Eco and Xho DHBV DNAs.Propagation and purification ofthe recombi-nants, as well aspreparation ofthe DNAs, wereperformed aspreviously described(3, 10, 14).

Containment. Containment conditions were as recom-mended by the French National Control Committee. The culture ofrecombinantbacteriophage wasdone under L3B1 conditions.

DNAnucleotidesequence. Sequence analyses were deter-mined bythe Maxam and Gilbert method(19). Usually, ca. 10 pmol of EcoDHBV DNA (20

Rg)

wasfullydigested each time with a given restriction enzyme. Fragments were

de-phosphorylatedand labeled with

[y-32P]ATP

and

polynucle-otide kinase as described previously (14). To separate the twolabeled ends, fragments were denatured by heating to 920C in the presence of 30%dimethylsulfoxide and fraction-ated by electrophoresis in acrylamide gel (19). Fragments larger than 600 base pairs were hydrolyzed with another

restriction enzyme. Under some circumstances, fragments

with a recessed 3' end were labeled with an

ot32P-labeled

nucleotidetriphosphate of choice and DNApolymeraseIas describedby Hartley and Donelson (13).

on November 10, 2019 by guest

http://jvi.asm.org/

(2)

DHBV GENOME NUCLEOTIDE SEQUENCE 783

f-J

FIG. 1. Diagram of analyzed DNA fragments. Vertical bars correspond to the positions of the labeled ends of restriction fragments used. Length of the arrows is relative to the number of analyzednucleotides.Row 1, HaeIII;row2. Hinfl;row3. Aliil;row

4, Sau3a; row5, RsaI; row 6, Mspl +BgllI. Sau3a fragments are

labeled at their 3'ends; the othersare labeledat their 5' ends.

RESULTS AND DISCUSSION

The complete nucleotide sequence of the EcoRI DHBV DNA cloned in Escheric-hia coli was determined

by

the

method of Maxam and Gilbert (19),

using

five different chemical reactions giving specific bands for

G,

AG, CT, C, and AC. The sequence was derived from a

large

number of

overlappingfragments sothatboth strands were

entirely

and independently analyzed, and all starting restriction sites

were analyzed within an overlapping fragment (Fig. 1). For

reasons explained below, the EcoRI site used for

cloning

was also overlapped by using an

XlzoI

DHBV clone,

elim-inating the risk of theloss ofa small DNA fragment located

betweentwoputative Ec oRI sites withinthe DHBVgenome. The DHBVsequence shown in Fig. 2 is 3,021 nucleotides

long, compared with 3,182for the HBV DNA and 3,308 for the WHV DNA. Restriction enzyme data obtained from virion DNA enabled us to define which strand was the L

strand and which was the S strand in both HBV and WHV (9, 10). We do not have such data for DHBV, but the position of theopen reading frames aswell asthe nucleotide

sequence homologies found within these three genomes indicate that the nucleotide sequence presented in Fig. 2 is

complementary to the L strand, alsocalled the minusstrand by Summers and Mason (31).

Location of the open reading frames. The number and distribution ofstop codons within the L strand is suchthat, as previously noticed with HBV and WHV, it is difficult to locate agene which could be transcribed from the S strand (Fig. 3). Only one region spanning from nucleotide 1397 to nucleotide 835 could have a substantial coding capacity. However, (i) the first in-phase ATG in this region

(open

reading frame 1) is toward the end, at position 904. very closetothestop

codon;

(ii) thereis noobviousacceptor site at the beginning ofopen reading frame 1 to which an exon

containing an ATG could be spliced; and (iii) neither HBV

nor WHV has corresponding open reading frames which

could give a protein homologous to the putative protein

made from open reading frames1 of DHBV. These suggest

the absence of a coding function for the S strand, as

previouslyobservedfor the HBVand WHVgenomes (9,

10).

On the contrary, the number and distribution of stop codons in the S strandleave open several regions, indicating

that mRNA could arise by transcription of the L strand. However, only three open reading frames are displayed

along the DHBV genome instead of the four previously noticed for the HBV and WHV genomes. These three

regions span from nucleotides 14 to 2527, 684 to 1784, and 2515 to 411. Their relative positions (Fig. 4) as well as

information derived fromnucleotide

aidd

aminoacid

compar-isons of HBV and WHV sequences indicate that (i) the

largest

open

reading

frame (nucleotides 14 to

2527)

corre-sponds

togene 6 ofHBV;

(ii)

theopen

reading

frame which is

overlapped by

region

6and goes fromnucleotides 684 to 1784

corresponds

to gene

7,

also called S (for surface

protein):

and

(iii)

regions 5 and 8 found in HBV and WHV

genomesarefusedintoa

single region

in

DHBV,

as

suggest-ed

by

the size and

position

ofthethird open

reading

frame

(nucleotides 2515 to

411).

Therefore,

we call it

region

5/8. Nucleotide sequence

comparisons.

By

using a computer

program

developed

by

Staden

(29),

thenucleotide sequence ofthe DHBVgenome was

compared

withthose of the HBV and WHVgenomes.

Although

nucleotidesequence

compari-sons show a

large degree

of

homology

between HBV and

WHV,

62 to

70% along

the genomes except in two small

regions

(9),

amuchsmaller level of

homology

wasobserved with the DHBV sequence.

Figure

5 shows several

graphs

which

demonstrate

this result. For the

largest

part of the DNA sequence, the

degree

of

homology

was below

40%.

Four

regions

located between nucleotides 100 and

200,

300 and

500.

600 and

800,

and 1300 and

2100

were more conserved and had about

50%

homology

with the homolo-gous

regions

found in the HBV and WHV genomes. The

region

between nucleotides 1300 and 1400 had the

highest

homology

(70%).

This

highly

conserved sequence is part of

regions

6and7, which could be read intwodifferent frames.

This

remarkably high degree

of

homology

is

probably

dueto the

product

ofgene6, which is

slightly

moreconserved than the

product

of gene

7,

as

explained

below. As for

HBV,

viruses found innature vary fromonetoanother(2,

28,

32).

The two DHBV

clones,

cloned

through

the

EcoRI

or

XlioI

site, were

exactly

the same size as

demonstrated

by

their

comparative

electrophoretic mobility

butexhibiteda

slightly

different restrictionpattern. In agreement withthis observa-tion, the nucleotide sequence we

determined

around the EcoRIsiteoftheX/loDHBVclone

diverged

from that of the

Eco DHBV clone to an extent of ca.

3%.

None of these nucleotide

changes

alters the

coding capacity

of this

se-quence.

Amino acidsequence

comparisons.

(i)

Region

6.Inthethree

viruses,

this

region

covered ca.

80%

of the genome. In

DHBV it went from

nucleotides

14 to

2528.

From the first ATG(residue

20)

uptothe TAA

stop

codon

(residue

2528),

a

protein of 836 amino acids can be

predicted,

as

compared

with 879 and 838 for the

corresponding

WHV and HBV

proteins

(9,

10).

These

proteins

have not been

identified

so

far,

but the size of region 6 leaves little doubt that it does havea

coding

function.A

comparison

ofthe

predicted

amino

acid sequences ofthese three proteins

clearly

shows that

they

are related to each other

(Fig.

6and

7).

Although

theDNAsequencesallow thegene6

products

of

DHBVandHBVtohave

nearly

identical

lengths

(836amino

acids,

compared with

838),

it is difficult to

align

these two

proteins

starting

from the first methionine in each region. If

regions

of amino acid

homology

are

paired,

then the best coincidenceisobtained

by

assuming

that theDHBV

region

6 protein starts at the second

in-phase

ATG. This would reduce the DHBV

protein

6 to a size of 786 amino acids.

Supporting

that

hypothesis,

thefirst

in-phase

ATG ofregion

6 is

preceded

at

position

-3

by

a

pyrimidine,

whereas the

second

in-phase

ATGis

preceded

at

position

-3

by

apurine

(16).

Because of its size, which is

comparable

to various

polymerases,

ithasbeen

proposed

that theDNA

polymerase

found

within

the virion would be coded

by

gene 6(10). An experiment

performed

by

Summers and Mason (31) has

I r

2 IJ!2 1 ;

3 rt

4 1.,

5 r j

6 iJ W

VOL. 49. 1984

on November 10, 2019 by guest

http://jvi.asm.org/

[image:2.612.83.277.75.146.2]
(3)

784 MANDART, KAY, AND GALIBERT

revealed the existence ofanRNAtranscript complementary

to the L strand during theprocess of DNA replication. We were therefore interested in comparing the amino acid sequencepredicted from the largeopenregions (region 6) of

the three known hepatitis viruseswith several viral encoded DNA polymerases orreversetranscriptases.

Theamino acid comparisonswere made withacomputer

program established by B. Caudron (personal

communica-tion) which is abletodetect small stretches of 20 amino acids with homology above 25%. With this program, the amino

acidsequences predicted fromgenes 6 werescored against theamino acidsequences of the DNA polymerase encoded

in the adenovirus 2 genome (1, 12) and the amino acid sequences of the avian and murine reverse transcriptases

(25, 27). Several stretches of amino acids with homologies

71

141

211

over25%werefound in all comparisons. However,

homolo-gies between adenovirus 2 DNA polymerase and the hepati-tis B virus region6products generally involved only oneof

the viruses at a time, and the positions of the various

homologies indicatednoparticularpattern.Onthecontrary, when the reverse transcriptases were compared with the

threehepatitis B virus region6products,at leastone setof amino acidsequencehomologies appearedtogiveacoherent

pattern. These sequences are indicated in Table 1. When DHBV and HBV gene 6 proteins and Rous sarcoma virus reversetranscriptase (RSVRT)arealigned by pairing thisset ofsequencehomologies, other smaller stretches of homolo-gy become apparent (Fig. 7). (The same is alsotrue when WHV region 6 protein is included.) Also, in a significant

number ofcasesinwhich therewas anamino aciddifference

Start of region 6 (frame 2)

CATGCTCATTTGAAAGCTTATGCAAAAATTAACGAGGAATCACTGGATAGGGCTAGGAGATTGCTTTGGT GGCATTACAACTGTTTACTGTGGGGAGAAGCTCAAGTTACTAACTATATTTCTCGTTTGCGTACTTGGTT GTCAACTCCTGAGAAATATAGAGGTAGAGATGCCCCGACCATTGAAGCAATCACTAGACCAATCCAGGTG GCTCAGGGAGGCAGAAAAACAACTACGGGTACTAGAAAACCTCGTGGACTCGAACCTAGAAGAAGAAAAG 281 TTAAAACCACAGTTGTCTATGGGAGAAGACGTTCAAAGTCCCGGGAAAGGAGAGCCCCTACACCCCAACG

End of the fused 5 + 8 region-1

351 TGCGGGCTCCCCTCTCCCACGTAGTTCGAGCAGCCACCATAGATCTCCCTCGCCTAGGAAATAAATTACC 421 TGCTAGGCATCACTTAGGTAAATTGTCAGGACTATATCAAATGAAGGGCTGTACTTTTAACCCAGAATGG 491 AAAGTACCAGATATTTCGGATACTCATTTTAATTTAGATGTAGTTAATGAGTGCCCTTCCCGAAATTGGA 561 AATATTTGACTCCAGCCAAATTCTGGCCCAAGAGCATTTCCTACTTTCCTGTCCAGGTAGGGGTTAAACC

Start of region 7 (frame 3)

631 AAAGTATCCTGACAATGTGATGCAACATGAATCAATAGTAGGTAAATATTTAACCAGGCTCTATGAAGCA 701 GGAATCCTTTATAAGCGGATATCTAAACATTTGGTCACATTTAAAGGTCAGCCTTATAATTGGGAACAGC 771 AACACCTTGTCAATCAACATCACATTTATGATGGGGCAACATCCAGCAAAATCAATGGACGTCAGACGGA 841 TAGAAGGAGGAGAAATACTGTTAAACCAACTTGCCGGAAGGATGATCCCAAAAGGGACTTTGACATGGTC 911 AGGCAAGTTTCCAACACTAGATCACGTGTTAGACCATGTGCAAACAATGGAGGAGATAAACACCCTCCAG 981 AATCAGGGAGCTTGGCCTGCTGGGGCGGGAAGGAGAGTAGGATTATCAAATCCGACTCCTCAAGAGATTC 1051

1121

1191

1261

1331 1401 1471

CTCAGCCCCAGTGGACTCCCGAGGAAGACCAAAAAGCACGCGAAGCTTTTCGCCGTTATCAAGAAGAAAG

ACCACCGGAAACCACCACCATTCCTCCGTCTTCCCCTCCTCAGTGGAAGCTACAACCCGGGGACGATCCA CTCCTGGGAAATCAGTCTCTCCTCGAGACTCATCCGCTATACCAGTCAGAACCAGCGGTGCCAGTGATAA AAACTCCCCCCTTGAAGAAGAAAATGTCTGGTACCTTCGGGGGAATACTAGCTGGCCTAATCGGATTACT GGTAAGCTTTTTCTTGTTGATAAAAATTCTAGAAATACTGAGGAGGCTAGATTGGTGGTGGATTTCTCTC AGTTCTCCAAAGGGAAAAATGCAATGCGCTTTCCAAGATACTGGAGCCCAAATCTCTCCACATTACGTAG GATCTTGCCCGTGGGGATGCCCAGGATTTCTTTGGACCTATCTCAGGCTTTTTATCATCTTCCTCTTAAT

FIG. 2 9

J. VIROL.

on November 10, 2019 by guest

http://jvi.asm.org/

[image:3.612.155.468.246.730.2]
(4)

DHBV GENOME NUCLEOTIDE SEQUENCE 785

between the DHBV 6 protein and RSVRT, the HBV 6 proteins and RSVRT were identical. Finally, if one takes into account that severaloftheobserved differences corre-spond to amino acids of the same family, involving a so-calledconservativechange (6), then one can suggest that the homologies observed between the hepatitis gene 6 proteins and the avian reverse transcriptase did not occur only by chance. However, more sophisticated computer work is

needed to establish more definitely what kind ofstructural

relationship exists between region 6 proteins and reverse transcriptases and whether this is due to evolution from a common ancestor or to convergence to fulfill a similar

function.

(ii) Region 7. This open reading frame went from nucleo-tide 684 in frame 3 to stop codon TAG 1785. It has been

1541 1611 1681

17 51 1821 1891 1901 2031 2101 2171 2241 2311

shown that the homologous HBV open reading frame codes for the surface protein, and according to DNA sequence homology and amino acid sequence homology, an identical result has been inferred for WHV. As we have already pointed out, significant DNA homologybetween DHBV and the other twoviruses was found in this region. This suggests that the open region 7 codes for the duck viral surface (DHBs) protein. Due to the presence of numerous ATGs in thereading frame, translation could start at several different places. From the first encountered ATG at position 693, a protein of 364 amino acids can be predicted. By using different sources ofdata, such as the molecular weight and theknown amino acid sequence of the N-terminal portion of the human viral surface(HBs) protein, it has been deduced that theN-methionine of the HBsprotein is not coded by the

CCTGCTAGTAGCAGCAGGCTTGCTGTATCTGACGGACAACGGGTCTACTATTTTAGGAAAGCTCCAATGG GCGTCGGTCTCAGCCCTTTTCTCCTCCATCTCTTCACTACTGCCCTCGGATCCGAAATCTCTCGTCGCTT TAACGTTTGGACTTTCACTTATATGGATGACTTCCTCCTCTGCCACCCAAACGCTCGTCACCTTAACGCA

End of region 7 ;

ATTAGCCACGCTGTCTGCTCTTTTTTACAAGAGTTAGGAATAAGAATAAACTTTG ACAAAACCACGCCTT CTCCGGTGAATGAAATAAGATTCCTCGGTTACCAGATTGATGAAAATTTCATGAAGATTG AAGAAAGCAG ATGGAAAGAATTAAGGACTGTAATCAAGAAAATAAAAGTAGGAGAATGGTATGACTGGAAATGTATTCAA AGATTTGTCGGGCATTTGAATTTTGTTTTGCCTTTTACTAAAGGTAATATTGAAATGTTAAAACCAATGT ATGCTGCTATTACTAACCAAGTAAACTTTAGCTTCTCTTCATCCTATAGGACTTTGTTATATAAACTAAC AATGGGTGTGTGTAAATTAAGAATAAAGCCAAAGTCCTCTGTACCTTTGCCACGTGTTAGCTACAGATGCT ACCCCAACACATGGCGCAATATCCCATATCACCGGCGGGAGCGCAGTGTTTGCTTTTTCAAAGGTCAGAC ATATACATGTTCAGGAACTATTGATGTCTTGTTTAGCCAAGATAATGATTAAACCACGTTGTCTCTTATC TGATTCAACTTTTGTTTGCCATAAGCGTTATCAGACGTTACCATGGCATTTTGCTATGTTGGCCAAACAA 2381 TTGCTCAAACCGATACAATTGTACTTTGTCCCGAGCAAATATAATCCTGCTGACGGCCCATCCAGGCACA

Start of the fused >s regions in frame 1 2451 AACCTCCTGATTGGACGGCTTTTCCATACACCCCTCTCTCGAAAGCAATATATATTCCACATAGGCTATG

End of region 6

2521 TGGAACTTAAGAATTACACCCCTCTCCTTCGGAGCTGCTTGCCAAGGTATCTTTACGTCTACATTGCTGT 2591 TGTCGTGTGTGACTGTACCTTTGGTATGTACCATTGTTTATCATTCTTGCTTATATATGGATATCAATGC 2661 TTCTAGAGCCTTAGCCAATGTGTATGATCTACCAGATGATTTCTTTCCAAAAATAGATGATCTTGTTAGA

2731 GATGCTAAAGACGCTTTAGAGCCTTATTGGAAATCAGATTCAATAAAGAAACATGTTTTGATTGCAACTC

2S01 ACTTTGTGGATCTCATTGAAGACTTCTGCCAGACTACACAGGGCATGCATGAAATAGCCGAATCATTAAG 2871 AGCTGTTATACCTCCCACTACTACTCCTGTTCCACCGGGTTATCTTATTCAGCACGAGGAAGCTGAAGAG

2941 ATACCTTTGGGAGATTTATTTAAACACCAAGAAGAAAGGATAGTAAGTTTCCAACCCGACTATCCGATTA

3011 CGGCTAGAATT

FIG. 2. Nucleotide sequence of the Eco DHBV DNA clone. The sequence shown is complementary to the viral L strand. VOL.49, 1984

on November 10, 2019 by guest

http://jvi.asm.org/

(5)

786 MANDART, KAY, AND GALIBERT

Chain S

2

3

2

5/8 *, 111, 5/8 6

toIon III 7 iCOtl

ChainL

100 o

1 1 "' 1 1 1 III i,

11iI 11111 dPlI3I3om 1..11- 1111111 1131Is

3 fJ'-3ap*fdkl3l mlII A 1 11 II 11 I II fill

FIG. 3. Diagram showing the localization of the nonsense co-donsonchainsSand L.Threereading framesweredefinedfrom the 5'end ofeachDNAstrand.Onchain S,frame1isdefinedby itsfirst tripletCAT, frame2isidentifiedbyATG,andframe 3 isidentified by TGC. On chain L, frame 1 is defined by its firsttriplet AAT, frame2isidentifiedbyATT, andframe3 isidentified by TTC. The viral DNA is circular, and its length in nucleotides (3,021) is a multiple of3; therefore, passing through the EcoRI site does not changethereadingframe. Upperverticalbarsindicate stop codons. Lowerverticalbarsrepresent ATG triplets. Numbers5/8,6,and 7 defineareas in which a viral gene has been located: region 5/8 goes from2515 to 411; region 6 goes from 14 to 2527; region 7 goes from 684 to 1784.

first in-phase ATG codon but by the third one (10). An

identical deduction has been made for the woodchuck viral surface (WHs) protein (9). According to these results,

trans-lationof DHBsprotein maystart with the secondor

follow-ing in-phaseATG.

A comparison ofthe amino acid sequence deduced from the DNA sequenceofopenreading frame7ofHBV, WHV, and DHBV showed no homology for the so-called pre-S region. On the other hand, a significant homology existed between the DHBV amino acid sequencestartingwith ATG 1284 and the N-terminal amino acid sequences of HBs and WHsproteins,suggesting,ifnotproving,that theN-terminal amino acid sequence for the DHBsprotein starts with ATG 1284, the seventhin-phaseATG within openreadingframe 7 (Fig. 8). From ATG 1284 up toTAG 1785,a protein of 167 amino acids with a molecular weight of 18,204 can be

predicted, ascomparedwith 25,645and25,422forthe WHs and HBs proteins,

respectively

(3, 8, 35). Electrophoretic results, showingan apparent molecularweightof17,000for

DHBs protein (Mason etal., personal communication), are ingood agreement withthe theoretical molecular weight.

A comparison ofthe amino acid sequences of the three surface proteins showed that afragment of about 50 amino acids, corresponding roughly to position 105 to 155 ofthe HBs protein sequence, was absent from the DHBs protein (Fig. 9). This deletion was also seen in the gene 6 protein (Fig. 9) andwasvisible atthe DNA level (Fig. 5). The most intriguing point about thisdeletion is that the main antigenic epitope of the HBs protein was tentatively located in that region byproteolytic digestion, amino acid comparison,and peptide synthesis (7, 8, 22, 23).

Several characteristic features of the HBs and WHs proteins have been previously noted, such as the existence of a very large hydrophobic sequence and of a

sequence,

Asn-X-Thr/Ser, knowntobe involved inglycosylation (30). The same is true for the DHBs protein (Fig. 9). There is a hydrophobic regionlimitedbyamino acids 79 and 97 with a hydrophobic index equalto3.31 (26). A potential glycosyla-tion site is also located at position 99. The homology observed between the DHBsproteinand the two others was much lower(35%)than between the HBs and WHsproteins (61%), and thecarboxylic regionsshowedverylittle

homolo-

gy-It isinterestingto notethat theDHBsproteinwaspartially encoded bythe sequence locatedbetween nucleotides 1300 to 1400 in which the largest sequencehomology (70%) was observed with the other two genomes. This sequence also coded for the gene 6 protein. This high percentage of homology was mainly due to the gene 6 product, which is moreconserved (Fig. 10).

Although we cannotformally prove thatATG 1284codes for the N-terminus of the DHBsprotein,it ishighly suggest-ed by the amino acid sequence homology observed among the three N-terminal sequences. However, the existence in all three viruses of alarge open readingframepreceding the putative N-termini of their surface proteins is most intrigu-ing. Various experiments have located the TATA boxand thecap site of the HBs mRNA 150 nucleotides ahead of the first ATG of region 7, suggesting that transcription of the pre-S region does occur (21). Recent experiments by Si mapping ofHBsprotein mRNA madein transfected mouse cells have detected a protected DNA fragment starting at position 3160 of the HBV genome (29a). Because of the locationatnucleotide2776 of a TATAbox,a moreprobable hypothesis is that HBs mRNA is spliced and that position 3160correspondstothe 5' end ofitsmainbody.Aconsensus

6

[image:5.612.84.277.68.208.2]

5 5

FIG. 4. Localizationof the openreadingframesonthe viralgenomes of HBV and WHV andcomparison with the DHBVgenome. The stripedareainregion7correspondstothepre-Ssequence. Arrows indicate thepositionof the first ATG found withinanopenreadingframe. Numbers 1to8refertothe variousopenreadingframesas defined in thetextandin Galibert etal.(9, 10).

J. VIROL.

l

on November 10, 2019 by guest

http://jvi.asm.org/

[image:5.612.106.507.569.694.2]
(6)

DHBV GENOME NUCLEOTIDE SEQUENCE 787

w

8-//

/

i ~~~~H

170 970

a080

t890

V09 1890

D

81

/

//

/

§/ H

170 970

/

8

10910 t890

/

/ 1890

waslostduring cloning through the EcoRI site by sequencing thisregion on anindependent XhoI DHBV clone.Therefore, the fused 5/8 region represents the actual structure of the DHBV genome. Because the nucleotide sequencehomology was too poorin this region, comparisonsdid not allowus to infer which part of the sequence was lost (or acquired) during evolution.

Acomparison of the amino acid sequence of DHBV region

5/8 withthecorresponding amino acidsequencesofregions5 and 8 of HBV and WHV showed a net degree ofhomology between thecarboxylic end of the molecule coded by regions 8 and 5/8. The same peculiar structure,involving repetition and increased amounts of basic aminoacids, was observed at thecarboxylicend,giving, upon comparison, a character-isticpicture (Fig. 12). This clearly suggests that this amino acid sequencecorrespondstothecoreprotein. The molecu-lar weight of the DHBV core protein has been estimated by gel electrophoresis to be 35,000 (W. S. Mason and J. Newbolt, personal communication). Thiswouldbe in good agreementwithaprotein made with the entire open reading

frame from ATG 2518 to TAA 412 for which a molecular weight of 34,986 can be predicted. However, some topologi-calproblem may prevent theuseof that ATG because of the

position ofthe nickon the minus strand, which was down-stream from ATG 2518. Iftranscription occurs on anicked genome, then transcription probably starts after the nick,

4

/

2620 300 2620

2/

300

FIG. 5. Nucleotide

sequence

comparisons. Sequences40 nucle-otideslong withahomology equalorsuperiorto50%werescored, andtheyareindicated by lines whosecoordinates

correspond

tothe position ofthat sequencewithin thetwocomparedgenomes. Left-hand row is acomparison ofWHV (W) and HBV (H) genomes; right-hand row is a comparison of DHBV (D) and HBV (H) genomes.

acceptor sequence at position 3171 in the

14BV

genome

supports this hypothesis. Donor and acceptor consensus sequences (20) are also found in the WHV and DHBV

sequences (Fig. 11).

The existence of splice and acceptor sequences at the suggested positionsraises thequestionastothechoice of the initiator ATG. In all three cases, the one defining the

N-terminal methionine is not the first encountered ATG, and neither is the first one with a purine at position -3 (16). Anotherquestionraisedbyaneventualspliceof HBsprotein mRNAis relativetothe conservationduringevolution ofan

open reading frame within the pre-S region. A working hypothesisis thatanotherproteincodedbythepre-S region alone orby thetotalityofregion7 is expressed.

(iii) Region5/8.Startingwithnucleotide2515,therewas an

open readingframe which continuedthrough the EcoRI site

uptostopcodon TAA 412. Because of its relative position, overlappingthe 3' and 5' ends ofregion 6 with its 5' and 3' ends, respectively, thisregionlooks likeaproductof fusion

of theformerlydefined 5 and 8regionsof HBV and WHV(9, 10).We eliminated thepossibilitythatasmall DNAfragment

H6

/

/

/

[image:6.612.86.275.71.426.2]

W6

FIG. 6. Comparison of the amino acid sequences of gene 6

proteins of HBV (H), DHBV (D), and WHV (W). Stretches 30

amino acids long with homology equal or superior to 20% are

indicated byaline.

H6

//

//

/

/

/

VOL. 49,1984

on November 10, 2019 by guest

http://jvi.asm.org/

[image:6.612.340.513.352.686.2]
(7)

D6 SerThrProGlyLysSerValSer o ArgAspSerSerAlaIleProValArgThrSerGlyAlaSerAspLysAsn RT ThrValAlaLeuHisLeuAlaIleProLeuLysTrpLysProAspHisThrProValTrpIleAspGlnTrpProLeu H6 CysTrpTrp o GlnPheArgAsnSerLysProCysSerAspTyrCysLeuSerLeuIleValAsnLeuLeuGluAsp SerProLeuGluGluGluAsnValTrpTyr o ArgGlyAsnThrSerTrpProAsnArg o ThrGlyLys o PheLeu

ProGlLysLeuValAlau6uThrGlnLeualGluLysGluLeuGlnLeuGlyHisIleGluProProLeuSers ys

TrpGlyProCysAlaGluHisGlyGluHisHislleArgIleProArgThrProSerArgValThrGlyGIyValPheLeu

ValAspLysAsnSerArgAsnThrGluGlu o ArgLeuValValAspPheSerGlnPheSerLysGlyLys o o Met TrpAsnThrProPhePheValIleArgLysAlaSerGlySerTyrArgLeuLeuHisAspLeuArgAlaValAsnAlaLys ValAspLysAsnProHisAsnThrAlaGluSerArgLeuValValAspPheSerGlnPheSer o GlyAsnTyrArgVal ArgPhe o ArgTyrTrpSerProAsnLeuSerThrLeuArgArgIle o o Val o Met o ArgIleSer o o

LeuValProPheGlyAlaValGlnGlnGlyAlaProValLeuSerAlaLeuProArgGlyTrpProLeuMetValLeuAsp SerTrp o LysPhe o o ProAsnLeuGlnSerLeuThrAsnLeu o SerSerAsnLeuSerTrpLeuSer o o

0 SerGlnAla o TyrHisLeu o o AsnProAlaSerSerSerArgLeu o ValSerAspGlyGlnArgValTyr LeuLysAspCysPhePheSerIleProLeuAlaGluGlnAspArgGluAlaPheAlaPheThrLeuProSerValAsnAsn ValSerAlaAla o TyrHisLeu o o HisProAlaAlaMetProHisLeuLeuValGlySerSerGlyLeuSerArg Tyr PheArgLysAla o Met o ValGlyLeu o o PheLeu-LeuHisLeuPhe GlnAlaProAlaArgArgPheGlnTrpLysValLeuProGlnGlyMetThrCysSerProThrlleCysGInLeuValVal Tyr// 54 aa //PheArgLysIle o Met o ValGlyLeu o o PheLeu-LeuAlaGlnPhe ThrThrAla o GlySerGluIleSerArgArgPheAsnValTrp-ThrPheThr o o o o Phe o o Cys

GlyGlnValLeuGluProLeuArgLeuLysHisProSerLeuCys-MetLeuHisTyrMetAspAspLeuLeuLeuAla ThrSerAlaIleCysSerValValArgArgAlaPheProHis o LeuAlaPheSer o o o o ValVal o Gly HisProAsnAlaArgHis o Asn o IleSerHisAla o Cys o Phe o GlnGluLeu o IleArg o AsnPhe AlaSerSerHisAspGlyLeuGluAlaAlaGlyGluGluValIleSerThrLeuGluArgAlaGlyPheThrIleSerPro

o Lys o ValGlnHis o o SerLeuPheThrAla o ThrAsnPhe o LeuSerLeu o IleHisLeuAsn o

o o ThrThrProSer o ValAsnGluIleArgPhe o o o GlnIleAspGluAsnPheMetLysIleGluGlu AspLysValGlnArgGluProGlyValGlnTyr LeuGlyTyrLysLeuGlySerThrTyrValAlaProValGly Asn o ThrLys o TrpGlyTyrSerLeuAsn-PheMet o o ValIle o CysTyrGlySerLeu o GlnGlu SerArgTrpLys o Leu o ThrVal1leLysLysIleLysValGlyGluTrpTyrAspTrpLysCysIleGlnArgPhe LeuValAla- GluProArgIleAlaThrLeuTrpAspValGlnLysLeuValGlySerLeuGlnTrpLeuArgProAla HisIltIleGlnLyslleLysGluCysPheArgLysLeuProIleAsnArgProIleAspTrpLysValCysGlnArgIle Val o HisLeuAsnPheValLeuProPheThrLysGlyAsnIleGluMetLeuLys o MetTyr o AlaIleThr o

LeuGlylleProProArgLeuMetGlyProPheTyrGluGlnLeuArgGlySerAspProAsnGluAlaArgGluTrpAsn Val o LeuLeuGlyPheAlaAlaProPheThrGlnCysGlyTyrProAlaLeuMet o LeuTyr o CysIleGlnSer GlnValAsnPheSerPheSerSerSerTyrArgThr o LeuTyrLysLeuThrMetGlyValCysLysLeuArgIleLys LeuAspMetLysMetAlaTrpArgGluIleValArgLeuSerThr ThrAlaAlaLeu LysGlnAlaPheThrPheSerProThrTyrLysAlaPheLeuCysLysGlnTyrLeuAsnLeuTyrProVal o ArgGln

ProLysSerSerValPro o0 ArgValAlaThrAsp o ThrProThrHis o o o Ser -HisIleThr- GluArgTrpAspProAlaLeuProLeuGluGlyAlaValAlaArgCysGluGlnGlyAlaIleGly-ValLeuGly-ArgProGlyLeuCysGlnValPheAlaAspAlaThrPro-ThrGlyTrp-- o LeuValMetGlyHisGlnArgMet Gly o SerAlaValPheAlaPhe o LysVal o AspIleHisVal o GluLeuLeuMetSerCysLeuAlaLysIle GInGlyLeuPheThrHisProArgSerCysLeuArgLeuPheSerThrGlnProThrLysAlaPheThrAlaTrpLeuGlu

Arg o Thr Phe o AlaProLeuProlleHis o AlaGluLeuLeu o AlaCysPheAlaArgSer MetileLysProArgCysLeuLeuSerAspSerThrPhe o CysHisLysArgTyrGlnThrLeuProTrpHisPheAla ValLeuThrLeuLeuIleThrLysLeuArgAlaSerAlaValArgThrPheGlyLysGluValAspIleLeuLeuLeuPro ArgSerGlyAlaAsn o IleGlyThrAspAsn o Val o LeuSerArgLysTyrThrSerPheProTrp o o Gly MetLeuAlaLysGlnLeu o LysProIleGlnLeuTyrPheValProSer o TyrAsnProAlaAspGlyPro o Arg

AlaCysPheArgGluAspLeuProLeuProGluGlyIleLeuLeuAlaLeuLysGlyPheAlaGlyLysIleArgSerSer

CysAlaAlaAsnTrpIle o ArgGlyThrSerPheValTyrValProSerAlaLeuAsnProAlaAspAspPro o Arg HisLys o ProAspTrpThrAlaPheProTyrThrProLeu o

LysAlaIleTyrIleProHisArgLeuCysGlyThr-AspThrProSerIlePheAspIleAlaArgProLeuHisValSerLeuLysValArgValThrAspHisProValProGly GlyArgLeuGlyLeuSerArgProLeuLeuArg o ProPheArgProThrThrGlyArg o SerLeuTyrAlaAspSer Stop

ProThrValPheThrAspAlaSerSerSerThrHisLysGlyValValValTrpArgGluGlyProArgTrpGluIle..RT

o Ser o ProSerHisLeuProAspArgVal o PheAlaSerProLeuHisValAlaTrpArgProPro Stop H6

FIG. 7. Amino acid sequencecomparison between RSVRT (RT, middle line), DHBV gene 6protein (D6, top line), and HBV gene 6 protein(H6, bottom line).RSVRT is takenasreference.Amino acids in either D6orH6 thatareidenticaltothecorresponding amino acid in RTarerepresentedbyanopencircle. Amino acids whichare notidentical butbelongtothesamefamilyareshaded.The amino acid sequence ofD6startswith the390th aminoacidafterthe first ATG of thereading frame,and H6startswithaminoacid 312. RTstartswith thefirst

ami-noacidofthemature protein.

788

on November 10, 2019 by guest

http://jvi.asm.org/

(8)
[image:8.612.58.554.93.187.2]

DHBV GENOME NUCLEOTIDE SEQUENCE 789

TABLE 1. Sequence comparisonshowinganonapeptide ofverysimilarsequencethat isfound within the threehepatitisBgene6 proteins and thetworeversetranscriptases

Peptide

Protein position Aminoacid

within sequence

RSVRT 180 Tyr Met Asp Asp Leu Leu Leu Ala Ala

MLVRT 342 a Valb _

WHs 583 Valb Glyb

DHBs 561 - Phe Cys His

HBs 538 - Valb Valb - Glyb

-a

_,

Aminoacids identicaltothose shownfor RSVRT.

bAmino acidsof the same group asthe amino acids shown for RSVRT.

and theinitiationcodonforthe5/8 gene product will then be ATG 2647, which is notfavored by the presence ofaT at position 2644 (16). On the other hand, the transcriptional template may not have a nick, allowing translation to start with ATG 2518.

We found nohomology between the N-terminal end ofthe DHBV5/8 region andregion5ofHBVorWHV.Therefore,

although we believe that the third open reading frame in DHBVdoes represent afusionof the 5 and 8regions,we can drawnofirmconclusions. However, it shouldbenoted that there was far less homology between HBV and WHV with region 5 than there was with the rest of the genome.

Atthe presenttime, noproteincodedbyregion5of HBV or WHV has been identified, and no function has been

clearlyproposed.However,theexistence ofageneencoded inthisregionhasbeensupported bycomparative analysisof

amino acid and nucleotide sequences (9). Therefore, the questionarises whether thefunction oftheprotein codedby HBV and WHVregion5has thesamefunctionastheprotein

codedby region 5/8 ofthe DHBV genome. In otherwords,

does the protein made by region 5/8 have two roles, one devotedtothecoreproteinandonedevotedtogeneprotein 5, or is region 5/8, either through RNA splicing or protein

processing,making twodifferent proteins?

Replicationorigin. The genomeof thehepatitisBvirus isa noncovalentcircle (15). Theinterruptionin the L strand has beenassignedin HBV and WHVtotheonlyregiondevoid of

codingcapacity,between theopenreadingframes 5 and 8in the vicinity of a hairpin structure (9). A similar hairpin structure, although different in sequence, was observed in the DHBVgenome,startingwith nucleotide 2504 andending withnucleotide 2525. This hairpinwaslocalized at the end of region 6 and overlappedthe beginningofregion 5/8. It was surrounded by two direct repeats (ACACCCCTCTC) at position 2478 and 2536, which are reminiscent of the short direct repeats foundatboth ends of retroviruses.

Nucleotidesequencecomparisonsof theHBV, WHV,and DHBV genomes atthe molecular level clearly showed that these viruses belong to the same family and were derived from a common ancestor. Whereas HBV and WHV share between 60 to 70% nucleotide sequence homology (9), DHBVshowed much lesshomology (around or below40%) foralargepartof the genome andshowedbetween 50 to 55% homology for the remaining part, except for a small se-quenceof 100 nucleotides which reached 70%. This indicates thatduringevolution the threevirusesdid not separate from each other at the same time, but that DHBV separated earlierfrom the ancestorof the twoothers. This is probably inconjunction with the fact that DHBV infects birds and that birds started toevolveseparately from mammals 250 million years ago.Inturn, this also couldindicate that, at least inthe caseofthe hepatitisvirusfamilybut possibly for all kinds of

viruses, viruses appeared very early during evolution and evolved in a fashion parallel to their target. Unfortunately,

SerLeuLeuGluThrHisProLeuTyrGlnSerGluProAlaValProValIleLysThrPro

DHBs 1206 TCTCTCCTCGAGACTCATCCGCTATACCAGTCAGAACCAGCGGTGCCAGTGATAAAAACTCCC HBs 79 GGAACAGTAAACCCTGTTCTGACTACTGCCTCTCCCTTATCGTCAATCTTCTCGAGGATTGGG

GlyThrValAsnProValLeuThrThrAla o ProLeuSerSerlIePheSerArgIleGly

ProLeuLysLysLysMMetSerGlyThrPheGlyGlylleLeuAlaGlyLeulieGlyLeuLeu

DHBs 1269 CCCTTGAAGAAGAAAATGTCTGGTACCTTCCGGGGAATACTAGCTGGCCTAATCGGATTACTG HBs 142 GACCCTGCGCTGAACATGGAGAACATCACATCAGGATTCCTAGGACCCCTTCTCGTGTTACAG AspProAlaLeuAsn o GluAsnIleThrSer o Phe o GlyPro o LeuVal o Gln

ValSerPhePheLeuLeuIleLyslIeLeuGluIleLeuArgArgLeuAspTrpTrpTrplle

DHBs 1332 GTAAGCTTTTTCTTGTTGATAAAAATTCTAGAAATACTGAGGAGGCTAGATTGGTGGTGGATT HBs 205 GCGGGGTTTTTCTTGTTGACAAGAATCCTCACAATACCGCAGAGTCTAGACTCGTGGTGGACT AlaGly o o o o ThrArg o o Thr o ProGinSer o o Ser o o Thr

SerLeuSerSerProLysGlyLysMetGlnCysAlaPheGlnAspThrGlyAlaGInlleSer

DHBs 1395 TCTCTCAGTTCTCCAAAGGGAAAAATGCAATGCGCTTTCCAAGATACTGGAGCCCAAATCTCT

HBs 268 TCTCTCAATTTTCTAGGGCGAACTACCGTGTGTCTTGGCCAAAATTCGCAGTCCCCAACC

o o AsnPheLeuGly o ThrThrVal o LeuGly o AsnSerGlnSerProThr

FIG. 8. Nucleotide and amino acid sequence comparison around ATG 1285 for DHBs protein and ATG 157 for HBs protein, which

probablycodefor the N-terminal methionine of the surfaceantigen. Althoughthereisnohomology upstream from theseATGs,numerous

identicalamino acids downstream from these positions can be observed. Shading indicates ATG 1285 and ATG 157. VOL. 49,1984

on November 10, 2019 by guest

http://jvi.asm.org/

(9)

DHBs 1 MetSerGlyThrPheGlyGlyIleLeuAlaGlyLeuIleGlyLeuLeuValSerPhePheLeuLeuIleLys

HBs 1 o GluAsnIleThrSer o Phe o GlyPro o LeuVal o GlnAlaGly o o o o ThrArg DHBs 25 IleLeuGluIleLeuArgArgLeuAspTrpTrpTrpIleSerLeuSerSerProLysGlyLysMetGlnCys

HBs 25 o o Thr o ProGlnSer o o Ser o o Thr o o AsnPheLeuGly o ThrThrVal o

DHBs 49 AlaPheGlnAspThrGlyAlaGlnIleSerProHisTyrValGlySerCysProTrpGlyCysProGlyPhe HBs 49 LeuGly o AsnSerGlnSerProThrSerAsn o SerProThr o o o ProThr o o o Tyr

6 //AspLeuSerGlnAlaPheTyrHisLeuProLeuAsnProAlaSerSerSerArgLeuAlaValSer

s 73 LeuTrpThrTyrLeuArgLeupheIleIlePheLeuLeuIleLeuLeuValAlaAlaGlyLeuLeuTyrLeu

DHBV 1506 /7 ACCTATCTCAGGCTTTTTATCATCTTCCTCTTAATCCTGCTAGTAGCAGCAGGCTTGCTGTATCTG

***.... ... 5*...X^... . ...

..-HBV 379 7/ ATGTGTCTGCGGCGTTTTATCATCTTCCTCTTCATCCTGCTGCTATGCCTCATCTTCTTGTTGGTT s 73 Arg o MetCys o o Arg.o o 0 oo Phe o o o LeuCysLeuIlePhe o LeuVal

6 // o Val o Ala o o o o o o o His o o AlaMetProHis o Leu o Gly 6 AspGlyGlnArgValTyr

s 97 ThrAspAsnGlySerThr

DHBV ACGGACAACGGGTCTACTA

HBV CTTCTGGACTATCAAGGTATGTTGCCCGTTTGTCCTCTAATTCCAGGATCCTCAACAACCAGCACGGGACCA

s 97 LeuLeuAspTyrGlnGlyMetLeuProValCysProLeuIleProGlySerSerThrThrSerThrGlyPro 6 SerSerGlyLeuSerArgTyrValAlaArgLeuSerSerAsnSerArgIleLeuAsnAsnGlnHisGlyThr HBV TGCCGGACCTGCATGACTACTGCTCAAGGAACCTCTATGTATCCCTCCTGTTGCTGTACCAAACCTTCGGAC

s 97 CysArgThrCysMetThrThrAlaGlnGlyThrSerMetTyrProSerCysCysCysThrLysProSerAsp 6 MetProAspLeuHisAspTyrCysSerArgAsnLeuTyrValSerLeuLeuLeuLeuTyrGlnThrPheGly

6 TyrPheArgLysAlaProMetGlyValGlyLeuSer

s 103 IleLeuGlyLysLeuGlnTrpAlaSerValSerAla

DHBV TTTTAGGAAAGCTCCAATGGGCGTCGGTCTCAGCC

HBV GGAAATTGCACCTGTATTCCCATCCCATCATCCTGGGCTTTCGGAAAATTCCTATGGGAGTGGGCCTCAGCC

s 121 GlyAsnCysThrCysIleProIleProSerSerTrpAlaPhe o o PheLeu o GluTrpAla o o

6 ArgLysLeuHisLeuTyrSerHisProIleIleLeuGly o o o Ile o o o o o o o

6 ProPheLeuLeuHisLeuPheThrThrAlaLeuGlySerGluIleSerArgArgPhe-AsnValTrpThr

s 115 LeuPheSerSerIleSerSerLeuLeuProSerAspProLysSerLeuValAlaLeu-ThrPheGlyLeu

DHBV CTTTTCTCCTCCATCTCTTCACTACTGCCCTCGGATCCGAAATCTCTCGTCGCTTT-AACGTTTGGACTT

HBV CGTTTCTCCTGGCTCAGTTTACTAGTGCCATTTGTTCAGTGGTTCGTAGGGCTTTCCCCCACTGTTTGGCTT s 145 Arg o o TrpLeu o Leu o Val o PheValGlnTrpLeuValGlyLeuSerPro o ValTrp o

6 o o o o AlaGin o o Ser 0 IleCys o ValValArg o Ala o ProHisCysLeuAla 6 PheThr//

s 138 SerLeuIleTrpMetThrSerSerSerAlaThrGlnThrLeuValThrLeuThrGln-LeuAlaThrLeu

DHBV TCACTTATATGGATGACTTCCTCCTCTGCCACCCAAACGCTCGTCACCTTAACGCA-ATTAGCCACGCTG

... ...i... ... ... . . . ... ..

HBV TCAGTTATATGGATGATGTGGTATTGGGGOCCAAGTCTGTACAGCATCTTGAGTCCCTTTTTACCGCTGTTA

s i69 o ValIleTrpMetMetTrpTyrTrpGlyProSerLeuTyrSerIleLeuSerProPheLeuProLeuLeu

6 0 Ser//

s 161

SerAlaLeuPheTyrLysSer-DHBV TCTGCTCTTTTTTACAAGAGTTAG// HBV CCAATTTTCTTTTGTCTTTGGGTATACATTTAA//

s 193

ProIlePhePheCysLeuTrpValTyrIle-FIG. 9. Amino acid and nucleotide sequencecomparisons ofgene7. The positionof a deletion within the DHBVgenomeaffectingthe

surfaceantigenandgene 6proteinhasbeen determinedbycomparingtheDHBsand HBsproteinsequences, the DHBV and HBVprotein6 sequences, and the DNA sequences. Toclarifythefigure,the gene 6protein sequences and the DNA sequencesareonlyshown inpart, delimitedby11.TheDHBs and HBsproteinsequencesareshownin full. Dotsindicatehomologybetween the DNAsequences. An open circle in theHBVproteinsequencesindicatesthattheaminoacidatthatpositionisthesame asthecorrespondingamino acid in thehomologous DHBVprotein.

s GlyPheLeuGlyProLeuLeuValLeuGinAlaGlyPhePheLeuLeuThrArgIleLeu

6 ArgIleProArgThrProSerArgValThrGlyGLyValPheLeuValAspLysAsnPro

HBV 174 AGGATTCCTAGGACCCCTTCTCGTGTTACAGGCGGGGTTTTTCTTGTTGACAAGAATCCT

DHBV 1300GGGAATACTAGCTGGCCTAATCGGATTACTGGTAAGCTTTTTCTTGTTGATAAAAATTCT 6 GlyAsnThrSerTrp o Asn o Ile o o LysLeu o o o o o 0 Ser

s o Ile o AlaGly o IleGly o LeuValSer o o o o IleLys o o

s ThrIleProGlnSerLeuAspSerTrpTrpThrSerLeuAsnPheLeuGlyGlyThr

6 HisAsnThrSerGluSerArgLeuValValAspPheSerGlnPheSerArgGlyAsnTyr

HBV 258 CACAATACCGCAGAGTCTAGACTCGTGGTGGACTTCTCTCAATTTTCTAGGGGGAACTA

DHBV 1384 AGAAATACTGAGGAGGCTAGATTGGTGGTGGATTTCTCTCAGTTCTCCAAAGGGAAAAA 6 Arg o o Glu o Ala o o o o o o o o o Lys o LysAsn s Glu o LeuArgArg o o Trp o o Ile o o SerSerProLy3 o Lys

FIG. 10. Amino acid and nucleotide sequencecomparison of thehighlyconserved sequence locatedinDHBV between nucleotide1300

and1400. As canbeseen, the gene 6proteinsare moreconservedin thisregionthan the surfaceantigens. Out of 40 aminoacids,24are

identi-calfor protein 6, whereasonly 18outof39areidentical in the surfaceantigen. Symbols arethe same asthose in Fig. 9. 790

on November 10, 2019 by guest

http://jvi.asm.org/

[image:9.612.126.482.35.526.2]
(10)

DHBV GENOME NUCLEOTIDE SEQUENCE 791

Cap Donor

II

HBV I TATATAA

2776

Acceptor N-methionine

ACTCATCCTCAG=AUCAC....

... AAC JAT

3170/3171 157

Cap Donor WHY several TA

A

TAAAGIGTAAC ....

rich region2949/2s95

Acceptor CTTTTCATCTCCAG ....

141/142

ACAAT1 CCTATGOAC....GA\ ATU TCA

193 19,' 2(0

Cap Donor DHBV several TA I CCAGCCTAGG...

Irich regionsl 616/617

Donor consens

Acceptor

TCTTCCCCTCCTCACGGA ...

1163/1164

CA

A

thishypothesiscannotbe testedfor themomentsinceHBV, WHV, and DHBV so far represent the only examples of viruses which belong tothe same family but infect widely

different hosts and whose genomes have been entirely se-quenced.

HO

7/w

W8

FIG. 12. Amino acid sequence comparison of thecore antigen.

Stretches of 20amino acidswithhomology equalorabove 20%are

indicatedby lines.

ACKNOWLEDGMENTS

We are very gratefulto B. Masson andJ. Summers for helpful discussions and forthegift of the DNA cloned recombinants.

This workwassupported inpartbyInstitut National de la Santeet

dela Recherche MddicalethroughgrantSC15.

LITERATURE CITED

1. Alestrom, P., G. Akusjarvi, U. Pettersson, and M. Pettersson. 1982.DNAsequenceanalysisof theregion encodingthe termi-nalproteinand thehypotheticalN. Geneproductofadenovirus

type2. J. Biol. Chem. 257:13492-13498.

2. Burrel, C. J.,P.Mackay,P.J. Greenaway,P. H.Hofschneider, andK.Murray.1979.ExpressioninEscherichia coli ofhepatitis B virus DNA cloned in plasmid pBR 322. Nature (London) 279:43-47.

3. Charnay, P.,E.Mandart,A.Hampe,F.Fitoussi,P.Tiollais,and F.Galibert. 1979.Localizationontheviralgenomeand nucleo-tidesequenceofthegenecoding for thetwomajorpolypeptides ofthehepatitisBsurfaceantigen (HBs Ag).NucleicAcids Res. 7:335-346.

4. Charnay, P., C. Pourcel,A.Louise, A.Fritsch,and P.Tiollais. 1979. Cloning in Escherichia coli and physical structure of hepatitisB virionDNA. Proc.Natl. Acad. Sci. U.S.A. 76:2222-2226.

5. Cummings,I.W.,J.K.Browne,W.A.Salser,G. V.Tyler,R. L.

Snyder, J. M.Smolec, andJ. Summers. 1980. Isolation charac-terization andcomparison of recombinant DNAsderived from the humanhepatitis B and woodchuck hepatitis virusgenome.

Proc.Natl. Acad. Sci. U.S.A. 77:1842-1846.

6. Dayhoff, M. O., R. V. Eck, and C. M. Park. 1972. A model of evolutionary change in proteins, p. 89-99. In M. 0. Dayhoff

(ed.), Atlas of protein sequence and structure 1972, vol. 5.

NationalBiomedical Research Foundation, Washington, D.C. 7. Dreesman,G. R., Y.Sandrez, I. Ionescu-Matin, J. T. Sparrow,

H. R.Six, D. L. Peterson, F. B. Hollinger,andJ. L. Melnick. 1982. Antibody to hepatitis B surface antigen after a single

inoculation ofuncoupled synthetic HBs Ag peptides. Nature (London) 295:158-160.

8. Galibert, F., T. N. Chen, and E. Mandart. 1981. Localization and nucleotidesequenceof thegenescoding for the woodchuck hepatitisvirussurfaceantigen: comparison with thegenecoding

AAA ArC TCT 12'4

Acceptor consensus (C)IINTAC/C

FIG. 11. Comparison ofputative controlsequencesfortheexpression ofgene 7. VOL.49, 1984

on November 10, 2019 by guest

http://jvi.asm.org/

[image:10.612.90.532.68.317.2] [image:10.612.72.275.424.693.2]
(11)

792 MANDART, KAY, AND GALIBERT

for the human hepatitis B virus surface antigen. Proc. Natl. Acad. Sci. U.S.A. 78:5315-5319.

9. Galibert, F., T. N. Chen, and E. Mandart. 1982. Nucleotide sequence of a clonedwoodchuckhepatitisvirus genome: com-parisonwith thehepatitisBvirus sequence. J.Virol. 41:51-65. 10. Galibert, F., E. Mandart, F. Fitoussi, P. Charnay, and F. Galibert. 1979. Nucleotide sequence of the hepatitis B virus genome (subtype ayw) cloned in E. coli. Nature (London) 281:646-650.

11. Gerlich, W. H., M. A. Feitelson, P. L. Marion, and W. S. Robinson. 1980. Structural relationships between the surface antigens of ground squirrel hepatitis virus and humanhepatitisB virus. J. Virol. 36:787-795.

12. Gingeras, T. R., D. Sciaky,R. E.Gelinas,J. Bing-Dong,C. E. Yen, M. M. Kelly, P. A. Bullock, B. L. Parsons, K. E.O'Neill, andR. J. Roberts. 1982.Nucleotidesequencesfromthe adeno-virus-2 genome. J. Biol. Chem. 257:13475-13491.

13. Hartley, J.L., and J. E. Donelson. 1980.Nucleotidesequence of the yeastplasmid. Nature (London)286:860-864.

14. Herisse, J., G. Courtois, and F. Galibert. 1980. Nucleotide sequence of the EcoRI D fragment of adenovirus 2 genome. NucleicAcids Res. 8:2173-2191.

15. Hruska, J. F., D. A. Clayton, J. L. R. Rubenstein, and W. S. Robinson. 1977. Structure of hepatitis B Dane particle DNA beforeandafter theDaneparticleDNApolymerase reaction.J. Virol.21:666-672.

16. Kozak, M. 1981. Possible role offlanking nucleotides in recogni-tionof AUG initiatorcodon byeukaryoticribosomes. Nucleic Acids Res.9:5233-5252.

17. Marion, P. L., L. S. Oshiro, D. C. Regnery, G. H. Scullard, and W.S. Robinson. 1980. AvirusinBeechey ground squirrels that isrelated tohepatitisBvirus on humans. Proc. Natl. Acad. Sci. U.S.A. 77:2941-2945.

18. Mason, W. S., G. Seal, and J. Summers. 1980. Virus of Pekin ducks with structural and biological relatedness to human hepatitisB virus.J. Virol. 36:829-836.

19. Maxam, A., and W. Gilbert. 1980.Sequencing end labeled DNA with base specific chemical cleavage. Methods Enzymol. 65:499-560.

20. Mount, S. M. 1982. A catalogue ofsplice junction sequences. Nucleic Acids Res. 10:459-472.

21. Pourcel,C., A.Louise,M.Gervais,N.Chenciner,M.-F.Dubois, and P. Tiollais. 1982. Transcription of the hepatitis B surface antigengeneinmousecells transformed with clonedviralDNA. J. Virol. 42:100-105.

22. Prince, A. M., H. Ikram, and T. P. Hopp. 1982. Hepatitis B virusvaccine: Identification ofHBsAg/aand HBsAg/dbutnot HBsAg/y subtypeantigenic determinantson asynthetic immu-nogenicpeptide. Proc. Natl. Acad. Sci. U.S.A.79:579-582. 23. Rao,K.R., and G. N.Vyas.1976. Biochemicalcharacterization

ofhepatitisB surfaceantigeninrelationtoserologicalactivity.

J. Biol. Stand. 4:295-304.

24. Robinson, W. S., P. L. Marion, M. A. Feitelson, and A. A. Siddiqui. 1981. the hepadna virus group: hepatitis B andrelated viruses, p. 57-58. In W. Szmuness, H. J. Alter, and J. E. Maynard (ed.), Proceedings of the International Symposium on ViralHepatitis. Franklin Institute Press, Philadelphia. 25. Schwartz, D., R. Tizard, and W. Gilbert. 1983. Nucleotide

sequenceofRous sarcomavirus. Cell32:853-869.

26. Segrest, J. P., and R. J. Feldmann. 1974. Membrane proteins: amino acid sequence and membrane penetration. J. Mol. Biol. 87:853-858.

27. Shinnick, T. M., R. A. Lerner, and J. G. Sutcliffe. 1981. Nucleotide sequence of Moloney murine leukaemia virus. Na-ture(London) 293:543-548.

28. Sninsky, J. J., A. Siddiqui, W. S. Robinson, and S. N. Cohen. 1979.Cloningandendonuclease mapping of the hepatitis Bviral genome.Nature (London) 279:346-348.

29. Staden,R. 1977.Sequence data handling by computer. Nucleic Acids Res.4:4037-4051.

29a.Stenlund, A., D. Lamy, J.Moreno-Lopez, H. Ahola, U. Petters-son, and P. Tiollais. 1983. Secretion of the hepatitis B virus surface antigen from mouse cells using an extra-chromosomal eucaryotic vector. EMBO J. 5:669-673.

30. Struck, D. K., W. J. Lennarz, and K. Brew. 1978. Primary structuralrequirements for the enzymatic formation of the N-glycosidic bond in glycoproteins. J. Biol. Chem.253:5784-5786. 31. Summers,J., and W.S. Mason. 1982.Replication ofthe genome of ahepatitis-B-like virus by reversetranscriptionof an RNA intermediate. Cell29:403-415.

32. Summers, J., A. O'Connell,and 1. Millman. 1975. Genome of hepatitisB virus: restrictionenzymecleavageand structureof DNA extracted from Dane particles. Proc. Natl. Acad. Sci. U.S.A. 72:4797-4801.

33. Summers,J., J. M.Smolec, and R.Snyder. 1978. Avirus similar tohumanhepatitisBvirus associated withhepatitis and hepato-ma in woodchucks. Proc. Natl. Acad. Sci. U.S.A. 75:4533-4537.

34. Summers,J., J. M.Smolec, B. G. Werner, Jr., T.J.Kelly, G. V. Tyler, and R. L.Snyder. 1980.HepatitisBvirusandwoodchuck hepatitis virus are membersofa novel class ofDNA viruses. Viruses in naturally occurring cancers. Cold Spring Harbor Conf. Cell Proliferation 7:459-470.

35. Valenzuala, P., P. Gray, M. Quiroza, J. Zaldivar, H. M. Goodman, and W. J. Rutter. 1979. Nucleotide sequence of the gene coding for the major protein of hepatitis B virus surface antigen. Nature (London) 280:815-819.

36. Valenzuala, P., M. Quiroga, J. Zaldivar, P. Gray, and W. J. Rutter. 1981.The nucleotide sequence of thehepatitis Bviral genome and theidentification ofthe majorviral genes. In B. Fields, R.Jalnisch, and C. F. Fox (ed.), Animal virus genetics. Academic Press, Inc., NewYork.

J. VIROL.

on November 10, 2019 by guest

http://jvi.asm.org/

Figure

FIG.1.labeled4,analyzedfragmentscorrespond Sau3a; Diagramof analyzed DNAfragments.Verticalbarstothepositionsof thelabeledendsof restriction used
FIG. 29
FIG. 4.Numbersstriped Localization of the open reading frames on the viral genomes of HBV and WHV and comparison with the DHBV genome
FIG. 5.otidesandhandright-handposition Nucleotide sequence comparisons. Sequences 40 nucle- long with a homology equal or superior to 50% were scored, they are indicated by lines whose coordinates correspond to the of that sequence within the two compared
+4

References

Related documents

cell polymerases as well as the viral polymerase and because in vivo viral density DNA could not be detected in wild-type HCMV in the presence of PAA (see above), much of the in

Filamentous virions became readily visible on the surface of phage-producing cells, whereas virus particles were rarely seen on freeze-etched nonproducing cells, for example,

Consequently, amorphous aggregates have a di ff erent rotamer pro fi le to monomers, and the proto fi bril β -sheet structures reveal two new rotamer states that appear to be

Our analysis of the four different business models also enables us to make sense of something often observed but not always explained, namely that the world is rather full of small

Identification of a dyslexic profile among the signing participants was more complex as different phonological measures were used that did not rely on speech

Social Cognitive Theory, Intervention, Experiment, Control, Student, College , and University.. The search process was conducted by entering combinations of keywords

methods based on the frequency theory of probability.. impulse responses or variance decompositions) involving models such as structural VARs. Note that this equation allows for

 Assess the impact of a community mobilization intervention through women groups on home care, health care seeking behaviour and maternal and infant mortality in the?.