Vol.49, No. 3 JOURNALOFVIROLOGY, Mar. 1984, p.782-792
0022-538X/84/030782-11$02.00/0
Copyright © 1984, American Society for Microbiology
Nucleotide Sequence of
a
Cloned
Duck Hepatitis
B
Virus Genome:
Comparison with Woodchuck and Human Hepatitis
B
Virus
Sequences
ELISABETH MANDART, ALAN KAY, AND FRANCIS GALIBERT*
Laboratoire d'Hematologie Experimentale, Centre Hayem, HopitalSaint-Louis, 75475 ParisCe'dex 10, France Received 16 June 1983/Accepted 6 October 1983
The nucleotide sequenceofanEcoRI duckhepatitis B virus (DHBV) clone waselucidated by usingthe
Maxam and Gilbert method. This sequence, which is 3,021 nucleotides long, was comparedwiththe two previously analyzed hepatitis B-like viruses (human andwoodchuck). From this comparison, itwas shown
that DHBV is derived from an ancestor common to the two others but has a slightly different genomic organization. There wasnointergenic region between genes5 and 8, whichwerefused intoa singleopen
reading frame in DHBV. Genes for the surface andcoreproteinswereassignedtoopenreading frames 7 and
5/8. Amino acid comparisons showed some structural relationship between gene 6 product and avian reverse transcriptase, suggesting either evolution from a common ancestor or convergence to some
particular structure tofulfill aspecific function. This should be correlated with the synthesis ofan RNA
intermediate duringDNAreplication. This is also takenas anargumentinfavor of thehypothesis thatgene
6 codesfor the DNApolymerase that isfound within the virion. DNA sequencecomparison also showed that thetwomammalian hepatitis B viruses are more homologous to each other thanthey areto DHBV, indicating thatDHBVstarts toevolveonitsownearlier than the twootherviruses,asdo birds compared
with mammals. Fromthis it is proposed that the virusesevolved inafashion parallel to the speciesthey infect.
Duck hepatitis B virus (DHBV), which was recently
isolated
(18), is the fourth member ofa new viral family called Hepadnaviridae, whose prototypeishuman hepatitisBvirus(HBV) (24). The fourmembersofthisgrowing family
have several characteristics in common, such as ultrastruc-tire, antigenic makeup, DNA size, and structure.
Similar-ities inthe
pathological
field have been describedaswell (11,17, 33, 34).
In the past few years, by molecular cloning (4, 5) and
nucleotide sequence analysis (3, 8-10, 35, 36), our
knowl-edge of the biology of these viruses has
increased.
Two genes, one coding forthecoat protein and the othercoding
for thecoreprotein,
havebeen identifiedonthe genomes ofHBV andwoodchuck hepatitis B virus (WHV). Twoother
possible coding regions have been
mapped,
but their func-tion and their products have not, asyet, been determined.In spite ofourgrowing knowledge aboutthe structure of
various viral components, the
biological study
of theseviruses is still complicated by the absence of cell cultures
susceptible toinfection.
Some ofthese
complications
have been overcomeby
thediscovery ofDHBV (18), which is not only able to infect ducks, an animaleasily kept in colony, but can alsoinfect andmultiplyinembryonic eggs. These
findings
caused usto undertake the elucidation of the nucleotide sequence of the DHBV genome. Duringthecourseofthisstudy, the useful-nessof the duck-DHBV modelwasprovenbythefinding
of Summers and Mason (31), who, using an in vitro system, found evidenceimplicating an RNA intermediate in DHBV DNA replication.Inthis paper,wereport thecompletenucleotide sequence of the genome of DHBV and compare its
primary
structure with thatof the HBV and WHVgenomes.*Correspondingauthor.
782
MATERIALS AND METHODS
Enzymes and chemicals. Restriction endonucleases came from NewEngland Biolabs and were used as recommended bythe manufacturer. DNA polymerase I canie from Boeh-ringerMannheimBiochemicals, andbacterialalkaline phos-phatase and polynucleotide kinase were from P. L. Biochemicals.
Chemicals used for nucleotide sequence analysiswere as described, previously (14). [-y-32P]ATP (specific activity,
>2,500 Ci/mmol) and
ot32P-labeled
nucleotide triphosphate (specificactivity, >3,000Ci/mmol)werefrom New England Nuclear Corp.Preparation ofEcoandXho DHBV DNAs. X-DHBV recom-binants wereconstructed andgiven to us by W. S. Mason et al. The cloned DNAs were referred to as Eco and Xho DHBV DNAs.Propagation and purification ofthe recombi-nants, as well aspreparation ofthe DNAs, wereperformed aspreviously described(3, 10, 14).
Containment. Containment conditions were as recom-mended by the French National Control Committee. The culture ofrecombinantbacteriophage wasdone under L3B1 conditions.
DNAnucleotidesequence. Sequence analyses were deter-mined bythe Maxam and Gilbert method(19). Usually, ca. 10 pmol of EcoDHBV DNA (20
Rg)
wasfullydigested each time with a given restriction enzyme. Fragments werede-phosphorylatedand labeled with
[y-32P]ATP
andpolynucle-otide kinase as described previously (14). To separate the twolabeled ends, fragments were denatured by heating to 920C in the presence of 30%dimethylsulfoxide and fraction-ated by electrophoresis in acrylamide gel (19). Fragments larger than 600 base pairs were hydrolyzed with another
restriction enzyme. Under some circumstances, fragments
with a recessed 3' end were labeled with an
ot32P-labeled
nucleotidetriphosphate of choice and DNApolymeraseIas describedby Hartley and Donelson (13).
on November 10, 2019 by guest
http://jvi.asm.org/
DHBV GENOME NUCLEOTIDE SEQUENCE 783
f-J
FIG. 1. Diagram of analyzed DNA fragments. Vertical bars correspond to the positions of the labeled ends of restriction fragments used. Length of the arrows is relative to the number of analyzednucleotides.Row 1, HaeIII;row2. Hinfl;row3. Aliil;row
4, Sau3a; row5, RsaI; row 6, Mspl +BgllI. Sau3a fragments are
labeled at their 3'ends; the othersare labeledat their 5' ends.
RESULTS AND DISCUSSION
The complete nucleotide sequence of the EcoRI DHBV DNA cloned in Escheric-hia coli was determined
by
themethod of Maxam and Gilbert (19),
using
five different chemical reactions giving specific bands forG,
AG, CT, C, and AC. The sequence was derived from alarge
number ofoverlappingfragments sothatboth strands were
entirely
and independently analyzed, and all starting restriction siteswere analyzed within an overlapping fragment (Fig. 1). For
reasons explained below, the EcoRI site used for
cloning
was also overlapped by using anXlzoI
DHBV clone,elim-inating the risk of theloss ofa small DNA fragment located
betweentwoputative Ec oRI sites withinthe DHBVgenome. The DHBVsequence shown in Fig. 2 is 3,021 nucleotides
long, compared with 3,182for the HBV DNA and 3,308 for the WHV DNA. Restriction enzyme data obtained from virion DNA enabled us to define which strand was the L
strand and which was the S strand in both HBV and WHV (9, 10). We do not have such data for DHBV, but the position of theopen reading frames aswell asthe nucleotide
sequence homologies found within these three genomes indicate that the nucleotide sequence presented in Fig. 2 is
complementary to the L strand, alsocalled the minusstrand by Summers and Mason (31).
Location of the open reading frames. The number and distribution ofstop codons within the L strand is suchthat, as previously noticed with HBV and WHV, it is difficult to locate agene which could be transcribed from the S strand (Fig. 3). Only one region spanning from nucleotide 1397 to nucleotide 835 could have a substantial coding capacity. However, (i) the first in-phase ATG in this region
(open
reading frame 1) is toward the end, at position 904. very closetothestop
codon;
(ii) thereis noobviousacceptor site at the beginning ofopen reading frame 1 to which an exoncontaining an ATG could be spliced; and (iii) neither HBV
nor WHV has corresponding open reading frames which
could give a protein homologous to the putative protein
made from open reading frames1 of DHBV. These suggest
the absence of a coding function for the S strand, as
previouslyobservedfor the HBVand WHVgenomes (9,
10).
On the contrary, the number and distribution of stop codons in the S strandleave open several regions, indicating
that mRNA could arise by transcription of the L strand. However, only three open reading frames are displayed
along the DHBV genome instead of the four previously noticed for the HBV and WHV genomes. These three
regions span from nucleotides 14 to 2527, 684 to 1784, and 2515 to 411. Their relative positions (Fig. 4) as well as
information derived fromnucleotide
aidd
aminoacidcompar-isons of HBV and WHV sequences indicate that (i) the
largest
openreading
frame (nucleotides 14 to2527)
corre-sponds
togene 6 ofHBV;(ii)
theopenreading
frame which isoverlapped by
region
6and goes fromnucleotides 684 to 1784corresponds
to gene7,
also called S (for surfaceprotein):
and(iii)
regions 5 and 8 found in HBV and WHVgenomesarefusedintoa
single region
inDHBV,
assuggest-ed
by
the size andposition
ofthethird openreading
frame(nucleotides 2515 to
411).
Therefore,
we call itregion
5/8. Nucleotide sequencecomparisons.
By
using a computerprogram
developed
by
Staden(29),
thenucleotide sequence ofthe DHBVgenome wascompared
withthose of the HBV and WHVgenomes.Although
nucleotidesequencecompari-sons show a
large degree
ofhomology
between HBV andWHV,
62 to70% along
the genomes except in two smallregions
(9),
amuchsmaller level ofhomology
wasobserved with the DHBV sequence.Figure
5 shows severalgraphs
which
demonstrate
this result. For thelargest
part of the DNA sequence, thedegree
ofhomology
was below40%.
Fourregions
located between nucleotides 100 and200,
300 and500.
600 and800,
and 1300 and2100
were more conserved and had about50%
homology
with the homolo-gousregions
found in the HBV and WHV genomes. Theregion
between nucleotides 1300 and 1400 had thehighest
homology
(70%).
Thishighly
conserved sequence is part ofregions
6and7, which could be read intwodifferent frames.This
remarkably high degree
ofhomology
isprobably
dueto theproduct
ofgene6, which isslightly
moreconserved than theproduct
of gene7,
asexplained
below. As forHBV,
viruses found innature vary fromonetoanother(2,28,
32).
The two DHBV
clones,
clonedthrough
theEcoRI
orXlioI
site, were
exactly
the same size asdemonstrated
by
theircomparative
electrophoretic mobility
butexhibitedaslightly
different restrictionpattern. In agreement withthis observa-tion, the nucleotide sequence wedetermined
around the EcoRIsiteoftheX/loDHBVclonediverged
from that of theEco DHBV clone to an extent of ca.
3%.
None of these nucleotidechanges
alters thecoding capacity
of thisse-quence.
Amino acidsequence
comparisons.
(i)
Region
6.Inthethreeviruses,
thisregion
covered ca.80%
of the genome. InDHBV it went from
nucleotides
14 to2528.
From the first ATG(residue20)
uptothe TAAstop
codon(residue
2528),
aprotein of 836 amino acids can be
predicted,
ascompared
with 879 and 838 for the
corresponding
WHV and HBVproteins
(9,
10).
Theseproteins
have not beenidentified
sofar,
but the size of region 6 leaves little doubt that it does haveacoding
function.Acomparison
ofthepredicted
aminoacid sequences ofthese three proteins
clearly
shows thatthey
are related to each other(Fig.
6and7).
Although
theDNAsequencesallow thegene6products
ofDHBVandHBVtohave
nearly
identicallengths
(836aminoacids,
compared with
838),
it is difficult toalign
these twoproteins
starting
from the first methionine in each region. Ifregions
of amino acidhomology
arepaired,
then the best coincidenceisobtainedby
assuming
that theDHBVregion
6 protein starts at the secondin-phase
ATG. This would reduce the DHBVprotein
6 to a size of 786 amino acids.Supporting
thathypothesis,
thefirstin-phase
ATG ofregion6 is
preceded
atposition
-3by
apyrimidine,
whereas thesecond
in-phase
ATGispreceded
atposition
-3by
apurine(16).
Because of its size, which is
comparable
to variouspolymerases,
ithasbeenproposed
that theDNApolymerase
found
within
the virion would be codedby
gene 6(10). An experimentperformed
by
Summers and Mason (31) hasI r
2 IJ!2 1 ;
3 rt
4 1.,
5 r j
6 iJ W
VOL. 49. 1984
on November 10, 2019 by guest
http://jvi.asm.org/
[image:2.612.83.277.75.146.2]784 MANDART, KAY, AND GALIBERT
revealed the existence ofanRNAtranscript complementary
to the L strand during theprocess of DNA replication. We were therefore interested in comparing the amino acid sequencepredicted from the largeopenregions (region 6) of
the three known hepatitis viruseswith several viral encoded DNA polymerases orreversetranscriptases.
Theamino acid comparisonswere made withacomputer
program established by B. Caudron (personal
communica-tion) which is abletodetect small stretches of 20 amino acids with homology above 25%. With this program, the amino
acidsequences predicted fromgenes 6 werescored against theamino acidsequences of the DNA polymerase encoded
in the adenovirus 2 genome (1, 12) and the amino acid sequences of the avian and murine reverse transcriptases
(25, 27). Several stretches of amino acids with homologies
71
141
211
over25%werefound in all comparisons. However,
homolo-gies between adenovirus 2 DNA polymerase and the hepati-tis B virus region6products generally involved only oneof
the viruses at a time, and the positions of the various
homologies indicatednoparticularpattern.Onthecontrary, when the reverse transcriptases were compared with the
threehepatitis B virus region6products,at leastone setof amino acidsequencehomologies appearedtogiveacoherent
pattern. These sequences are indicated in Table 1. When DHBV and HBV gene 6 proteins and Rous sarcoma virus reversetranscriptase (RSVRT)arealigned by pairing thisset ofsequencehomologies, other smaller stretches of homolo-gy become apparent (Fig. 7). (The same is alsotrue when WHV region 6 protein is included.) Also, in a significant
number ofcasesinwhich therewas anamino aciddifference
Start of region 6 (frame 2)
CATGCTCATTTGAAAGCTTATGCAAAAATTAACGAGGAATCACTGGATAGGGCTAGGAGATTGCTTTGGT GGCATTACAACTGTTTACTGTGGGGAGAAGCTCAAGTTACTAACTATATTTCTCGTTTGCGTACTTGGTT GTCAACTCCTGAGAAATATAGAGGTAGAGATGCCCCGACCATTGAAGCAATCACTAGACCAATCCAGGTG GCTCAGGGAGGCAGAAAAACAACTACGGGTACTAGAAAACCTCGTGGACTCGAACCTAGAAGAAGAAAAG 281 TTAAAACCACAGTTGTCTATGGGAGAAGACGTTCAAAGTCCCGGGAAAGGAGAGCCCCTACACCCCAACG
End of the fused 5 + 8 region-1
351 TGCGGGCTCCCCTCTCCCACGTAGTTCGAGCAGCCACCATAGATCTCCCTCGCCTAGGAAATAAATTACC 421 TGCTAGGCATCACTTAGGTAAATTGTCAGGACTATATCAAATGAAGGGCTGTACTTTTAACCCAGAATGG 491 AAAGTACCAGATATTTCGGATACTCATTTTAATTTAGATGTAGTTAATGAGTGCCCTTCCCGAAATTGGA 561 AATATTTGACTCCAGCCAAATTCTGGCCCAAGAGCATTTCCTACTTTCCTGTCCAGGTAGGGGTTAAACC
Start of region 7 (frame 3)
631 AAAGTATCCTGACAATGTGATGCAACATGAATCAATAGTAGGTAAATATTTAACCAGGCTCTATGAAGCA 701 GGAATCCTTTATAAGCGGATATCTAAACATTTGGTCACATTTAAAGGTCAGCCTTATAATTGGGAACAGC 771 AACACCTTGTCAATCAACATCACATTTATGATGGGGCAACATCCAGCAAAATCAATGGACGTCAGACGGA 841 TAGAAGGAGGAGAAATACTGTTAAACCAACTTGCCGGAAGGATGATCCCAAAAGGGACTTTGACATGGTC 911 AGGCAAGTTTCCAACACTAGATCACGTGTTAGACCATGTGCAAACAATGGAGGAGATAAACACCCTCCAG 981 AATCAGGGAGCTTGGCCTGCTGGGGCGGGAAGGAGAGTAGGATTATCAAATCCGACTCCTCAAGAGATTC 1051
1121
1191
1261
1331 1401 1471
CTCAGCCCCAGTGGACTCCCGAGGAAGACCAAAAAGCACGCGAAGCTTTTCGCCGTTATCAAGAAGAAAG
ACCACCGGAAACCACCACCATTCCTCCGTCTTCCCCTCCTCAGTGGAAGCTACAACCCGGGGACGATCCA CTCCTGGGAAATCAGTCTCTCCTCGAGACTCATCCGCTATACCAGTCAGAACCAGCGGTGCCAGTGATAA AAACTCCCCCCTTGAAGAAGAAAATGTCTGGTACCTTCGGGGGAATACTAGCTGGCCTAATCGGATTACT GGTAAGCTTTTTCTTGTTGATAAAAATTCTAGAAATACTGAGGAGGCTAGATTGGTGGTGGATTTCTCTC AGTTCTCCAAAGGGAAAAATGCAATGCGCTTTCCAAGATACTGGAGCCCAAATCTCTCCACATTACGTAG GATCTTGCCCGTGGGGATGCCCAGGATTTCTTTGGACCTATCTCAGGCTTTTTATCATCTTCCTCTTAAT
FIG. 2 9
J. VIROL.
on November 10, 2019 by guest
http://jvi.asm.org/
[image:3.612.155.468.246.730.2]DHBV GENOME NUCLEOTIDE SEQUENCE 785
between the DHBV 6 protein and RSVRT, the HBV 6 proteins and RSVRT were identical. Finally, if one takes into account that severaloftheobserved differences corre-spond to amino acids of the same family, involving a so-calledconservativechange (6), then one can suggest that the homologies observed between the hepatitis gene 6 proteins and the avian reverse transcriptase did not occur only by chance. However, more sophisticated computer work is
needed to establish more definitely what kind ofstructural
relationship exists between region 6 proteins and reverse transcriptases and whether this is due to evolution from a common ancestor or to convergence to fulfill a similar
function.
(ii) Region 7. This open reading frame went from nucleo-tide 684 in frame 3 to stop codon TAG 1785. It has been
1541 1611 1681
17 51 1821 1891 1901 2031 2101 2171 2241 2311
shown that the homologous HBV open reading frame codes for the surface protein, and according to DNA sequence homology and amino acid sequence homology, an identical result has been inferred for WHV. As we have already pointed out, significant DNA homologybetween DHBV and the other twoviruses was found in this region. This suggests that the open region 7 codes for the duck viral surface (DHBs) protein. Due to the presence of numerous ATGs in thereading frame, translation could start at several different places. From the first encountered ATG at position 693, a protein of 364 amino acids can be predicted. By using different sources ofdata, such as the molecular weight and theknown amino acid sequence of the N-terminal portion of the human viral surface(HBs) protein, it has been deduced that theN-methionine of the HBsprotein is not coded by the
CCTGCTAGTAGCAGCAGGCTTGCTGTATCTGACGGACAACGGGTCTACTATTTTAGGAAAGCTCCAATGG GCGTCGGTCTCAGCCCTTTTCTCCTCCATCTCTTCACTACTGCCCTCGGATCCGAAATCTCTCGTCGCTT TAACGTTTGGACTTTCACTTATATGGATGACTTCCTCCTCTGCCACCCAAACGCTCGTCACCTTAACGCA
End of region 7 ;
ATTAGCCACGCTGTCTGCTCTTTTTTACAAGAGTTAGGAATAAGAATAAACTTTG ACAAAACCACGCCTT CTCCGGTGAATGAAATAAGATTCCTCGGTTACCAGATTGATGAAAATTTCATGAAGATTG AAGAAAGCAG ATGGAAAGAATTAAGGACTGTAATCAAGAAAATAAAAGTAGGAGAATGGTATGACTGGAAATGTATTCAA AGATTTGTCGGGCATTTGAATTTTGTTTTGCCTTTTACTAAAGGTAATATTGAAATGTTAAAACCAATGT ATGCTGCTATTACTAACCAAGTAAACTTTAGCTTCTCTTCATCCTATAGGACTTTGTTATATAAACTAAC AATGGGTGTGTGTAAATTAAGAATAAAGCCAAAGTCCTCTGTACCTTTGCCACGTGTTAGCTACAGATGCT ACCCCAACACATGGCGCAATATCCCATATCACCGGCGGGAGCGCAGTGTTTGCTTTTTCAAAGGTCAGAC ATATACATGTTCAGGAACTATTGATGTCTTGTTTAGCCAAGATAATGATTAAACCACGTTGTCTCTTATC TGATTCAACTTTTGTTTGCCATAAGCGTTATCAGACGTTACCATGGCATTTTGCTATGTTGGCCAAACAA 2381 TTGCTCAAACCGATACAATTGTACTTTGTCCCGAGCAAATATAATCCTGCTGACGGCCCATCCAGGCACA
Start of the fused >s regions in frame 1 2451 AACCTCCTGATTGGACGGCTTTTCCATACACCCCTCTCTCGAAAGCAATATATATTCCACATAGGCTATG
End of region 6
2521 TGGAACTTAAGAATTACACCCCTCTCCTTCGGAGCTGCTTGCCAAGGTATCTTTACGTCTACATTGCTGT 2591 TGTCGTGTGTGACTGTACCTTTGGTATGTACCATTGTTTATCATTCTTGCTTATATATGGATATCAATGC 2661 TTCTAGAGCCTTAGCCAATGTGTATGATCTACCAGATGATTTCTTTCCAAAAATAGATGATCTTGTTAGA
2731 GATGCTAAAGACGCTTTAGAGCCTTATTGGAAATCAGATTCAATAAAGAAACATGTTTTGATTGCAACTC
2S01 ACTTTGTGGATCTCATTGAAGACTTCTGCCAGACTACACAGGGCATGCATGAAATAGCCGAATCATTAAG 2871 AGCTGTTATACCTCCCACTACTACTCCTGTTCCACCGGGTTATCTTATTCAGCACGAGGAAGCTGAAGAG
2941 ATACCTTTGGGAGATTTATTTAAACACCAAGAAGAAAGGATAGTAAGTTTCCAACCCGACTATCCGATTA
3011 CGGCTAGAATT
FIG. 2. Nucleotide sequence of the Eco DHBV DNA clone. The sequence shown is complementary to the viral L strand. VOL.49, 1984
on November 10, 2019 by guest
http://jvi.asm.org/
786 MANDART, KAY, AND GALIBERT
Chain S
2
3
2
5/8 *, 111, 5/8 6
toIon III 7 iCOtl
ChainL
100 o
1 1 "' 1 1 1 III i,
11iI 11111 dPlI3I3om 1..11- 1111111 1131Is
3 fJ'-3ap*fdkl3l mlII A 1 11 II 11 I II fill
FIG. 3. Diagram showing the localization of the nonsense co-donsonchainsSand L.Threereading framesweredefinedfrom the 5'end ofeachDNAstrand.Onchain S,frame1isdefinedby itsfirst tripletCAT, frame2isidentifiedbyATG,andframe 3 isidentified by TGC. On chain L, frame 1 is defined by its firsttriplet AAT, frame2isidentifiedbyATT, andframe3 isidentified by TTC. The viral DNA is circular, and its length in nucleotides (3,021) is a multiple of3; therefore, passing through the EcoRI site does not changethereadingframe. Upperverticalbarsindicate stop codons. Lowerverticalbarsrepresent ATG triplets. Numbers5/8,6,and 7 defineareas in which a viral gene has been located: region 5/8 goes from2515 to 411; region 6 goes from 14 to 2527; region 7 goes from 684 to 1784.
first in-phase ATG codon but by the third one (10). An
identical deduction has been made for the woodchuck viral surface (WHs) protein (9). According to these results,
trans-lationof DHBsprotein maystart with the secondor
follow-ing in-phaseATG.
A comparison ofthe amino acid sequence deduced from the DNA sequenceofopenreading frame7ofHBV, WHV, and DHBV showed no homology for the so-called pre-S region. On the other hand, a significant homology existed between the DHBV amino acid sequencestartingwith ATG 1284 and the N-terminal amino acid sequences of HBs and WHsproteins,suggesting,ifnotproving,that theN-terminal amino acid sequence for the DHBsprotein starts with ATG 1284, the seventhin-phaseATG within openreadingframe 7 (Fig. 8). From ATG 1284 up toTAG 1785,a protein of 167 amino acids with a molecular weight of 18,204 can be
predicted, ascomparedwith 25,645and25,422forthe WHs and HBs proteins,
respectively
(3, 8, 35). Electrophoretic results, showingan apparent molecularweightof17,000forDHBs protein (Mason etal., personal communication), are ingood agreement withthe theoretical molecular weight.
A comparison ofthe amino acid sequences of the three surface proteins showed that afragment of about 50 amino acids, corresponding roughly to position 105 to 155 ofthe HBs protein sequence, was absent from the DHBs protein (Fig. 9). This deletion was also seen in the gene 6 protein (Fig. 9) andwasvisible atthe DNA level (Fig. 5). The most intriguing point about thisdeletion is that the main antigenic epitope of the HBs protein was tentatively located in that region byproteolytic digestion, amino acid comparison,and peptide synthesis (7, 8, 22, 23).
Several characteristic features of the HBs and WHs proteins have been previously noted, such as the existence of a very large hydrophobic sequence and of a
sequence,
Asn-X-Thr/Ser, knowntobe involved inglycosylation (30). The same is true for the DHBs protein (Fig. 9). There is a hydrophobic regionlimitedbyamino acids 79 and 97 with a hydrophobic index equalto3.31 (26). A potential glycosyla-tion site is also located at position 99. The homology observed between the DHBsproteinand the two others was much lower(35%)than between the HBs and WHsproteins (61%), and thecarboxylic regionsshowedverylittle
homolo-
gy-It isinterestingto notethat theDHBsproteinwaspartially encoded bythe sequence locatedbetween nucleotides 1300 to 1400 in which the largest sequencehomology (70%) was observed with the other two genomes. This sequence also coded for the gene 6 protein. This high percentage of homology was mainly due to the gene 6 product, which is moreconserved (Fig. 10).
Although we cannotformally prove thatATG 1284codes for the N-terminus of the DHBsprotein,it ishighly suggest-ed by the amino acid sequence homology observed among the three N-terminal sequences. However, the existence in all three viruses of alarge open readingframepreceding the putative N-termini of their surface proteins is most intrigu-ing. Various experiments have located the TATA boxand thecap site of the HBs mRNA 150 nucleotides ahead of the first ATG of region 7, suggesting that transcription of the pre-S region does occur (21). Recent experiments by Si mapping ofHBsprotein mRNA madein transfected mouse cells have detected a protected DNA fragment starting at position 3160 of the HBV genome (29a). Because of the locationatnucleotide2776 of a TATAbox,a moreprobable hypothesis is that HBs mRNA is spliced and that position 3160correspondstothe 5' end ofitsmainbody.Aconsensus
6
[image:5.612.84.277.68.208.2]5 5
FIG. 4. Localizationof the openreadingframesonthe viralgenomes of HBV and WHV andcomparison with the DHBVgenome. The stripedareainregion7correspondstothepre-Ssequence. Arrows indicate thepositionof the first ATG found withinanopenreadingframe. Numbers 1to8refertothe variousopenreadingframesas defined in thetextandin Galibert etal.(9, 10).
J. VIROL.
l
on November 10, 2019 by guest
http://jvi.asm.org/
[image:5.612.106.507.569.694.2]DHBV GENOME NUCLEOTIDE SEQUENCE 787
w
8-//
/
i ~~~~H
170 970
a080
t890V09 1890
D
81
/
//
/
§/ H
170 970
/
8
10910 t890
/
/ 1890
waslostduring cloning through the EcoRI site by sequencing thisregion on anindependent XhoI DHBV clone.Therefore, the fused 5/8 region represents the actual structure of the DHBV genome. Because the nucleotide sequencehomology was too poorin this region, comparisonsdid not allowus to infer which part of the sequence was lost (or acquired) during evolution.
Acomparison of the amino acid sequence of DHBV region
5/8 withthecorresponding amino acidsequencesofregions5 and 8 of HBV and WHV showed a net degree ofhomology between thecarboxylic end of the molecule coded by regions 8 and 5/8. The same peculiar structure,involving repetition and increased amounts of basic aminoacids, was observed at thecarboxylicend,giving, upon comparison, a character-isticpicture (Fig. 12). This clearly suggests that this amino acid sequencecorrespondstothecoreprotein. The molecu-lar weight of the DHBV core protein has been estimated by gel electrophoresis to be 35,000 (W. S. Mason and J. Newbolt, personal communication). Thiswouldbe in good agreementwithaprotein made with the entire open reading
frame from ATG 2518 to TAA 412 for which a molecular weight of 34,986 can be predicted. However, some topologi-calproblem may prevent theuseof that ATG because of the
position ofthe nickon the minus strand, which was down-stream from ATG 2518. Iftranscription occurs on anicked genome, then transcription probably starts after the nick,
4
/
2620 300 2620
2/
300FIG. 5. Nucleotide
sequence
comparisons. Sequences40 nucle-otideslong withahomology equalorsuperiorto50%werescored, andtheyareindicated by lines whosecoordinatescorrespond
tothe position ofthat sequencewithin thetwocomparedgenomes. Left-hand row is acomparison ofWHV (W) and HBV (H) genomes; right-hand row is a comparison of DHBV (D) and HBV (H) genomes.acceptor sequence at position 3171 in the
14BV
genomesupports this hypothesis. Donor and acceptor consensus sequences (20) are also found in the WHV and DHBV
sequences (Fig. 11).
The existence of splice and acceptor sequences at the suggested positionsraises thequestionastothechoice of the initiator ATG. In all three cases, the one defining the
N-terminal methionine is not the first encountered ATG, and neither is the first one with a purine at position -3 (16). Anotherquestionraisedbyaneventualspliceof HBsprotein mRNAis relativetothe conservationduringevolution ofan
open reading frame within the pre-S region. A working hypothesisis thatanotherproteincodedbythepre-S region alone orby thetotalityofregion7 is expressed.
(iii) Region5/8.Startingwithnucleotide2515,therewas an
open readingframe which continuedthrough the EcoRI site
uptostopcodon TAA 412. Because of its relative position, overlappingthe 3' and 5' ends ofregion 6 with its 5' and 3' ends, respectively, thisregionlooks likeaproductof fusion
of theformerlydefined 5 and 8regionsof HBV and WHV(9, 10).We eliminated thepossibilitythatasmall DNAfragment
H6
/
/
/
[image:6.612.86.275.71.426.2]W6
FIG. 6. Comparison of the amino acid sequences of gene 6
proteins of HBV (H), DHBV (D), and WHV (W). Stretches 30
amino acids long with homology equal or superior to 20% are
indicated byaline.
H6
//
//
/
/
/
VOL. 49,1984
on November 10, 2019 by guest
http://jvi.asm.org/
[image:6.612.340.513.352.686.2]D6 SerThrProGlyLysSerValSer o ArgAspSerSerAlaIleProValArgThrSerGlyAlaSerAspLysAsn RT ThrValAlaLeuHisLeuAlaIleProLeuLysTrpLysProAspHisThrProValTrpIleAspGlnTrpProLeu H6 CysTrpTrp o GlnPheArgAsnSerLysProCysSerAspTyrCysLeuSerLeuIleValAsnLeuLeuGluAsp SerProLeuGluGluGluAsnValTrpTyr o ArgGlyAsnThrSerTrpProAsnArg o ThrGlyLys o PheLeu
ProGlLysLeuValAlau6uThrGlnLeualGluLysGluLeuGlnLeuGlyHisIleGluProProLeuSers ys
TrpGlyProCysAlaGluHisGlyGluHisHislleArgIleProArgThrProSerArgValThrGlyGIyValPheLeu
ValAspLysAsnSerArgAsnThrGluGlu o ArgLeuValValAspPheSerGlnPheSerLysGlyLys o o Met TrpAsnThrProPhePheValIleArgLysAlaSerGlySerTyrArgLeuLeuHisAspLeuArgAlaValAsnAlaLys ValAspLysAsnProHisAsnThrAlaGluSerArgLeuValValAspPheSerGlnPheSer o GlyAsnTyrArgVal ArgPhe o ArgTyrTrpSerProAsnLeuSerThrLeuArgArgIle o o Val o Met o ArgIleSer o o
LeuValProPheGlyAlaValGlnGlnGlyAlaProValLeuSerAlaLeuProArgGlyTrpProLeuMetValLeuAsp SerTrp o LysPhe o o ProAsnLeuGlnSerLeuThrAsnLeu o SerSerAsnLeuSerTrpLeuSer o o
0 SerGlnAla o TyrHisLeu o o AsnProAlaSerSerSerArgLeu o ValSerAspGlyGlnArgValTyr LeuLysAspCysPhePheSerIleProLeuAlaGluGlnAspArgGluAlaPheAlaPheThrLeuProSerValAsnAsn ValSerAlaAla o TyrHisLeu o o HisProAlaAlaMetProHisLeuLeuValGlySerSerGlyLeuSerArg Tyr PheArgLysAla o Met o ValGlyLeu o o PheLeu-LeuHisLeuPhe GlnAlaProAlaArgArgPheGlnTrpLysValLeuProGlnGlyMetThrCysSerProThrlleCysGInLeuValVal Tyr// 54 aa //PheArgLysIle o Met o ValGlyLeu o o PheLeu-LeuAlaGlnPhe ThrThrAla o GlySerGluIleSerArgArgPheAsnValTrp-ThrPheThr o o o o Phe o o Cys
GlyGlnValLeuGluProLeuArgLeuLysHisProSerLeuCys-MetLeuHisTyrMetAspAspLeuLeuLeuAla ThrSerAlaIleCysSerValValArgArgAlaPheProHis o LeuAlaPheSer o o o o ValVal o Gly HisProAsnAlaArgHis o Asn o IleSerHisAla o Cys o Phe o GlnGluLeu o IleArg o AsnPhe AlaSerSerHisAspGlyLeuGluAlaAlaGlyGluGluValIleSerThrLeuGluArgAlaGlyPheThrIleSerPro
o Lys o ValGlnHis o o SerLeuPheThrAla o ThrAsnPhe o LeuSerLeu o IleHisLeuAsn o
o o ThrThrProSer o ValAsnGluIleArgPhe o o o GlnIleAspGluAsnPheMetLysIleGluGlu AspLysValGlnArgGluProGlyValGlnTyr LeuGlyTyrLysLeuGlySerThrTyrValAlaProValGly Asn o ThrLys o TrpGlyTyrSerLeuAsn-PheMet o o ValIle o CysTyrGlySerLeu o GlnGlu SerArgTrpLys o Leu o ThrVal1leLysLysIleLysValGlyGluTrpTyrAspTrpLysCysIleGlnArgPhe LeuValAla- GluProArgIleAlaThrLeuTrpAspValGlnLysLeuValGlySerLeuGlnTrpLeuArgProAla HisIltIleGlnLyslleLysGluCysPheArgLysLeuProIleAsnArgProIleAspTrpLysValCysGlnArgIle Val o HisLeuAsnPheValLeuProPheThrLysGlyAsnIleGluMetLeuLys o MetTyr o AlaIleThr o
LeuGlylleProProArgLeuMetGlyProPheTyrGluGlnLeuArgGlySerAspProAsnGluAlaArgGluTrpAsn Val o LeuLeuGlyPheAlaAlaProPheThrGlnCysGlyTyrProAlaLeuMet o LeuTyr o CysIleGlnSer GlnValAsnPheSerPheSerSerSerTyrArgThr o LeuTyrLysLeuThrMetGlyValCysLysLeuArgIleLys LeuAspMetLysMetAlaTrpArgGluIleValArgLeuSerThr ThrAlaAlaLeu LysGlnAlaPheThrPheSerProThrTyrLysAlaPheLeuCysLysGlnTyrLeuAsnLeuTyrProVal o ArgGln
ProLysSerSerValPro o0 ArgValAlaThrAsp o ThrProThrHis o o o Ser -HisIleThr- GluArgTrpAspProAlaLeuProLeuGluGlyAlaValAlaArgCysGluGlnGlyAlaIleGly-ValLeuGly-ArgProGlyLeuCysGlnValPheAlaAspAlaThrPro-ThrGlyTrp-- o LeuValMetGlyHisGlnArgMet Gly o SerAlaValPheAlaPhe o LysVal o AspIleHisVal o GluLeuLeuMetSerCysLeuAlaLysIle GInGlyLeuPheThrHisProArgSerCysLeuArgLeuPheSerThrGlnProThrLysAlaPheThrAlaTrpLeuGlu
Arg o Thr Phe o AlaProLeuProlleHis o AlaGluLeuLeu o AlaCysPheAlaArgSer MetileLysProArgCysLeuLeuSerAspSerThrPhe o CysHisLysArgTyrGlnThrLeuProTrpHisPheAla ValLeuThrLeuLeuIleThrLysLeuArgAlaSerAlaValArgThrPheGlyLysGluValAspIleLeuLeuLeuPro ArgSerGlyAlaAsn o IleGlyThrAspAsn o Val o LeuSerArgLysTyrThrSerPheProTrp o o Gly MetLeuAlaLysGlnLeu o LysProIleGlnLeuTyrPheValProSer o TyrAsnProAlaAspGlyPro o Arg
AlaCysPheArgGluAspLeuProLeuProGluGlyIleLeuLeuAlaLeuLysGlyPheAlaGlyLysIleArgSerSer
CysAlaAlaAsnTrpIle o ArgGlyThrSerPheValTyrValProSerAlaLeuAsnProAlaAspAspPro o Arg HisLys o ProAspTrpThrAlaPheProTyrThrProLeu o
LysAlaIleTyrIleProHisArgLeuCysGlyThr-AspThrProSerIlePheAspIleAlaArgProLeuHisValSerLeuLysValArgValThrAspHisProValProGly GlyArgLeuGlyLeuSerArgProLeuLeuArg o ProPheArgProThrThrGlyArg o SerLeuTyrAlaAspSer Stop
ProThrValPheThrAspAlaSerSerSerThrHisLysGlyValValValTrpArgGluGlyProArgTrpGluIle..RT
o Ser o ProSerHisLeuProAspArgVal o PheAlaSerProLeuHisValAlaTrpArgProPro Stop H6
FIG. 7. Amino acid sequencecomparison between RSVRT (RT, middle line), DHBV gene 6protein (D6, top line), and HBV gene 6 protein(H6, bottom line).RSVRT is takenasreference.Amino acids in either D6orH6 thatareidenticaltothecorresponding amino acid in RTarerepresentedbyanopencircle. Amino acids whichare notidentical butbelongtothesamefamilyareshaded.The amino acid sequence ofD6startswith the390th aminoacidafterthe first ATG of thereading frame,and H6startswithaminoacid 312. RTstartswith thefirst
ami-noacidofthemature protein.
788
on November 10, 2019 by guest
http://jvi.asm.org/
DHBV GENOME NUCLEOTIDE SEQUENCE 789
TABLE 1. Sequence comparisonshowinganonapeptide ofverysimilarsequencethat isfound within the threehepatitisBgene6 proteins and thetworeversetranscriptases
Peptide
Protein position Aminoacid
within sequence
RSVRT 180 Tyr Met Asp Asp Leu Leu Leu Ala Ala
MLVRT 342 a Valb _
WHs 583 Valb Glyb
DHBs 561 - Phe Cys His
HBs 538 - Valb Valb - Glyb
-a
_,
Aminoacids identicaltothose shownfor RSVRT.bAmino acidsof the same group asthe amino acids shown for RSVRT.
and theinitiationcodonforthe5/8 gene product will then be ATG 2647, which is notfavored by the presence ofaT at position 2644 (16). On the other hand, the transcriptional template may not have a nick, allowing translation to start with ATG 2518.
We found nohomology between the N-terminal end ofthe DHBV5/8 region andregion5ofHBVorWHV.Therefore,
although we believe that the third open reading frame in DHBVdoes represent afusionof the 5 and 8regions,we can drawnofirmconclusions. However, it shouldbenoted that there was far less homology between HBV and WHV with region 5 than there was with the rest of the genome.
Atthe presenttime, noproteincodedbyregion5of HBV or WHV has been identified, and no function has been
clearlyproposed.However,theexistence ofageneencoded inthisregionhasbeensupported bycomparative analysisof
amino acid and nucleotide sequences (9). Therefore, the questionarises whether thefunction oftheprotein codedby HBV and WHVregion5has thesamefunctionastheprotein
codedby region 5/8 ofthe DHBV genome. In otherwords,
does the protein made by region 5/8 have two roles, one devotedtothecoreproteinandonedevotedtogeneprotein 5, or is region 5/8, either through RNA splicing or protein
processing,making twodifferent proteins?
Replicationorigin. The genomeof thehepatitisBvirus isa noncovalentcircle (15). Theinterruptionin the L strand has beenassignedin HBV and WHVtotheonlyregiondevoid of
codingcapacity,between theopenreadingframes 5 and 8in the vicinity of a hairpin structure (9). A similar hairpin structure, although different in sequence, was observed in the DHBVgenome,startingwith nucleotide 2504 andending withnucleotide 2525. This hairpinwaslocalized at the end of region 6 and overlappedthe beginningofregion 5/8. It was surrounded by two direct repeats (ACACCCCTCTC) at position 2478 and 2536, which are reminiscent of the short direct repeats foundatboth ends of retroviruses.
Nucleotidesequencecomparisonsof theHBV, WHV,and DHBV genomes atthe molecular level clearly showed that these viruses belong to the same family and were derived from a common ancestor. Whereas HBV and WHV share between 60 to 70% nucleotide sequence homology (9), DHBVshowed much lesshomology (around or below40%) foralargepartof the genome andshowedbetween 50 to 55% homology for the remaining part, except for a small se-quenceof 100 nucleotides which reached 70%. This indicates thatduringevolution the threevirusesdid not separate from each other at the same time, but that DHBV separated earlierfrom the ancestorof the twoothers. This is probably inconjunction with the fact that DHBV infects birds and that birds started toevolveseparately from mammals 250 million years ago.Inturn, this also couldindicate that, at least inthe caseofthe hepatitisvirusfamilybut possibly for all kinds of
viruses, viruses appeared very early during evolution and evolved in a fashion parallel to their target. Unfortunately,
SerLeuLeuGluThrHisProLeuTyrGlnSerGluProAlaValProValIleLysThrPro
DHBs 1206 TCTCTCCTCGAGACTCATCCGCTATACCAGTCAGAACCAGCGGTGCCAGTGATAAAAACTCCC HBs 79 GGAACAGTAAACCCTGTTCTGACTACTGCCTCTCCCTTATCGTCAATCTTCTCGAGGATTGGG
GlyThrValAsnProValLeuThrThrAla o ProLeuSerSerlIePheSerArgIleGly
ProLeuLysLysLysMMetSerGlyThrPheGlyGlylleLeuAlaGlyLeulieGlyLeuLeu
DHBs 1269 CCCTTGAAGAAGAAAATGTCTGGTACCTTCCGGGGAATACTAGCTGGCCTAATCGGATTACTG HBs 142 GACCCTGCGCTGAACATGGAGAACATCACATCAGGATTCCTAGGACCCCTTCTCGTGTTACAG AspProAlaLeuAsn o GluAsnIleThrSer o Phe o GlyPro o LeuVal o Gln
ValSerPhePheLeuLeuIleLyslIeLeuGluIleLeuArgArgLeuAspTrpTrpTrplle
DHBs 1332 GTAAGCTTTTTCTTGTTGATAAAAATTCTAGAAATACTGAGGAGGCTAGATTGGTGGTGGATT HBs 205 GCGGGGTTTTTCTTGTTGACAAGAATCCTCACAATACCGCAGAGTCTAGACTCGTGGTGGACT AlaGly o o o o ThrArg o o Thr o ProGinSer o o Ser o o Thr
SerLeuSerSerProLysGlyLysMetGlnCysAlaPheGlnAspThrGlyAlaGInlleSer
DHBs 1395 TCTCTCAGTTCTCCAAAGGGAAAAATGCAATGCGCTTTCCAAGATACTGGAGCCCAAATCTCT
HBs 268 TCTCTCAATTTTCTAGGGCGAACTACCGTGTGTCTTGGCCAAAATTCGCAGTCCCCAACC
o o AsnPheLeuGly o ThrThrVal o LeuGly o AsnSerGlnSerProThr
FIG. 8. Nucleotide and amino acid sequence comparison around ATG 1285 for DHBs protein and ATG 157 for HBs protein, which
probablycodefor the N-terminal methionine of the surfaceantigen. Althoughthereisnohomology upstream from theseATGs,numerous
identicalamino acids downstream from these positions can be observed. Shading indicates ATG 1285 and ATG 157. VOL. 49,1984
on November 10, 2019 by guest
http://jvi.asm.org/
DHBs 1 MetSerGlyThrPheGlyGlyIleLeuAlaGlyLeuIleGlyLeuLeuValSerPhePheLeuLeuIleLys
HBs 1 o GluAsnIleThrSer o Phe o GlyPro o LeuVal o GlnAlaGly o o o o ThrArg DHBs 25 IleLeuGluIleLeuArgArgLeuAspTrpTrpTrpIleSerLeuSerSerProLysGlyLysMetGlnCys
HBs 25 o o Thr o ProGlnSer o o Ser o o Thr o o AsnPheLeuGly o ThrThrVal o
DHBs 49 AlaPheGlnAspThrGlyAlaGlnIleSerProHisTyrValGlySerCysProTrpGlyCysProGlyPhe HBs 49 LeuGly o AsnSerGlnSerProThrSerAsn o SerProThr o o o ProThr o o o Tyr
6 //AspLeuSerGlnAlaPheTyrHisLeuProLeuAsnProAlaSerSerSerArgLeuAlaValSer
s 73 LeuTrpThrTyrLeuArgLeupheIleIlePheLeuLeuIleLeuLeuValAlaAlaGlyLeuLeuTyrLeu
DHBV 1506 /7 ACCTATCTCAGGCTTTTTATCATCTTCCTCTTAATCCTGCTAGTAGCAGCAGGCTTGCTGTATCTG
***.... ... 5*...X^... . ...
..-HBV 379 7/ ATGTGTCTGCGGCGTTTTATCATCTTCCTCTTCATCCTGCTGCTATGCCTCATCTTCTTGTTGGTT s 73 Arg o MetCys o o Arg.o o 0 oo Phe o o o LeuCysLeuIlePhe o LeuVal
6 // o Val o Ala o o o o o o o His o o AlaMetProHis o Leu o Gly 6 AspGlyGlnArgValTyr
s 97 ThrAspAsnGlySerThr
DHBV ACGGACAACGGGTCTACTA
HBV CTTCTGGACTATCAAGGTATGTTGCCCGTTTGTCCTCTAATTCCAGGATCCTCAACAACCAGCACGGGACCA
s 97 LeuLeuAspTyrGlnGlyMetLeuProValCysProLeuIleProGlySerSerThrThrSerThrGlyPro 6 SerSerGlyLeuSerArgTyrValAlaArgLeuSerSerAsnSerArgIleLeuAsnAsnGlnHisGlyThr HBV TGCCGGACCTGCATGACTACTGCTCAAGGAACCTCTATGTATCCCTCCTGTTGCTGTACCAAACCTTCGGAC
s 97 CysArgThrCysMetThrThrAlaGlnGlyThrSerMetTyrProSerCysCysCysThrLysProSerAsp 6 MetProAspLeuHisAspTyrCysSerArgAsnLeuTyrValSerLeuLeuLeuLeuTyrGlnThrPheGly
6 TyrPheArgLysAlaProMetGlyValGlyLeuSer
s 103 IleLeuGlyLysLeuGlnTrpAlaSerValSerAla
DHBV TTTTAGGAAAGCTCCAATGGGCGTCGGTCTCAGCC
HBV GGAAATTGCACCTGTATTCCCATCCCATCATCCTGGGCTTTCGGAAAATTCCTATGGGAGTGGGCCTCAGCC
s 121 GlyAsnCysThrCysIleProIleProSerSerTrpAlaPhe o o PheLeu o GluTrpAla o o
6 ArgLysLeuHisLeuTyrSerHisProIleIleLeuGly o o o Ile o o o o o o o
6 ProPheLeuLeuHisLeuPheThrThrAlaLeuGlySerGluIleSerArgArgPhe-AsnValTrpThr
s 115 LeuPheSerSerIleSerSerLeuLeuProSerAspProLysSerLeuValAlaLeu-ThrPheGlyLeu
DHBV CTTTTCTCCTCCATCTCTTCACTACTGCCCTCGGATCCGAAATCTCTCGTCGCTTT-AACGTTTGGACTT
HBV CGTTTCTCCTGGCTCAGTTTACTAGTGCCATTTGTTCAGTGGTTCGTAGGGCTTTCCCCCACTGTTTGGCTT s 145 Arg o o TrpLeu o Leu o Val o PheValGlnTrpLeuValGlyLeuSerPro o ValTrp o
6 o o o o AlaGin o o Ser 0 IleCys o ValValArg o Ala o ProHisCysLeuAla 6 PheThr//
s 138 SerLeuIleTrpMetThrSerSerSerAlaThrGlnThrLeuValThrLeuThrGln-LeuAlaThrLeu
DHBV TCACTTATATGGATGACTTCCTCCTCTGCCACCCAAACGCTCGTCACCTTAACGCA-ATTAGCCACGCTG
... ...i... ... ... . . . ... ..
HBV TCAGTTATATGGATGATGTGGTATTGGGGOCCAAGTCTGTACAGCATCTTGAGTCCCTTTTTACCGCTGTTA
s i69 o ValIleTrpMetMetTrpTyrTrpGlyProSerLeuTyrSerIleLeuSerProPheLeuProLeuLeu
6 0 Ser//
s 161
SerAlaLeuPheTyrLysSer-DHBV TCTGCTCTTTTTTACAAGAGTTAG// HBV CCAATTTTCTTTTGTCTTTGGGTATACATTTAA//
s 193
ProIlePhePheCysLeuTrpValTyrIle-FIG. 9. Amino acid and nucleotide sequencecomparisons ofgene7. The positionof a deletion within the DHBVgenomeaffectingthe
surfaceantigenandgene 6proteinhasbeen determinedbycomparingtheDHBsand HBsproteinsequences, the DHBV and HBVprotein6 sequences, and the DNA sequences. Toclarifythefigure,the gene 6protein sequences and the DNA sequencesareonlyshown inpart, delimitedby11.TheDHBs and HBsproteinsequencesareshownin full. Dotsindicatehomologybetween the DNAsequences. An open circle in theHBVproteinsequencesindicatesthattheaminoacidatthatpositionisthesame asthecorrespondingamino acid in thehomologous DHBVprotein.
s GlyPheLeuGlyProLeuLeuValLeuGinAlaGlyPhePheLeuLeuThrArgIleLeu
6 ArgIleProArgThrProSerArgValThrGlyGLyValPheLeuValAspLysAsnPro
HBV 174 AGGATTCCTAGGACCCCTTCTCGTGTTACAGGCGGGGTTTTTCTTGTTGACAAGAATCCT
DHBV 1300GGGAATACTAGCTGGCCTAATCGGATTACTGGTAAGCTTTTTCTTGTTGATAAAAATTCT 6 GlyAsnThrSerTrp o Asn o Ile o o LysLeu o o o o o 0 Ser
s o Ile o AlaGly o IleGly o LeuValSer o o o o IleLys o o
s ThrIleProGlnSerLeuAspSerTrpTrpThrSerLeuAsnPheLeuGlyGlyThr
6 HisAsnThrSerGluSerArgLeuValValAspPheSerGlnPheSerArgGlyAsnTyr
HBV 258 CACAATACCGCAGAGTCTAGACTCGTGGTGGACTTCTCTCAATTTTCTAGGGGGAACTA
DHBV 1384 AGAAATACTGAGGAGGCTAGATTGGTGGTGGATTTCTCTCAGTTCTCCAAAGGGAAAAA 6 Arg o o Glu o Ala o o o o o o o o o Lys o LysAsn s Glu o LeuArgArg o o Trp o o Ile o o SerSerProLy3 o Lys
FIG. 10. Amino acid and nucleotide sequencecomparison of thehighlyconserved sequence locatedinDHBV between nucleotide1300
and1400. As canbeseen, the gene 6proteinsare moreconservedin thisregionthan the surfaceantigens. Out of 40 aminoacids,24are
identi-calfor protein 6, whereasonly 18outof39areidentical in the surfaceantigen. Symbols arethe same asthose in Fig. 9. 790
on November 10, 2019 by guest
http://jvi.asm.org/
[image:9.612.126.482.35.526.2]DHBV GENOME NUCLEOTIDE SEQUENCE 791
Cap Donor
II
HBV I TATATAA2776
Acceptor N-methionine
ACTCATCCTCAG=AUCAC....
... AAC JAT3170/3171 157
Cap Donor WHY several TA
A
TAAAGIGTAAC ....rich region2949/2s95
Acceptor CTTTTCATCTCCAG ....
141/142
ACAAT1 CCTATGOAC....GA\ ATU TCA
193 19,' 2(0
Cap Donor DHBV several TA I CCAGCCTAGG...
Irich regionsl 616/617
Donor consens
Acceptor
TCTTCCCCTCCTCACGGA ...
1163/1164
CA
Athishypothesiscannotbe testedfor themomentsinceHBV, WHV, and DHBV so far represent the only examples of viruses which belong tothe same family but infect widely
different hosts and whose genomes have been entirely se-quenced.
HO
7/w
W8
FIG. 12. Amino acid sequence comparison of thecore antigen.
Stretches of 20amino acidswithhomology equalorabove 20%are
indicatedby lines.
ACKNOWLEDGMENTS
We are very gratefulto B. Masson andJ. Summers for helpful discussions and forthegift of the DNA cloned recombinants.
This workwassupported inpartbyInstitut National de la Santeet
dela Recherche MddicalethroughgrantSC15.
LITERATURE CITED
1. Alestrom, P., G. Akusjarvi, U. Pettersson, and M. Pettersson. 1982.DNAsequenceanalysisof theregion encodingthe termi-nalproteinand thehypotheticalN. Geneproductofadenovirus
type2. J. Biol. Chem. 257:13492-13498.
2. Burrel, C. J.,P.Mackay,P.J. Greenaway,P. H.Hofschneider, andK.Murray.1979.ExpressioninEscherichia coli ofhepatitis B virus DNA cloned in plasmid pBR 322. Nature (London) 279:43-47.
3. Charnay, P.,E.Mandart,A.Hampe,F.Fitoussi,P.Tiollais,and F.Galibert. 1979.Localizationontheviralgenomeand nucleo-tidesequenceofthegenecoding for thetwomajorpolypeptides ofthehepatitisBsurfaceantigen (HBs Ag).NucleicAcids Res. 7:335-346.
4. Charnay, P., C. Pourcel,A.Louise, A.Fritsch,and P.Tiollais. 1979. Cloning in Escherichia coli and physical structure of hepatitisB virionDNA. Proc.Natl. Acad. Sci. U.S.A. 76:2222-2226.
5. Cummings,I.W.,J.K.Browne,W.A.Salser,G. V.Tyler,R. L.
Snyder, J. M.Smolec, andJ. Summers. 1980. Isolation charac-terization andcomparison of recombinant DNAsderived from the humanhepatitis B and woodchuck hepatitis virusgenome.
Proc.Natl. Acad. Sci. U.S.A. 77:1842-1846.
6. Dayhoff, M. O., R. V. Eck, and C. M. Park. 1972. A model of evolutionary change in proteins, p. 89-99. In M. 0. Dayhoff
(ed.), Atlas of protein sequence and structure 1972, vol. 5.
NationalBiomedical Research Foundation, Washington, D.C. 7. Dreesman,G. R., Y.Sandrez, I. Ionescu-Matin, J. T. Sparrow,
H. R.Six, D. L. Peterson, F. B. Hollinger,andJ. L. Melnick. 1982. Antibody to hepatitis B surface antigen after a single
inoculation ofuncoupled synthetic HBs Ag peptides. Nature (London) 295:158-160.
8. Galibert, F., T. N. Chen, and E. Mandart. 1981. Localization and nucleotidesequenceof thegenescoding for the woodchuck hepatitisvirussurfaceantigen: comparison with thegenecoding
AAA ArC TCT 12'4
Acceptor consensus (C)IINTAC/C
FIG. 11. Comparison ofputative controlsequencesfortheexpression ofgene 7. VOL.49, 1984
on November 10, 2019 by guest
http://jvi.asm.org/
[image:10.612.90.532.68.317.2] [image:10.612.72.275.424.693.2]792 MANDART, KAY, AND GALIBERT
for the human hepatitis B virus surface antigen. Proc. Natl. Acad. Sci. U.S.A. 78:5315-5319.
9. Galibert, F., T. N. Chen, and E. Mandart. 1982. Nucleotide sequence of a clonedwoodchuckhepatitisvirus genome: com-parisonwith thehepatitisBvirus sequence. J.Virol. 41:51-65. 10. Galibert, F., E. Mandart, F. Fitoussi, P. Charnay, and F. Galibert. 1979. Nucleotide sequence of the hepatitis B virus genome (subtype ayw) cloned in E. coli. Nature (London) 281:646-650.
11. Gerlich, W. H., M. A. Feitelson, P. L. Marion, and W. S. Robinson. 1980. Structural relationships between the surface antigens of ground squirrel hepatitis virus and humanhepatitisB virus. J. Virol. 36:787-795.
12. Gingeras, T. R., D. Sciaky,R. E.Gelinas,J. Bing-Dong,C. E. Yen, M. M. Kelly, P. A. Bullock, B. L. Parsons, K. E.O'Neill, andR. J. Roberts. 1982.Nucleotidesequencesfromthe adeno-virus-2 genome. J. Biol. Chem. 257:13475-13491.
13. Hartley, J.L., and J. E. Donelson. 1980.Nucleotidesequence of the yeastplasmid. Nature (London)286:860-864.
14. Herisse, J., G. Courtois, and F. Galibert. 1980. Nucleotide sequence of the EcoRI D fragment of adenovirus 2 genome. NucleicAcids Res. 8:2173-2191.
15. Hruska, J. F., D. A. Clayton, J. L. R. Rubenstein, and W. S. Robinson. 1977. Structure of hepatitis B Dane particle DNA beforeandafter theDaneparticleDNApolymerase reaction.J. Virol.21:666-672.
16. Kozak, M. 1981. Possible role offlanking nucleotides in recogni-tionof AUG initiatorcodon byeukaryoticribosomes. Nucleic Acids Res.9:5233-5252.
17. Marion, P. L., L. S. Oshiro, D. C. Regnery, G. H. Scullard, and W.S. Robinson. 1980. AvirusinBeechey ground squirrels that isrelated tohepatitisBvirus on humans. Proc. Natl. Acad. Sci. U.S.A. 77:2941-2945.
18. Mason, W. S., G. Seal, and J. Summers. 1980. Virus of Pekin ducks with structural and biological relatedness to human hepatitisB virus.J. Virol. 36:829-836.
19. Maxam, A., and W. Gilbert. 1980.Sequencing end labeled DNA with base specific chemical cleavage. Methods Enzymol. 65:499-560.
20. Mount, S. M. 1982. A catalogue ofsplice junction sequences. Nucleic Acids Res. 10:459-472.
21. Pourcel,C., A.Louise,M.Gervais,N.Chenciner,M.-F.Dubois, and P. Tiollais. 1982. Transcription of the hepatitis B surface antigengeneinmousecells transformed with clonedviralDNA. J. Virol. 42:100-105.
22. Prince, A. M., H. Ikram, and T. P. Hopp. 1982. Hepatitis B virusvaccine: Identification ofHBsAg/aand HBsAg/dbutnot HBsAg/y subtypeantigenic determinantson asynthetic immu-nogenicpeptide. Proc. Natl. Acad. Sci. U.S.A.79:579-582. 23. Rao,K.R., and G. N.Vyas.1976. Biochemicalcharacterization
ofhepatitisB surfaceantigeninrelationtoserologicalactivity.
J. Biol. Stand. 4:295-304.
24. Robinson, W. S., P. L. Marion, M. A. Feitelson, and A. A. Siddiqui. 1981. the hepadna virus group: hepatitis B andrelated viruses, p. 57-58. In W. Szmuness, H. J. Alter, and J. E. Maynard (ed.), Proceedings of the International Symposium on ViralHepatitis. Franklin Institute Press, Philadelphia. 25. Schwartz, D., R. Tizard, and W. Gilbert. 1983. Nucleotide
sequenceofRous sarcomavirus. Cell32:853-869.
26. Segrest, J. P., and R. J. Feldmann. 1974. Membrane proteins: amino acid sequence and membrane penetration. J. Mol. Biol. 87:853-858.
27. Shinnick, T. M., R. A. Lerner, and J. G. Sutcliffe. 1981. Nucleotide sequence of Moloney murine leukaemia virus. Na-ture(London) 293:543-548.
28. Sninsky, J. J., A. Siddiqui, W. S. Robinson, and S. N. Cohen. 1979.Cloningandendonuclease mapping of the hepatitis Bviral genome.Nature (London) 279:346-348.
29. Staden,R. 1977.Sequence data handling by computer. Nucleic Acids Res.4:4037-4051.
29a.Stenlund, A., D. Lamy, J.Moreno-Lopez, H. Ahola, U. Petters-son, and P. Tiollais. 1983. Secretion of the hepatitis B virus surface antigen from mouse cells using an extra-chromosomal eucaryotic vector. EMBO J. 5:669-673.
30. Struck, D. K., W. J. Lennarz, and K. Brew. 1978. Primary structuralrequirements for the enzymatic formation of the N-glycosidic bond in glycoproteins. J. Biol. Chem.253:5784-5786. 31. Summers,J., and W.S. Mason. 1982.Replication ofthe genome of ahepatitis-B-like virus by reversetranscriptionof an RNA intermediate. Cell29:403-415.
32. Summers, J., A. O'Connell,and 1. Millman. 1975. Genome of hepatitisB virus: restrictionenzymecleavageand structureof DNA extracted from Dane particles. Proc. Natl. Acad. Sci. U.S.A. 72:4797-4801.
33. Summers,J., J. M.Smolec, and R.Snyder. 1978. Avirus similar tohumanhepatitisBvirus associated withhepatitis and hepato-ma in woodchucks. Proc. Natl. Acad. Sci. U.S.A. 75:4533-4537.
34. Summers,J., J. M.Smolec, B. G. Werner, Jr., T.J.Kelly, G. V. Tyler, and R. L.Snyder. 1980.HepatitisBvirusandwoodchuck hepatitis virus are membersofa novel class ofDNA viruses. Viruses in naturally occurring cancers. Cold Spring Harbor Conf. Cell Proliferation 7:459-470.
35. Valenzuala, P., P. Gray, M. Quiroza, J. Zaldivar, H. M. Goodman, and W. J. Rutter. 1979. Nucleotide sequence of the gene coding for the major protein of hepatitis B virus surface antigen. Nature (London) 280:815-819.
36. Valenzuala, P., M. Quiroga, J. Zaldivar, P. Gray, and W. J. Rutter. 1981.The nucleotide sequence of thehepatitis Bviral genome and theidentification ofthe majorviral genes. In B. Fields, R.Jalnisch, and C. F. Fox (ed.), Animal virus genetics. Academic Press, Inc., NewYork.
J. VIROL.