0
settings
الوضع الليلي
moon
انماط الصفحة الرئيسية arrow
EN
1
المرجع الالكتروني للمعلوماتية

النبات

مواضيع عامة في علم النبات

الجذور - السيقان - الأوراق

النباتات الوعائية واللاوعائية

البذور (مغطاة البذور - عاريات البذور)

الطحالب

النباتات الطبية

الحيوان

مواضيع عامة في علم الحيوان

علم التشريح

التنوع الإحيائي

البايلوجيا الخلوية

الأحياء المجهرية

البكتيريا

الفطريات

الطفيليات

الفايروسات

علم الأمراض

الاورام

الامراض الوراثية

الامراض المناعية

الامراض المدارية

اضطرابات الدورة الدموية

مواضيع عامة في علم الامراض

الحشرات

التقانة الإحيائية

مواضيع عامة في التقانة الإحيائية

التقنية الحيوية المكروبية

التقنية الحيوية والميكروبات

الفعاليات الحيوية

وراثة الاحياء المجهرية

تصنيف الاحياء المجهرية

الاحياء المجهرية في الطبيعة

أيض الاجهاد

التقنية الحيوية والبيئة

التقنية الحيوية والطب

التقنية الحيوية والزراعة

التقنية الحيوية والصناعة

التقنية الحيوية والطاقة

البحار والطحالب الصغيرة

عزل البروتين

هندسة الجينات

التقنية الحياتية النانوية

مفاهيم التقنية الحيوية النانوية

التراكيب النانوية والمجاهر المستخدمة في رؤيتها

تصنيع وتخليق المواد النانوية

تطبيقات التقنية النانوية والحيوية النانوية

الرقائق والمتحسسات الحيوية

المصفوفات المجهرية وحاسوب الدنا

اللقاحات

البيئة والتلوث

علم الأجنة

اعضاء التكاثر وتشكل الاعراس

الاخصاب

التشطر

العصيبة وتشكل الجسيدات

تشكل اللواحق الجنينية

تكون المعيدة وظهور الطبقات الجنينية

مقدمة لعلم الاجنة

الأحياء الجزيئي

مواضيع عامة في الاحياء الجزيئي

علم وظائف الأعضاء

الغدد

مواضيع عامة في الغدد

الغدد الصم و هرموناتها

الجسم تحت السريري

الغدة النخامية

الغدة الكظرية

الغدة التناسلية

الغدة الدرقية والجار الدرقية

الغدة البنكرياسية

الغدة الصنوبرية

مواضيع عامة في علم وظائف الاعضاء

الخلية الحيوانية

الجهاز العصبي

أعضاء الحس

الجهاز العضلي

السوائل الجسمية

الجهاز الدوري والليمف

الجهاز التنفسي

الجهاز الهضمي

الجهاز البولي

المضادات الميكروبية

مواضيع عامة في المضادات الميكروبية

مضادات البكتيريا

مضادات الفطريات

مضادات الطفيليات

مضادات الفايروسات

علم الخلية

الوراثة

الأحياء العامة

المناعة

التحليلات المرضية

الكيمياء الحيوية

مواضيع متنوعة أخرى

الانزيمات

قم بتسجيل الدخول اولاً لكي يتسنى لك الاعجاب والتعليق.

Heterochromatin DNA and Transposon Repeats

المؤلف:  Strachan, T., & Read, A.

المصدر:  Human molecular genetics

الجزء والصفحة:  5th E, P306-311

2026-10-05

68

+

-

20

As listed below, two main DNA sequence classes are present in very high copy numbers in the human genome:

 • Tandemly repeated short DNA sequences located at centromeres and other regions of constitutive heterochromatin (which remains condensed throughout the cell cycle);

• Interspersed transposon repeats that are distributed across the nuclear genome and vary in size from hundreds of base pairs to several kilobases in length.

Constitutive heterochromatin is very largely defined by long arrays of tandem DNA repeats

Constitutive heterochromatin is the highly condensed chromatin that is usually located at the periphery of the nucleus, attached to the nuclear membrane (the euchromatic DNA tends to be concentrated in the center of the nucleus where it can be actively transcribed). The underlying DNA accounts for around 200 Mb (~6.5%) of the human genome (see Table 1) and encompasses megabase-sized regions at the centromeres, multiple kilobases of DNA at the telomeres of all chromosomes, and large additional regions on some chromosomes, including: the majority of the Y chromosome; most of the short arms of the acrocentric chromosomes (13, 14, 15, 21, and 22); and very substantial regions of pericentromeric heterochromatin, notably on chromosomes 1, 9, 16, and 19 (see the chromosome banding image on the inside back cover).

Table1. REFERENCE SEQUENCES AND SOME CHARACTERISTICS OF THE 24 HUMAN CHROMOSOMAL DNA MOLECULES

Table1. (Continued) REFERENCE SEQUENCES AND SOME CHARACTERISTICS OF THE 24 HUMAN CHROMOSOMAL DNA MOLECULES

The DNA of the telomeres is very highly conserved in sequence: in vertebrates (and in many other eukaryotes) it consists of long arrays of tandem repeats of a hexanucleotide sequence TTAGGG (or CCCTAA on the opposing strand); according to the length of the arrays, the telomere repeats belong to the minisatellite class of noncoding tandem repeats (see Table 2). Each array extends over several kilobases but the length varies with age, being reduced each time the DNA of a cell replicates in preparation for cell division (because of the chromosome end-replication problem). The telomere hexanucleotides are not represented in the GRCh38.p12 reference sequence); instead, telomere DNA is simply acknowledged by an arbitrary length of 10 kb, with the nucleotides simply represented by the letter N. (In human chromosome 1, for example, the location of telomere sequences in the reference sequence is marked by the following coordinates: short arm—chr1:1–10,000; and long arm—chr1: 248,946,423–248,956,422.) Immediately proximal to the telomere repeats are simple-sequence sub-telomeric DNA repeats.

Table2. MAJOR CLASSES OF HIGH-COPY-NUMBER TANDEMLY REPEATED HUMAN DNA

Table2. (Continued) MAJOR CLASSES OF HIGH-COPY-NUMBER TANDEMLY REPEATED HUMAN DNA

The DNA underlying constitutive heterochromatin at the centromeres and other regions mostly consists of very long arrays of high-copy-number tandemly repeated DNA sequences, known as satellite DNA. Large tracts of heterochromatin are typically com posed of a mosaic of different satellite DNA sequences that are occasionally interrupted by transposon repeats. There are different satellite DNA organizations, and the repeated unit may be a very simple sequence (less than 10 nucleotides long) or a moderately com plex sequence extending to over 100 nucleotides long; see Table 2 for a classification of human satellite DNA and other high-copy-number tandem repeats.

Various satellite DNA families are associated with human centromeres (Figure 1), but only the α-satellite is known to be present at all human centromeres, and its repeat units often contain a binding site for a specific centromere protein, CENPB. Cloned α-satellite arrays have been shown to seed de novo centromeres in human cells, indicating that the α-satellite must have an important role in centromere function.

Fig1. Human centromere and α-satellite DNA organization. α-satellite DNA (alphoid DNA) is a prominent component of the centromere of all human chromosomes. It is composed of repeats of a 171 bp sequence but significant sequence differences are found between different 171 bp repeats (represented here by different colored arrows at bottom; the monomers can vary at up to 30–40% of nucleotide positions). Higher-order repeats (HOR) mark regions where a series of similarly orientated monomer repeats has been tandemly repeated to form long multimer arrays with very high levels of sequence identity (97–100%) between the multiple HORs in an array. In this hypothetical example, we imagine two arrays composed respectively of HOR-1 repeats (each consisting of a sequence of eight monomers) and HOR-2 repeats (with a nine-monomer sequence). Clusters of simple monomers that can be in different orientations are typically found in the interval between the multimer arrays. Outside the α-satellite HORs, the centromere sequences often include α-satellite monomer clusters (αM) and simple sequence (SS) satellite DNAs. For a specific example—the centromere of human chromosome 10—see Figure 4 in the paper of Aldrup-Macdonald & Sullivan (2014) (PMID 24683489) listed under Further Reading.

Transcription of heterochromatin DNA

We typically think of heterochromatin DNA sequences as being transcriptionally inactive, but RNA transcripts can be produced from the tandemly repeated DNA sequences underlying constitutive heterochromatin. Telomeric repeat-containing RNA (TERRA) transcripts of variable length (100 bp–9 kb) are notably produced in the G1 phase of the cell cycle and contain both subtelomeric sequences and C-rich telomere hexanucleotide repeats. TERRA transcripts have multiple roles, including regulation of telomere length and telomere capping and replication. Satellite DNA sequences in pericentromeric and centromeric regions can also be transcribed. The output varies according to developmental stage and cell type, but is amplified in response to cellular stresses, such as heat shock, exposure to hazardous chemicals and heavy metals, and so on.

Transposon-derived repeats make up the majority of the human genome and arose very largely through retrotransposition

 The majority of the human genome is made up of interspersed repetitive noncoding DNA sequences derived from transposons (also called transposable elements), mobile DNA sequences that can migrate to different regions of the genome. About 45% of the genome can readily be seen to be made up of transposon repeats, but much of the remaining “unique” DNA is likely to have been derived from ancient transposon copies that have diverged extensively over long evolutionary timescales. (The most sensitive computer programs indicate that at least two-thirds of the human genome arose in this way.)

In humans, and other mammals, only a tiny minority of transposon repeats are actively transposing. Both the frequency of transposition and the percentage of full-length repeats in a repeat family depend on the family’s evolutionary age: recently evolved repeats have a comparatively high percentage of full-length repeats and of actively transposing members, and conversely, more ancient transposon repeats are frequently truncated copies or have inactivating mutations. Transposons that can trans pose independently are described as autonomous. Others are nonautonomous: they can transpose only with the help of an autonomous transposon.

DNA transposons versus retrotransposons

 A small minority of human transposon repeats originated from the DNA transposon class. Transposons of this type have terminal inverted repeats and migrate directly without any copying of the sequence using a “cut-and-paste” mechanism (they make a transposase to excise the DNA sequence, which then re-inserts elsewhere in the genome). There are two major human DNA transposon superfamilies (Figure 2) plus a variety of less frequent families. Although these repeats were actively transposing in the past, there is much less evidence of recent transposition; they are often described, therefore, as transposon fossils.

Fig2. Human transposon repeat classes. Note that full-length repeats are rare; most are truncated copies. Some full length LINEs (of the LINE-1 subfamily) are able to transpose independently because they can make a functional reverse transcriptase. SINEs are nonautonomous: they need a reverse transcriptase to be supplied. LTR transposons resemble retroviruses and have long terminal repeats (LTR) characteristic of retroviruses. They include endogenous retroviruses (with gag, pol, and env genes), plus truncated LTR elements that have lost key retroviral sequences. DNA transposons use a cut-and-paste transposition, but human DNA transposon repeats seem to be unable to transpose, being truncated or having a mutated transposase gene. Various human DNA transposon superfamilies exist, of which the most numerous are the hAT and the Tc1/mariner superfamilies (see PMID 17339369). LINE, long interspersed nuclear element; SINE, short interspersed nuclear element; HERV, human endogenous retrovirus. (Adapted from The International Human Genome Sequencing Consortium [2001] Nature 409:860–921; PMID 11237011. With permission from Springer Nature. Copyright © 2001.)

The great majority of human transposon repeats belong to the retrotransposon class (also called retroposons). They can transpose using a reverse transcriptase to convert an RNA transcript into a cDNA copy that then integrates into the genomic DNA at a different location (a “copy-and-paste” mechanism). There are three major types of mammalian retrotransposon repeat, as listed below and illustrated in Figure 2.

• LINEs (long interspersed nuclear elements; over 6 kb when full length) have a comparatively long evolutionary history: equivalent sequences are present in other mammals, such as mice. Human LINEs, consisting of three distantly related families (LINE-1, LINE-2, and LINE-3), are located primarily in euchromatic regions, preferentially in the dark AT-rich G-bands of metaphase chromosomes. The LINE-1 (or L1) family is the predominant LINE family, accounting for 17% of the genome, and is detailed below. It continues to have actively transposing members that are the only autonomous transposon repeats in the human genome.

• SINEs (short interspersed nuclear elements; full-length members are less than 400 nucleotides long). SINEs cannot transpose independently. However, SINEs and LINEs share sequences at their 3′ end, and SINEs have been shown to be mobilized by neighboring LINE repeats. By parasitizing on the LINE transposition machinery, SINEs can attain high copy numbers. The primate-specific Alu repeat is the most abundant SINE in the human genome, and is detailed below; the next most common human SINE family are mammalian-wide interspersed repeats, known as MIR elements.

• LTR transposons (repeats that resemble retroviruses, minimally having the long terminal repeats [LTR] characteristic of retroviruses). There are two subclasses: human endogenous retroviruses (HERVs) and LTR elements. HERVs have the gag, pol, and env genes of retroviruses, and arose when infectious retroviruses repeatedly entered the germ line of hosts over many tens of millions of years. LTR elements are effectively truncated HERVs: they retain LTR sequences but have lost key retroviral sequences.

In addition to the major classes above, some repeats are composites of different classes, notably the SVA family, as described below.

The LINE-1 (L1) family

Full-length functional LINE-1 elements make two proteins: an RNA-binding protein, p40, and a protein with both endonuclease and reverse transcriptase activities (Figure 3). Full-length copies bring with them their own promoter (located in the 5′ untranslated region) that can be used after integration in a permissive region of the genome. After translation, the LINE-1 RNA assembles with its own encoded proteins and moves to the nucleus.

Fig3. Structure of human LINE-1, Alu, and SVA repeats. The LINE-1 p40 protein is an RNA-binding protein with a nucleic acid chaperone activity. Converging arrows mark potential transcription from a bidirectional internal protein within the 5′ UTR (untranslated region) of LINE-1 elements. At the other end is an An /Tn sequence, often described as the 3′ poly(A) tail (pA). The LINE-1 endonuclease cuts one strand of a DNA duplex, preferably within the sequence TTTT↓A, and the reverse transcriptase uses the released 3′-OH end to prime cDNA synthesis. New insertion sites are flanked by a small target-site duplication (flanking black arrowheads). Alu repeats often consist of two monomer repeats that have similar sequences terminating in an A-rich or An /Tn sequence (oligo A) but differ in size because of the insertion of a 32 bp element within the larger repeat. The smaller repeat has internal components of an RNA polymerase III promoter (boxes A and B). Nonautonomous SVA repeats are usually more than 2 kb in length and have both an Alu-like sequence and a 3′ HERV fragment (called SINE-R), separated by a VNTR sequence. They may often be transcribed from promoters in flanking DNA sequence. VNTR, variable number of tandem repeats.

To integrate into genomic DNA, the LINE-1 endonuclease cuts a DNA duplex on one strand, leaving a free 3′ OH group that serves as a primer for reverse transcription from the 3′ end of the LINE RNA. The endonuclease’s preferred cleavage site is TTTT↓A; hence the preference for integrating into AT-rich regions. During integration, the reverse transcription often fails to proceed to the 5′ end, resulting in truncated, nonfunctional insertions. Accordingly, only about 1 in 100 copies are full length, and most LINE-derived repeats are short (the average size for all LINE-1 copies is 900 bp).

The LINE-1 machinery is responsible for reverse transcription of all retroelements in the genome, including nonautonomous SINEs and SVA repeats, and also copies of mRNA transcripts that integrate in the genome to create processed pseudogenes and retro genes. Of the 6000 or so full-length LINE-1 sequences, about 60–100 are still capable of transposing, and they occasionally cause disease as a result of aberrant gene expression after insertion.

Alu repeats

The human Alu repeat is the most abundant sequence in the human genome. The full-length repeat is about 280 bp long and consists of two tandem repeats, each about 120 bp in length followed by an A-rich or An /Tn sequence. Monomers, containing only one of the two tandem repeats, and various truncated versions of dimers and monomers are also common, giving a genome-wide average of 230 bp. Alu repeats are primate specific, but subfamilies of different evolutionary ages can be identified, of which the Y and S subfamilies contain the most mobile Alu sequences.

Alu repeats have a relatively high GC content and, although dispersed mainly throughout the euchromatic regions of the genome, are preferentially located in the GC-rich and gene-rich R chromosome bands, in striking contrast to the preferential location of LINEs in AT-rich DNA. When located within genes they are, like LINE-1 elements, almost always confined to introns and the untranslated regions.

Like other mammalian SINEs, Alu repeats originated from cDNA copies of small RNAs transcribed by RNA polymerase III that re-integrated into germ-line DNA. Genes transcribed by RNA polymerase III often have internal promoters, and so the cDNA copies of transcripts carry with them their own promoter sequences that can activate transcription of the cDNA copy in a permissive chromatin location. Both the Alu repeat and, independently, the mouse B1 repeat originated from cDNA copies of 7SL RNA, the short RNA that is a component of the signal recognition particle, using a retrotransposition mechanism. Some other SINEs, such as MIR elements and the mouse B2 repeat, are known to be retrotransposed copies of tRNA sequences.

SVA repeats

SVA repeats are composite repeats, having both an Alu-like sequence and a truncated HERV component (the 3′ end of the env gene plus an LTR) that was originally named SINE-R. These two elements are separated by a VNTR region with a variable number of tandem copies of a 35–50 bp sequence (the name SVA is short for SINE-R–VNTR–Alu). The Alu sequence is preceded by tandem CCCTCT repeats, and the HERV component is followed by an (A/T)n tail, called a poly(A) sequence. This recently evolved, hominid-specific family contains mostly full-length sequences (~2 kb in length), and although there are only 2700 human SVA repeats, they are the third most actively transposing human transposon repeat sequence (after Alu repeats and LINE-1 repeats).

The double-edged nature of transposons: both friends and foes

As detailed in Chapter 13, transposons have been crucially important in genome evolution, and copies of transposons have been valuable sources of novel functional sequences, including not just new regulatory sequences, but also new exons, and, very occasionally, even new genes. Transposon repeats may also assist gene and exon duplication (which are also important in genome evolution) by stabilizing local mispairing of chromatids, and they can alter expression of host genes in different ways, such as by offering new regulatory sequences or new splice sites. Retrotransposons also actively transpose during neurogenesis, creating an additional level of genetic diversity that may be valuable in promoting neuron diversity.

While transposons offer many advantages, active transposons pose a threat to the genome. Effectively, they are mutagens and potentially harmful, and excessive mobilization of transposons can result in chaos. To contain this threat, the first line of defense is epigenetic regulation: the transcription of active retrotransposons is down-regulated by setting epigenetic marks to alter the chromatin state (typical epigenetic marks are DNA methylation and histone modifications). However, epigenetic marks across the genome are erased in early development (in preparation for re-setting of epigenetic marks, including imprinting marks); that allows an opportunity for transient transcription of the retrotransposons, leading to their rapid proliferation.

Because of the importance of the germ line, and because the chromatin state of germ-line cells may be more inherently permissive to retrotransposon activity than that of somatic cells, a variety of different genome defense strategies have evolved to limit the potentially dangerous spread of retrotransposons in the germ line. Many of these mechanisms are often used also to combat virus infections (viruses originated from ancient transposons). Interested readers wishing to explore this topic further might like to consult the reviews by Molaro & Malik (2016) (PMID 26821364) and Goodier (2016) (PMID 27525044) listed under Further Reading. We briefly describe below one important method of containing the threat posed by retrotransposons, which is based on RNA silencing.

Controlling retrotransposons using piRNA-mediated RNA silencing

 Both piRNAs and, to a lesser extent, endogenous short interfering RNAs (siRNAs) are important in transposon control in the germ line using RNA silencing.  piRNA-based silencing is similar in some respects, but different in others. piRNAs, resembling a type of repeat-associated siRNA, are typically 24–31 nucleotides long, slightly longer than siRNAs, and they have a dis tinct preference for a U at the +1 position. The name is a contraction of PiWi protein interacting RNAs (they were first discovered in Drosophila where they bind to the PiWi protein and two related proteins; the equivalent mammalian proteins are the cytoplasmic proteins MILI and MIWI and a nuclear protein MIWI2).

Whereas both siRNAs and miRNAs are produced by cleavage of double-stranded RNA precursors into functional small RNAs, piRNAs are produced from single-stranded RNAs, and predominantly in germ cells. The majority of piRNAs originate from long (50–100 kb) RNA polymerase II transcripts produced from numerous regions of the genome with a high density of truncated transposon repeats. The capped and poly adenylated transcripts are exported to the cytoplasm where they are each subject to processing and ultimately the production of thousands of piRNAs from transcripts.

Like siRNAs, piRNAs act as guide RNAs, recognizing and binding to complementary DNA and RNA sequences, and like siRNAs, their job is to bind executive silencing proteins and deliver them to the target sequences. For example, in the embryonic mouse male germ line, piRNAs originating from transposons can bind the MIWI2 protein and transport it to the sites of complementary sequences within transposons across the genome; the deposited MIWI2 proteins recruit transcriptional silencing complexes to the target sequences, often repressing their transcription by CpG methylation. Similarly, the cytoplasmic MILI protein can be bound by piRNAs that then bind to transposon transcripts with the complementary sequence, whereupon the slicer endonuclease cleaves the transcript, leading to its degradation, a form of post-transcriptional silencing.

اخر الاخبار

اشترك بقناتنا على التلجرام ليصلك كل ما هو جديد