• No results found

Optim.ization of BLAS

In document dtj v06 03 1994 pdf (Page 47-49)

BLAS 2 rou t ines operate on matrix, vector. and scalar data. The data st ructu res are larger and more complex than the BLAS 1 data structures and the operations more compl icated . Accordi ngly, these rou tines lend themse lves to more sophisticated optim ization techniques.

Optimized DX.,\1l BLAS 2 routines are typica l ly 20 percent to 100 percent faster than the public domain rou ti nes. Figure 2 il lustrates this performance improvement for the matrix-vector mu ltiply routine. DGE.\1V. and the triangu lar solve routine. DTRSY.H

The DXML DGEMY uses a data-blocking technique that asymptotical l y performs two floati ng-poi nt operat i o ns for each mem ory access, compared to the publ ic domain ve rsion, which performs two floating-poi n t operations for every three memory accesses. 19 This tec hnique is designed to m i n i mize translation bu ffer and data cache misses and m axi­ m ize the use of floating-point registers u' IH 2o The same data prefetch considerations used on the BLAS 1 routines are also used on the BLAS 2 rou ti nes.

The DXJ\'IL version of the DTRSY rou tine partitions the problem such that a sma.l l triangu lar solve oper­ ation is performed. The resu l t of this solve opera­ tion is then used in a DGEM V operation to u pdate the remainder of the vector. The process is repeated unt i l the final triangu lar update comp letes the operation. Thus the DTRSV rout ine rel ies heavily on the optim izati o ns used in the DGEMY rou tine.

Scientific Computing Optimizations for AJpha 1 40 1 20 I 1 00 (f) 0:: 80 0 -' u. 2 60 40 ' ' ' ' ' ' ' - . ' .._ "'.-. - - -.- ..,. _ _ _ _ 20 0 KEY: 200 400 600 BOO ORDER OF V ECTORS/MATRICES --BLAS D G E M V - --- D X M L D G E M V · · · · · BLAS DTRSV -·- · - DXML DTRSV 1 000

Figure 2 Performance of BLAS 2 Rou tines DGEMV and DTRSV

As with BLAS 1 routines, BLAS 2 rout ines be nefit greatly from data cache. Although the effect is less dramatic for the BLAS 2 routines, Figure 2 clearly shows the three-step profile observed in Figure 1 .

Best performance is achieved when both matrix and vector fit in the primary cache. Performance is lower but flat over the region where the data fits on the secondary board level cache. The final per­ formance plateau is reached when data resides entirely in memory.

Optimization of BLAS 3

BLAS 3 rou tines operate primarily on matrices. The operations and data structures are more compl i­ cated that those of BLAS 1 and BLAS 2 routines. Typically, BLAS 3 routines perform many computa­ tions on each data element. These routines exhibit a great deal of data reuse and thus naturally lend them­ selves to sophisticated optimization techniques.

DXML BLAS 3 rou t i nes are general ly two to ten times faster than their public domain counterparts. The plots in Figure 3 show these performance dif­ ferences for the ma trix-matrix mul tiply routine,

DGEMM, and the triangular solve routine with multi­

ple right -hand sides, DTRSM 9

Al l performance optimization techniques used for the DXML B LAS 1 and BLAS 2 routines are used on the DXM L BLAS 3 routines. In particul ar, data­ blocking techniques are used extensively. Portions

50 1 80 1 60 1 40 1 20 (f) 0:: 1 00 0 -' I I

80 I 60 40 20 0 KEY: / - - I I 200 -- BLAS D G E M M - - - - DXML DGEMM · BLAS DTRSM - · - - DXML DTRSM 400 600 O R D E R OF MATRICES 800 1 000

Figure 3 Performance of BIAS 3 Routines DGEMM and DTRSM

of mat rices are copied to page-aligned work areas where secondary cache and translation bu ffer misses are eliminated and primary cache misses are absolutely minimized.

As a n example, within the primary compute loop of the DXJVIL DGEMM rou tine, there are no transla­ tion bu ffer misses, no secondary cache misses, and, on average, only one pri mary cache miss for every 42 floating-point operations. Performance within this key loop is also enhanced by carefu l ly using floating-point registers so that fou r floating-point operations are performed for each memory read access. Much of the DXML BLAS 3 performance advantage over the publ ic domain routi nes is a direct consequence of a greatly improved ratio of floating-point operations per memory access.

The DXML DTRSM routine is optimized in a man­

ner similar to its BLAS 2 counterpart, DTRSV. A small triangu lar system is solved . The resu l ti ng matrix is then used by DGEMM to update the remainder of the right-hand-side matrix. Consequently, most of the DXML DTRSM performance is directly attrib­ utable to the DXML DGEMM routine. In fact, the tech­ n iques used in DGEMM pervade DXM L BLAS 3

routines.

Figure 3 i l lustrates a key feature of DXML BLAS 3 routines. Whereas the performance of public domain rou tines degrades sign ificantly as the matrices become too large to fi t in caches, DXJ\1 L

DXML: A High-performance Scient1jic Subroutine Library

routines are relatively insensit ive to array size, shape , or orientation." 9 The performance of a DXML BLAS 3 routine typical ly reaches an asymptote and remains there regard less of problem size.

In document dtj v06 03 1994 pdf (Page 47-49)