Home
Alewife
AlphaServer SC
AP 3000
AV 25000
The Challenge
CM-2
CM-5
Dash
E10000
Exemplar V21000
Flash
Gamma II Plus
Hector
ILLIAC IV
iPSC
KSR-2
MP-1216
NCUBE/ten
NUMAchine
Paragon
RS/6000 SP
SR8000
SX-5
T3E
VPP5000
                        The HP Exemplar V2600

Machine type: RISC-based distributed-memory multi-processor
Models: Exemplar V2600.
Operating system: HP-UX (HP's Unix flavour)
Connection structure: Ring
Compilers: Fortran 77, Fortran 90, Parallel Fortran, HPF, C, C++
The V2600 is the latest in the series of Exemplar systems that have been offered first by Convex and later by HP since 1995. The architecture, however, has not radically changed: up to 32 PA-RISC 8600 chips are clustered via a crossbar to form an SMP node. The PA-RISC 8600 CPUs run at a clock cycle of 1.81 ns. As a CPU contains 2 floating-point units that are able do execute a combined floating multiply-add instruction, in favourable circumstances four flops/cycle can be achieved and a Theoretical Peak Performance of 2.27 Gflop/s per CPU can be attained. Per SMP node the peak speed is 72.7 Gflop/s. Up to four SMP nodes can be coupled by a so-called SCA HyperLink, uni-directional SCI rings with an aggregate bandwidth of 3.84 GB/s, while the aggregate bandwidth within an SMP node is 15.36 GB/s. The HyperLinks tolerate multiple outstanding requests and, in addition, there is a "HyperLinkcache" that both help in hiding the communication latency in inter-node communication. As in the former systems a shared memory parallel model is supported. HP is a partner in the OpenMP organisation and will therefore make available this style of shared-memory parallel programming in addition to (and later on instead of) its proprietary parallel model. The shared-memory parallelism is not confined to the SMP nodes: a multi-node system can be addressed globally making the Exemplar a ccNUMA system. The memory latency within and between nodes differs by about a factor of 3--3.5. Measured Performances:
For the new V2600 system no results are available yet. However for the V2500 which is using PA-RISC 8500 processors with a 2.27 ns cycle time but which has essentially the same architecture in a speed of 31.59 Gflop/s is reported for a 1-cabinet, 32 processor system when solving a 41,000-order dense linear system, an efficiency of 56% on this problem.
Model                                       Exemplar V2600    
Clock cycle                                1.81 ns                 
Theor. peak performance:
Per proc. (64-bits)                      2.27 Gflop/s          
Maximal (64-bits)                       291 Gflop/s           
Main memory
Memory/node                            <=1 GB                
Memory/maximal                       <= 128 GB            
No. of processors                       16-128                  
Communication bandwidth:
aggregate (per cabinet)                15.36 GB/s           
aggregate (inter-cabinet)              3.84 GB/s             
System Parameters:
1