Determination of performance characteristics of scientific applications on IBM Blue Gene/Q

C Evangelinos and RE Walkup and V Sachdeva and KE Jordan and H Gahvari and IH Chung and MP Perrone and L Lu and LK Liu and K Magerlein, IBM JOURNAL OF RESEARCH AND DEVELOPMENT, 57, 9 (2013).

DOI: 10.1147/JRD.2012.2229901

The IBM Blue Gene (R)/Q platform presents scientists and engineers with a rich set of hardware features such as 16 cores per chip sharing a Level 2 cache, a wide SIMD (single-instruction, multiple-data) unit, a five-dimensional torus network, and hardware support for collective operations. An especially important feature is that the cores have four "hardware threads," which makes it possible to hide latencies and obtain a high fraction of the peak issue rate from each core. All of these hardware resources present unique performance-tuning opportunities on Blue Gene/Q. We provide an overview of several important applications and solvers and study them on Blue Gene/Q using performance counters and Message Passing Interface profiles. We discuss how Blue Gene/Q tools help us understand the interaction of the application with the hardware and software layers and provide guidance for optimization. On the basis of our analysis, we discuss code improvement strategies targeting Blue Gene/Q. Information about how these algorithms map to the Blue Gene (R) architecture is expected to have an impact on future system design as we move to the exascale era.

Return to Publications page