Skip to main navigation Skip to search Skip to main content

A multilevel parallelization framework for high-order stencil computations

  • Hikmet Dursun
  • , Ken Ichi Nomura
  • , Liu Peng
  • , Richard Seymour
  • , Weiqiang Wang
  • , Rajiv K. Kalia
  • , Aiichiro Nakano
  • , Priya Vashishta

Research output: Chapter in Book/Report/Conference proceedingConference contribution

35 Scopus citations

Abstract

Stencil based computation on structured grids is a common kernel to broad scientific applications. The order of stencils increases with the required precision, and it is a challenge to optimize such high-order stencils on multicore architectures. Here, we propose a multilevel parallelization framework that combines: (1) inter-node parallelism by spatial decomposition; (2) intra-chip parallelism through multithreading; and (3) data-level parallelism via single-instruction multiple-data (SIMD) techniques. The framework is applied to a 6th order stencil based seismic wave propagation code on a suite of multicore architectures. Strong-scaling scalability tests exhibit superlinear speedup due to increasing cache capacity on Intel Harpertown and AMD Barcelona based clusters, whereas weak-scaling parallel efficiency is 0.92 on 65,536 BlueGene/P processors. Multithreading+SIMD optimizations achieve 7.85-fold speedup on a dual quad-core Intel Clovertown, and the data-level parallel efficiency is found to depend on the stencil order. © 2009 Springer.
Original languageEnglish
Title of host publicationLecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
Pages642-653
Number of pages12
Volume5704 LNCS
DOIs
StatePublished - Nov 9 2009
Externally publishedYes

Fingerprint

Dive into the research topics of 'A multilevel parallelization framework for high-order stencil computations'. Together they form a unique fingerprint.

Cite this