TY - GEN
T1 - Accelerating Quantum Light-Matter Dynamics on Graphics Processing Units
AU - Razakh, Taufeq Mohammed
AU - Linker, Thomas
AU - Luo, Ye
AU - Kalia, Rajiv K.
AU - Nomura, Ken Ichi
AU - Vashishta, Priya
AU - Nakano, Aiichiro
N1 - Publisher Copyright:
© 2024 IEEE.
PY - 2024
Y1 - 2024
N2 - To study light-matter interaction, we have developed a linear-scaling DC-MESH (divide-and-conquer Maxwell-Ehrenfest-surface hopping) simulation algorithm, where our globally-sparse and locally-dense electronic solvers, multiple time-scale splitting, and shadow dynamics achieve high scalability and allow the most compute-intensive quantum dynamics kernel based on time-dependent density functional theory to reside on GPU with minimal CPU-GPU data transfer. GPU computation based on OpenMP target constructs is accelerated by: (i) data and loop reordering for better memory access patterns; (ii) hierarchical GPU offloading using teams-distribute and parallel constructs, respectively, for coarse and fine computations; (iii) algebraic 'BLASification' of the nonlocal computational bottleneck; and (iv) GPU-resident data structures facilitated by custom C++ class initializer and destructor based on OpenMP target data constructs. We have thereby achieved 644-fold speedup on Nvidia A100 GPU over AMD EPYC 7543 CPU of the Polaris computer at Argonne Leadership Computing Facility. In addition, the DC-MESH code exhibits a weak-scaling parallel efficiency of 96.73% on 256 nodes (or 1,024 GPUs) of Polaris for 5,120-atom PbTiO3 material. This enables the study of light-induced topological switching for future ultrafast and ultralow-power ferroelectric topotronics applications.
AB - To study light-matter interaction, we have developed a linear-scaling DC-MESH (divide-and-conquer Maxwell-Ehrenfest-surface hopping) simulation algorithm, where our globally-sparse and locally-dense electronic solvers, multiple time-scale splitting, and shadow dynamics achieve high scalability and allow the most compute-intensive quantum dynamics kernel based on time-dependent density functional theory to reside on GPU with minimal CPU-GPU data transfer. GPU computation based on OpenMP target constructs is accelerated by: (i) data and loop reordering for better memory access patterns; (ii) hierarchical GPU offloading using teams-distribute and parallel constructs, respectively, for coarse and fine computations; (iii) algebraic 'BLASification' of the nonlocal computational bottleneck; and (iv) GPU-resident data structures facilitated by custom C++ class initializer and destructor based on OpenMP target data constructs. We have thereby achieved 644-fold speedup on Nvidia A100 GPU over AMD EPYC 7543 CPU of the Polaris computer at Argonne Leadership Computing Facility. In addition, the DC-MESH code exhibits a weak-scaling parallel efficiency of 96.73% on 256 nodes (or 1,024 GPUs) of Polaris for 5,120-atom PbTiO3 material. This enables the study of light-induced topological switching for future ultrafast and ultralow-power ferroelectric topotronics applications.
KW - algebraic BLASijication
KW - GPU acceleration
KW - light-matter interaction
KW - quantum dynamics
KW - time-dependent density functional theory
UR - https://www.scopus.com/pages/publications/85200723536
U2 - 10.1109/IPDPSW63119.2024.00176
DO - 10.1109/IPDPSW63119.2024.00176
M3 - Conference contribution
AN - SCOPUS:85200723536
T3 - 2024 IEEE International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2024
SP - 1057
EP - 1066
BT - 2024 IEEE International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2024
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2024 IEEE International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2024
Y2 - 27 May 2024 through 31 May 2024
ER -