Academic Journal

Efficient and Accurate Selection of Optimal Collective Communication Algorithms Using Analytical Performance Modeling

التفاصيل البيبلوغرافية
العنوان: Efficient and Accurate Selection of Optimal Collective Communication Algorithms Using Analytical Performance Modeling
المؤلفون: Emin Nuriyev, Alexey Lastovetsky
المصدر: IEEE Access, Vol 9, Pp 109355-109373 (2021)
بيانات النشر: IEEE, 2021.
سنة النشر: 2021
المجموعة: LCC:Electrical engineering. Electronics. Nuclear engineering
مصطلحات موضوعية: Message passing, collective communication algorithms, communication performance modeling, MPI, Electrical engineering. Electronics. Nuclear engineering, TK1-9971
الوصف: The performance of collective operations has been a critical issue since the advent of Message Passing Interface (MPI). Many algorithms have been proposed for each MPI collective operation but none of them proved optimal in all situations. Different algorithms demonstrate superior performance depending on the platform, the message size, the number of processes, etc. MPI implementations perform the selection of the collective algorithm empirically, executing a simple runtime decision function. While efficient, this approach does not guarantee the optimal selection. As a more accurate but equally efficient alternative, the use of analytical performance models of collective algorithms for the selection process was proposed and studied. Unfortunately, the previous attempts in this direction have not been successful. We revisit the analytical model-based approach and propose two innovations that significantly improve the selective accuracy of analytical models: (1) We derive analytical models from the code implementing the algorithms rather than from their high-level mathematical definitions. This results in more detailed and relevant models. (2) We estimate model parameters separately for each collective algorithm and include the execution of this algorithm in the corresponding communication experiment. We experimentally demonstrate the accuracy and efficiency of our approach using Open MPI broadcast and gather algorithms and two different Grid’5000 clusters and one supercomputer.
نوع الوثيقة: article
وصف الملف: electronic resource
اللغة: English
تدمد: 2169-3536
Relation: https://ieeexplore.ieee.org/document/9502598/; https://doaj.org/toc/2169-3536
DOI: 10.1109/ACCESS.2021.3101689
URL الوصول: https://doaj.org/article/3728205aff604d8aa13e8ae25273bbe7
رقم الانضمام: edsdoj.3728205aff604d8aa13e8ae25273bbe7
قاعدة البيانات: Directory of Open Access Journals
الوصف
تدمد:21693536
DOI:10.1109/ACCESS.2021.3101689