Algorithm for automatic loop parallelization for graphics processing units

Parallelization of loop operators is a long standing problem of parallel programming. The widespread use of graphics processing units for computational tasks has resulted in the new statement of the mentioned problem for this class of multicore systems. The purpose of this work is to improve the mec...

Ausführliche Beschreibung

Gespeichert in:
Bibliographische Detailangaben
Veröffentlicht in:PROBLEMS IN PROGRAMMING
Datum:2018
Heft:4
Сторінки:28-36
ISSN:1727-4907
Автори та афіліації:
  • А.Yu. Doroshenko — Institute of Software Systems NAS of Ukraine
  • O.A. Yatsenko — Institute of Software Systems NAS of Ukraine
  • O.G. Beketov — Institute of Software Systems NAS of Ukraine
Ключові слова:обчислення загального призначення на графічних процесорах, проектування і синтез програм, проектування та синтез програм, оптимізація циклів, синтез програм, автоматизація оптимізації паралельних програм, методи паралелізації, паралельна програма, синхронізація паралельних програм, автоматизоване проектування програм
Hauptverfasser: Doroshenko, А.Yu., Yatsenko, O.A., Beketov, O.G.
Format: Artikel
Sprache:Ukrainisch
Veröffentlicht: PROBLEMS IN PROGRAMMING 2018
Schlagworte:
Online Zugang:https://pp.isofts.kiev.ua/index.php/ojs1/article/view/308
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
Назва журналу:Problems in programming
Завантажити файл: Pdf

Institution

Problems in programming
Beschreibung
Zusammenfassung:Parallelization of loop operators is a long standing problem of parallel programming. The widespread use of graphics processing units for computational tasks has resulted in the new statement of the mentioned problem for this class of multicore systems. The purpose of this work is to improve the mechanism of transformation of cyclic operators for loop parallelization for execution on a graphics processing unit. Software tool for computation optimization that allows to parallelize cyclic operators semi‑automatically was developed. Data bufferization synchronized with main loop execution was implemented, and the software tool using the rewriting rules system TermWare was built and integrated with the toolkit for design and synthesis of programs IDS. The developed system was tested using heterogeneous multicore cluster. The advantages of the developed system in comparison with well-known parallelization system Par4All consist in processing speed and the possibility of processing of data amounts exceeding the amount of memory of a graphics processing unit, and also the ability to use several graphics processing units simultaneously. The developed system was applied for parallelization of a serial loop, which is the part of a numerical weather forecasting program.Problems in programming 2017; 4: 028-036
ISSN:1727-4907
DOI:10.15407/pp2017.04.028