High Performance Fortran: A Practical Analysis

Allan Knies; Matthew O'keefe; Tom Macdonald

doi:10.1155/1994/150306

High Performance Fortran: A Practical Analysis

Scientific Programming ◽

10.1155/1994/150306 ◽

1994 ◽

Vol 3 (3) ◽

pp. 187-199 ◽

Cited By ~ 7

Author(s):

Allan Knies ◽

Matthew O'keefe ◽

Tom Macdonald

Keyword(s):

High Performance ◽

Parallel Machines ◽

Production Quality ◽

Efficient Production ◽

Data Parallel ◽

High Performance Fortran ◽

Multiple Data ◽

Computing Industry ◽

Application Developers ◽

Important Design

The recently released high performance Fortran forum (HPFF) proposal has stirred much interest in the high performance computing industry. HPFF's most important design goal is to create a language that has source code portability and that achieves high performance on single instruction multiple data (SIMD), distributed-memory multiple instruction multiple data (MIMD), and shared-memory MIMD architectures. The HPFF proposal brings to the forefront many questions about design of portable and efficient languages for parallel machines. In this article, we discuss issues that need to be addressed before an efficient production quality compiler will be available for any such language. We examine some specific issues that are related to HPF's model of computation and analyze several implementation issues. We also provide some results from another data parallel compiler to help gain insight on some of the implementation issues that are relevant to HPF. Finally, we provide a summary of options currently available for application developers in industry.

Download Full-text

PGHPF – An Optimizing High Performance Fortran Compiler for Distributed Memory Machines

Scientific Programming ◽

10.1155/1997/705102 ◽

1997 ◽

Vol 6 (1) ◽

pp. 29-40 ◽

Cited By ~ 9

Author(s):

Zeki Bozkus ◽

Larry Meadows ◽

Steven Nakamoto ◽

Vincent Schuster ◽

Mark Young

Keyword(s):

High Performance ◽

Distributed Memory ◽

Parallel Machines ◽

High Efficiency ◽

Memory Systems ◽

Production Quality ◽

Distributed Memory Machines ◽

High Performance Fortran ◽

Application Developers ◽

Efficient Software

High Performance Fortran (HPF) is the first widely supported, efficient, and portable parallel programming language for shared and distributed memory systems. HPF is realized through a set of directive-based extensions to Fortran 90. It enables application developers and Fortran end-users to write compact, portable, and efficient software that will compile and execute on workstations, shared memory servers, clusters, traditional supercomputers, or massively parallel processors. This article describes a production-quality HPF compiler for a set of parallel machines. Compilation techniques such as data and computation distribution, communication generation, run-time support, and optimization issues are elaborated as the basis for an HPF compiler implementation on distributed memory machines. The performance of this compiler on benchmark programs demonstrates that high efficiency can be achieved executing HPF code on parallel architectures.

Download Full-text

A Linear Algebra Framework for Static High Performance Fortran Code Distribution

Scientific Programming ◽

10.1155/1997/195689 ◽

1997 ◽

Vol 6 (1) ◽

pp. 3-27 ◽

Cited By ~ 22

Author(s):

Corinne Ancourt ◽

Fabien Coelho ◽

FranÇois Irigoin ◽

Ronan Keryell

Keyword(s):

Linear Algebra ◽

High Performance ◽

Address Space ◽

Data Parallel ◽

High Performance Fortran ◽

Multiple Data ◽

Fortran Code ◽

Code Distribution ◽

Overlap Analysis ◽

Data Parallel Programming

High Performance Fortran (HPF) was developed to support data parallel programming for single-instruction multiple-data (SIMD) and multiple-instruction multiple-data (MIMD) machines with distributed memory. The programmer is provided a familiar uniform logical address space and specifies the data distribution by directives. The compiler then exploits these directives to allocate arrays in the local memories, to assign computations to elementary processors, and to migrate data between processors when required. We show here that linear algebra is a powerful framework to encode HPF directives and to synthesize distributed code with space-efficient array allocation, tight loop bounds, and vectorized communications forINDEPENDENTloops. The generated code includes traditional optimizations such as guard elimination, message vectorization and aggregation, and overlap analysis. The systematic use of an affine framework makes it possible to prove the compilation scheme correct.

Download Full-text

Opus: A Coordination Language for Multidisciplinary Applications

Scientific Programming ◽

10.1155/1997/632908 ◽

1997 ◽

Vol 6 (4) ◽

pp. 345-362 ◽

Cited By ~ 25

Author(s):

Barbara Chapman ◽

Matthew Haines ◽

Piyush Mehrotra ◽

Hans Zima ◽

John Van Rosendale

Keyword(s):

High Performance ◽

Data Repository ◽

Task Parallelism ◽

Central Concept ◽

Coordination Language ◽

Data Parallel ◽

High Performance Fortran ◽

Parallel Languages ◽

Wide Range ◽

And Task

Data parallel languages, such as High Performance Fortran, can be successfully applied to a wide range of numerical applications.However, many advanced scientific and engineering applications are multidisciplinary and heterogeneous in nature, and thus do not fit well into the data parallel paradigm. In this paper we present Opus, a language designed to fill this gap. The central concept of Opus is a mechanism called ShareD Abstractions (SDA). An SDA can be used as a computation server, i.e., a locus of computational activity, or as a data repository for sharing data between asynchronous tasks. SDAs can be internally data parallel, providing support for the integration of data and task parallelism as well as nested task parallelism. They can thus be used to express multidisciplinary applications in a natural and efficient way. In this paper we describe the features of the language through a series of examples and give an overview of the runtime support required to implement these concepts in parallel and distributed environments.

Download Full-text

Performance Issues in High Performance Fortran Implementations of Sensor-Based Applications

Scientific Programming ◽

10.1155/1997/372831 ◽

1997 ◽

Vol 6 (1) ◽

pp. 59-72 ◽

Cited By ~ 1

Author(s):

David R. O'hallaron ◽

Jon Webb ◽

Jaspal Subhlok

Keyword(s):

High Performance ◽

Parallel Machines ◽

Radar Imaging ◽

Synthetic Aperture ◽

Application Domain ◽

Resonance Imaging ◽

High Performance Fortran ◽

Intel Paragon ◽

Independent Loops ◽

Tracking Radar

Applications that get their inputs from sensors are an important and often overlooked application domain for High Performance Fortran (HPF). Such sensor-based applications typically perform regular operations on dense arrays, and often have latency and through put requirements that can only be achieved with parallel machines. This article describes a study of sensor-based applications, including the fast Fourier transform, synthetic aperture radar imaging, narrowband tracking radar processing, multibaseline stereo imaging, and medical magnetic resonance imaging. The applications are written in a dialect of HPF developed at Carnegie Mellon, and are compiled by the Fx compiler for the Intel Paragon. The main results of the study are that (1) it is possible to realize good performance for realistic sensor-based applications written in HPF and (2) the performance of the applications is determined by the performance of three core operations: independent loops (i.e., loops with no dependences between iterations), reductions, and index permutations. The article discusses the implications for HPF implementations and introduces some simple tests that implementers and users can use to measure the efficiency of the loops, reductions, and index permutations generated by an HPF compiler.

Download Full-text

Experiences in Data-Parallel Programming

Scientific Programming ◽

10.1155/1997/260463 ◽

1997 ◽

Vol 6 (1) ◽

pp. 153-158 ◽

Cited By ~ 1

Author(s):

Terry W. Clark ◽

Reinhard v. Hanxleden ◽

Ken Kennedy

Keyword(s):

High Performance ◽

Parallel Applications ◽

Scientific Application ◽

Data Parallel ◽

High Performance Fortran ◽

Engineering Data ◽

Source Program ◽

Parallel Compiler ◽

Fortran D ◽

Near Future

To efficiently parallelize a scientific application with a data-parallel compiler requires certain structural properties in the source program, and conversely, the absence of others. A recent parallelization effort of ours reinforced this observation and motivated this correspondence. Specifically, we have transformed a Fortran 77 version of GROMOS, a popular dusty-deck program for molecular dynamics, into Fortran D, a data-parallel dialect of Fortran. During this transformation we have encountered a number of difficulties that probably are neither limited to this particular application nor do they seem likely to be addressed by improved compiler technology in the near future. Our experience with GROMOS suggests a number of points to keep in mind when developing software that may at some time in its life cycle be parallelized with a data-parallel compiler. This note presents some guidelines for engineering data-parallel applications that are compatible with Fortran D or High Performance Fortran compilers.

Download Full-text

High Performance Fortran, Version 2

Parallel Processing Letters ◽

10.1142/s0129626497000425 ◽

1997 ◽

Vol 07 (04) ◽

pp. 437-449 ◽

Cited By ~ 3

Author(s):

Robert Schreiber

Keyword(s):

High Performance ◽

Data Parallelism ◽

Data Mapping ◽

Parallel Language ◽

Data Parallel ◽

High Performance Fortran ◽

Concurrent Execution ◽

Parallel Tasks ◽

New Ideas ◽

Data Parallel Programming

This paper introduces the ideas that underly the data-parallel language High Performance Fortran (HPF) and the new ideas in version 2 of HPF. It first reviews HPF's key language elements. It discusses the meaning of data parallelism and the limitations of HPF version 1 as a data-parallel programming language. The second part of the paper is a review of the development of version 2 of HPF. The extended language, under development in 1996, includes a richer data mapping capability; an extension to the independent loop that allows reduction operations in the loop range; a means for directing the mapping of computation as well as data; and a way to specify concurrent execution of several parallel tasks on disjoint subsets of processors.

Download Full-text

Implementing O(N) N-Body Algorithms Efficiently in Data-Parallel Languages

Scientific Programming ◽

10.1155/1996/425936 ◽

1996 ◽

Vol 5 (4) ◽

pp. 337-364 ◽

Cited By ~ 7

Author(s):

Yu Hu ◽

S. Lennart Johnsson

Keyword(s):

High Performance ◽

Data Distribution ◽

Optimization Techniques ◽

Total Efficiency ◽

Total Execution Time ◽

Data Parallel ◽

High Performance Fortran ◽

Parallel Languages ◽

Connection Machine ◽

Body Method

The optimization techniques for hierarchical O(N) N-body algorithms described here focus on managing the data distribution and the data references, both between the memories of different nodes and within the memory hierarchy of each node. We show how the techniques can be expressed in data-parallel languages, such as High Performance Fortran (HPF) and Connection Machine Fortran (CMF). The effectiveness of our techniques is demonstrated on an implementation of Anderson's hierarchical O(N) N-body method for the Connection Machine system CM-5/5E. Of the total execution time, communication accounts for about 10–20% of the total time, with the average efficiency for arithmetic operations being about 40% and the total efficiency (including communication) being about 35%. For the CM-5E, a performance in excess of 60 Mflop/s per node (peak 160 Mflop/s per node) has been measured.

Download Full-text

An Introduction to High Performance Fortran

Scientific Programming ◽

10.1155/1995/612973 ◽

1995 ◽

Vol 4 (2) ◽

pp. 87-113 ◽

Cited By ~ 11

Author(s):

John Merlin ◽

Anthony Hey

Keyword(s):

Parallel Computation ◽

High Performance ◽

Data Distribution ◽

Parallel Architectures ◽

Data Parallel ◽

High Performance Fortran ◽

Concurrent Execution ◽

Fortran 90

High Performance Fortran (HPF) is an informal standard for extensions to Fortran 90 to assist its implementation on parallel architectures, particularly for data-parallel computation. Among other things, it includes directives for specifying data distribution across multiple memories, and concurrent execution features. This article provides a tutorial introduction to the main features of HPF.

Download Full-text

OPUS: HETEROGENEOUS COMPUTING WITH DATA PARALLEL TASKS

Parallel Processing Letters ◽

10.1142/s0129626499000256 ◽

1999 ◽

Vol 09 (02) ◽

pp. 275-289 ◽

Cited By ~ 3

Author(s):

ERWIN LAURE ◽

PIYUSH MEHROTRA ◽

HANS ZIMA

Keyword(s):

High Performance ◽

Heterogeneous Computing ◽

Coarse Grain ◽

Tool Support ◽

Task Parallelism ◽

Coordination Language ◽

Data Parallel ◽

High Performance Fortran ◽

Parallel Tasks ◽

Object Based

The coordination language Opus is an object-based extension of High Performance Fortran (HPF) that supports the integration of coarse-grain task parallelism with HPF-style data parallelism. In this paper we discuss Opus in the Context of multidisciplinary applications (MDAs) which execute in a heterogencous environment. After outlining the major properties of such applications and a number of different approaches towards providing language and tool support for MDAs we describe the salientfeatures of Opus and its implementation, emphasizing the issues related to the coordination of data-parallel HPF programs in a heterogencous environment.

Download Full-text

Data parallel programming: The promises and limitations of high performance fortran

Parallel Computation - Lecture Notes in Computer Science ◽

10.1007/3-540-57314-3_10 ◽

1993 ◽

pp. 114-114

Author(s):

Piyush Mehrotra

Keyword(s):

Parallel Programming ◽

High Performance ◽

Data Parallel ◽

High Performance Fortran ◽

Data Parallel Programming

Download Full-text