Griffith HPC User Guide

Griffith HPC User Guide

Contents

1. Introduction

Gowonda is a 792 core HPC cluster and consists of a mixture of SGI Altix XE and SGI® Rackable™ C2114-4TY14 servers. It is managed by the eResearch Services unit in Information Services. Gowonda was made possible by a grant from the Queensland Cyber Infrastructure Foundation (QCIF) and funding from the University Electronic Infrastructure Capital Fund.

Gowonda will be used to run computations that require large amount of computing resources. All Griffith University researchers and researchers from QCIF affiliated institutions will be able to get access to Gowonda. Similarly Griffith researchers can call upon additional resources from QCIF affiliated institutions if required. Gowonda is a major increase in capability for Griffith University and its partners.
It formally came into operation on Aug 1st 2011.

There is no plan to charge for legitimate research usage of gowonda. However, we will need to meet expectations of the stakeholders. Consequently, we will need to be able to account for all usage on the cluster to satisfy our stakeholders. Information is provided below about how you can assist in this.

This document is loosely modeled after “A Beginner's Guide to the Barrine Linux HPC Cluster” written by Dr. David Green (HPC manager, UQ) , ICE Cluster User Guide written by Bryan Hughes and Wiki pages of the City University of New York (CUNY) HPC Center (see reference for further details).

2. Gowonda Overview

2.1 Hardware

2.1.1 Hardware (2024 upgrade) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

2.1.2 Hardware (2019 upgrade) . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

Node Names 

Total Number 

Mem per node

Cores per node

Processor Type

GPU card 

 

 

gc-prd-hpcn001

gc-prd-hpcn002

gc-prd-hpcn003

gc-prd-hpcn004

gc-prd-hpcn005

gc-prd-hpcn006

 

6

192GB

72

2x Intel Xeon 6140

 

 

 

n061 (gpu node)

1

500 GB

96

2X AMD EPYC 7413 

24-Core Processor

5 X A100

NVIDIA A100 80GB PCIe

 

 

n060 (gpu node)

1

380 GB

72

Intel(R) Xeon(R) Gold 6140 CPU @ 2.30GHz

8 X V100

NVIDIA V100-PCIE-32GB

 

 

2.2 Old Hardware . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

Node Type

Total Number of Cores

Total Amount of Memory (GB)

Compute Nodes

Cores Per node

Mem per Node (GB)

Memory per Core

Processor Type

Small Memory Nodes

 

 

48

48

(n001-n004)

4

12

1

Intel(R) Xeon(R) CPU X5650 @ 2.67GHz

Medium Memory Nodes

108

216

9

(n005-n009,n010-n012,n019)

12

24

2

Intel(R) Xeon(R) CPU X5650 @ 2.67GHz

Large Memory Nodes

72

288

6

(n013-n018)

12

48

4

Intel(R) Xeon(R) CPU X5650 @ 2.67GHz

Extra Large Memory Nodes with GPU (see table below for more details about GPU)

48

384

4

(n020-n023)

12

96

8

Intel(R) Xeon(R) CPU X5650 @ 2.67GHz
NVIDIA Tesla C2050

Extra Large Memory Nodes (no GPU)

96

768

8

(n031-n038)

12

96

8

Intel(R) Xeon(R) CPU X5650 @ 2.67GHz

Special Nodes (no GPU)

64

128

16

(n039-n042)

16

32

2

Intel(R) Xeon(R) CPU E5-2670 0 @ 2.60GHz

Special Large Memory Nodes (no GPU)

64

1024

8

(n044-n047)

16

256

16

Intel(R) Xeon(R) CPU E5-2670 0 @ 2.60GHz

Special Large Memory Nodes (no GPU)

192

1536

24

(aspen01-aspen12)

16

128

8

Intel(R) Xeon(R) CPU E5-2670 0 @ 2.60GHz

Please note that each of the Extra Large nodes (n020,n021, n022 and n023) have 2 nvidia tesla C-2050 GPU cards.

Node Type

Programming Model

Total Number of CUDA Cores

Total Amount of Memory (GB)

Compute Nodes

CUDA Cores Per node

CUDA cards per node

Mem per Node (GB)

Memory per Core

Processor Type

Extra Large Memory Nodes with GPU

GPU

4X2X448

384

4

(n020-n023)

2X448

2

96

8

Intel(R) Xeon(R) CPU X5650 @ 2.67GHz
NVIDIA Tesla C2050

Special Administrative Nodes (Not used for computing purposes)

Node Type

Node Name

Total Number of Cores

Total Amount of Memory (GB)

Mem per Node (GB)

Memory per Core

Processor Type

File Servers

n024,n025

24

96G

48GB

4GB

Intel(R) Xeon(R) CPU X5650 @ 2.67GHz

Test Node

testhpc

12

24G

24GB

2GB

Intel(R) Xeon(R) CPU X5650 @ 2.67GHz

login Node

gowonda

12

48G

48GB

4GB

Intel(R) Xeon(R) CPU X5650 @ 2.67GHz

Admin Node

n030 (admin)

12

24G

24GB

2GB

Intel(R) Xeon(R) CPU X5650 @ 2.67GHz

More information about the Intel(R) Xeon(R) CPU X5650 processor can be obtained here.

In addition to the above, there is a special windows HPC node.

Node type

Total Number of Cores

Total Amount of Memory (GB)

Compute Nodes

Cores Per node

Mem per Node (GB)

Memory per Core

Processor Type

OS

Windows 2008 Large Memory Node

12

48

1

n029

12

48

4

Intel(R) Xeon(R) CPU X5650 @ 2.67GHz

Windows 2008 R2 with windows HPC pack

Instructions for using the Windows HPC is given in a separate user guide

2.3 Software . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

The operating system on gowonda is RedHat Enterprise Linux (RHEL) 6.1 updated with SGI foundation-2.4 suite and accelerate-1.2 support package. The queuing system used is PBS Pro 11 . The exception is the Windows HPC node which runs Windows 2008 R2 with Windows HPC pack and Windows HPCS job scheduler. Gowonda has the following compilers and parallel library software. Much more detail on each can be found below.

  • GNU C, C++ and Fortran compilers;

  • Portland Group, Inc. optimizing C, C++, and Fortran compilers (currently awaiting to be installed);

  • The Intel Cluster Studio including the Intel C, C++ and Fortran compilers, Math and Kernel Library;

  • Intel MPI , SGI's proprietary MPT and OpenMPI

  • Oracle Solaris Studio Compiler (formerly called Sun Studio Compiler)

The following third party applications are currently installed or will be installed shortly. The Gowonda HPC Center staff will be happy to work with any user interested in installing additional applications, subject to meeting that application's license requirements.

Software

Version

Usage

Status

AutoDOCK

4.2.3

Module load autodock423 autodockvina112

Installed

Bioperl

 

 

TBI (To be Installed)

Blast

 

 

Installed

CUDA

4.0

module load cuda/4.0

Installed)

Gaussian03

 

module load gaussian/g03

Installed

Gaussian09

 

module load gaussian/g09

Installed

Gromacs

 

 

Installed

gromos

1.0.0

module load gromos/1.0.0

Installed

MATLAB

2009b,2011a

module load matlab/2009b,module load matlab/2011a

Installed

MrBayes

 

 

TBI (To be Installed)

NAMD

 

module load NAMD/NAMD28b1

Installed

numpy

1.5.1

module load python/2.7.1

Installed

PyCogent

-

module load python/2.7.1

Installed

qiime

 

 

To be Installed

R

 

module load R/2.13.0

Installed

SciPy

0.9.0

module load python/2.7.1

Installed

VASP

-

-

TBI

The following graphics, IO, and scientific libraries are also supported.

Software

Version

Usage

Status

Atlas

3.9.39

module load ATLAS/3.9.39

Installed

FFTW

3.2.2.,3.3a

module load fftw/3.3-alpha-intel

Installed

GSL-

1.09,1.15

module load module load gsl/gsl-1.15

Installed

LAPACK-

3.3.0

-

-

NETCDF-

3.6.2,3.6.3,4.0,4.1.1,4.1.2

e.g. module load NetCDF/4.1.2

Installed

3 Support

3.1 Hours of Operation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . ..

The fourth Thursday mornings in the month from 8:00AM to 12PM are normally reserved (but not always used) for scheduled maintenance. Please plan accordingly. Unplanned maintenance to remedy system related problems may be scheduled as needed. Reasonable attempts will be made to inform users running on those systems when these needs arise.

3.2 User Support . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

Users are encouraged to read this Wiki carefully. In particular, the sections on compiling and running parallel programs, and the section on the PBS Pro batch queuing system will give you the essential knowledge needed to use the gowonda cluster.

Gowonda cluster staff along with outside vendors also offer some courses to the Griffith HPC community in parallel programming techniques, HPC computing architecture, and the essentials of using our systems. Please follow our mailings on the subject and feel free to inquire about such courses. We can schedule training visits and classes at the various Griffith campuses. Please let us know if such are training visit is of interest.

Users with further questions or requiring immediate assistance in use of the systems should submit a ticket here

 

Support staff may be contacted via:
Griffith Library and IT help 3735 5555 or X55555
email support: Submit form 
You can log cases on service desk (category: eResearch services.HPC)

eResearch Services, Griffith University
Phone: +61 - 7 - 373 56649 (GMT +10 Hours)
Email:  Submit a Question to the HPC Support Team
Web: griffith.edu.au/eresearch-services

3.3 Service Alerts . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .

Any major disruptions/events will be communicated through emails. These emails will be logged on the following web pages as well.