CARL
contact
- Please contact us, preferably by email, at one of the addresses listed below. Please contact our support address in the first instance.
- If you would prefer to contact us by telephone, please feel free to call us on the numbers below, or arrange a call via email.
- Our usual office hours are between 9.00 am and 5.30 pm. If we are unavailable, you can always reach us by email
HPC Support
Address
CARL
CARL (named after Carl von Ossietzky)
CARL, funded by the German Research Foundation (DFG) and the Ministry of Science and Culture (MWK) of the State of Lower Saxony, is a multi-purpose cluster designed to meet the needs of compute-intensive and data-driven research projects in the main areas of
- Quantum Chemistry and Quantum Dynamics,
- Theoretical Physics,
- the neurosciences (including hearing research),
- Oceanic and Marine Research,
- biodiversity, and
- Computer Science
Like its sister cluster EDDY, CARL is operated by the IT Services of the University of Oldenburg. The system is used by more than 20 research groups from the Faculty of Mathematics and Science, and a couple of research groups from the Department of Computing Science within the School of Computing Science, Business Administration, Economics and Law.
Hardware Overview
- 327 compute nodes (7,640 CPU cores, 77 TB of main memory (RAM), 271 TFlop/s theoretical peak)
- 158 ‘standard’ nodes
- LENOVO NeXtScale nx360 M5 HPC Node (Model 5465-FT1)
- 2x Intel Xeon CPU E5-2650 v4 12-core at 2.2 GHz
- 256 GB main memory (RAM)
- 1 TB 7.2K 6 Gbps HDD (used for local storage)
- 128 "low-memory" nodes
- LENOVO NeXtScale nx360 M5 HPC Node (Model 5465-FT1)
- 2x Intel Xeon E5-2650 v4 12-core CPUs at 2.2 GHz
- 128 GB main memory (RAM)
- 1 TB 7.2K 6 Gbps HDD (used for local storage)
- 30 "high-memory" nodes
- LENOVO System x3650 M5 Node (Model 8871-FT)
- 2x Intel Xeon E5-2667 v4 8-core CPUs at 3.2 GHz
- 512 GB main memory (RAM)
- 2 ‘pre- and post-processing’ nodes
- Lenovo System x3850 X6 Node (Model 6241-FT1)
- 4x Intel Xeon E7-8891 v4 10-core CPUs at 2.8 GHz
- 2048 GB main memory (RAM)
- 9 ‘GPU’ nodes
- LENOVO NeXtScale nx360 M5 HPC Node (Model 5465-FT1)
- 2x Intel Xeon E5-2650 v4 12-core CPUs at 2.2 GHz
- 256 GB main memory (RAM)
- 1 TB 7.2K 6 Gbps HDD (used for local storage)
- NVIDIA GPU
- 158 ‘standard’ nodes
- Management and Login Nodes
- 2 administration nodes in an active/passive high-availability (HA) configuration
- LENOVO System x3550 M5 Node Model 8869-FT1
- 2x Intel Xeon CPU E5-2650 v4 12-core at 2.2GHz
- 256 GB main memory (RAM) @ 2400 MHz
- 4x 1.2TB 7.2K 12Gbps SAS HDD (configured as a hardware RAID 10)
- Connect-IB single-port card
The master nodes are shared between CARL and its sister cluster EDDY and run all vital cluster services (node provisioning, DHCP, DNS, LDAP, NFS, Job Management System, etc.). They also provide monitoring functions for both clusters (with automated alerting). Monitoring covers hardware components (health status of all servers, temperature, power consumption, etc.) as well as basic cluster services (with automated restart should a service fail)
- 2 login nodes for user access to the system, software development (programming environment), and job submission and control. These nodes have the same specifications as the administration nodes described above.
- 2 administration nodes in an active/passive high-availability (HA) configuration
- Internal networks
- An InfiniBand network comprising 2 spine switches and 11 leaf switches with an 8:1 blocking factor. Consequently, each leaf switch is connected to 32 MPC nodes. The maximum data transfer rate is 56.250 Gb/s (4x FDR).
- Secondly, a physically separate Gigabit Ethernet network (‘base network’) for vital cluster services (node provisioning, DHCP, DNS, LDAP, NFS, Job Management System, etc.)
- A 10 Gb Ethernet backbone network connecting the management and login nodes, the storage system, and the Gigabit Ethernet (MPI and base network) leaf switches
- An IPMI network for hardware monitoring and control, including access to the VGA console (KVM functionality), enabling full remote management of the cluster
- Storage System
- General Parallel File System (GPFS) with 1.392 PB of total storage space, of which approximately 926 TB is usable. We are using four declustered arrays to ensure the availability of our data. The failure of a single hard drive goes unnoticed; even if two hard drives fail, the ‘critical rebuild’ would take only about 45 minutes. The high-memory and pre- and post-processing nodes have additional local storage space (up to 1 TB).
- Enterprise-class scalable NAS cluster (manufacturer: EMC Isilon). This is where the home directories are stored. All data saved on the Isilon is backed up, and it is possible to work with snapshots. The data is accessible via a 10GbE connection.
The Isilon serves as the central storage for IT Services, which is why it is also used for HPC. Disk space is allocated to the two clusters depending on the proportion of the storage system’s hardware that was funded by the FLOW and HERO project funds, respectively.
SYSTEM SOFTWARE AND MIDDLEWARE
- All cluster nodes are running Red Hat Enterprise Linux (RHEL)
- Cluster management: Bright Cluster Manager
- Commercial compilers, debuggers and profilers, and performance libraries: Intel Cluster Studio, PGI
- Workload management: SLURM
SELECTED APPLICATIONS RUNNING ON CARL
(Due to licensing restrictions, some applications are only accessible to specific users or research groups.)
- Quantum chemistry packages: Gaussian 09, MOLCAS, MOLPRO, VASP
- LEDA – a C++ class library for efficient data types and algorithms
- MATLAB, including the Parallel Computing Toolbox
- FVCOM – The Unstructured Grid Finite Volume Coastal Ocean Model
PICTURES


More pictures can be found here.