HPC Cluster Footprint
This page provides the active hardware footprint, configuration details, and capacity metrics for the High Performance Computing (HPC) clusters managed by Columbia University Information Technology (CUIT).
Insomnia Cluster Footprint
Insomnia went live in February 2024 and initially was a joint purchase by 21 research groups and departments. Unlike its predecessors, this new high performance computing design allows researchers to buy not only a node but half or even a quarter of a node on the cluster.
Insomnia is faculty-governed by the cross-disciplinary SRCPAC and is administered and supported by CUIT’s High Performance Computing team.
Insomnia is a new type of design intended to allow indefinite expansion to Columbia's shared high performance computing cluster, adding new hardware and capabilities as needed. For Insomnia, the nodes are expected to remain in production for the duration of their operational lifecycle and will only be decommissioned due to hardware failure, end-of-life status, or data center capacity constraints related to power, cooling, or available rack space.
All of Insomnia's high performance computing servers are equipped with Dual Intel Xeon Platinum 8460Y processors (2 GHz):
90
Active Compute Nodes
142
Individual Accelerators
2.282PB
GPFS Filesystem
HDR Infiniband
RHEL 9.x
Slurm Scheduler
Node Type
- Standard Nodes
Number of Nodes
- 41
Node Type
- High Memory Nodes
Number of Nodes
- 19
Node Type
- GPU Nodes
Number of Nodes
- 30
GPU Node Type
- NVIDIA A6000 8 GPU modules per server
Number of Nodes
- 13
GPU Node Type
- NVIDIA A6000 4 GPU modules per server
Number of Nodes
- 2
GPU Node Type
- NVIDIA H100 2 GPU modules per server
Number of Nodes
- 3
GPU Node Type
- NVIDIA L40 2 GPU modules per server
Number of Nodes
- 3
GPU Node Type
- NVIDIA L40s 2 GPU modules per server
Number of Nodes
- 9
Total Cluster Memory:
~58.3 TiB
Total Physical Cores:
7,144
Total Logical CPUs:
14,288
Includes hyperthreading logical configuration metrics.
GPU Comparison Table
- GPU Type
- NVIDIA RTX PRO 6000
- VRAM
- 48GB GDDR6
- Best Use Case
- Visualization, AI inference, rendering, and medium-sized LLMs
- Strengths
- Excellent graphics and visualization performance,
efficient power usage, and strong mixed-workload
support - Considerations
- Lower AI training performance
compared to H100/H200
- GPU Type
- NVIDIA L40
- VRAM
- 48GB GDDR6
- Best Use Case
- AI inference, graphics, Omniverse, VDI, and rendering
- Strengths
- Balanced GPU for AI and visualization workloads
with strong inference performance - Considerations
- Not optimized for very large-scale AI
training
- GPU Type
- NVIDIA L40S
- VRAM
- 48GB GDDR6
- Best Use Case
- Generative AI, inference, fine-tuning, and mixed HPC/AI workloads
- Strengths
- Higher AI and tensor performance than L40 with
strong inference throughput - Considerations
- Still lower raw training performance
than H100/H200
- GPU Type
- NVIDIA RTX PRO 4500
- VRAM
- 32GB GDDR7
- Best Use Case
- Generative AI inference, fine-tuning, rendering, and mixed professional workloads
- Strengths
- Blackwell architecture with next-gen tensor cores, faster GDDR7 memory (896 GB/s bandwidth), and strong AI inference throughput — effective L40S successor
- Considerations
- Lower VRAM (32GB vs 48GB) than L40s; less suited for very large model training
- GPU Type
- NVIDIA H100
- VRAM
- 80GB GDDR6
- Best Use Case
- Large-scale AI training, HPC simulations, and distributed deep learning
- Strengths
- Extremely high tensor performance, NVLink support,
and optimized for AI/HPC - Considerations
- Higher cost and demand
- GPU Type
- NVIDIA H200
- VRAM
- 141GB HBM3e
- Best Use Case
- Very large AI models, memory-intensive HPC workloads, and LLM fine-tuning
- Strengths
- Larger memory capacity and bandwidth than H100,
ideal for memory-heavy workloads - Considerations
- Premium pricing and availability
constraints
Other & Retired Clusters
Ginsburg high performance computing went live in February 2021 and is a joint purchase by 33 research groups and departments.
The cluster is faculty-governed by the cross-disciplinary SRCPAC and is administered and supported by CUIT’s High Performance Computing team.
Tentative retirement dates
Ginsburg Phase 1 retirement: February 2026COMPLETED- Ginsburg Phase 2 retirement: March 2027
- Ginsburg Phase 3 retirement: December 2027
Specifications
286 nodes with a total of 9152 cores (32 cores per node):
All of Ginburg's high performance computing servers are equipped with Dual Intel Xeon Gold 6226R processors (2.9 GHz):
- 191 Standard Nodes (192 GB)
- 56 High Memory Nodes (768 GB)
- 18 NVIDIA RTX 8000 GPU nodes (2 GPUs modules per server)
- 4 NVIDIA V100S GPU nodes (2 GPU modules per server)
- 8 NVIDIA A100 GPU Nodes (2 GPU modules per server)
- 9 NVIDIA A40 GPU Nodes (2 GPU modules per server)
- 1PB of DDN ES7790 Lustre storage
- HDR Infiniband
- Red Hat Enterprise Linux 8
- Slurm job scheduler
The Terremoto high-performance computing cluster was officially retired in March 2025. However, it is currently operating in a reactivated state for specific user groups. PI's who previously invested in Terremoto have the opportunity to reactivate their retired nodes. These nodes will remain available for use until they experience hardware failure or the Data Center requires the physical space to accommodate new incoming clusters.
Specifications
86 reactivated nodes with a total of 2064 cores:
- 68 Standard Nodes (192 GB)
- 8 High Memory Nodes (754 GB)
- 10 NVIDIA V100 GPU Nodes (2 GPU modules per server)
- NFS Storage
- HDR Infiniband
- Red Hat Enterprise Linux 9
- Slurm job scheduler
Habanero Shared HPC Cluster
The Habanero high performance computing cluster, retired in December 2023, was located in the Shared Research Computing Facility (SRCF), a dedicated portion of the university data center on the Morningside campus.
Yeti Shared HPC Cluster
The Yeti high performance computing cluster retired in 2019, was located in the Shared Research Computing Facility (SRCF), a dedicated portion of the university data center on the Morningside campus.
