Adding a VM to a GPU cluster
In GPU clusters, you can create VMs with 8 GPUs based on one of the following platforms:
- AMD EPYC™ with NVIDIA® Ampere® A100 (
gpu-standard-v3). - Gen2 (
gpu-standard-v3i). - GPU PLATFORM V4 (
gpu-standard-v4).
Such VMs must be deployed from a dedicated image with NVIDIA drivers.
Note
You can host your GPU cluster in one of these availability zones: ru-central1-a, ru-central1-b, and ru-central1-d. The VM must be created within the same availability zone as the cluster.
-
In the management console
, select the folder where you want to create the VM. -
Navigate to Compute Cloud.
-
In the left-hand panel, select
Virtual machines and click Create virtual machine. -
Under Boot disk image, select an image with pre-installed NVIDIA drivers.
-
In the Availability zone field, select the availability zone where your GPU cluster resides.
-
Under Computing resources, navigate to the Custom tab and specify:
-
Under General information, specify the VM name.
-
Click Create VM.
If you do not have the Yandex Cloud CLI yet, install and initialize it.
The folder used by default is the one specified when creating the CLI profile. To change the default folder, use the yc config set folder-id <folder_ID> command. You can also specify a different folder for any command using --folder-name or --folder-id. If you access a resource by its name, the search will be limited to the default folder. If you access a resource by its ID, the search will be global, i.e., through all folders based on access permissions.
In the terminal, run this command:
export YC_GPU_CLUSTER=$(yc compute gpu-cluster list --format=json | jq -r .[].id)
export YC_ZONE="ru-central1-a"
export SUBNET_NAME="my-subnet-name"
export SUBNET_ID=$(yc vpc subnet get --name=$SUBNET_NAME --format=json | jq -r .id)
yc compute instance create \
--name node-gpu-test \
--create-boot-disk size=64G,image-id=<image_ID_with_drivers>,type=network-ssd \
--ssh-key=$HOME/.ssh/id_rsa.pub \
--gpus 8 \
--cores 224 \
--memory=952G \
--zone $YC_ZONE \
--network-interface subnet-id=$SUBNET_ID,nat-ip-version=ipv4 \
--platform gpu-standard-v3 \
--gpu-cluster-id=$YC_GPU_CLUSTER
Where:
--name: VM name.--create-boot-disk: VM disk properties.--ssh-key: Path to the file with the public SSH key.--gpus: Number of GPUs.--cores: Number of vCPUs.--memory: Amount of RAM.--zone: Availability zone.--network-interface: VM network interface settings.--platform: Platform ID.--gpu-cluster-id: GPU cluster ID.
If you do not have Terraform yet, install it and configure the Yandex Cloud provider.
To manage infrastructure using Terraform under a service account or user accounts (a Yandex account, a federated account, or a local user), authenticate using the appropriate method.
-
In the Terraform configuration file, describe the resource you want to create:
provider "yandex" { zone = "ru-central1-a" } resource "yandex_compute_disk" "boot-disk" { name = "<disk_name>" type = "<disk_type>" zone = "ru-central1-a" size = "<disk_size>" image_id = "<image_ID_with_drivers>" } resource "yandex_compute_instance" "default" { name = "vm-gpu" platform_id = "gpu-standard-v3" zone = "ru-central1-a" gpu_cluster_id = "<GPU_cluster_ID>" resources { cores = "224" memory = "952" gpus = "8" } boot_disk { disk_id = yandex_compute_disk.boot-disk.id } network_interface { subnet_id = "${yandex_vpc_subnet.subnet-1.id}" nat = true } metadata = { user-data = "#cloud-config\nusers:\n - name: <username>\n groups: sudo\n shell: /bin/bash\n sudo: 'ALL=(ALL) NOPASSWD:ALL'\n ssh_authorized_keys:\n - ${file("<path_to_public_SSH_key>")}" } } resource "yandex_vpc_network" "network-1" { name = "network1" } resource "yandex_vpc_subnet" "subnet-1" { name = "subnet1" zone = "<availability_zone>" v4_cidr_blocks = ["192.168.10.0/24"] network_id = "${yandex_vpc_network.network-1.id}" }Where:
-
yandex_compute_disk: Boot disk description, whereimage_idis the ID of the image with the drivers. -
gpu_cluster_id: GPU cluster ID. This is a required setting. -
yandex_vpc_network: Cloud network description. -
yandex_vpc_subnet: Description of the subnet to create your VM in.Note
If you already have suitable resources, such as a cloud network and subnet, you do not need to redefine them. Specify their names and IDs in the appropriate parameters.
For more information about
yandex_compute_instanceproperties, see this Terraform provider guide.
-
-
Under
metadata, specify your username and path to the public SSH key. For more information, see VM metadata. -
Create the resources:
-
In the terminal, navigate to the configuration file directory.
-
Make sure the configuration is correct using this command:
terraform validateIf the configuration is valid, you will get this message:
Success! The configuration is valid. -
Run this command:
terraform planYou will see a list of resources and their properties. No changes will be made at this step. Terraform will show any errors in the configuration.
-
Apply the configuration changes:
terraform apply -
Type
yesand press Enter to confirm the changes.
-
This will create a VM in the specified GPU cluster. You can check the new VM and its configuration using the management console
yc compute instance get <VM_name>
To create a VM in a GPU cluster, use the create REST API method for the Instance resource or the InstanceService/Create gRPC API call. In the request body, specify the GPU cluster ID in the gpuClusterId field.