G P U Cluster Service

Create a GPU cluster

POST
/compute/clusters

Create an Instant Cluster on Together's high-performance GPU clusters. With features like on-demand scaling, long-lived resizable high-bandwidth shared DC-local storage, Kubernetes and Slurm cluster flavors, a REST API, and Terraform support, you can run workloads flexibly without complex infrastructure management.

Authorization

bearerAuth
AuthorizationBearer <token>

In: header

Request Body

application/json

GPU Cluster create request

GPU Cluster create request

cluster_type?string

Type of cluster to create.

Value in

  • "KUBERNETES"
  • "SLURM"
region?string

Region to create the GPU cluster in. Usable regions can be found from client.clusters.list_regions()

gpu_type?string

Type of GPU to use in the cluster

Value in

  • "H100_SXM"
  • "H200_SXM"
  • "RTX_6000_PCI"
  • "L40_PCIE"
  • "B200_SXM"
  • "H100_SXM_INF"
  • "B300_SXM"
num_gpus?integer

Number of GPUs to allocate in the cluster. This must be multiple of 8. For example, 8, 16 or 24

cluster_name?string

Name of the GPU cluster.

duration_days?integer

Duration in days to keep the cluster running.

shared_volume?

Inline configuration to create a shared volume with the cluster creation.

volume_id?string

ID of an existing volume to use with the cluster creation.

billing_type?string

RESERVED billing types allow you to specify the duration of the cluster reservation via the duration_days field. ON_DEMAND billing types will give you ownership of the cluster until you delete it. SCHEDULED_CAPACITY billing types allow you to reserve capacity for a scheduled time window. You must specify the reservation_start_time and reservation_end_time with this request.

Value in

  • "RESERVED"
  • "ON_DEMAND"
  • "SCHEDULED_CAPACITY"
auto_scaled?boolean
Deprecated

Whether GPU cluster should be auto-scaled based on the workload. By default, it is not auto-scaled.

Defaultfalse
auto_scale_max_gpus?integer

Maximum number of GPUs to which the cluster can be auto-scaled up. This field is required if auto_scaled is true.

slurm_shm_size_gib?integer

Shared memory size in GiB for Slurm cluster. This field is required if cluster_type is SLURM.

capacity_pool_id?string

ID of the capacity pool to use for the cluster. This field is optional and only applicable if the cluster is created from a capacity pool.

reservation_start_time?string

Reservation start time of the cluster. This field is required for SCHEDULED billing to specify the reservation start time for the cluster. If not provided, the cluster provisions immediately.

Formatdate-time
reservation_end_time?string

Reservation end time of the cluster. This field is required for SCHEDULED billing to specify the reservation end time for the cluster.

Formatdate-time
install_traefik?boolean

Whether to install Traefik ingress controller in the cluster. This field is only applicable for Kubernetes clusters and is false by default.

Defaultfalse
cuda_version?string

Legacy CUDA selector for this cluster. Bare semantic values such as 12.5 select ubuntu-22.04; existing OS-suffixed values remain accepted for compatibility. Must be paired with nvidia_driver_version. Prefer nvidia_version_id for new integrations.

nvidia_driver_version?string

Legacy NVIDIA driver selector for this cluster. For example, 550. Must be paired with cuda_version. Prefer nvidia_version_id for new integrations.

nvidia_version_id*string

Canonical region-specific NVIDIA version ID. If cuda_version and nvidia_driver_version are also set, they must resolve to the same catalog entry.

slurm_image?string

Custom Slurm image for Slurm clusters.

oidc_config?
project_id?string

Project ID for the cluster. If not set, the project from the request context is used.

acceptance_tests_params?

AcceptanceTestsParams groups all GPU acceptance test options when enabled is true.

cluster_config?
num_capacity_pool_gpus?integer

Number of GPUs to allocate from the capacity pool. Must be a multiple of 8 and not exceed num_gpus.

auto_scale?boolean

Whether to enable auto-scaling for the cluster. If true, the cluster will automatically scale the number of GPU worker nodes between num_gpus and auto_scale_max_gpus based on the workload.

num_preemptible_gpus?integer

Number of preemptible GPUs to request alongside on-demand capacity. Must be a multiple of 8. Preemptible nodes are cheaper but may be reclaimed when on-demand capacity is needed elsewhere; the system fulfills this asynchronously and surfaces the actual count in allocated_preemptible_gpus.

num_reserved_gpus?integer

Number of prepaid (PLG) reserved GPUs for this cluster. When omitted for RESERVED billing on create, the server defaults this to num_gpus.

add_ons?array<>

Add-ons to enable on the cluster at creation time.

Response Body

application/json

curl -X POST "https://example.com/compute/clusters" \  -H "Content-Type: application/json" \  -d '{    "region": "string",    "gpu_type": "H100_SXM",    "num_gpus": 0,    "cluster_name": "string",    "billing_type": "RESERVED"  }'
{  "cluster_id": "string",  "cluster_type": "KUBERNETES",  "region": "string",  "gpu_type": "H100_SXM",  "cluster_name": "string",  "duration_hours": 0,  "volumes": [    {      "volume_id": "string",      "volume_name": "string",      "size_tib": 0,      "status": "string"    }  ],  "status": "WaitingForControlPlaneNodes",  "control_plane_nodes": [    {      "node_id": "string",      "status": "string",      "host_name": "string",      "num_cpu_cores": 0,      "memory_gib": 0,      "network": "string",      "phase_transitions": [        {          "phase": "NODE_PHASE_PENDING",          "transition_time": "2019-08-24T14:15:22Z"        }      ],      "public_ipv4": "string"    }  ],  "gpu_worker_nodes": [    {      "node_id": "string",      "status": "string",      "host_name": "string",      "num_cpu_cores": 0,      "num_gpus": 0,      "memory_gib": 0,      "networks": [        "string"      ],      "instance_id": "string",      "latest_remediation": {        "id": "string",        "cluster_id": "string",        "instance_id": "string",        "mode": "REMEDIATION_MODE_VM_ONLY",        "trigger": "REMEDIATION_TRIGGER_MANUAL",        "state": "PENDING_APPROVAL",        "reason": "string",        "active_health_check_run_id": "string",        "passive_health_check_event_id": "string",        "requested_by": "string",        "create_time": "2019-08-24T14:15:22Z",        "reviewed_by": "string",        "review_time": "2019-08-24T14:15:22Z",        "review_comment": "string",        "start_time": "2019-08-24T14:15:22Z",        "end_time": "2019-08-24T14:15:22Z",        "error_message": "string",        "update_time": "2019-08-24T14:15:22Z",        "instance_name": "string",        "linked_alerts": [          {            "passive_health_check_alert_id": "string",            "instance_id": "string",            "cluster_id": "string",            "target_vm": "string",            "alert_name": "string",            "severity": "PHC_SEVERITY_INFO",            "annotations": {              "property1": "string",              "property2": "string"            },            "started_at": "2019-08-24T14:15:22Z",            "resolved_at": "2019-08-24T14:15:22Z",            "node_remediation_intent_id": "string",            "annotation": {              "title": "string",              "description": "string",              "summary_line": "string",              "xid": {                "events": [                  {                    "xid_code": "string",                    "mnemonic": "string",                    "count": 0                  }                ]              },              "slurm_node_unavailable": {                "reason": "string"              }            }          }        ]      },      "slurm_worker_hostname": "string",      "phase_transitions": [        {          "phase": "NODE_PHASE_PENDING",          "transition_time": "2019-08-24T14:15:22Z"        }      ],      "marked_for_deletion": true,      "public_ipv4": "string",      "ib_hca_type": "string",      "ib_hca_count": 0,      "nvswitch_count": 0,      "nvswitch_type": "string",      "ephemeral_storage": "string",      "auto_remediation_enabled": true,      "deleted_at": "2019-08-24T14:15:22Z"    }  ],  "kube_config": "string",  "num_gpus": 0,  "slurm_shm_size_gib": 0,  "capacity_pool_id": "string",  "reservation_start_time": "2019-08-24T14:15:22Z",  "reservation_end_time": "2019-08-24T14:15:22Z",  "install_traefik": true,  "cuda_version": "string",  "nvidia_driver_version": "string",  "created_at": "2019-08-24T14:15:22Z",  "oidc_config": {    "issuer_url": "string",    "client_id": "string",    "username_claim": "string",    "username_prefix": "string",    "group_claim": "string",    "group_prefix": "string",    "ca_cert": "string"  },  "project_id": "string",  "cluster_config": {    "load_balancer": "NONE",    "kubernetes_dashboard_enabled": true,    "jumphost_enabled": true,    "slurm_startup_scripts": {      "worker_prolog": "string",      "worker_epilog": "string",      "controller_prolog": "string",      "controller_epilog": "string",      "login_init_script": "string",      "nodeset_init_script": "string",      "extra_slurm_conf": "string"    },    "ingress": {      "enabled": true    },    "observability": {      "enabled": true    },    "gpu_operator_version": "string",    "network_operator_version": "string",    "ssh_ca_enabled": true  },  "num_cpu_workers": 0,  "phase_transitions": [    {      "phase": "CLUSTER_PHASE_QUEUED",      "transition_time": "2019-08-24T14:15:22Z"    }  ],  "desired_preemptible_gpus": 0,  "allocated_preemptible_gpus": 0,  "billing_type": "RESERVED",  "add_ons": [    {      "name": "string",      "add_on_type": "string",      "config": {        "dashboard": {          "enabled": true        },        "ingress": {          "enabled": true        },        "torchpass": {          "enabled": true        },        "slurm_web": {          "enabled": true        },        "headlamp": {          "enabled": true        }      },      "state": {        "dashboard": {},        "ingress": {},        "torchpass": {},        "slurm_web": {},        "headlamp": {}      }    }  ],  "machine_cluster_id": "string",  "first_ready_at": "2019-08-24T14:15:22Z",  "is_in_substrate": true,  "control_plane_ready": true,  "ums_project_id": "string",  "ums_org_id": "string",  "os_image": "string",  "nvidia_driver_version_id": "string",  "num_capacity_pool_gpus": 0,  "num_reserved_gpus": 0,  "deleted_gpu_worker_nodes": [    {      "node_id": "string",      "status": "string",      "host_name": "string",      "num_cpu_cores": 0,      "num_gpus": 0,      "memory_gib": 0,      "networks": [        "string"      ],      "instance_id": "string",      "latest_remediation": {        "id": "string",        "cluster_id": "string",        "instance_id": "string",        "mode": "REMEDIATION_MODE_VM_ONLY",        "trigger": "REMEDIATION_TRIGGER_MANUAL",        "state": "PENDING_APPROVAL",        "reason": "string",        "active_health_check_run_id": "string",        "passive_health_check_event_id": "string",        "requested_by": "string",        "create_time": "2019-08-24T14:15:22Z",        "reviewed_by": "string",        "review_time": "2019-08-24T14:15:22Z",        "review_comment": "string",        "start_time": "2019-08-24T14:15:22Z",        "end_time": "2019-08-24T14:15:22Z",        "error_message": "string",        "update_time": "2019-08-24T14:15:22Z",        "instance_name": "string",        "linked_alerts": [          {            "passive_health_check_alert_id": "string",            "instance_id": "string",            "cluster_id": "string",            "target_vm": "string",            "alert_name": "string",            "severity": "PHC_SEVERITY_INFO",            "annotations": {              "property1": "string",              "property2": "string"            },            "started_at": "2019-08-24T14:15:22Z",            "resolved_at": "2019-08-24T14:15:22Z",            "node_remediation_intent_id": "string",            "annotation": {              "title": "string",              "description": "string",              "summary_line": "string",              "xid": {                "events": [                  {                    "xid_code": "string",                    "mnemonic": "string",                    "count": 0                  }                ]              },              "slurm_node_unavailable": {                "reason": "string"              }            }          }        ]      },      "slurm_worker_hostname": "string",      "phase_transitions": [        {          "phase": "NODE_PHASE_PENDING",          "transition_time": "2019-08-24T14:15:22Z"        }      ],      "marked_for_deletion": true,      "public_ipv4": "string",      "ib_hca_type": "string",      "ib_hca_count": 0,      "nvswitch_count": 0,      "nvswitch_type": "string",      "ephemeral_storage": "string",      "auto_remediation_enabled": true,      "deleted_at": "2019-08-24T14:15:22Z"    }  ],  "node_lifecycle_events": [    {      "node_id": "string",      "reason": "string",      "message": "string",      "timestamp": "2019-08-24T14:15:22Z"    }  ]}