Deployments

Create a new deployment

POST
/deployments

Create a new deployment with specified configuration

Authorization

bearerAuth
AuthorizationBearer <token>

In: header

Request Body

application/json

Deployment configuration

args?array<string>

Args overrides the container's CMD. Provide as an array of arguments (e.g., ["python", "app.py"])

autoscaling?||

Autoscaling configuration. Example: {"metric": "QueueBacklogPerWorker", "target": 1.01} to scale based on queue backlog. Omit or set to null to disable autoscaling

capacity_type?string

Controls how replicas above reserved capacity behave. stable replicas stay running after scale-up; preemptible replicas may be evicted during capacity contention.

Value in

  • "stable"
  • "preemptible"
command?array<string>

Command overrides the container's ENTRYPOINT. Provide as an array (e.g., ["/bin/sh", "-c"])

cpu?number

CPU is the number of CPU cores to allocate per container instance (e.g., 0.1 = 100 milli cores)

Range0.1 <= value
description?string

Description is an optional human-readable description of your deployment

environment_variables?array<>

EnvironmentVariables is a list of environment variables to set in the container. Each must have a name and either a value or value_from_secret

gpu_count?integer

GPUCount is the number of GPUs to allocate per container instance. Defaults to 0 if not specified

gpu_type*string

GPUType specifies the GPU hardware to use (e.g., "h100-80gb").

Value in

  • "h100-80gb"
  • "h100-40gb-mig"
  • "h200-140gb"
  • "b200-192gb"
health_check_path?string

HealthCheckPath is the HTTP path for health checks (e.g., "/health"). If set, the platform checks this endpoint to determine container health.

image*string

Image is the container image to deploy from registry.together.ai.

max_replicas?integer

MaxReplicas is the maximum number of container instances. Defaults to MinReplicas if not set.

memory?number

Memory is the amount of RAM to allocate per container instance in GiB (e.g., 0.5 = 512MiB)

Rangevalue <= 1000
min_replicas?integer

MinReplicas is the minimum number of container instances to run. Defaults to 1 if not specified

model_mounts?array<>

Model weights to preload from Together's model registry into the container. At most one mount is supported, and it cannot be used with volumes.

name*string

Name is the unique identifier for your deployment. Must contain lowercase letters, numbers, or hyphens, start with a lowercase letter or number, and be 4-63 characters. It cannot be changed.

Length4 <= length <= 63
port?integer

Port is the container port your application listens on (e.g., 8080 for web servers). Required if your application serves traffic

Range1 <= value <= 65535
storage?integer

Storage is the amount of ephemeral disk storage to allocate per container instance (e.g., 10 = 10GiB)

Rangevalue <= 400
termination_grace_period_seconds?integer

TerminationGracePeriodSeconds is the time in seconds to wait for graceful shutdown before forcefully terminating the replica

volumes?array<>

Volumes is a list of volume mounts to attach to the container. Each mount must reference an existing volume by name

Response Body

application/json

application/json

application/json

curl -X POST "https://example.com/deployments" \  -H "Content-Type: application/json" \  -d '{    "gpu_type": "h100-80gb",    "image": "string",    "name": "string"  }'
{  "args": [    "string"  ],  "autoscaling": {    "metric": "HTTPTotalRequests",    "target": 0,    "time_interval_minutes": 0  },  "capacity_type": "stable",  "command": [    "string"  ],  "cpu": 0,  "created_at": "2019-08-24T14:15:22Z",  "description": "string",  "desired_replicas": 0,  "environment_variables": [    {      "name": "string",      "value": "string",      "value_from_secret": "string"    }  ],  "gpu_count": 0,  "gpu_type": "h100-80gb",  "health_check_path": "string",  "id": "string",  "image": "string",  "max_replicas": 0,  "memory": 0,  "min_replicas": 0,  "model_mounts": [    {      "model_id": "string",      "mount_path": "string",      "revision_id": "string"    }  ],  "name": "string",  "object": "deployment",  "port": 0,  "ready_replicas": 0,  "replica_events": {    "property1": {      "image": "string",      "replica_ready_since": "string",      "replica_status": "string",      "replica_status_message": "string",      "replica_status_reason": "string",      "revision_id": "string",      "volume_preload_completed_at": "string",      "volume_preload_started_at": "string",      "volume_preload_status": "string"    },    "property2": {      "image": "string",      "replica_ready_since": "string",      "replica_status": "string",      "replica_status_message": "string",      "replica_status_reason": "string",      "revision_id": "string",      "volume_preload_completed_at": "string",      "volume_preload_started_at": "string",      "volume_preload_status": "string"    }  },  "status": "Updating",  "storage": 0,  "termination_grace_period_seconds": 0,  "updated_at": "2019-08-24T14:15:22Z",  "volumes": [    {      "mount_path": "string",      "name": "string",      "version": 0    }  ]}