Deployments

Update a deployment

PATCH
/deployments/{id}

Update an existing deployment configuration

Authorization

bearerAuth
AuthorizationBearer <token>

In: header

Path Parameters

id*string

Deployment ID or name

Request Body

application/json

Updated deployment configuration

args?array<string>

Args overrides the container's CMD. Provide as an array of arguments (e.g., ["python", "app.py"])

autoscaling?||

Autoscaling configuration for the deployment. Set to {} to disable autoscaling

capacity_type?string

Controls how replicas above reserved capacity behave. stable replicas stay running after scale-up; preemptible replicas may be evicted during capacity contention.

Value in

  • "stable"
  • "preemptible"
command?array<string>

Command overrides the container's ENTRYPOINT. Provide as an array (e.g., ["/bin/sh", "-c"])

cpu?number

CPU is the number of CPU cores to allocate per container instance (e.g., 0.1 = 100 milli cores)

Range0.1 <= value
description?string

Description is an optional human-readable description of your deployment

environment_variables?array<>

EnvironmentVariables is a list of environment variables to set in the container. Replaces all existing environment variables.

gpu_count?integer

GPUCount is the number of GPUs to allocate per container instance

gpu_type?string

GPUType specifies the GPU hardware to use (e.g., "h100-80gb")

Value in

  • "h100-80gb"
  • "h100-40gb-mig"
  • "h200-140gb"
  • "b200-192gb"
health_check_path?string

HealthCheckPath is the HTTP path for health checks (e.g., "/health"). Set to empty string to disable health checks

image?string

Image is the container image to deploy from registry.together.ai.

max_replicas?integer

MaxReplicas is the maximum number of replicas that can be scaled up to.

memory?number

Memory is the amount of RAM to allocate per container instance in GiB (e.g., 0.5 = 512MiB)

Rangevalue <= 1000
min_replicas?integer

MinReplicas is the minimum number of replicas to run

model_mounts?array<>

Replacement model weights to mount into the deployment. At most one mount is supported, and it cannot be used with volumes.

name?string

Name is the new unique identifier for your deployment. Must contain only alphanumeric characters, underscores, or hyphens (1-100 characters)

Length1 <= length <= 100
port?integer

Port is the container port your application listens on (e.g., 8080 for web servers)

Range1 <= value <= 65535
storage?integer

Storage is the amount of ephemeral disk storage to allocate per container instance (e.g., 10 = 10GiB)

Rangevalue <= 400
termination_grace_period_seconds?integer

TerminationGracePeriodSeconds is the time in seconds to wait for graceful shutdown before forcefully terminating the replica

volumes?array<>

Volumes is a list of volume mounts to attach to the container. Replaces all existing volumes.

Response Body

application/json

application/json

application/json

application/json

curl -X PATCH "https://example.com/deployments/string" \  -H "Content-Type: application/json" \  -d '{}'
{  "args": [    "string"  ],  "autoscaling": {    "metric": "HTTPTotalRequests",    "target": 0,    "time_interval_minutes": 0  },  "capacity_type": "stable",  "command": [    "string"  ],  "cpu": 0,  "created_at": "2019-08-24T14:15:22Z",  "description": "string",  "desired_replicas": 0,  "environment_variables": [    {      "name": "string",      "value": "string",      "value_from_secret": "string"    }  ],  "gpu_count": 0,  "gpu_type": "h100-80gb",  "health_check_path": "string",  "id": "string",  "image": "string",  "max_replicas": 0,  "memory": 0,  "min_replicas": 0,  "model_mounts": [    {      "model_id": "string",      "mount_path": "string",      "revision_id": "string"    }  ],  "name": "string",  "object": "deployment",  "port": 0,  "ready_replicas": 0,  "replica_events": {    "property1": {      "image": "string",      "replica_ready_since": "string",      "replica_status": "string",      "replica_status_message": "string",      "replica_status_reason": "string",      "revision_id": "string",      "volume_preload_completed_at": "string",      "volume_preload_started_at": "string",      "volume_preload_status": "string"    },    "property2": {      "image": "string",      "replica_ready_since": "string",      "replica_status": "string",      "replica_status_message": "string",      "replica_status_reason": "string",      "revision_id": "string",      "volume_preload_completed_at": "string",      "volume_preload_started_at": "string",      "volume_preload_status": "string"    }  },  "status": "Updating",  "storage": 0,  "termination_grace_period_seconds": 0,  "updated_at": "2019-08-24T14:15:22Z",  "volumes": [    {      "mount_path": "string",      "name": "string",      "version": 0    }  ]}