Update a deployment
Update an existing deployment configuration
Authorization
bearerAuth In: header
Path Parameters
Deployment ID or name
Request Body
application/json
Updated deployment configuration
Args overrides the container's CMD. Provide as an array of arguments (e.g., ["python", "app.py"])
Autoscaling configuration for the deployment. Set to {} to disable autoscaling
Controls how replicas above reserved capacity behave. stable replicas stay running after scale-up; preemptible replicas may be evicted during capacity contention.
Value in
- "stable"
- "preemptible"
Command overrides the container's ENTRYPOINT. Provide as an array (e.g., ["/bin/sh", "-c"])
CPU is the number of CPU cores to allocate per container instance (e.g., 0.1 = 100 milli cores)
0.1 <= valueDescription is an optional human-readable description of your deployment
EnvironmentVariables is a list of environment variables to set in the container. Replaces all existing environment variables.
GPUCount is the number of GPUs to allocate per container instance
GPUType specifies the GPU hardware to use (e.g., "h100-80gb")
Value in
- "h100-80gb"
- "h100-40gb-mig"
- "h200-140gb"
- "b200-192gb"
HealthCheckPath is the HTTP path for health checks (e.g., "/health"). Set to empty string to disable health checks
Image is the container image to deploy from registry.together.ai.
MaxReplicas is the maximum number of replicas that can be scaled up to.
Memory is the amount of RAM to allocate per container instance in GiB (e.g., 0.5 = 512MiB)
value <= 1000MinReplicas is the minimum number of replicas to run
Replacement model weights to mount into the deployment. At most one mount is supported, and it cannot be used with volumes.
Name is the new unique identifier for your deployment. Must contain only alphanumeric characters, underscores, or hyphens (1-100 characters)
1 <= length <= 100Port is the container port your application listens on (e.g., 8080 for web servers)
1 <= value <= 65535Storage is the amount of ephemeral disk storage to allocate per container instance (e.g., 10 = 10GiB)
value <= 400TerminationGracePeriodSeconds is the time in seconds to wait for graceful shutdown before forcefully terminating the replica
Volumes is a list of volume mounts to attach to the container. Replaces all existing volumes.
Response Body
application/json
application/json
application/json
application/json
curl -X PATCH "https://example.com/deployments/string" \ -H "Content-Type: application/json" \ -d '{}'{ "args": [ "string" ], "autoscaling": { "metric": "HTTPTotalRequests", "target": 0, "time_interval_minutes": 0 }, "capacity_type": "stable", "command": [ "string" ], "cpu": 0, "created_at": "2019-08-24T14:15:22Z", "description": "string", "desired_replicas": 0, "environment_variables": [ { "name": "string", "value": "string", "value_from_secret": "string" } ], "gpu_count": 0, "gpu_type": "h100-80gb", "health_check_path": "string", "id": "string", "image": "string", "max_replicas": 0, "memory": 0, "min_replicas": 0, "model_mounts": [ { "model_id": "string", "mount_path": "string", "revision_id": "string" } ], "name": "string", "object": "deployment", "port": 0, "ready_replicas": 0, "replica_events": { "property1": { "image": "string", "replica_ready_since": "string", "replica_status": "string", "replica_status_message": "string", "replica_status_reason": "string", "revision_id": "string", "volume_preload_completed_at": "string", "volume_preload_started_at": "string", "volume_preload_status": "string" }, "property2": { "image": "string", "replica_ready_since": "string", "replica_status": "string", "replica_status_message": "string", "replica_status_reason": "string", "revision_id": "string", "volume_preload_completed_at": "string", "volume_preload_started_at": "string", "volume_preload_status": "string" } }, "status": "Updating", "storage": 0, "termination_grace_period_seconds": 0, "updated_at": "2019-08-24T14:15:22Z", "volumes": [ { "mount_path": "string", "name": "string", "version": 0 } ]}