Create a deployment
Creates a model deployment under an endpoint. The deployment provisions asynchronously; monitor its status before routing live traffic to it.
Authorization
bearerAuth In: header
Path Parameters
ID of the project that owns the endpoint.
ID of the endpoint that will contain the deployment.
Query Parameters
When true, validates the request without creating or provisioning a deployment.
Request Body
application/json
Configuration for creating a deployment that binds a model and immutable config to an endpoint.
Name for the deployment within its endpoint. Returned as a fully-qualified endpoint string.
Deprecated. Use model. Model identifier to serve, accepted when model is unset.
Deprecated. Use model with a /revisions/{revisionId} segment. If omitted, the latest revision is resolved at creation.
Deprecated. Use config. Config revision identifier to deploy, accepted when config is unset.
Model resource name in the form projects/{projectId}/models/{modelId}[/revisions/{revisionId}]. Omit the revision segment to pin the latest revision at creation time.
Autoscaling configuration for the deployment.
Immutable config revision in the form projects/{projectId}/configs/{configRevisionId}. The config must be compatible with the model.
Inactive timeout in minutes. Use 0 or omit to disable automatic stopping; otherwise accepted values are 30 through 1440.
Maximum number of inference requests that may be in flight to a single replica. If omitted, the platform uses one less than the config's per-replica concurrency limit to reserve a health-check slot. Values above that maximum are reduced on create; 0 means unlimited when the config limit is 1 or less.
Placement controls where a deployment is scheduled.
Response Body
application/json
application/json
curl -X POST "https://example.com/projects/string/endpoints/string/deployments" \ -H "Content-Type: application/json" \ -d '{ "name": "string", "autoscaling": {} }'{ "id": "string", "projectId": "string", "endpointId": "string", "name": "string", "createdAt": "2019-08-24T14:15:22Z", "updatedAt": "2019-08-24T14:15:22Z", "modelId": "string", "modelRevisionId": "string", "model": "string", "autoscaling": { "minReplicas": 0, "maxReplicas": 0, "scaleDownWindow": "string", "scaleUpWindow": "string", "scaleToZeroWindow": "string", "scalingMetrics": [ { "name": "active_sessions", "type": "METRIC_TARGET_TYPE_VALUE", "target": 0, "percentile": "string" } ], "scaleUp": { "policies": [ { "type": "SCALING_POLICY_TYPE_PODS", "value": 1, "periodSeconds": 1 } ], "selectPolicy": "SCALING_POLICY_SELECT_MAX" }, "scaleDown": { "policies": [ { "type": "SCALING_POLICY_TYPE_PODS", "value": 1, "periodSeconds": 1 } ], "selectPolicy": "SCALING_POLICY_SELECT_MAX" } }, "configId": "string", "config": "string", "speculatorId": "string", "speculatorRevisionId": "string", "speculator": "string", "estimatedEffectiveTrafficShare": 0.1, "maxConcurrentRequestsPerReplica": "string", "inactiveTimeout": 0, "etag": "string", "hardware": "string", "trafficMode": "TRAFFIC_MODE_LIVE", "runtimeInfo": { "engineType": "string", "engineVersion": "string", "functionCallingSupported": true, "structuredOutputSupported": true }, "desiredReplicas": 0, "status": { "state": "DEPLOYMENT_STATE_PROVISIONING", "readyReplicas": 0, "message": "string", "scheduledReplicas": 0, "details": { "region": [ { "region": "string", "scheduledReplicas": 0, "readyReplicas": 0 } ] } }, "placement": { "inline": { "regions": [ "string" ], "constraint": "ENFORCEMENT_REQUIRED", "compliancePolicy": { "hipaa": true } } }}