Weights sync
Submits a weights-sync operation that makes the session's current trained parameters available for sampling. Call this after optim-step when you want subsequent samples to use the updated policy.
Authorization
bearerAuth In: header
Path Parameters
Training session ID
Header Parameters
Required key that makes retries return the original operation; use a new key for changed request bodies.
Request Body
application/json
Request body for publishing updated policy parameters for sampling.
How updated parameters are made available for sampling. See WeightSyncType for accepted values.
Value in
- "WEIGHT_SYNC_TYPE_SYNCHRONOUS"
- "WEIGHT_SYNC_TYPE_BACKGROUND_PUBLISH"
- "WEIGHT_SYNC_TYPE_PIPELINE"
Response Body
application/json
application/json
curl -X POST "https://example.com/rl/training-sessions/string/operations/weights-sync" \ -H "Idempotency-Key: string" \ -H "Content-Type: application/json" \ -d '{ "weight_sync_type": "WEIGHT_SYNC_TYPE_SYNCHRONOUS" }'{ "id": "string", "status": "TRAINING_OPERATION_STATUS_UNSPECIFIED", "output": { "weights_version": "string" }, "error": { "code": "TRAINING_OPERATION_ERROR_CODE_UNSPECIFIED", "message": "string" }}Optimizer step POST
Submits an optimizer step operation that will asynchronously apply accumulated gradients to update model parameters. Does not make the updated parameters available for sampling; call `weights-sync` afterwards when you want subsequent samples to use the updated policy.
Sample POST
Submits a sample operation that will asynchronously generate text completions with logprobs.