Optimizer step
Submits an optimizer step operation that will asynchronously apply accumulated gradients to update model parameters. Does not make the updated parameters available for sampling; call weights-sync afterwards when you want subsequent samples to use the updated policy.
Authorization
bearerAuth In: header
Path Parameters
Training session ID
Header Parameters
Required key that makes retries return the original operation; use a new key for changed request bodies.
Request Body
application/json
Request body for an optimizer step.
Adam optimizer overrides for this step.
Muon optimizer overrides for this step.
Response Body
application/json
application/json
curl -X POST "https://example.com/rl/training-sessions/string/operations/optim-step" \ -H "Idempotency-Key: string" \ -H "Content-Type: application/json" \ -d '{}'{ "id": "string", "status": "TRAINING_OPERATION_STATUS_UNSPECIFIED", "output": { "step": "string" }, "error": { "code": "TRAINING_OPERATION_ERROR_CODE_UNSPECIFIED", "message": "string" }}Get custom forward-backward operation GET
Retrieves the current status and result of a custom forward-backward operation.
Weights sync POST
Submits a weights-sync operation that makes the session's current trained parameters available for sampling. Call this after `optim-step` when you want subsequent samples to use the updated policy.