R L

Optimizer step

POST
/rl/training-sessions/{session_id}/operations/optim-step

Submits an optimizer step operation that will asynchronously apply accumulated gradients to update model parameters. Does not make the updated parameters available for sampling; call weights-sync afterwards when you want subsequent samples to use the updated policy.

Authorization

bearerAuth
AuthorizationBearer <token>

In: header

Path Parameters

session_id*string

Training session ID

Header Parameters

Idempotency-Key*string

Required key that makes retries return the original operation; use a new key for changed request bodies.

Request Body

application/json

Request body for an optimizer step.

adam_params?

Adam optimizer overrides for this step.

muon_params?

Muon optimizer overrides for this step.

Response Body

application/json

application/json

curl -X POST "https://example.com/rl/training-sessions/string/operations/optim-step" \  -H "Idempotency-Key: string" \  -H "Content-Type: application/json" \  -d '{}'
{  "id": "string",  "status": "TRAINING_OPERATION_STATUS_UNSPECIFIED",  "output": {    "step": "string"  },  "error": {    "code": "TRAINING_OPERATION_ERROR_CODE_UNSPECIFIED",    "message": "string"  }}