Sample
Submits a sample operation that will asynchronously generate text completions with logprobs.
Authorization
bearerAuth In: header
Path Parameters
Training session ID
Header Parameters
Required key that makes retries return the original operation; use a new key for changed request bodies.
Request Body
application/json
Model inputs to sample from
Optional sampling parameters
Number of completions to generate per prompt
1When true, also compute teacher-forced log-probabilities for the model input tokens and return them in SampleResult.prompt_logprobs.
falseNumber of most likely alternative tokens to return per model input token in SampleResult.topk_prompt_logprobs. 0 disables top-k prompt log-probabilities. Maximum 20.
0 <= value <= 200When true, enable reuse of the expert selections from sampled sequences during training. Only supported for mixture-of-experts models; ignored for other models.
falseResponse Body
application/json
application/json
curl -X POST "https://example.com/rl/training-sessions/string/operations/sample" \ -H "Idempotency-Key: string" \ -H "Content-Type: application/json" \ -d '{ "model_inputs": [ { "chunks": [ { "encoded_text": { "tokens": [ "string" ] } } ] } ] }'{ "id": "string", "status": "TRAINING_OPERATION_STATUS_UNSPECIFIED", "output": { "results": [ { "sequences": [ { "tokens": [ "string" ], "logprobs": [ 0 ], "stop_reason": "STOP_REASON_LENGTH", "prompt_cache_hit_tokens": 0, "routed_experts_key": "string" } ], "prompt_logprobs": [ 0.1 ], "topk_prompt_logprobs": [ { "token_ids": [ 0 ], "logprobs": [ 0.1 ] } ], "policy_segments": [ { "version": 0, "start_token": 0 } ] } ] }, "error": { "code": "TRAINING_OPERATION_ERROR_CODE_UNSPECIFIED", "message": "string" }}Weights sync POST
Submits a weights-sync operation that makes the session's current trained parameters available for sampling. Call this after `optim-step` when you want subsequent samples to use the updated policy.
Create inference checkpoint POST
Submits an operation that will asynchronously save the current LoRA adapter as an inference checkpoint and upload it to object storage.