Fine tuning

Preview tokenized data

POST
/fine-tunes/preview

Preview how sampled rows from a fine-tuning training file will be tokenized before packing.

Authorization

bearerAuth
AuthorizationBearer <token>

In: header

Request Body

application/json

Request body for previewing tokenized fine-tuning data.

model*string

Name of the base model whose tokenizer and chat template will be used.

training_file*string

File-ID of the uploaded JSONL training file to sample for preview.

training_method?"sft"

Fine-tuning method to preview. Only supervised fine-tuning is currently supported.

Default"sft"

Value in

  • "sft"
train_on_inputs?boolean

Whether prompt or user-message tokens should contribute to training loss in the preview.

Defaultfalse
top_k?integer

Maximum number of rows from the start of the training file to tokenize.

Range1 <= value <= 50
Default5

Response Body

application/json

application/json

application/json

application/json

application/json

application/json

application/json

curl -X POST "https://example.com/fine-tunes/preview" \  -H "Content-Type: application/json" \  -d '{    "model": "string",    "training_file": "string"  }'
{  "model": "string",  "dataset_format": "general",  "max_seq_length": 0,  "train_on_inputs": true,  "rows": [    {      "input_ids": [        0      ],      "tokens": [        "string"      ],      "labels": [        0      ],      "trained_spans": [        [          0,          0        ]      ],      "num_tokens": 0,      "num_trained_tokens": 0,      "truncated": true    }  ]}