Supported Models Service
List supported models
Lists Together-hosted base models that can be deployed for dedicated inference, together with their capabilities and certified deployment profiles.
Authorization
bearerAuth AuthorizationBearer <token>
In: header
Query Parameters
modality?string
Filter models by input modality.
Value in
- "MODALITY_TEXT"
- "MODALITY_IMAGE"
- "MODALITY_AUDIO"
- "MODALITY_VIDEO"
product?string
Filter models by product surface.
Value in
- "PRODUCT_SERVERLESS"
- "PRODUCT_DEDICATED"
- "PRODUCT_FINE_TUNING"
search?string
Case-insensitive search across model IDs, names, and descriptions.
limit?integer
Maximum number of models to return.
after?string
Cursor from a previous supported-model list response.
Response Body
application/json
application/json
curl -X GET "https://example.com/supported-models"{ "data": [ { "id": "string", "name": "string", "displayName": "string", "description": "string", "inputModalities": [ "MODALITY_TEXT" ], "outputModalities": [ "MODALITY_TEXT" ], "products": [ "PRODUCT_SERVERLESS" ], "features": [ "FEATURE_TOOL_CALLING" ], "capabilities": [ "CAPABILITY_CHAT" ], "architecture": "string", "contextLength": "string", "publisher": "string", "status": "SUPPORTED_MODEL_STATUS_RECOMMENDED", "tags": [ "string" ], "inputFormat": "string", "outputFormat": "string", "serverlessEndpoint": "string", "pricing": { "input": 0.1, "output": 0.1, "cachedInput": 0.1 }, "familyId": "string", "displayType": "string", "baseModelId": "string", "baseModel": "string", "deploymentProfiles": [ { "profileId": "string", "certifiedConfigRevisionId": "string", "certifiedModelRevisionId": "string", "gpuType": "string", "gpuCount": 0, "quantization": "string", "tensorParallelSize": 0, "performanceBenchmarks": { "decodingSpeedTps": 0.1, "timeToFirstTokenMs": 0, "maxContextLength": "string" }, "config": "string", "model": "string", "parallelism": "string", "modelName": "string" } ], "createdAt": "2019-08-24T14:15:22Z", "updatedAt": "2019-08-24T14:15:22Z" } ], "next_cursor": "string", "object": "list"}Get an inference instance type GET
Retrieves the GPU resources, pricing, regional availability, and best-effort capacity headroom for one inference instance type.
Get a supported model GET
Retrieves a Together-hosted base model and the certified model, configuration, hardware, and performance profiles available for deployment.