Supported Models Service

List supported models

GET
/supported-models

Lists Together-hosted base models that can be deployed for dedicated inference, together with their capabilities and certified deployment profiles.

Authorization

bearerAuth
AuthorizationBearer <token>

In: header

Query Parameters

modality?string

Filter models by input modality.

Value in

  • "MODALITY_TEXT"
  • "MODALITY_IMAGE"
  • "MODALITY_AUDIO"
  • "MODALITY_VIDEO"
product?string

Filter models by product surface.

Value in

  • "PRODUCT_SERVERLESS"
  • "PRODUCT_DEDICATED"
  • "PRODUCT_FINE_TUNING"
search?string

Case-insensitive search across model IDs, names, and descriptions.

limit?integer

Maximum number of models to return.

after?string

Cursor from a previous supported-model list response.

Response Body

application/json

application/json

curl -X GET "https://example.com/supported-models"
{  "data": [    {      "id": "string",      "name": "string",      "displayName": "string",      "description": "string",      "inputModalities": [        "MODALITY_TEXT"      ],      "outputModalities": [        "MODALITY_TEXT"      ],      "products": [        "PRODUCT_SERVERLESS"      ],      "features": [        "FEATURE_TOOL_CALLING"      ],      "capabilities": [        "CAPABILITY_CHAT"      ],      "architecture": "string",      "contextLength": "string",      "publisher": "string",      "status": "SUPPORTED_MODEL_STATUS_RECOMMENDED",      "tags": [        "string"      ],      "inputFormat": "string",      "outputFormat": "string",      "serverlessEndpoint": "string",      "pricing": {        "input": 0.1,        "output": 0.1,        "cachedInput": 0.1      },      "familyId": "string",      "displayType": "string",      "baseModelId": "string",      "baseModel": "string",      "deploymentProfiles": [        {          "profileId": "string",          "certifiedConfigRevisionId": "string",          "certifiedModelRevisionId": "string",          "gpuType": "string",          "gpuCount": 0,          "quantization": "string",          "tensorParallelSize": 0,          "performanceBenchmarks": {            "decodingSpeedTps": 0.1,            "timeToFirstTokenMs": 0,            "maxContextLength": "string"          },          "config": "string",          "model": "string",          "parallelism": "string",          "modelName": "string"        }      ],      "createdAt": "2019-08-24T14:15:22Z",      "updatedAt": "2019-08-24T14:15:22Z"    }  ],  "next_cursor": "string",  "object": "list"}