Sail does not requantize weights. Each row links to the weights Sail
currently serves.
Core models
| Model | Model ID | Image | LoRA | Reasoning |
|---|---|---|---|---|
moonshotai/Kimi-K3 | ||||
zai-org/GLM-5.3 | ||||
zai-org/GLM-5.3-Flash | ||||
deepseek-ai/DeepSeek-V4.1-Flash | ||||
deepseek-ai/DeepSeek-V4-Pro-0813 | ||||
deepseek-ai/DeepSeek-V4-Flash-0731 | ||||
moonshotai/Kimi-K2.6 | ||||
google/gemma-4-31B-it | ||||
nvidia/Gemma-4-31B-IT-NVFP4 | ||||
google/gemma-4-12B-it | ||||
openai/gpt-oss-120b |
Flex-only models
Models served with theflex completion window exclusively.
Notes
- If we offer multiple quantizations of a model, we list them as separate model IDs.
- Use
GET /v1/modelsto confirm runtime availability for your API key. - For per-model rates by completion window, see Pricing.