Caleta Private AI Appliance / Model catalogue
The measured model catalogue
Every model the appliance ships, with download sizes, measured speeds on real Azure VM sizes, and the deployment pairings we have proven. Measured, not promised.
Be up and running in minutes: deploy a box into your own subscription, in the region that is cheapest for it today, then serve any model on this page from the appliance menu.
How we measure
Every figure on this page was measured on the exact appliance image we ship, with the serve settings it ships with. Tokens per second are single-stream decode speeds. "Caleta Score" is our in-house infrastructure-as-code benchmark, where higher is better. Deployment pairings marked measured ran on that hardware in that region; estimated pairings are engineering estimates awaiting a measured run.
Deployment pairings and spot prices below are from our latest catalogue snapshot. Spot prices move; the appliance's own spot placement menu action ranks regions from data refreshed daily. The menu on the box is authoritative and this page may lag it.
Pick a model, get a proven size and the best regions to run it
Recommendations come only from configurations we have run on the shipping image; region rankings come from spot data we measure daily.
Pick a model
GLM-5.2 on Azure
| Variant | Type | Download | Caleta Score | Measured |
|---|---|---|---|---|
| GLM-5.2 IQ2_M | MoE, 40B active | 239 GB | 90 | NC96 4xA100: 28 t/s · E64: 7 t/s |
| GLM-5.2 IQ3_S | MoE, 40B active | 288 GB | 98 | NC96: 26 t/s |
| GLM-5.2 (744B) (Azure Cobalt Arm) | MoE frontier, 40B active | 467 GB | - | ~9.5-11.7 t/s Cobalt CPU · frontier flagship; E64ps_v6 only |
Proven deployments
| Variant | VM size | Engine | Mode | All-in USD/hr* | Status |
|---|---|---|---|---|---|
| GLM-5.2 (UD-Q4_K_XL) | Standard_ND96is_MI300X_v5 | llamacpp | interactive | $11.49 (francecentral) | measured |
| GLM-5.2 (Q8_0) | Standard_ND96is_MI300X_v5 | llamacpp | interactive | $11.49 (francecentral) | measured |
| GLM-5.2 (UD-IQ2_M) | Standard_NC96ads_A100_v4 | llamacpp | interactive | $3.93 (italynorth) | measured |
| GLM-5.2 (UD-IQ3_S) | Standard_NC96ads_A100_v4 | llamacpp | interactive | $3.93 (italynorth) | measured |
DeepSeek-R1-0528 on Azure
| Variant | Type | Download | Caleta Score | Measured |
|---|---|---|---|---|
| DeepSeek-R1-0528 (Q2) | MoE, 37B active | 234 GB | unscored | E64: 9 t/s |
Proven deployments
| Variant | VM size | Engine | Mode | All-in USD/hr* | Status |
|---|---|---|---|---|---|
| DeepSeek-R1 (Q8_0) | Standard_ND96is_MI300X_v5 | llamacpp | interactive | $11.49 (francecentral) | estimated |
| DeepSeek-R1 (UD-Q2_K_XL) | Standard_E64ads_v7 | llamacpp | batch | $0.98 (eastus) | measured |
Devstral-Small-2 on Azure
| Variant | Type | Download | Caleta Score | Measured |
|---|---|---|---|---|
| Devstral-Small-2 (Q4) | dense 24B | 14 GB | 84 | NC80: 92 t/s |
Proven deployments
| Variant | VM size | Engine | Mode | All-in USD/hr* | Status |
|---|---|---|---|---|---|
| Devstral-Small-2 (UD-Q4_K_XL) | Standard_NC24ads_A100_v4 | llamacpp | interactive | $1.00 (italynorth) | measured |
| Devstral-Small-2 (UD-Q4_K_XL) | Standard_NC80adis_H100_v5 | llamacpp | interactive | $3.64 (indonesiacentral) | measured |
Gemma-4-26B-A4B on Azure
| Variant | Type | Download | Caleta Score | Measured |
|---|---|---|---|---|
| Gemma-4-26B-A4B (Q4) | MoE, 4B active | 15 GB | 82 | NC80: 190 t/s |
Proven deployments
| Variant | VM size | Engine | Mode | All-in USD/hr* | Status |
|---|---|---|---|---|---|
| Gemma-4-26B (UD-Q4_K_XL) | Standard_NC24ads_A100_v4 | llamacpp | interactive | $1.00 (italynorth) | measured |
| Gemma-4-26B (UD-Q4_K_XL) | Standard_NC80adis_H100_v5 | llamacpp | interactive | $3.64 (indonesiacentral) | measured |
GLM-4.7-Flash on Azure
| Variant | Type | Download | Caleta Score | Measured |
|---|---|---|---|---|
| GLM-4.7-Flash (Q4) | MoE, 3B active | 18 GB | 56 | NC80: 170 t/s |
Proven deployments
| Variant | VM size | Engine | Mode | All-in USD/hr* | Status |
|---|---|---|---|---|---|
| GLM-4.7-Flash (Q4_K_M) | Standard_NC24ads_A100_v4 | llamacpp | interactive | $1.00 (italynorth) | measured |
| GLM-4.7-Flash (Q4_K_M) | Standard_NC80adis_H100_v5 | llamacpp | interactive | $3.64 (indonesiacentral) | measured |
gpt-oss-20b on Azure
| Variant | Type | Download | Caleta Score | Measured |
|---|---|---|---|---|
| gpt-oss-20b (MXFP4) | MoE, 4B active | 12 GB | 60 | NC80: 306 t/s |
| gpt-oss-20b (Azure Cobalt Arm) | MoE reasoner, 4B active | 12 GB | - | ~75 t/s Cobalt CPU · fastest strong reasoner; fits E8ps_v6 |
Proven deployments
| Variant | VM size | Engine | Mode | All-in USD/hr* | Status |
|---|---|---|---|---|---|
| gpt-oss-20b (MXFP4) | Standard_NC24ads_A100_v4 | llamacpp | interactive | $1.00 (italynorth) | measured |
| gpt-oss-20b (MXFP4) | Standard_NC80adis_H100_v5 | llamacpp | interactive | $3.64 (indonesiacentral) | measured |
Kimi-K2.7-Code on Azure
| Variant | Type | Download | Caleta Score | Measured |
|---|---|---|---|---|
| Kimi-K2.7-Code (Q2) | MoE, 32B active | 339 GB | unscored | E64: 9.4 t/s |
Proven deployments
| Variant | VM size | Engine | Mode | All-in USD/hr* | Status |
|---|---|---|---|---|---|
| Kimi-K2.7-Code (UD-Q4_K_XL) | Standard_ND96is_MI300X_v5 | llamacpp | interactive | $11.49 (francecentral) | measured |
| Kimi-K2.7-Code (UD-Q2_K_XL) | Standard_E64ads_v7 | llamacpp | batch | $0.98 (eastus) | measured |
Phi-4 on Azure
| Variant | Type | Download | Caleta Score | Measured |
|---|---|---|---|---|
| Phi-4 (Q4) | dense 14B | 9 GB | 74 | NC80: 152 t/s |
Proven deployments
| Variant | VM size | Engine | Mode | All-in USD/hr* | Status |
|---|---|---|---|---|---|
| Phi-4 (Q4_K_M) | Standard_NC24ads_A100_v4 | llamacpp | interactive | $1.00 (italynorth) | measured |
| Phi-4 (Q4_K_M) | Standard_NC80adis_H100_v5 | llamacpp | interactive | $3.64 (indonesiacentral) | measured |
Qwen2.5-Coder-32B on Azure
| Variant | Type | Download | Caleta Score | Measured |
|---|---|---|---|---|
| Qwen2.5-Coder-32B (Q4) | dense 32B | 19 GB | 86 | NC80: 68.4 t/s |
Proven deployments
| Variant | VM size | Engine | Mode | All-in USD/hr* | Status |
|---|---|---|---|---|---|
| Qwen2.5-Coder-32B (Q4_K_M) | Standard_NC24ads_A100_v4 | llamacpp | interactive | $1.00 (italynorth) | measured |
| Qwen2.5-Coder-32B (Q4_K_M) | Standard_NC80adis_H100_v5 | llamacpp | interactive | $3.64 (indonesiacentral) | measured |
Qwen3-235B on Azure
| Variant | Type | Download | Caleta Score | Measured |
|---|---|---|---|---|
| Qwen3-235B (Q4) | MoE, 22B active | 126 GB | 86 | NC80: 82.5 t/s · E64: 13 t/s |
| Qwen3-235B-A22B (Azure Cobalt Arm) | MoE big, 22B active | 142 GB | - | ~15 t/s Cobalt CPU · big brain, RAM-heavy |
Proven deployments
| Variant | VM size | Engine | Mode | All-in USD/hr* | Status |
|---|---|---|---|---|---|
| Qwen3-235B (UD-Q4_K_XL) | Standard_NC80adis_H100_v5 | llamacpp | interactive | $3.64 (indonesiacentral) | measured |
| Qwen3-235B-BF16 (BF16) | Standard_ND96is_MI300X_v5 | llamacpp | interactive | $11.49 (francecentral) | measured |
Qwen3-Coder-30B on Azure
| Variant | Type | Download | Caleta Score | Measured |
|---|---|---|---|---|
| Qwen3-Coder-30B (Q4), the trial | MoE, 3B active | 18 GB | 80 | NC80: 263 t/s · E32: 35 t/s |
| Qwen3-Coder-30B-A3B (Azure Cobalt Arm) | MoE coder, 3B active | 19 GB | - | ~56-64 t/s Cobalt CPU · fast mature coder; safe default |
Proven deployments
| Variant | VM size | Engine | Mode | All-in USD/hr* | Status |
|---|---|---|---|---|---|
| Qwen3-Coder-30B (UD-Q4_K_XL) | Standard_NC24ads_A100_v4 | llamacpp | interactive | $1.00 (italynorth) | measured |
| Qwen3-Coder-30B (UD-Q4_K_XL) | Standard_NC80adis_H100_v5 | llamacpp | interactive | $3.64 (indonesiacentral) | measured |
Qwen3-Coder-Next-80B on Azure
| Variant | Type | Download | Caleta Score | Measured |
|---|---|---|---|---|
| Qwen3-Coder-Next-80B (Q4) | MoE, 3B active | 46 GB | 78 | NC80: 169 t/s |
Proven deployments
| Variant | VM size | Engine | Mode | All-in USD/hr* | Status |
|---|---|---|---|---|---|
| Qwen3-Coder-Next-80B (UD-Q4_K_XL) | Standard_NC24ads_A100_v4 | llamacpp | interactive | $1.00 (italynorth) | measured |
| Qwen3-Coder-Next-80B (UD-Q4_K_XL) | Standard_NC80adis_H100_v5 | llamacpp | interactive | $3.64 (indonesiacentral) | measured |
Qwen3.6-27B on Azure
| Variant | Type | Download | Caleta Score | Measured |
|---|---|---|---|---|
| Qwen3.6-27B (Q4) | dense 27B | 17 GB | 90 | NC80 2xH100: 72.7 t/s |
Proven deployments
| Variant | VM size | Engine | Mode | All-in USD/hr* | Status |
|---|---|---|---|---|---|
| Qwen3.6-27B (UD-Q4_K_XL) | Standard_NC24ads_A100_v4 | llamacpp | interactive | $1.00 (italynorth) | measured |
| Qwen3.6-27B (UD-Q4_K_XL) | Standard_NC80adis_H100_v5 | llamacpp | interactive | $3.64 (indonesiacentral) | measured |
Qwen3.6-35B-A3B on Azure
| Variant | Type | Download | Caleta Score | Measured |
|---|---|---|---|---|
| Qwen3.6-35B-A3B (Q4) | MoE, 3B active | 21 GB | 76 | NC80: 205 t/s |
| Qwen3.6-35B-A3B (Azure Cobalt Arm) | MoE coder, 3B active | 22 GB | - | ~34 t/s Cobalt CPU · best coder quality (SWE-bench 73.4) |
Proven deployments
| Variant | VM size | Engine | Mode | All-in USD/hr* | Status |
|---|---|---|---|---|---|
| Qwen3.6-35B (UD-Q4_K_XL) | Standard_NC24ads_A100_v4 | llamacpp | interactive | $1.00 (italynorth) | measured |
| Qwen3.6-35B (UD-Q4_K_XL) | Standard_NC80adis_H100_v5 | llamacpp | interactive | $3.64 (indonesiacentral) | measured |
DeepSeek-V4-Flash-0731 on Azure
| Variant | Type | Download | Caleta Score | Measured |
|---|---|---|---|---|
| DeepSeek-V4-Flash-0731 (MXFP4) | MoE, 13B active | 155 GB | unscored | NC80: 54.6 t/s · E64: 9.7 t/s |
Proven deployments
| Variant | VM size | Engine | Mode | All-in USD/hr* | Status |
|---|---|---|---|---|---|
| DeepSeek-V4-Flash-0731 (MXFP4) | Standard_NC80adis_H100_v5 | llamacpp | interactive | $3.64 (indonesiacentral) | measured |
gpt-oss-120b on Azure
| Variant | Type | Download | Caleta Score | Measured |
|---|---|---|---|---|
| gpt-oss-120b (MXFP4) | MoE, 5B active | 61 GB | unscored | NC24 1xA100: 140 t/s · E64: 48 t/s |
| gpt-oss-120b (Azure Cobalt Arm) | MoE reasoner, 5B active | 63 GB | - | ~53 t/s Cobalt CPU · 120B brain at 5B-active speed; value pick |
Proven deployments
| Variant | VM size | Engine | Mode | All-in USD/hr* | Status |
|---|---|---|---|---|---|
| gpt-oss-120b (MXFP4) | Standard_NC24ads_A100_v4 | llamacpp | interactive | $1.00 (italynorth) | measured |
Qwen3-Coder-480B on Azure
Proven deployments
| Variant | VM size | Engine | Mode | All-in USD/hr* | Status |
|---|---|---|---|---|---|
| Qwen3-Coder-480B (Q8_0) | Standard_ND96is_MI300X_v5 | llamacpp | interactive | $11.49 (francecentral) | measured |
Qwen3.8-27B on Azure
| Variant | Type | Download | Caleta Score | Measured |
|---|---|---|---|---|
| Qwen3.8-27B (Q4) | dense 27B | 18 GB | unscored | NC24 A100: 47.5 t/s |
Proven deployments
| Variant | VM size | Engine | Mode | All-in USD/hr* | Status |
|---|---|---|---|---|---|
| Qwen3.8-27B (UD-Q4_K_XL) | Standard_NC24ads_A100_v4 | llamacpp | interactive | $1.00 (italynorth) | measured |
ERNIE-4.5-21B-A3B on Azure
| Variant | Type | Download | Caleta Score | Measured |
|---|---|---|---|---|
| ERNIE-4.5-21B-A3B (Azure Cobalt Arm) | MoE general, 3B active | 13 GB | - | ~67 t/s Cobalt CPU · tiny, Apache-2.0; fits E8ps_v6 |
GLM-4.5-Air on Azure
| Variant | Type | Download | Caleta Score | Measured |
|---|---|---|---|---|
| GLM-4.5-Air (Azure Cobalt Arm) | MoE mid, 12B active | 73 GB | - | ~25 t/s Cobalt CPU · best mid-tier, MIT; needs E16ps_v6 or larger |
Qwen3-30B-A3B-Instruct on Azure
| Variant | Type | Download | Caleta Score | Measured |
|---|---|---|---|---|
| Qwen3-30B-A3B-Instruct (Azure Cobalt Arm) | MoE general, 3B active | 19 GB | - | ~63 t/s Cobalt CPU · fast general daily-driver |
Qwen3-Next-80B-A3B on Azure
| Variant | Type | Download | Caleta Score | Measured |
|---|---|---|---|---|
| Qwen3-Next-80B-A3B (Azure Cobalt Arm) | MoE general, 3B active | 49 GB | - | ~28 t/s Cobalt CPU · smart, not a speed pick |
*All-in prices are the VM spot price as measured in the listed region for our latest snapshot, plus the appliance software fee for that size as read from the Marketplace price grid. Both land on your Azure bill. Spot moves; treat these as the shape of the economics, not a quote.
Get notified when the catalogue changes
Leave your email and we will tell you when the rankings or the model catalogue change.
Every one of these runs on the same Marketplace image.