Caleta Private AI Appliance / Model catalogue

The measured model catalogue

Every model the appliance ships, with download sizes, measured speeds on real Azure VM sizes, and the deployment pairings we have proven. Measured, not promised.

DEPLOY NOW

Be up and running in minutes: deploy a box into your own subscription, in the region that is cheapest for it today, then serve any model on this page from the appliance menu.

How we measure

Every figure on this page was measured on the exact appliance image we ship, with the serve settings it ships with. Tokens per second are single-stream decode speeds. "Caleta Score" is our in-house infrastructure-as-code benchmark, where higher is better. Deployment pairings marked measured ran on that hardware in that region; estimated pairings are engineering estimates awaiting a measured run.

Deployment pairings and spot prices below are from our latest catalogue snapshot. Spot prices move; the appliance's own spot placement menu action ranks regions from data refreshed daily. The menu on the box is authoritative and this page may lag it.

Pick a model, get a proven size and the best regions to run it

Recommendations come only from configurations we have run on the shipping image; region rankings come from spot data we measure daily.

1

Pick a model

GLM-5.2 on Azure

VariantTypeDownloadCaleta ScoreMeasured
GLM-5.2 IQ2_MMoE, 40B active239 GB90NC96 4xA100: 28 t/s · E64: 7 t/s
GLM-5.2 IQ3_SMoE, 40B active288 GB98NC96: 26 t/s
GLM-5.2 (744B) (Azure Cobalt Arm)MoE frontier, 40B active467 GB-~9.5-11.7 t/s Cobalt CPU · frontier flagship; E64ps_v6 only

Proven deployments

VariantVM sizeEngineModeAll-in USD/hr*Status
GLM-5.2 (UD-Q4_K_XL)Standard_ND96is_MI300X_v5llamacppinteractive$11.49 (francecentral)measured
GLM-5.2 (Q8_0)Standard_ND96is_MI300X_v5llamacppinteractive$11.49 (francecentral)measured
GLM-5.2 (UD-IQ2_M)Standard_NC96ads_A100_v4llamacppinteractive$3.93 (italynorth)measured
GLM-5.2 (UD-IQ3_S)Standard_NC96ads_A100_v4llamacppinteractive$3.93 (italynorth)measured

DeepSeek-R1-0528 on Azure

VariantTypeDownloadCaleta ScoreMeasured
DeepSeek-R1-0528 (Q2)MoE, 37B active234 GBunscoredE64: 9 t/s

Proven deployments

VariantVM sizeEngineModeAll-in USD/hr*Status
DeepSeek-R1 (Q8_0)Standard_ND96is_MI300X_v5llamacppinteractive$11.49 (francecentral)estimated
DeepSeek-R1 (UD-Q2_K_XL)Standard_E64ads_v7llamacppbatch$0.98 (eastus)measured

Devstral-Small-2 on Azure

VariantTypeDownloadCaleta ScoreMeasured
Devstral-Small-2 (Q4)dense 24B14 GB84NC80: 92 t/s

Proven deployments

VariantVM sizeEngineModeAll-in USD/hr*Status
Devstral-Small-2 (UD-Q4_K_XL)Standard_NC24ads_A100_v4llamacppinteractive$1.00 (italynorth)measured
Devstral-Small-2 (UD-Q4_K_XL)Standard_NC80adis_H100_v5llamacppinteractive$3.64 (indonesiacentral)measured

Gemma-4-26B-A4B on Azure

VariantTypeDownloadCaleta ScoreMeasured
Gemma-4-26B-A4B (Q4)MoE, 4B active15 GB82NC80: 190 t/s

Proven deployments

VariantVM sizeEngineModeAll-in USD/hr*Status
Gemma-4-26B (UD-Q4_K_XL)Standard_NC24ads_A100_v4llamacppinteractive$1.00 (italynorth)measured
Gemma-4-26B (UD-Q4_K_XL)Standard_NC80adis_H100_v5llamacppinteractive$3.64 (indonesiacentral)measured

GLM-4.7-Flash on Azure

VariantTypeDownloadCaleta ScoreMeasured
GLM-4.7-Flash (Q4)MoE, 3B active18 GB56NC80: 170 t/s

Proven deployments

VariantVM sizeEngineModeAll-in USD/hr*Status
GLM-4.7-Flash (Q4_K_M)Standard_NC24ads_A100_v4llamacppinteractive$1.00 (italynorth)measured
GLM-4.7-Flash (Q4_K_M)Standard_NC80adis_H100_v5llamacppinteractive$3.64 (indonesiacentral)measured

gpt-oss-20b on Azure

VariantTypeDownloadCaleta ScoreMeasured
gpt-oss-20b (MXFP4)MoE, 4B active12 GB60NC80: 306 t/s
gpt-oss-20b (Azure Cobalt Arm)MoE reasoner, 4B active12 GB-~75 t/s Cobalt CPU · fastest strong reasoner; fits E8ps_v6

Proven deployments

VariantVM sizeEngineModeAll-in USD/hr*Status
gpt-oss-20b (MXFP4)Standard_NC24ads_A100_v4llamacppinteractive$1.00 (italynorth)measured
gpt-oss-20b (MXFP4)Standard_NC80adis_H100_v5llamacppinteractive$3.64 (indonesiacentral)measured

Kimi-K2.7-Code on Azure

VariantTypeDownloadCaleta ScoreMeasured
Kimi-K2.7-Code (Q2)MoE, 32B active339 GBunscoredE64: 9.4 t/s

Proven deployments

VariantVM sizeEngineModeAll-in USD/hr*Status
Kimi-K2.7-Code (UD-Q4_K_XL)Standard_ND96is_MI300X_v5llamacppinteractive$11.49 (francecentral)measured
Kimi-K2.7-Code (UD-Q2_K_XL)Standard_E64ads_v7llamacppbatch$0.98 (eastus)measured

Phi-4 on Azure

VariantTypeDownloadCaleta ScoreMeasured
Phi-4 (Q4)dense 14B9 GB74NC80: 152 t/s

Proven deployments

VariantVM sizeEngineModeAll-in USD/hr*Status
Phi-4 (Q4_K_M)Standard_NC24ads_A100_v4llamacppinteractive$1.00 (italynorth)measured
Phi-4 (Q4_K_M)Standard_NC80adis_H100_v5llamacppinteractive$3.64 (indonesiacentral)measured

Qwen2.5-Coder-32B on Azure

VariantTypeDownloadCaleta ScoreMeasured
Qwen2.5-Coder-32B (Q4)dense 32B19 GB86NC80: 68.4 t/s

Proven deployments

VariantVM sizeEngineModeAll-in USD/hr*Status
Qwen2.5-Coder-32B (Q4_K_M)Standard_NC24ads_A100_v4llamacppinteractive$1.00 (italynorth)measured
Qwen2.5-Coder-32B (Q4_K_M)Standard_NC80adis_H100_v5llamacppinteractive$3.64 (indonesiacentral)measured

Qwen3-235B on Azure

VariantTypeDownloadCaleta ScoreMeasured
Qwen3-235B (Q4)MoE, 22B active126 GB86NC80: 82.5 t/s · E64: 13 t/s
Qwen3-235B-A22B (Azure Cobalt Arm)MoE big, 22B active142 GB-~15 t/s Cobalt CPU · big brain, RAM-heavy

Proven deployments

VariantVM sizeEngineModeAll-in USD/hr*Status
Qwen3-235B (UD-Q4_K_XL)Standard_NC80adis_H100_v5llamacppinteractive$3.64 (indonesiacentral)measured
Qwen3-235B-BF16 (BF16)Standard_ND96is_MI300X_v5llamacppinteractive$11.49 (francecentral)measured

Qwen3-Coder-30B on Azure

VariantTypeDownloadCaleta ScoreMeasured
Qwen3-Coder-30B (Q4), the trialMoE, 3B active18 GB80NC80: 263 t/s · E32: 35 t/s
Qwen3-Coder-30B-A3B (Azure Cobalt Arm)MoE coder, 3B active19 GB-~56-64 t/s Cobalt CPU · fast mature coder; safe default

Proven deployments

VariantVM sizeEngineModeAll-in USD/hr*Status
Qwen3-Coder-30B (UD-Q4_K_XL)Standard_NC24ads_A100_v4llamacppinteractive$1.00 (italynorth)measured
Qwen3-Coder-30B (UD-Q4_K_XL)Standard_NC80adis_H100_v5llamacppinteractive$3.64 (indonesiacentral)measured

Qwen3-Coder-Next-80B on Azure

VariantTypeDownloadCaleta ScoreMeasured
Qwen3-Coder-Next-80B (Q4)MoE, 3B active46 GB78NC80: 169 t/s

Proven deployments

VariantVM sizeEngineModeAll-in USD/hr*Status
Qwen3-Coder-Next-80B (UD-Q4_K_XL)Standard_NC24ads_A100_v4llamacppinteractive$1.00 (italynorth)measured
Qwen3-Coder-Next-80B (UD-Q4_K_XL)Standard_NC80adis_H100_v5llamacppinteractive$3.64 (indonesiacentral)measured

Qwen3.6-27B on Azure

VariantTypeDownloadCaleta ScoreMeasured
Qwen3.6-27B (Q4)dense 27B17 GB90NC80 2xH100: 72.7 t/s

Proven deployments

VariantVM sizeEngineModeAll-in USD/hr*Status
Qwen3.6-27B (UD-Q4_K_XL)Standard_NC24ads_A100_v4llamacppinteractive$1.00 (italynorth)measured
Qwen3.6-27B (UD-Q4_K_XL)Standard_NC80adis_H100_v5llamacppinteractive$3.64 (indonesiacentral)measured

Qwen3.6-35B-A3B on Azure

VariantTypeDownloadCaleta ScoreMeasured
Qwen3.6-35B-A3B (Q4)MoE, 3B active21 GB76NC80: 205 t/s
Qwen3.6-35B-A3B (Azure Cobalt Arm)MoE coder, 3B active22 GB-~34 t/s Cobalt CPU · best coder quality (SWE-bench 73.4)

Proven deployments

VariantVM sizeEngineModeAll-in USD/hr*Status
Qwen3.6-35B (UD-Q4_K_XL)Standard_NC24ads_A100_v4llamacppinteractive$1.00 (italynorth)measured
Qwen3.6-35B (UD-Q4_K_XL)Standard_NC80adis_H100_v5llamacppinteractive$3.64 (indonesiacentral)measured

DeepSeek-V4-Flash-0731 on Azure

VariantTypeDownloadCaleta ScoreMeasured
DeepSeek-V4-Flash-0731 (MXFP4)MoE, 13B active155 GBunscoredNC80: 54.6 t/s · E64: 9.7 t/s

Proven deployments

VariantVM sizeEngineModeAll-in USD/hr*Status
DeepSeek-V4-Flash-0731 (MXFP4)Standard_NC80adis_H100_v5llamacppinteractive$3.64 (indonesiacentral)measured

gpt-oss-120b on Azure

VariantTypeDownloadCaleta ScoreMeasured
gpt-oss-120b (MXFP4)MoE, 5B active61 GBunscoredNC24 1xA100: 140 t/s · E64: 48 t/s
gpt-oss-120b (Azure Cobalt Arm)MoE reasoner, 5B active63 GB-~53 t/s Cobalt CPU · 120B brain at 5B-active speed; value pick

Proven deployments

VariantVM sizeEngineModeAll-in USD/hr*Status
gpt-oss-120b (MXFP4)Standard_NC24ads_A100_v4llamacppinteractive$1.00 (italynorth)measured

Qwen3-Coder-480B on Azure

Proven deployments

VariantVM sizeEngineModeAll-in USD/hr*Status
Qwen3-Coder-480B (Q8_0)Standard_ND96is_MI300X_v5llamacppinteractive$11.49 (francecentral)measured

Qwen3.8-27B on Azure

VariantTypeDownloadCaleta ScoreMeasured
Qwen3.8-27B (Q4)dense 27B18 GBunscoredNC24 A100: 47.5 t/s

Proven deployments

VariantVM sizeEngineModeAll-in USD/hr*Status
Qwen3.8-27B (UD-Q4_K_XL)Standard_NC24ads_A100_v4llamacppinteractive$1.00 (italynorth)measured

ERNIE-4.5-21B-A3B on Azure

VariantTypeDownloadCaleta ScoreMeasured
ERNIE-4.5-21B-A3B (Azure Cobalt Arm)MoE general, 3B active13 GB-~67 t/s Cobalt CPU · tiny, Apache-2.0; fits E8ps_v6

GLM-4.5-Air on Azure

VariantTypeDownloadCaleta ScoreMeasured
GLM-4.5-Air (Azure Cobalt Arm)MoE mid, 12B active73 GB-~25 t/s Cobalt CPU · best mid-tier, MIT; needs E16ps_v6 or larger

Qwen3-30B-A3B-Instruct on Azure

VariantTypeDownloadCaleta ScoreMeasured
Qwen3-30B-A3B-Instruct (Azure Cobalt Arm)MoE general, 3B active19 GB-~63 t/s Cobalt CPU · fast general daily-driver

Qwen3-Next-80B-A3B on Azure

VariantTypeDownloadCaleta ScoreMeasured
Qwen3-Next-80B-A3B (Azure Cobalt Arm)MoE general, 3B active49 GB-~28 t/s Cobalt CPU · smart, not a speed pick

*All-in prices are the VM spot price as measured in the listed region for our latest snapshot, plus the appliance software fee for that size as read from the Marketplace price grid. Both land on your Azure bill. Spot moves; treat these as the shape of the economics, not a quote.

Get notified when the catalogue changes

Leave your email and we will tell you when the rankings or the model catalogue change.

Only catalogue and ranking updates, no marketing. Unsubscribe any time. See our Privacy Policy.

Every one of these runs on the same Marketplace image.