Azure AI Foundry and Self-Hosted LLMs, With Cost Discipline
We help Azure shops run AI properly. Foundry deployment in your tenant, self-hosting built on our own Private AI Appliance, and a FinOps overlay across both. The same cost discipline we apply to Azure infrastructure, applied to AI.
AI Foundry, Deployed Properly
Azure AI Foundry inside your tenant, on your security boundary, in the right region for your workload and budget. Region choice and SKU selection set the cost ceiling for everything that follows, so we get this part right first.
Foundry Deployment in Your Tenant
Deploy Azure AI Foundry behind private endpoints with RBAC, content filtering, Key Vault integration, and monitoring on day one. Same security boundary as the rest of your Azure estate.
Region & Region-Pricing Guidance
Foundry pricing varies sharply between regions and SKUs. We map your model and quota requirements to the cheapest region that still meets your latency and data-residency needs, often well under the default choice.
Multi-Model Routing
Route expensive prompts to capable models, simple ones to cheaper models. GPT-5 family, Mistral, Llama, DeepSeek: all run on Azure compute, all under one RBAC and one bill. Anthropic Claude available too, but routes to Anthropic infrastructure (worth knowing if data residency is your driver).
Self-Hosted Spot GPU Lab on Azure
Batch inference, fine-tuning, embeddings, evals. Once a workload is sustained and high-volume, managed-API pricing breaks down fast. A spot GPU pool running our Private AI Appliance behind its OpenAI-compatible API can cost a fraction of the managed equivalent, and the data stays inside your VNET. Our break-even analysis shows you the actual numbers for your workload.
GPU Right-Sizing
A100 vs T4 vs L4 vs MI300X. We benchmark against your actual workload and pick the right SKU instead of the biggest one. Often the cheaper card wins.
Spot GPU Lab on Azure
Run sustained-throughput workloads (batch inference, fine-tuning, evals, embeddings) on Azure Spot GPU VMs for up to 90% discount. Checkpoint-based pipelines that handle eviction gracefully.
Self-Hosted: the Caleta Appliance
For self-hosting we deploy our own Private AI Appliance: llama.cpp, an OpenAI-compatible API and a private chat UI, every model measured on the exact image. For bespoke builds we also set up vLLM where sustained throughput calls for it. Useful when token volume makes managed APIs painful, or when you need data to stay in your VNET.
Not always the right answer. Spiky, low-volume workloads usually belong on managed APIs. Steady, high-volume workloads usually don't. Our break-even analysis tells you which side of the line you're on before you commit.
FinOps Overlay for AI Workloads
The same blueprints approach as our Azure FinOps service, applied to AI. Break-even analysis, scheduling, pricing gotchas, and the dashboards that catch the £4,000 agentic-loop bill before month-end.
Break-Even Analysis
At what monthly token volume does self-hosting beat managed APIs? We model managed-API pricing vs spot GPU vs reserved GPU and show you the crossover point before you commit to either path.
Working-Hours Scheduling
Dev and staging GPU instances left running around the clock are the new "forgotten VMs". Auto-shutdown policies, scale-to-zero where possible, and on-demand resume. Non-prod GPU hours drop to the hours actually worked.
Pricing Gotcha Audit
Per-token pricing differences between regions, hidden Foundry hosting fees, PTU vs PAYG break-even, egress on RAG pipelines, log ingestion volume from agentic workflows. The bills that surprise people.
Token Consumption Dashboard
Which teams, which models, how many tokens, what cost, by week. Anomaly alerts when an agentic loop goes wrong and racks up £4,000 overnight.
Sustained-Throughput Routing
Steady, high-volume workloads belong on reserved or self-hosted infrastructure. Spiky, low-volume workloads belong on PAYG. We split your traffic correctly.
Regional Arbitrage
Latency-insensitive AI workloads routed to cheaper Azure regions. UK South vs East US2 vs Sweden Central pricing differences are real, and Foundry SKU availability varies between them too.
Why Us, Not an AI Specialist?
AI specialists optimise the model. We optimise the infrastructure and the bill.
| AI Specialists | Caleta |
|---|---|
| Focuses on the model and the use case | Focuses on the infrastructure and the bill |
| Demo-driven proof-of-concept first | Production deployment in your tenant, on your security boundary |
| "We'll figure out cost later" | Cost model on day one. Break-even and routing decisions made up front |
| Managed APIs only, vendor lock-in baked in | Managed APIs, self-hosted spot GPU, or hybrid. Whichever the maths supports |
| Separate AI silo bolted onto the org | AI deployed as part of your existing Azure infrastructure with FinOps overlay |
How We Engage
Start with a focused review. No long pitch, no obligation.
Book a 30-Min Review
We talk through your current AI plans, workloads, and Azure environment.
We Audit
Read-only access to your tenant. Foundry config, GPU SKUs, region choice, token spend, governance gaps. 3-5 working days.
Receive the Report
Cost projections, region recommendations, break-even analysis, gotcha list, and a 90-day roadmap.
Build or Hand Over
Implement yourself with the report, or engage us for the deployment and FinOps overlay.
Works With Our Other Services
Azure FinOps
AI cost work usually surfaces broader Azure overspend. Our FinOps service tackles waste across the whole estate.
Smart Hands
Deploying GPU hardware in a Slough data centre? Our Smart Hands team provides same-day installation and support across 38+ facilities.
Data Shuttle
Moving large training datasets or model artefacts between cloud and on-prem? Data Shuttle transfers terabytes in hours, not weeks.
AI & Infrastructure Insights
Practical guidance on AI infrastructure and cost discipline
4.3 Million Downloads: The Version of Qwen People Actually Run
Qwen3.8-27B's community build passed 4.3 million monthly downloads on Hugging Face, four times the official release. Across the ecosystem, 12.3 of 14.4 million downloads are community builds. Self-hosting isn't the niche. It's the mainstream.
You don't need $10,000 of hardware to run DeepSeek's new model
All 284 billion parameters at 16.2 tokens a second on a $0.44 an hour Azure box with no GPU in it. Five machines, one engine build, three measured numbers each: how fast it writes, how fast it reads your prompt, and how much of the memory bus it actually uses. Observed spot prices with region and date on every figure, and no derived economics anywhere.
A private AI server for every person on the team, from $16 a month
Not a seat on someone else's service: a whole private server each, inside your own Azure tenancy, from $16 a month in compute. Measured on Microsoft's Arm boxes: 0.26s first token, a cache that makes every request after the first instant, a server that switches itself off when idle, and seven models resident on one machine.
I ran a 744 billion parameter AI model with no GPU, for under a dollar an hour
GLM-5.2, one of the most capable open weight models in the world, running entirely in RAM on a GPU-less Azure box at 94 cents an hour on spot. Slow for one person, bleeding edge software, and early numbers that point somewhere interesting: batch output around $9.60 per million tokens.
We help Azure shops control AI costs
Get in touch for a 30-minute review. No obligation, no long pitch. Just a focused conversation about your AI plans and where the costs are likely to land.