A new October 10 Microsoft engineering case study examines serving Qwen3.8-27B on an Azure H100 Spot VM with vLLM, Application Insights, and Grafana. Its useful lesson is the difference between an impressive busy interval and the result across the full operating period. This is a published implementation study, not a new Spot service launch. Microsoft case study.
For teams considering self-hosted inference, I would turn that distinction into the first pilot requirement: account for the time and resources needed to deliver the workload, including periods when the model is not serving useful requests.
Keep the Comparison Honest
The author reports about 46 hours of VM runtime and roughly $98 in compute costs during the first week, compared with about $86 of token usage valued at the selected model's published API rates. The model-server uptime metric covered only 37.9 hours. Setup, loading, and early runs helped explain the difference. Reported measurement window.
That token valuation is a comparison benchmark, not revenue earned by the VM or proof of savings on your own bill. My recommendation is to compare the proposed service with the alternative your organization would actually use, under its applicable prices and workload requirements.
Keep workload quality in the comparison as well. A cheaper serving configuration is only a suitable alternative if its results meet the application's acceptance criteria.
Measure More Than Serving Uptime
I would maintain a pilot ledger with allocation time, initialization time, model-ready time, useful request volume, failed requests, and recovery periods. Reconcile the operational dashboard with the billed resources rather than assuming both clocks cover the same interval.
Use the same observation window for every side of the comparison. Include quieter periods if the application's normal traffic is uneven; selecting only a busy demonstration can hide the shape of the real workload.
Track supporting resources separately. Storage, networking, observability, and operator effort should be visible to the person approving the hosting approach, even when compute remains the largest item.
Make the dashboard's gaps explicit. If a metric stops arriving, distinguish an unavailable model server from an unavailable telemetry pipeline before interpreting the missing data as zero activity.
Spot Changes the Recovery Contract
Microsoft Learn states that Spot VMs have no SLA or high-availability guarantee and can be evicted when Azure needs the capacity. Scheduled Events notifications are best effort, up to 30 seconds before eviction. With the Deallocate policy, disks can keep incurring storage charges, and successful reallocation is not guaranteed. Spot availability and eviction behavior.
For an inference pilot, I would rehearse interruption with a safe workload. Record how requests are handled during shutdown, how the client learns that the service is unavailable, and what happens to work that was accepted but not completed.
Time the recovery until the application can successfully complete a request again. A running VM is only one milestone; the operator needs evidence that the serving path is ready for the intended workload.
Assign an owner to the fallback decision. Some applications can queue work, while others need a separately planned capacity option. Define that choice before using the pilot to support a production hosting proposal.
Treat the Hourly Rate as a Dated Input
The current Spot guidance describes prices as variable by region and VM size, and allocation depends on available capacity and quota. Historical eviction data can help inform a placement review, but it is not a reservation of future capacity. Pricing and capacity considerations.
I would record the region, SKU, price observation date, and quota assumptions with every comparison. Recheck those inputs at the point of deployment instead of copying a headline hourly figure into a long-lived budget.
Before considering another region, review the application's data-location and network requirements. A rate comparison should remain inside the locations the organization is willing and permitted to operate.
Practical Cloud Engineer Takeaway
Compare the complete operating period, including initialization and recovery.
Treat token-rate valuation as a benchmark, not earned revenue.
Reconcile serving metrics with billed runtime and supporting resources.
Rehearse interrupted requests and recovery using safe test data.
Name the fallback owner and preserve the application's quality requirements.
Date every price, capacity, and quota assumption.
Who Should Care?
Azure infrastructure engineers, AI platform teams, and application owners evaluating GPU-hosted inference for workloads that can tolerate interruptions.
Make the Hosting Decision Reproducible
The useful outcome is a pilot report that another engineer can understand and repeat: its workload, measurement window, resource costs, quality results, and recovery evidence should all be clear.
The evaluation steps above are recommendations. The measurements belong to the linked Microsoft case study; I have not reproduced its benchmark or deployed an H100 for this article.
Sources
Stay radical, stay curious, and keep pushing the boundaries of what is possible in the cloud.
Chriz
Beyond Cloud with Chriz
Comments