F6 - Foundry development options, PTUs

				MODEL
	  ┌──────────┴──────────┐
	  ▼                                                              ▼
SERVERLESS API                     MANAGED COMPUTE
     │                                                               │
 Microsoft-managed                              dedicated GPU
   model serving                                              capacity
serverless API

  • Microsoft hosts and manages the model-serving infrastructure and exposes an inference API

provisioned throughput

  • deployment reserves model-processing capacity for the application using Provisioned Throughput Units (PTUs)

estimate workload
      ↓
estimate PTUs
      ↓
benchmark
      ↓
observe utilization/latency
      ↓
adjust
managed compute

  • managed compute hosts supported open-source, partner and custom models on dedicated GPU capacity, while Microsoft manages the serving runtime and underlying infrastructure

serving path billing
serverless standard tokens
serverless provisioned PTUs
managed compute accelerator/GPU time