VCF 9.1 Private AI Foundation

Private AI Sizer

Size your GPU infrastructure for multi-model LLM inference on VMware Cloud Foundation. All calculations follow VCF 9.1 design recommendations.

GPU Selection

Cluster Configuration

2= 16 total GPUs

Model Stack

1 model
Model 1
Prompt length512 tokens
Replicas1
1 GPU
Request

Cluster GPU Allocation

Free
Llama 3.1 8B(1)
GPUs Used
1 / 16
GPUs Free
15
Utilization
6%

1Llama 3.1 8BFP16 · 4K ctx

1 replica · 1 GPUs
114
Concurrent
225
Tokens/sec
8 ms
TTFT
14.9 GB
Model Size
Weights 21%KV Cache 79%Free 0%

Host DRAM & Compute Requirements

Per Host
VM RAM (DRAM)
1.3 TB
2× GPU memory
Server RAM
1.6 TB
2.5× GPU memory
vCPUs
64
8 per GPU
Total Cluster (2 hosts)
Total VM RAM
2.5 TB
Total Server RAM
3.1 TB
Total vCPUs
128
Per VCF 9.1: VM RAM 1-2× GPU memory (PAIF-ACC-RCMD-008) to support vLLM runtime, model loading buffers, and KV cache overhead. Server RAM 2-3× (PAIF-ACC-RCMD-009) for hypervisor overhead. vCPUs 4-8 per GPU (PAIF-ACC-RCMD-011) for tokenization and request handling.

Methodology

All sizing formulas follow the VMware Cloud Foundation 9.1 Private AI Compute Detailed Design (PAIF-ACC-RCMD-001 through PAIF-ACC-RCMD-011). The 90% GPU memory utilization reflects the vLLM runtime default (gpu_memory_utilization=0.90), not a VCF specification.

Throughput and TTFT are theoretical single-request estimates. Production performance varies based on batching, concurrent load, model implementation, and vGPU vs. DirectPath I/O configuration. For production deployments, benchmark with your actual workloads.