meta-llama/Llama-3.3-70B-Instruct
Llama 3.3 70B dense model with NVIDIA FP8/FP4 quantized variants for Hopper and Blackwell GPUs
Guide
Overview
Llama 3.3 70B Instruct is Meta's 70-billion parameter dense language model. NVIDIA provides FP8 and FP4 quantized variants optimized for Hopper (H100/H200) and Blackwell (B200/GB200) GPUs. FP4 is Blackwell-only and provides the best VRAM efficiency.
TPU support is provided through vLLM TPU with a recipe for Trillium.
Prerequisites
- Hardware: 1x H100/H200 (FP8), 1x B200 (FP4), 1x MI300X/MI325X/MI355X, 2x GPUs, 8x Xeon6/Xeon5 NUMA nodes, or 4x Intel Arc Pro B70 (online FP8)
- vLLM >= 0.12.0
- CUDA Driver >= 575 for GPUs
- Docker with NVIDIA Container Toolkit (recommended) for GPUs
AMD ROCm
The ROCm wheel requires Python 3.12, ROCm 7.0+, and glibc >= 2.35. Use the ROCm Docker image when your host environment does not meet those requirements.
uv venv --python 3.12
source .venv/bin/activate
uv pip install vllm --extra-index-url https://wheels.vllm.ai/rocm
Launch on 1x MI300X / MI325X / MI355X with AITER enabled:
export VLLM_ROCM_USE_AITER=1
vllm serve meta-llama/Llama-3.3-70B-Instruct \
--tensor-parallel-size 1
The first AITER launch may take several minutes while ROCm kernels are compiled and cached.
pip (Intel Xeon 6 CPUs)
For Intel and AMD x86 CPUs, follow the CPU pre-built wheels installation instructions.
Docker (Intel Xeon 6 CPUs)
docker pull vllm/vllm-openai-cpu:latest-x86_64 # For Intel Xeon 6
Docker (Cloud TPU — Trillium)
TPU uses the separate vllm/vllm-tpu image (no pip wheel). Pull the tag specified by the upstream Trillium recipe, then run:
docker run -itd --name llama33-tpu \
--privileged --network host --shm-size 16G \
-v /dev/shm:/dev/shm -e HF_TOKEN=$HF_TOKEN \
vllm/vllm-tpu:latest \
--model meta-llama/Llama-3.3-70B-Instruct \
--tensor-parallel-size 8 \
--max-model-len 16384 \
--host 0.0.0.0 --port 8000
Trillium requires a 4-chip slice minimum.
Intel Xeon 6 Deployment via Docker
Launch the x86 CPU vLLM Docker container for meta-llama/Llama-3.3-70B-Instruct:
docker run \
--privileged --ipc=host -p 8000:8000 \
-v ~/.cache/huggingface:/root/.cache/huggingface \
vllm/vllm-openai-cpu:latest-x86_64 meta-llama/Llama-3.3-70B-Instruct \
--tensor-parallel-size 8
TP/DP remain user/deployment choices rather than recipe defaults.
The following validation settings are not model requirements and are intentionally not portable CPU recipe defaults:
--max-num-batched-tokensand--max-num-seqs: scheduler batch/concurrency tuning for workload shape, latency targets, and platform capacity.--gpu-memory-utilization: platform memory-budget tuning.--no-enable-prefix-caching: workload/benchmark cache-policy tuning.VLLM_ENGINE_ITERATION_TIMEOUT_S: runtime operational timeout tuning.
Hardware-specific overrides may still use these settings when they are part of a validated platform configuration.
Intel Arc Pro B70 (XPU)
Validated on 4x Intel Arc Pro B70 (32 GB per card) with the official vLLM XPU
image vllm/vllm-openai-xpu:latest, TP=4. BF16 weights (~140 GB) exceed the
128 GB node, so the validated configuration quantizes the base checkpoint
online to FP8 (--quantization fp8) with an FP8 KV cache.
docker run --device /dev/dri \
-v /dev/dri/by-path:/dev/dri/by-path --shm-size=16g \
--privileged --ipc=host -p 8000:8000 \
-v ~/.cache/huggingface:/root/.cache/huggingface \
-e HF_TOKEN=$HF_TOKEN \
-e PYTHONUNBUFFERED=1 \
-e TORCH_LLM_ALLREDUCE=1 \
-e VLLM_ALLOW_LONG_MAX_MODEL_LEN=1 \
-e VLLM_WORKER_MULTIPROC_METHOD=spawn \
vllm/vllm-openai-xpu:latest \
meta-llama/Llama-3.3-70B-Instruct \
--tensor-parallel-size 4 \
--block-size 64 \
--enforce-eager \
--no-enable-prefix-caching \
--disable-sliding-window \
--max-model-len 9472 \
--max-num-batched-tokens 8192 \
--gpu-memory-utilization 0.9 \
--quantization fp8 \
--kv-cache-dtype fp8
Client Usage
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="unused")
response = client.chat.completions.create(
model="nvidia/Llama-3.3-70B-Instruct-FP8",
messages=[{"role": "user", "content": "Hello, how are you?"}],
)
print(response.choices[0].message.content)
Troubleshooting
FP4 variant not loading: FP4 is only supported on Blackwell (compute capability 10.0). Use FP8 on Hopper.
OOM with BF16 on single GPU: Use the FP8 variant (~70 GB) or FP4 variant (~40 GB) to fit on a single GPU.