In this guide, we show how to serve Qwen3-4B and Qwen3-32B.
You can reproduce this experiment from your dev environment (e.g. your laptop). You need to install gcloud locally to complete this tutorial.
To install gcloud cli please follow this guide: Install the gcloud CLI
Once it is installed, you can login to GCP from your terminal with this command: gcloud auth login.
We create a single VM. For Qwen3-4B, 1 chip is sufficient and for the 32B model, at least 4 chips are required. If you need a different number of chips, you can set a different value for --topology such as 1x1, 2x4, etc.
To learn more about topologies: v6e VM Types.
Note: Acquiring on-demand TPUs can be challenging due to high demand. We recommend using Queued Resources to ensure you get the required capacity.
This command attempts to create a TPU VM immediately.
export TPU_NAME=your-tpu-name
export ZONE=your-tpu-zone
export PROJECT=your-tpu-project
# this command creates a tpu vm with 4 Trillium (v6e) chips - adjust it to suit your needs
gcloud alpha compute tpus tpu-vm create $TPU_NAME \
--type v6e --topology 2x2 \
--project $PROJECT --zone $ZONE --version v2-alpha-tpuv6eWith Queued Resources, you submit a request for TPUs and it gets fulfilled when capacity is available.
export TPU_NAME=your-tpu-name
export ZONE=your-tpu-zone
export PROJECT=your-tpu-project
export QR_ID=your-queued-resource-id # e.g. my-qr-request
# This command requests a v6e-4 (4 chips). Adjust accelerator-type for different sizes.
# For 1 chip (Qwen3-4B), use --accelerator-type v6e-1.
gcloud alpha compute tpus queued-resources create $QR_ID \
--node-id $TPU_NAME \
--project $PROJECT --zone $ZONE \
--accelerator-type v6e-4 \
--runtime-version v2-alpha-tpuv6eYou can check the status of your request with:
gcloud alpha compute tpus queued-resources list --project $PROJECT --zone $ZONEOnce the state is ACTIVE, your TPU VM is ready and you can proceed to the next steps.
gcloud compute tpus tpu-vm ssh $TPU_NAME --project $PROJECT --zone=$ZONEexport DOCKER_URI=vllm/vllm-tpu:latestsudo docker run -it --rm --name $USER-vllm --privileged --net=host \
-v /dev/shm:/dev/shm \
--shm-size 100gb \
--entrypoint /bin/bash ${DOCKER_URI}Note: 100GB should be sufficient for the 32B model. For the 4B model allocate at least 10GB for the weights. See this guide for attaching durable block storage to TPUs.
Export your hugging face token along with other environment variables inside the container.
export HF_HOME=/dev/shm
export HF_TOKEN=<your HF token>Now we start the vllm server. Make sure you keep this terminal open for the entire duration of this experiment.
export MAX_MODEL_LEN=4096
export TP=4 # number of chips
vllm serve Qwen/Qwen3-32B \
--seed 42 \
--disable-log-requests \
--gpu-memory-utilization 0.98 \
--max-num-batched-tokens 2048 \
--max-num-seqs 256 \
--tensor-parallel-size $TP \
--max-model-len $MAX_MODEL_LENFor the 4B model, we recommend --max-num-batched-tokens 1024 --max-num-seqs 128.
It takes a few minutes depending on the model size to prepare the server - once you see the below snippet in the logs, it means that the server is ready to serve requests or run benchmarks:
INFO: Started server process [7]
INFO: Waiting for application startup.
INFO: Application startup complete.
INFO: Uvicorn running on http://0.0.0.0:8000 (Press CTRL+C to quit)Open a new terminal to test the server and run the benchmark (keep the previous terminal open).
First, we ssh into the TPU vm via the new terminal:
export TPU_NAME=your-tpu-name
export ZONE=your-tpu-zone
export PROJECT=your-tpu-project
gcloud compute tpus tpu-vm ssh $TPU_NAME --project $PROJECT --zone=$ZONEsudo docker exec -it $USER-vllm bashLet's submit a test request to the server. This helps us to see if the server is launched properly and we can see legitimate response from the model.
curl http://localhost:8000/v1/completions \
-H "Content-Type: application/json" \
-d '{
"model": "Qwen/Qwen3-32B",
"prompt": "I love the mornings, because ",
"max_tokens": 200,
"temperature": 0
}'You will need to install datasets as it's not available in the base vllm image.
pip install datasetsFinally, we are ready to run the benchmark:
export MAX_INPUT_LEN=1800
export MAX_OUTPUT_LEN=128
export HF_TOKEN=<your HF token>
cd /workspace/vllm
vllm bench serve \
--backend vllm \
--model "Qwen/Qwen3-32B" \
--dataset-name random \
--num-prompts 1000 \
--random-input-len=$MAX_INPUT_LEN \
--random-output-len=$MAX_OUTPUT_LEN \
--seed 100The snippet below is what you’d expect to see - the numbers vary based on the vllm version, the model size and the TPU instance type/size.
============ Serving Benchmark Result ============
Successful requests: xxxxxxx
Benchmark duration (s): xxxxxxx
Total input tokens: xxxxxxx
Total generated tokens: xxxxxxx
Request throughput (req/s): xxxxxxx
Output token throughput (tok/s): xxxxxxx
Total Token throughput (tok/s): xxxxxxx
---------------Time to First Token----------------
Mean TTFT (ms): xxxxxxx
Median TTFT (ms): xxxxxxx
P99 TTFT (ms): xxxxxxx
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): xxxxxxx
Median TPOT (ms): xxxxxxx
P99 TPOT (ms): xxxxxxx
---------------Inter-token Latency----------------
Mean ITL (ms): xxxxxxx
Median ITL (ms): xxxxxxx
P99 ITL (ms): xxxxxxx
==================================================