The surge in AI adoption is transforming everything from chatbots to content generation. Still, a common pain point remains: How can organizations confidently size GPU resources for inference workloads and optimize Total Cost of Ownership (TCO)? With a dizzying mix of latency targets, model choices, quirky traffic patterns, and budget constraints, it’s easy to feel lost in the weeds, even before you’ve deployed a single model.
Today’s inference landscape is shaped by more than just hardware specs or “tokens per second.” Teams face an ever-growing list of sizing decisions: What kind of latency actually matters, Time to First Token (TTFT) average, 99th percentile latency, intertoken latency, or something else? How will your use case’s token patterns drive GPU memory and compute needs? What’s the right balance between on-prem core capacity and cloud-based elasticity?
This post offers a practical framework for mapping your use case to the right GPU footprint, sizing inference GPU infrastructure around real workload behavior rather than guesswork. We’ll walk through the inputs that matter most, including use case, token patterns, latency targets, concurrency, cache hit rate, model choice, and deployment strategy. Along the way, developers and infrastructure teams will see how core-and-flex capacity planning, right-sized GPUs, and model optimization techniques such as quantization, pruning, and distillation can improve performance while lowering TCO.
Cutting through the noise starts with one deceptively simple question: What problem are you solving? Different use cases map to wildly different infrastructure footprints. At a high level, most inference workloads fall into one of these four buckets:
| Use Case | Cached Input Tokens | Input Tokens | Output Tokens | Example Scenario |
|---|---|---|---|---|
| AI Chatbots/Copilots: Long input / short output | 1,000 – 5,000 | 2,000 – 8,000 | 200 – 800 | Limited RAG, multi-turn conversations |
| AI Agents: Extreme long context | >128,000 | 500 – 1,000 | 200 – 300 | Deep research, extended RAG |
| Content Generation: Short input / long output | 50 – 300 | 200 – 1,000 | 1,000 – 4,000 | Email/story generation, search |
| Translation Apps | 50 – 250 | 200 – 1,000 | 200 – 1,000 | Language/code translation, refactoring |
After mapping your use case, build your sizing plan around these dimensions:
Don’t let unpredictable workloads drive up costs. Adopt a core-and-flex strategy:
This model strikes a balance between capital efficiency (capex) and operational agility (opex), ensuring you aren’t overprovisioning or stifling growth.
The following scenarios illustrate how different enterprise workloads can approach GPU sizing and TCO estimation. These examples are intended for demonstration purposes only; actual GPU counts, configurations, and cost outcomes will vary based on model type, workload complexity, concurrency, and performance targets.
1. Financial services – Copilot for relationship managers
A local credit union deploys an AI copilot for relationship managers, analyzing complex client emails (long input, short output) to deliver quick, tailored responses and knowledge retrieval. Each copilot session averages 5,000 input tokens and 500 output tokens per query.
2. Life sciences — AI agent for drug discovery
A pharmaceutical start-up lab uses an AI agent to support scientific research teams, processing full-text research articles (extremely long context) to surface insights and generate result summaries. Queries often involve 20,000 input tokens and 2,000 output tokens.
3. Media and marketing – Real-time content generator
A mid-size digital marketing agency builds a generative system that creates personalized emails and ad copy from short briefs (500 input, 2,000 output tokens). With campaign launches requiring bursts of increased simultaneous users and a strong need to balance creative output speed and cost efficiency during production cycles.
4. Technology consulting — Large scale translation platform
An enterprise IT firm develops a multilingual code and documentation translation tool for client deployments. Each request (1,000 tokens input/output) comes from globally distributed teams.
Optimizing for TCO often comes down to strategically reducing your model’s memory footprint, one of the highest-leverage moves available. A smaller footprint lets you serve on compact or lower-cost GPUs, and in many cases drop an entire GPU tier. There are three levers, roughly in order of engineering effort:
None of these are one-time efforts; revisit them as models and workloads evolve. At enterprise scale, the cumulative savings in hardware, power, and operational overhead justify the investment.
Models typically ship in 16-bit floating point (FP16/BF16) at two bytes per parameter. Quantization re-expresses the weights, and optionally activations and KV cache, in an 8-bit format (FP8 or INT8) at one byte, roughly halving weight memory. That freed memory lets you drop to a smaller GPU, or fit a larger batch or longer KV cache on the same GPU to raise throughput and lower cost per token.
It’s the “quick win” because it needs no retraining. Post-training quantization (PTQ) converts an already-trained model in place, using only a small set of representative prompts for a calibration pass that sets per-layer scaling factors mapping the FP16 range into 8-bit with minimal distortion.
FP8 is the recommended starting point – typically close to lossless for inference, with more headroom than INT8 or INT4. Accuracy tolerance varies by use case; validate against your workload before production. When PTQ loss exceeds acceptable thresholds, escalate to Quantization-Aware Training (QAT), which fine-tunes with quantization simulated in the forward pass so weights adapt to lower precision. Refer to ModelOpt for more information.
NVIDIA ModelOpt makes PTQ a few lines of code:
import torch
import modelopt.torch.quantization as mtq
from modelopt.torch.export import export_hf_checkpoint
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained(
"meta-llama/Llama-3.1-8B-Instruct", dtype=torch.float16, device_map="auto"
).eval()
tokenizer = AutoTokenizer.from_pretrained("meta-llama/Llama-3.1-8B-Instruct")
def calibration_loop(model):
for prompt in ["Summarize this client email:", "What are the key risks here?"]:
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
with torch.no_grad():
model(**inputs)
model = mtq.quantize(model, mtq.FP8_DEFAULT_CFG, forward_loop=calibration_loop)
export_hf_checkpoint(model, export_dir="./llama-3.1-8b-fp8")

As shown in Figure 1, above, FP8 quantization reduces Llama-3.1-8B weight memory from 16.06GB to 9.08GB — a 43.5% reduction with no retraining. For deeper walkthroughs, see Optimizing LLMs for Performance and Accuracy with Post-Training Quantization, Post-Training Quantization of LLMs with NVIDIA NeMo and TensorRT Model Optimizer, and Turning FP8 Checkpoints into High-Performance Inference Engines with TensorRT.
When quantization isn’t enough, pruning and distillation enable deeper compression. Pruning removes less critical components, including entire layers (depth pruning) or attention heads, FFN channels, and embedding dimensions (width pruning). Knowledge distillation recovers accuracy by training the pruned student against the original teacher. The one-time compute cost is offset by sustained savings in hardware utilization, power, and operational overhead.
The example below uses NVIDIA NeMo with Qwen3-8B as the teacher, targeting a ~6B student. Convert the Hugging Face model to NeMo checkpoint format and preprocess WikiText-103-v1 first, then prune. The pruning step is where the architecture is actually reshaped – depth pruning trims 36->24 layers, while width pruning shrinks ffn_hidden_size 12288->9216 and hidden_size 4096->3584, both yielding roughly 6B parameters:
Prerequisites: 2x NVIDIA H100 or A100 80GB GPUs, a Docker-enabled environment, and the NeMo container (nvcr.io/nvidia/nemo:25.11, nvidia-modelopt==0.37.0).
NEMO_TEACHER_PATH="./nemo_ckpt/Qwen3-8B.nemo" # original (un-pruned) Qwen3-8B .nemo
# Pruning runs single-GPU — FastNAS asserts tp_size=1 for both depth and width pruning
PRUNE_COMMON="--devices 1 --tp_size 1 --pp_size 1 \
--restore_path ${NEMO_TEACHER_PATH} --legacy_ckpt \
--seq_length ${SEQ_LENGTH} --num_train_samples ${NUM_TRAIN_SAMPLES} --mbs ${MICRO_BATCH_SIZE} \
--data_paths ${DATA_PATHS} --index_mapping_dir ${INDEX_MAPPING_DIR}"
PRUNE="torchrun --nproc_per_node 1 ${NEMO_ROOT}/scripts/llm/gpt_prune.py"
# Step 1a — Depth pruning: 36 → 24 layers (~6B model)
${PRUNE} ${PRUNE_COMMON} --save_path ${ROOT_DIR}/Qwen3-8B-nemo-depth-pruned \
--target_num_layers 24
# Step 1b — Width pruning: ffn_hidden_size 12288→9216, hidden_size 4096→3584
${PRUNE} ${PRUNE_COMMON} --save_path ${ROOT_DIR}/Qwen3-8B-nemo-width-pruned \
--target_ffn_hidden_size 9216 --target_hidden_size 3584
#Step 2 — Distillation: train each pruned student against the teacher
TRAIN="torchrun --nproc_per_node ${DEVICES} ${NEMO_ROOT}/scripts/llm/gpt_train.py"
# Step 2a — Depth-pruned student
${TRAIN} \
--name depth_distill --model_path ${ROOT_DIR}/Qwen3-8B-nemo-depth-pruned \
--teacher_path ${NEMO_TEACHER_PATH} --log_dir ${ROOT_DIR}/depth_distill_logs --legacy_ckpt \
--data_paths ${DATA_PATHS} --index_mapping_dir ${INDEX_MAPPING_DIR} --seq_length ${SEQ_LENGTH} \
--max_steps ${MAX_STEPS} --gbs ${GLOBAL_BATCH_SIZE} --mbs ${MICRO_BATCH_SIZE} \
--val_check_interval ${VAL_CHECK_INTERVAL} --precision bf16-mixed
# Step 2b — Width-pruned student (same as above, swap the two lines below)
# --name width_distill --model_path ${ROOT_DIR}/Qwen3-8B-nemo-width-pruned

Note: For demonstrative purposes, the dataset used here is comparatively small, so these numbers should be treated as illustrative rather than definitive. As Figure 2 shows, above, in this run width pruning reached a lower final validation loss (3.21 vs 3.60), while depth pruning converged faster. For scale reference, the NVIDIA published run produces a 6B depth-pruned model that is 30% faster than Qwen3-4B at higher MMLU accuracy (72.5 vs 70.0).
For the full pipeline, end-to-end documentation, and architectural recommendations, see LLM Model Pruning and Knowledge Distillation with NVIDIA NeMo Framework, Pruning and Distilling LLMs Using NVIDIA TensorRT Model Optimizer, and the Llama-3.1-to-Minitron blog.
GPU sizing is an ongoing optimization, not a fixed choice made at deployment. Quantization, pruning, and distillation give teams concrete levers to cut infrastructure cost without sacrificing performance. Start with quantization for an immediate footprint reduction, layer in pruning and distillation as workloads mature, and revisit as models evolve – producing smaller, faster models that expand AI’s reach across mobile, edge, and embedded applications.
To get started with model optimization, check out the NVIDIA Model Optimizer GitHub and Hugging Face with a deeper dive into model quantization on how to post-train quantization with Model Optimizer.