Inference Triage
A symptom-driven debugging guide for LLM inference in production. Pick your engine above, find your symptom on the left, and follow the checks in order — each step names the exact metric, command, or flag, and says what the answer means. Written from real on-call incidents on large GPU fleets.
The triage method#
Every latency problem in LLM serving decomposes into the same anatomy. A request's end-to-end time is:
E2E = queue wait + prefill (→ TTFT) + decode (ITL × output tokens)
- Queue wait — the scheduler hasn't admitted the request yet (not enough KV-cache blocks or batch slots).
- Prefill — processing the whole prompt in one shot. Compute-bound: cost grows with prompt length (plus an attention term). Ends at the first output token, so TTFT = queue + prefill.
- Decode — one token per engine step. Memory-bandwidth-bound: each step re-reads the weights and the growing KV cache. ITL (inter-token latency) rises with the size of the running batch.
This gives you the first fork of every investigation:
TTFT high, ITL normal → the problem is before decode: queueing or prefill. Go to §01.
ITL high (tokens trickle out) → decode is the bottleneck: batch too large, bandwidth, parallelism. Go to §02.
Only p99 is bad, median fine → intermittent contention: preemptions, head-of-line blocking, cold caches. Go to §03.
Three rules that save hours:
- Localize before you tune. Measure at the engine first (its own metrics), then at the gateway, then at the client. If the engine says TTFT is 200 ms but the user sees 4 s, the problem is in front of the engine — load balancer, router, gateway timeout, TLS, or client-side token counting. Don't touch engine flags for a network problem.
- One lever at a time. Almost every knob trades latency against throughput. Change one, re-measure the same workload, keep or revert.
- Know your traffic. Prompt-length and output-length distributions decide everything. A p99 prompt of 60k tokens explains "random" TTFT spikes that no config change will fix.
Everything below assumes you can read the engine's Prometheus metrics. vLLM: exposed at /metrics by default. SGLang: start with --enable-metrics. TensorRT-LLM: the Triton backend exposes nv_trt_llm_* on the metrics port (:8002/metrics by default); trtllm-serve also exposes Prometheus metrics on recent versions. If you have no metrics, your first fix is deploying them.
High TTFT#
TTFT = queue wait + prefill. The single most important question: is the request waiting, or is prefill itself slow? One metric answers it.
# is there a queue?
curl -s localhost:8000/metrics | grep -E 'num_requests_(waiting|running)'
# vllm:num_requests_waiting > 0 sustained → queueing
# also check queue time directly (newer versions):
curl -s localhost:8000/metrics | grep request_queue_time
# is there a queue? (requires --enable-metrics)
curl -s localhost:30000/metrics | grep -E 'num_(queue|running)_reqs'
# sglang:num_queue_reqs > 0 sustained → queueing
# the console log also prints: #running-req and #queue-req every few steps
# Triton backend — is there a queue?
curl -s localhost:8002/metrics | grep nv_trt_llm_request_metrics
# request_type="waiting" > 0 sustained → queueing
# request_type="scheduled" vs "active" shows what actually made it into the batch
Path A — requests are queueing
The scheduler can't admit requests. Admission is gated by two budgets: batch slots (max concurrent sequences) and free KV-cache blocks. Find which one is exhausted:
KV cache full? vllm:gpu_cache_usage_perc pinned near 1.0 → the KV pool is the limit, not compute.
Levers: raise --gpu-memory-utilization (default 0.9); use --kv-cache-dtype fp8 (halves KV size); lower --max-model-len if you don't need it (it bounds worst-case reservation); shrink weights via quantization to free VRAM for KV.
Batch slots full? vllm:num_requests_running flat at --max-num-seqs while cache usage < 0.9 → raise --max-num-seqs (watch ITL, see §02).
Neither full but still queueing? Look at vllm:num_preemptions_total — preempted requests re-enter the queue and re-prefill (see §03).
Genuinely at capacity → add replicas and load-balance; this engine is full.
KV pool full? sglang:token_usage near 1.0 → the token pool (KV) is the limit.
Levers: raise --mem-fraction-static; use --kv-cache-dtype fp8_e4m3; check sglang:cache_hit_rate — a high radix-cache hit rate stretches the same pool much further, so make sure you're not defeating it (random prompt prefixes, disabled cache).
Batch slots full? sglang:num_running_reqs flat at --max-running-requests → raise it (watch ITL).
Retractions? If decode runs out of KV, SGLang retracts requests back to the queue (log lines mention retract). Lower --schedule-conservativeness below 1.0 to admit less aggressively, or add KV headroom.
Genuinely at capacity → add replicas / data parallelism (--dp-size).
KV blocks exhausted? nv_trt_llm_kv_cache_block_metrics{kv_cache_block_type="free"} near 0 → KV is the limit.
Levers: raise kv_cache_free_gpu_mem_fraction (runtime config, default ~0.9); enable FP8 KV cache (quantized checkpoint / kv_cache_dtype); rebuild the engine with a smaller max_seq_len if over-provisioned.
Batch full? Active requests flat at the engine's max_batch_size → rebuild with a larger trtllm-build --max_batch_size (TRT-LLM bakes this into the engine) and check max_num_tokens too.
Scheduler policy: guaranteed_no_evict (default) admits conservatively — it reserves KV for each request's worst case. max_utilization admits more but can evict under pressure (tail risk, §03).
Genuinely at capacity → add replicas; Triton can also queue-and-reject with a queue policy instead of letting TTFT grow unbounded.
Path B — no queue, prefill itself is slow
Then TTFT is honest compute time. Three questions:
- How long are the prompts? Check the engine's prompt-token histograms or your gateway logs. Prefill cost is roughly linear in prompt tokens (quadratic attention term kicks in at long context). A 50k-token prompt legitimately takes seconds — the fix is caching, not tuning.
- Are you re-computing shared prefixes? System prompts, few-shot examples, and RAG scaffolding are identical across requests. Prefix caching turns those tokens into ~free.
- Is a long prefill blocking the batch? Without chunked prefill, one giant prompt monopolizes an engine step and everyone else's TTFT/ITL spikes with it.
- Prefix caching: on by default in V1 (
--enable-prefix-caching). Verify it's working:vllm:gpu_prefix_cache_hit_rate(or hits/queries counters). Low hit rate with a shared system prompt → check that requests share an exact token prefix (a timestamp in the system prompt kills it). - Chunked prefill: default-on in V1; on V0 set
--enable-chunked-prefill. Tune the prefill/decode mix with--max-num-batched-tokens— smaller values favor ITL, larger favor prefill throughput. - Multiple API server workers (
--api-server-count) if HTTP/tokenization is the bottleneck at high QPS (CPU-bound frontend, see §09).
- RadixAttention prefix caching is the default and is SGLang's superpower — watch
sglang:cache_hit_rate. Structure prompts so shared content is a literal prefix; route similar-prefix traffic to the same replica (cache-aware routing) to multiply hit rate. - Chunked prefill: tune
--chunked-prefill-size(lower = decode-friendly, higher = prefill-friendly). - Schedule policy:
--schedule-policy lpm(longest-prefix-match, default) reorders the queue for cache hits; if strict fairness matters more,fcfs.
- KV cache reuse (prefix caching): enable
enable_kv_cache_reuse(engine build + runtime). Requires paged KV cache (default in recent builds). - Chunked context: enable
enable_chunked_contextso long prompts don't monopolize inflight batching; requires the engine to be built with appropriatemax_num_tokens. - Engine build matters: prefill speed is baked at build time — check the engine was built for the right GPU arch and with enough
max_num_tokensto batch prefill efficiently.
Then it's not the workload: check GPU health (§09 — clock throttling makes prefill 2-4x slower), cold starts (first requests after deploy pay CUDA-graph capture / torch.compile warmup — always warm up before shifting traffic), or CPU-side tokenization (very long prompts + busy Python frontend).
Slow streaming — high ITL#
Decode is memory-bandwidth-bound: every engine step reads all weights plus the KV cache of every running sequence. The theoretical floor is bytes moved / memory bandwidth — on an H100 SXM (~3.3 TB/s) a 70B FP16 model (~140 GB) can't decode faster than ~24 tokens/s per sequence at batch 1, no matter what you tune. Everything above that floor is contention you can fight.
How big is the running batch? ITL scales with concurrent sequences — each step does more KV reads.
vLLM vllm:num_requests_running high and ITL (vllm:time_per_output_token_seconds) over SLO → cap --max-num-seqs lower. This is the fundamental latency/throughput dial.
SGLang sglang:num_running_reqs high and ITL over SLO → cap --max-running-requests. The log's gen throughput line shows aggregate tok/s — healthy aggregate + slow individual streams = batch is just large.
TensorRT-LLM active request count near max_batch_size → lower runtime max_batch_size (can be set below the engine's built max) or rebuild.
Is prefill stealing decode steps? Mixed batches share the token budget; heavy prompt traffic inflates ITL.
vLLM lower --max-num-batched-tokens so each step carries less prefill; chunked prefill (default V1) keeps chunks small.
SGLang lower --chunked-prefill-size.
TensorRT-LLM enable enable_chunked_context; consider a lower runtime max_num_tokens.
Shrink the bytes. Weight quantization (FP8/INT4) cuts the per-step weight read linearly; FP8 KV cache halves KV reads. Both are ITL wins beyond just fitting memory.
CUDA graphs on? Decode steps are launch-overhead-sensitive.
vLLM make sure you're NOT running --enforce-eager in production — it disables CUDA graphs and costs real ITL. (It's a debug flag; see §06.)
SGLang CUDA graphs are on by default; check --cuda-graph-max-bs covers your actual batch sizes — batches above it fall back to eager mode.
TensorRT-LLM kernels are compiled ahead of time — the equivalent failure is running an engine built for a different GPU/config, which silently underperforms.
Tensor parallelism tax. TP does an all-reduce every layer, every token. Over NVLink that's fine; over PCIe it dominates ITL. Rule of thumb: use the smallest TP that fits the model; prefer more replicas (data parallel) over wider TP for throughput. Verify links: nvidia-smi nvlink --status, and check the topology with nvidia-smi topo -m (you want NV#, not PHB/PIX, between TP peers).
Speculative decoding — if latency at low batch matters, a draft model / EAGLE-style speculation multiplies effective tokens per step. If it's already on and ITL got worse, check the acceptance rate (engine logs/metrics): a poorly matched draft wastes compute. It also helps least at high batch — the GPU is already busy.
Estimated ITL floor ≈ (weights bytes ÷ TP) / (memory BW × ~0.7 utilization) + KV read term. If measured ITL at batch 1 is already 2-3x above this on an idle GPU, suspect the GPU or the deployment (throttling, wrong arch build, eager mode), not the scheduler.
p99 tail latency#
Tails come from intermittent contention. The usual suspects, in the order they're likely:
1. Preemption / eviction / retraction
When the KV pool runs dry mid-flight, engines take a victim: the request loses its KV cache and later re-prefills from scratch — its E2E can double or worse. This is the #1 cause of mysterious p99.
Metric: vllm:num_preemptions_total climbing = smoking gun (also logged with a warning). Fixes: more KV headroom (--gpu-memory-utilization, FP8 KV, quantized weights), lower --max-num-seqs, and bound request lengths (--max-model-len, cap max_tokens per request at the gateway).
Look for retract events (metric/log). Fix: --schedule-conservativeness > 1.0 makes admission more conservative (fewer retractions, slightly lower utilization); add KV headroom as above.
With batch_scheduler_policy: max_utilization, eviction under pressure is by design — switch to guaranteed_no_evict (default) if tails matter more than peak utilization.
2. Head-of-line blocking by long prompts
One 80k-token prefill in the batch stalls every other stream for that step. Fix: chunked prefill (§01 Path B), and consider routing long-context traffic to a dedicated pool so interactive traffic never shares a batch with it. At fleet scale this becomes prefill/decode disaggregation (e.g. NVIDIA Dynamo): separate prefill workers feed KV to decode workers, so TTFT work never contends with ITL work.
3. Cold anything
- First requests after a deploy pay CUDA-graph capture,
torch.compile, and cache warmup — warm up before shifting traffic (send a few synthetic requests covering your batch shapes). - Prefix-cache misses after restart: hit rate takes minutes to recover; expect a TTFT bump and don't page on it.
- Autoscaling from zero: model load takes minutes (weights download + load + warmup) — keep a floor of warm replicas.
4. The fleet, not the engine
- One sick replica drags fleet p99 while every healthy node looks fine. Compare per-replica ITL/TTFT; a single outlier → check that GPU (§09: throttling, ECC, XID).
- Load balancing skew: least-connections beats round-robin for LLMs because request costs vary wildly; cache-aware routing (route same-prefix traffic to the same replica) is better still.
- Gateway timeouts/retries: a retry storm doubles load exactly when the system is slow. Check gateway retry policy against engine E2E p99.
Low throughput#
Throughput problems are budget problems. The engine can only run as many concurrent tokens as its scarcest budget allows. Find the binding constraint:
KV cache — usage ~100% with a queue → every extra GB of KV is direct throughput.
vLLM vllm:gpu_cache_usage_perc. Levers: --gpu-memory-utilization up (0.9 → 0.95 on a dedicated GPU), --kv-cache-dtype fp8, quantize weights, trim --max-model-len.
SGLang sglang:token_usage. Levers: --mem-fraction-static up, FP8 KV, and above all raise the radix-cache hit rate — shared-prefix tokens cost ~nothing.
TensorRT-LLM free KV blocks ~0. Levers: kv_cache_free_gpu_mem_fraction up, FP8 KV, max_utilization scheduling if tails allow.
Concurrency caps — running count flat at the configured max with KV to spare → raise --max-num-seqs--max-running-requestsmax_batch_size until ITL hits your SLO. Throughput rises with batch size until bandwidth saturates.
Token budget per step — --max-num-batched-tokens too small starves prefill throughput;--chunked-prefill-size too small starves prefill throughput;max_num_tokens too small starves prefill throughput; too large hurts ITL. Sweep it against your real traffic mix.
Not enough offered load — the embarrassing one. If running count is low and there's no queue, the engine is idle; the bottleneck is upstream (client concurrency, gateway limits, connection pool). Run a saturation sweep (e.g. vllm bench serve, guidellm, or your own harness) before blaming the engine.
Parallelism layout — for throughput, prefer the smallest TP that fits + more replicas. Wide TP spends bandwidth on all-reduce; MoE models often want expert/data parallel layouts instead. If you doubled GPUs with TP and got <1.5x, this is why.
Tokens/s can look great while users time out (long queues, aborted streams). Track request completion rate within SLO ("goodput") alongside raw throughput — scheduler changes that raise one often lower the other.
OOM & KV pressure#
First, know your memory map. Total VRAM = weights + KV cache pool + activations/workspace + CUDA graphs & framework overhead. KV per token = 2 × layers × kv_heads × head_dim × bytes (GQA shrinks it; MLA compresses it further).
OOM at startup
- vLLM pre-allocates the KV pool at startup to
--gpu-memory-utilization(default 0.9) of the GPU. If something else lives on that GPU (another process, a sidecar), startup OOMs. Fix: evict the tenant (checknvidia-smiprocesses) or lower the fraction. - Model simply too big →
--tensor-parallel-sizeup, or quantize (AWQ/GPTQ/FP8), or a smaller--max-model-len(shrinks activation workspace + profiling allocation). - CUDA graph capture takes memory at startup — as a diagnostic only,
--enforce-eagerconfirms whether graphs are what pushes you over.
--mem-fraction-static(default ~0.9) sizes weights + KV pool. OOM at startup → lower it; OOM during serving → usually the dynamic side (activations, CUDA graphs) → also lower it, or reduce--cuda-graph-max-bs.- Multimodal models spike activation memory on image encoding — leave more headroom than a text-only deployment.
- Too big →
--tp-sizeup or quantize.
- Engine build parameters are hard commitments:
max_batch_size×max_num_tokenssize the activation workspace at build time. OOM loading an engine that "should fit" → rebuild with smaller maxima. kv_cache_free_gpu_mem_fractioncontrols the KV pool grab after weights load — lower it if load succeeds but serving OOMs immediately.- Weights too big → build with more TP ranks or a quantized checkpoint (FP8/INT4 via Model Optimizer).
OOM after hours of serving
The dangerous kind — the pool math was fine until a rare shape showed up:
- A rare long request: one max-context request needs KV the profiler never planned for concurrently with peak batch. Bound it: cap context and
max_tokensat the gateway; don't serve 128k context "just in case." - Multimodal spikes: a batch of large images through the vision encoder is an activation spike far above text steady-state.
- Memory creep outside the pool: long-lived Python processes accumulate fragmentation; NCCL/communication buffers grow with topology changes. If a replica OOMs weekly, schedule rolling restarts and file the bug with a captured
torch.cuda.memory_summary(). - Host OOM (not CUDA): the container gets OOMKilled by cgroups — weights streaming through page cache, tokenizer memory, big Python heaps. Check
dmesgforoom_kill; raise container memory, not GPU anything.
nvidia-smi --query-gpu=memory.used,memory.total --format=csv per GPU, then subtract known weights size — the remainder is pool + workspace. Any process you don't recognize in nvidia-smi owns your missing gigabytes.
Errors, timeouts & crashes#
Classify by where the error is generated
- Client/gateway timeouts (504, stream cut mid-generation): compare gateway timeout vs engine E2E p99 — long generations legitimately run minutes. Fix the timeout and use streaming keep-alives; don't "fix" the engine.
- Engine rejects (4xx): usually context-length overflow (prompt +
max_tokens> model max) or malformed requests. These should be caught at the gateway with a clear message, not passed through as mystery 400s. - Engine 5xx / crash: read the last 100 lines before the restart — engines almost always log the real cause (CUDA OOM → §05; illegal memory access / NCCL / assertion → below).
Hangs (no response, no crash)
- Multi-GPU deadlock: one TP rank died or desynced and the rest wait on a collective forever. Check every rank's log, not just rank 0. NCCL watchdog timeouts in the log confirm it. Restart the whole group; investigate the rank that stopped logging first.
NCCL_DEBUG=INFO(orWARN) for the next occurrence. - CUDA-graph edge cases: a hang on a specific request shape that reproduces → capture it, then retry with
--enforce-eager:retry with CUDA graphs disabled (--disable-cuda-graph):try a debug build / different kernel selection: if eager mode fixes it, you've isolated a graph bug — pin the version and file upstream with the repro. - Hardware:
dmesg | grep -i xid. Xid 79 ("fell off the bus"), uncorrectable ECC (Xid 48), or row-remap events mean the GPU is the patient — drain the node (§09).
Crash loops on restart
- Same OOM every time → the config never fit; see §05 startup section.
- Crashes under traffic only → replay recent traffic shapes against a canary; usually one request shape (huge prompt, pathological sampling params, enormous
n) — add gateway validation. - After a driver/image change → CUDA/driver mismatch; engines fail at first kernel launch. Roll back first, debug second.
Gateways that retry failed/slow requests amplify overload exactly when the engine is sick. During an incident, check the retry rate before concluding "traffic spiked."
Bad output quality#
Quality regressions are config bugs far more often than model bugs. Work down this list — it's ordered by hit rate in real incidents:
Chat template. The #1 cause of "suddenly dumb" or leaking <|im_end|>-style tokens. The template must match the model exactly.
vLLM uses the template from the model repo; override with --chat-template. Diff what's actually applied: request with "echo": true is gone, so log the rendered prompt at the gateway, or call /tokenize with the chat messages and detokenize to inspect.
SGLang --chat-template flag; SGLang also has model-specific conversation templates — verify the right one is auto-detected for fine-tunes with renamed configs.
TensorRT-LLM templating typically happens in the client/frontend layer (or trtllm-serve) — the engine sees token IDs. Verify the frontend applies the model's template, not a default.
Sampling parameters. Compare what the client sends vs the model's recommended defaults (generation_config.json). Classic failures: temperature 0 exposing repetition, missing stop tokens (generation runs to max_tokens then truncates mid-sentence), a client library silently injecting top_k or penalties.
Tokenizer mismatch. Fine-tunes with added special tokens need the matching tokenizer. Symptom: rare glitch tokens, wrong stop behavior. Verify tokenizer files came from the same revision as the weights.
Quantization regression. If quality dropped when you shipped FP8/INT4: A/B the same prompts against the unquantized model. KV-cache INT4 in particular needs evaluation; FP8 KV is usually safe. MoE models are more quantization-sensitive in their routers/experts than dense models.
Long-context degradation. Ignoring instructions only on long prompts → check RoPE scaling config matches the model card (a wrong rope_scaling silently serves garbage past the native window), and that you're not truncating the middle of the prompt at the gateway.
Speculative decoding + structured output. If gibberish correlates with spec-decode being enabled, or JSON-mode/grammar constraints misbehave, disable that feature for the affected route and file the repro — these paths have had correctness bugs in every engine.
Reasoning/tool-call parsers. "Empty content" with reasoning models is often the parser eating the answer: the model emitted <think>-style blocks and the parser mis-split content vs reasoning. Check the raw completion before the parser (--reasoning-parser choice in vLLM) (--reasoning-parser / tool-call parser choice in SGLang) (frontend parser configuration).
Keep 10-20 fixed prompts with known-good outputs per model. Run them on every deploy (new engine version, new quantization, new flags). Five minutes of CI catches 90% of "the model got dumber" incidents before users do.
Won't start / no model served#
Is it still loading? Big models take minutes: download + load + warmup. Watch the log; don't restart a pod that's 80% through loading weights (crash-loop-by-impatience is real — set startup probes generous enough for model size).
Can it get the weights? Gated/private HF repos need HF_TOKEN; air-gapped clusters need a local model path or mirror. The error is usually an unambiguous 401/timeout in the first 30 log lines.
Does the server bind but serve nothing?
vLLM curl :8000/v1/models — empty/refused while the log shows a traceback → the API server survived an engine death. The root cause is above in the log (usually OOM or a bad flag combination).
SGLang curl :30000/get_model_info and /health_generate (does a real generation step — catches "HTTP up, engine wedged").
TensorRT-LLM curl :8000/v2/health/ready (Triton). Model shows unavailable → the backend failed to load the engine; Triton's log names which model instance and why.
Engine/config mismatch.
TensorRT-LLM the classic: engine built for a different GPU arch (SM90 engine on an A100), different TP world size, or older TRT-LLM version than the runtime. Engines are not portable across arch/TP/version — rebuild for the exact target.
vLLM flag combinations that don't support the model architecture fail at startup with a traceback — read it; the incompatible feature is named (e.g. a quantization method unsupported for that model).
SGLang same pattern — attention backend / quantization / model-arch combos are validated at startup; the log names the offender.
Behind a router/gateway? The engine can be perfectly healthy while the router never registers it. Check the router's worker list, and what the router's health probe expects — a gateway probing /health + /server_info + model-ID metadata won't register a backend that answers none of them (the exact failure mode documented in the SMG Local Lab: Router ready |workers: []).
Environment: nvidia-smi works inside the container? (device plugin / runtime class). CUDA driver ≥ what the image needs? Ports actually exposed? Thirty seconds of basics before deep debugging.
GPU-level checks#
The 2-minute health pass
# 1. who's on the GPU, how much memory, what utilization
nvidia-smi
# 2. live per-second view: power, util, mem, clocks, temp
nvidia-smi dmon -s pucvmet
# 3. is it throttling? anything other than "Not Active" is your answer
nvidia-smi -q -d PERFORMANCE | grep -A8 "Clocks Event Reasons\|Clocks Throttle Reasons"
# 4. hardware errors
dmesg -T | grep -i xid | tail
nvidia-smi -q -d ECC | grep -A3 "Aggregate"
# 5. interconnect (TP deployments)
nvidia-smi nvlink --status
nvidia-smi topo -m
Reading it right
utilization.gpulies to you. It means "at least one kernel was resident" — a memory-bandwidth-bound decode can show 90%+ "utilization" while SMs mostly wait. For truth use DCGM profiling metrics:DCGM_FI_PROF_SM_ACTIVE,PIPE_TENSOR_ACTIVE(are tensor cores busy? → prefill/compute),DRAM_ACTIVE(is bandwidth busy? → decode). Decode-heavy serving should show high DRAM_ACTIVE and moderate SM_ACTIVE — if both are low under load, the bottleneck is outside the GPU (scheduler, CPU, network).- Throttling silently halves performance.
SW Power Cap→ power limit set low (checknvidia-smi -q -d POWER);HW Thermal Slowdown→ cooling problem, clocks drop 30-50%. A single thermally-throttled GPU in a TP group drags every rank to its speed — and looks exactly like "the model got slower." - Xid codes worth memorizing: 31 (GPU memory page fault — usually a software bug, grab the stack), 48 (double-bit ECC — drain the node), 63/64 (row remapping — plan replacement), 79 (GPU fell off the bus — hardware/power, drain now). Any Xid correlating with your incident timeline → the GPU is the root cause, stop tuning software.
- Topology:
nvidia-smi topo -mbetween TP peers should show NV# (NVLink). PHB/PIX (PCIe) between TP ranks = your all-reduce is 10x slower than you think (see §02).
The host side (routinely forgotten)
- CPU throttling: tokenization, detokenization, HTTP, and Python scheduling are CPU work. In containers check
/sys/fs/cgroup/cpu.stat→nr_throttledclimbing = your CPU limit is strangling the frontend; symptoms mimic engine slowness at high QPS. - Host RAM / OOMKilled:
dmesg | grep -i oom_kill— the container died, not the GPU. - Disk for model loads: slow node-local disk or a saturated network filesystem turns 2-minute loads into 20 — matters for autoscaling and incident recovery time.
Cross-engine cheat sheet#
Metrics
| Concept | vLLM | SGLang | TensorRT-LLM (Triton) |
|---|---|---|---|
| Requests waiting (queue) | vllm:num_requests_waiting | sglang:num_queue_reqs | nv_trt_llm_request_metrics{request_type="waiting"} |
| Requests running | vllm:num_requests_running | sglang:num_running_reqs | ...{request_type="active"} |
| KV cache usage | vllm:gpu_cache_usage_perc | sglang:token_usage | nv_trt_llm_kv_cache_block_metrics{...="used"/"free"/"max"} |
| TTFT histogram | vllm:time_to_first_token_seconds | sglang:time_to_first_token_seconds | measure at frontend / Triton request duration metrics |
| ITL histogram | vllm:time_per_output_token_seconds | sglang:inter_token_latency_seconds | derive from per-request timing at frontend |
| E2E latency | vllm:e2e_request_latency_seconds | sglang:e2e_request_latency_seconds | nv_inference_request_duration_us |
| Preemption / eviction | vllm:num_preemptions_total | retract events (logs/metrics) | implied by policy max_utilization; watch KV free blocks |
| Prefix cache effectiveness | vllm:gpu_prefix_cache_hit_rate (or hits/queries) | sglang:cache_hit_rate | KV reuse stats (iteration stats / logs) |
Config levers
| Lever | vLLM | SGLang | TensorRT-LLM |
|---|---|---|---|
| Max concurrent sequences | --max-num-seqs | --max-running-requests | max_batch_size (build + runtime) |
| Token budget per step | --max-num-batched-tokens | --chunked-prefill-size | max_num_tokens (build + runtime) |
| GPU memory for weights+KV | --gpu-memory-utilization | --mem-fraction-static | kv_cache_free_gpu_mem_fraction |
| Chunked prefill | default on (V1) / --enable-chunked-prefill | --chunked-prefill-size | enable_chunked_context |
| Prefix caching | default on (V1) / --enable-prefix-caching | RadixAttention (default); --disable-radix-cache to turn off | enable_kv_cache_reuse |
| KV cache dtype | --kv-cache-dtype fp8 | --kv-cache-dtype fp8_e4m3 | FP8 KV via quantized checkpoint / config |
| Context length cap | --max-model-len | --context-length | max_seq_len (engine build) |
| Tensor parallel | --tensor-parallel-size | --tp-size | TP baked at engine build (--tp_size) |
| Scheduling behavior | (V1 scheduler; priority via --scheduling-policy) | --schedule-policy lpm|fcfs, --schedule-conservativeness | batch_scheduler_policy: guaranteed_no_evict | max_utilization |
| Disable CUDA graphs (debug) | --enforce-eager | --disable-cuda-graph | n/a (AOT-compiled engine) |
| Health / model check | GET /health, /v1/models | GET /health_generate, /get_model_info | GET /v2/health/ready (Triton) |
All three engines rename flags and metrics between releases (vLLM V0→V1 especially). Treat names here as the map, and --help / /metrics on your actual build as the territory. Corrections welcome — this guide gets updated as engines evolve.
Companion resources: SMG Mastery (what the gateway/router layer does above these engines) and SMG Local Lab (reproduce routing failures locally).