Choosing a GPU: L4, A100, or H100 for your workload
GPU selection is often driven by whichever card is available or most recently announced, rather than by the actual requirements of the workload. L4, A100, and H100 occupy distinct positions on the cost, memory, and throughput spectrum, and picking correctly avoids both overpaying and hitting a capacity wall later. This post covers how to match the workload to the card.
L4 for cost-efficient inference
The L4 is built around inference efficiency and cost per token rather than raw training throughput. For serving small to mid-sized models, especially with quantization applied, an L4 fleet can offer a lower total cost than a smaller number of higher-end cards for the same aggregate throughput.
The tradeoff is memory capacity. Larger models or those requiring long context windows at full precision may not fit comfortably on an L4, pushing you toward quantization or a larger card regardless of the cost advantage L4 offers per unit.
A100 as the general-purpose workhorse
The A100 remains a strong default for fine-tuning and mid-scale training workloads, offering a balance of memory capacity and compute throughput at a lower cost than H100. For many fine-tuning jobs on models up to tens of billions of parameters, an A100 or a small A100 cluster is sufficient without needing the latest generation hardware.
A100 availability is also generally better than H100, since demand for the newest hardware often outstrips supply. For workloads that are not throughput-constrained by the newest architecture's advantages, this availability difference can matter as much as the raw performance difference.
H100 for large-scale training and high-throughput inference
The H100 delivers substantially higher throughput for large model training and for inference workloads with demanding latency or concurrency requirements. It also supports newer numerical formats that improve training efficiency for models that can take advantage of them.
The cost premium over A100 is significant, so H100 is best justified when the workload is throughput-bound and the cost per unit of useful work actually improves, not simply when the workload could technically run on either card.
- Choose L4 for cost-sensitive inference of small to mid-sized, quantized models
- Choose A100 for general fine-tuning and moderate-scale training
- Choose H100 for large-scale training or inference workloads that are demonstrably throughput-bound
Memory bandwidth matters more than headline compute for many LLM workloads
LLM inference, particularly at low batch sizes, is frequently memory-bandwidth bound rather than compute bound, meaning the theoretical compute advantage of a newer card does not translate proportionally into faster inference. Benchmark on your actual model and batch size rather than relying on headline specification comparisons between cards.
This is one of the more common sources of disappointment when teams upgrade hardware expecting a proportional throughput increase and see a smaller gain than anticipated, because the bottleneck was never compute in the first place.
Right-sizing across the workload lifecycle
A workload's GPU requirements often change across its lifecycle: experimentation and fine-tuning may need an A100 or H100 briefly, while steady-state production inference of the resulting model may run comfortably and more cheaply on L4. Building infrastructure that supports moving a workload between card types as it matures avoids being locked into the hardware chosen for an earlier phase.
Revisit GPU allocation periodically rather than treating the initial choice as permanent, since model updates, quantization improvements, and changes in traffic volume can shift the right answer over time.
Want this applied to your data?
We scope private AI projects in one call and start with a pilot you can evaluate.