Skip to content
inference.academy

glossary/compute/tensor-cores

Tensor cores

The GPU units that do small dense matrix multiplies in one instruction, at far higher rate than the general-purpose cores. Nearly all of a transformer's arithmetic is matrix multiplication, so peak inference FLOPS means tensor core FLOPS. Each generation adds lower-precision formats with higher rates, which is where FP8 and FP4 speedups come from. A kernel that is not feeding tensor cores is leaving most of the chip idle.


312 versus 19.5 TFLOPS

The A100's dense BF16 rate on tensor cores against its FP32 rate on the general cores, a factor of 16.


See it happen


Related