Gemma 4 on a Tesla T4: QAT Weights Decode 1.79x Faster Than bf16 - DEV Community
A step by step deployment of Gemma 4 E2B with vLLM on a single Tesla T4 attached to a Compute Engine VM, and a measured comparison of the QAT w4a16 checkpoint against the bf16 reference on the same card.