Gemma 4 QAT models: Optimizing model compression for mobile and laptop efficiency We’re releasing Gemma 4 quantization-aware training checkpoints, reducing memory requirements and improving on-device performance. Read More Feed 6 Jun 2026
Introducing Gemma 4 12B: a unified, encoder-free multimodal model An overview of Gemma 4 12B, a model designed to bring high-performance multimodal intelligence directly to your laptop. Read More Feed 4 Jun 2026
Accelerating Gemma 4: faster inference with multi-token prediction drafters An overview of how Multi-Token Prediction (MTP) drafters are making Gemma 4 models up to 3x faster at inference. Read More Feed 6 May 2026