A technique that reduces the precision of a model's internal numbers to make it smaller and faster to run, with a small trade-off in accuracy.
Quantization is often what makes it possible to run a large open-weight model on a normal laptop instead of needing a data centre.
One of 60 free AI glossary terms
Plain-language definitions for the AI jargon you'll actually run into — no email needed, ever, for this section.
Browse the Full Glossary →