A model that generates output one piece at a time, using everything it has produced so far to predict what comes next.
This is how most LLMs write — token by token, each new token chosen based on the ones before it, not the whole answer planned out in advance.
One of 60 free AI glossary terms
Plain-language definitions for the AI jargon you'll actually run into — no email needed, ever, for this section.
Browse the Full Glossary →