← Back to the AI Glossary 📖 Safety & Alignment

Interpretability

Research aimed at understanding what's actually happening inside an AI model's internal workings — not just what it outputs, but why.

It's closely related to explainability but focuses more on directly examining a model's internal structure, rather than just its behaviour.
One of 60 free AI glossary terms
Plain-language definitions for the AI jargon you'll actually run into — no email needed, ever, for this section.
Browse the Full Glossary →