This is a mixture-of-experts model: it only reads part of itself for each token, so it runs much faster than its size suggests, but the whole file still has to fit in memory. Around 10 tokens/sec reads comfortably; under 5 feels slow. How we estimate.
Other gpt-oss sizes
Embed this table
Free for blogs, READMEs and model cards. Updates automatically.
<iframe src="https://vsmacs.com/embed/gpt-oss-120b" width="100%" height="420" style="border:0" loading="lazy" title="gpt-oss 120B speed by machine"></iframe>