Around 10 tokens/sec reads comfortably; under 5 feels slow. How we estimate.
Other codellama sizes
Embed this table
Free for blogs, READMEs and model cards. Updates automatically.
<iframe src="https://vsmacs.com/embed/codellama-70b" width="100%" height="420" style="border:0" loading="lazy" title="Code Llama 70B speed by machine"></iframe>