Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

5090 is plenty for the Q4_K_M quantized version of 3.6 27B with reduced context size.

I run it on a 3090(24GB) and 64k context using GGUF format and llama-cpp. Double 3090 gives you 128k, quad 3090 gets you to full context - 256k.



Or you can run quantized context, there's some degradation but it fits in a lot less.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: