Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Q8 or Q6_UD with no KV cache quantization. I swear it matters even more with small activated parameters MOE model despite the minimal KL divergence drop


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: