Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

If you want to do a 1-bit model you have to QAT at pre-training with way more data than chinchilla to compensate for the cliffs (like 50x). Quantization on an existing pre-trained model will almost always collapse at 1-bit
 help



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: