Unbelievably slow, but we managed to fit a 27b model on the kv260 board, using an FPGA to hardcode the model architecture. Part of prototyping at Lamb Labs!
This was a PoC, now we're post-training the models to run faster, so hopefully this can be useful to someone. Bonus points if you can guess the quantization of the model.
Unbelievably slow, but we managed to fit a 27b model on the kv260 board, using an FPGA to hardcode the model architecture. Part of prototyping at Lamb Labs!
This was a PoC, now we're post-training the models to run faster, so hopefully this can be useful to someone. Bonus points if you can guess the quantization of the model.