GGUF Quants
#1
by Darkknight535 - opened
🥲
I'm working on releasing an even better REAP in FP8, and from that I can make more downstream quants like Q4_K_M and dynamic I-quants. There's potential to support different weight formats and reduce bpw to balance decode speed while retaining model performance.
🥲
I'm working on releasing an even better REAP in FP8, and from that I can make more downstream quants like Q4_K_M and dynamic I-quants. There's potential to support different weight formats and reduce bpw to balance decode speed while retaining model performance.
Thanks 🔥 🫡