GGUF Quants

#1
by Darkknight535 - opened

🥲

I'm working on releasing an even better REAP in FP8, and from that I can make more downstream quants like Q4_K_M and dynamic I-quants. There's potential to support different weight formats and reduce bpw to balance decode speed while retaining model performance.

🥲

I'm working on releasing an even better REAP in FP8, and from that I can make more downstream quants like Q4_K_M and dynamic I-quants. There's potential to support different weight formats and reduce bpw to balance decode speed while retaining model performance.

Thanks 🔥 🫡

Sign up or log in to comment