Thanks for the model, how to serve?

#1
by nimishchaudhari - opened

Hello

Thanks a lot for the model, always keen on testing <2b models for cpu inference. Can you recommend a serving engine for this? For testing I suppose we need to use hf transformers library?

Thanks for yet another European entry to the AI land.

Danish Foundation Models org

HF transformers or llama cpp with patch (see: https://huggingface.co/sinimiini/HRM-Text-1B-GGUF)

Consider oMLX downloaders as first class users, #imjustsaying

Sign up or log in to comment