grimulkan/LimaRP-augmented
Viewer • Updated • 804 • 116 • 33
How to use mpasila/Mistral-LiPPA-12B with Transformers:
# Use a pipeline as a high-level helper
from transformers import pipeline
pipe = pipeline("text-generation", model="mpasila/Mistral-LiPPA-12B")
messages = [
{"role": "user", "content": "Who are you?"},
]
pipe(messages) # Load model directly
from transformers import AutoTokenizer, AutoModelForCausalLM
tokenizer = AutoTokenizer.from_pretrained("mpasila/Mistral-LiPPA-12B")
model = AutoModelForCausalLM.from_pretrained("mpasila/Mistral-LiPPA-12B", device_map="auto")
messages = [
{"role": "user", "content": "Who are you?"},
]
inputs = tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
tokenize=True,
return_dict=True,
return_tensors="pt",
).to(model.device)
outputs = model.generate(**inputs, max_new_tokens=40)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:]))How to use mpasila/Mistral-LiPPA-12B with vLLM:
# Install vLLM from pip:
pip install vllm
# Start the vLLM server:
vllm serve "mpasila/Mistral-LiPPA-12B"
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:8000/v1/chat/completions" \
-H "Content-Type: application/json" \
--data '{
"model": "mpasila/Mistral-LiPPA-12B",
"messages": [
{
"role": "user",
"content": "What is the capital of France?"
}
]
}'docker model run hf.co/mpasila/Mistral-LiPPA-12B
How to use mpasila/Mistral-LiPPA-12B with SGLang:
# Install SGLang from pip:
pip install sglang
# Start the SGLang server:
python3 -m sglang.launch_server \
--model-path "mpasila/Mistral-LiPPA-12B" \
--host 0.0.0.0 \
--port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
-H "Content-Type: application/json" \
--data '{
"model": "mpasila/Mistral-LiPPA-12B",
"messages": [
{
"role": "user",
"content": "What is the capital of France?"
}
]
}'docker run --gpus all \
--shm-size 32g \
-p 30000:30000 \
-v ~/.cache/huggingface:/root/.cache/huggingface \
--env "HF_TOKEN=<secret>" \
--ipc=host \
lmsysorg/sglang:latest \
python3 -m sglang.launch_server \
--model-path "mpasila/Mistral-LiPPA-12B" \
--host 0.0.0.0 \
--port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
-H "Content-Type: application/json" \
--data '{
"model": "mpasila/Mistral-LiPPA-12B",
"messages": [
{
"role": "user",
"content": "What is the capital of France?"
}
]
}'How to use mpasila/Mistral-LiPPA-12B with Docker Model Runner:
docker model run hf.co/mpasila/Mistral-LiPPA-12B
LoRA trained in 4-bit with 8k context using mistralai/Mistral-Nemo-Base-2407 as the base model for 1 epoch.
Dataset used is mpasila/LimaRP-PIPPA-Mix-8K-Context which was made using grimulkan/LimaRP-augmented and KaraKaraWitch/PIPPA-ShareGPT-formatted.
Merged from this LoRA: mpasila/Mistral-LiPPA-LoRA-12B
So uhh it does kinda work, maybe not the best datasets but uhh it's something.
Unsloth changed assistant to gpt and user to human.
This mistral model was trained 2x faster with Unsloth and Huggingface's TRL library.
Base model
mistralai/Mistral-Nemo-Base-2407