File size: 3,537 Bytes
1e205df
66a1574
 
 
 
ac0dcf1
 
66a1574
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
ac0dcf1
 
66a1574
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
ac0dcf1
66a1574
69f8645
66a1574
69f8645
66a1574
69f8645
66a1574
 
69f8645
66a1574
69f8645
66a1574
 
 
 
 
 
69f8645
66a1574
69f8645
66a1574
69f8645
66a1574
69f8645
66a1574
69f8645
66a1574
 
 
 
69f8645
66a1574
 
 
 
 
 
69f8645
66a1574
 
69f8645
66a1574
 
 
 
ac0dcf1
66a1574
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
---
license: apache-2.0
language:
  - en
library_name: transformers
pipeline_tag: text-generation
tags:
  - text-generation
  - causal-lm
  - conversational
  - qwen2
  - qwen2.5
  - transformers
  - safetensors
  - gguf
  - unsloth
  - llama.cpp
  - vllm
  - coding
  - mathematics
  - reasoning
  - lora
  - peft
base_model:
  - Qwen/Qwen2.5-7B-Instruct
  - Qwen/Qwen2.5-Coder-7B-Instruct
  - Qwen/Qwen2.5-Math-7B-Instruct
datasets:
  - HuggingFaceH4/ultrachat_200k
  - openai/gsm8k
  - tatsu-lab/alpaca
  - m-a-p/CodeFeedback-Filtered-Instruction
---

# Keefe-Discere

<p align="center">
  <strong>An 8B-class instruction-following language model focused on reasoning, coding, mathematics, and agentic tool use.</strong>
</p>

<p align="center">
  <a href="https://huggingface.co/KeefeBuild/Keefe-Discere">
    <img src="https://img.shields.io/badge/Hugging%20Face-Keefe--Discere-orange" alt="Hugging Face">
  </a>
  <img src="https://img.shields.io/badge/Parameters-~8B-blue" alt="Parameters">
  <img src="https://img.shields.io/badge/Precision-BF16-blue" alt="Precision">
  <img src="https://img.shields.io/badge/Context-32K-purple" alt="Context">
  <img src="https://img.shields.io/badge/License-Apache--2.0-green" alt="License">
</p>

<p align="center">
  <a href="#quick-start">Quick Start</a><a href="#model-information">Model Info</a><a href="#training-data">Training Data</a><a href="#roadmap">Roadmap</a><a href="#citation">Citation</a>
</p>

---

## Overview

**Keefe-Discere** is an independently developed language-model project by **KeefeBuild**, built on the Qwen2.5-7B instruction-tuned architecture and enhanced through a two-stage process:

1. **Model Merging (DARE-TIES):** Combining general, coding, and mathematics specialists into a single balanced 15GB checkpoint.
2. **Targeted Post-Training (QLoRA v1.1):** Fine-tuning a LoRA adapter on curated instruction, reasoning, and code-execution data to improve agentic tool use and mathematical reliability.

The project is designed as a general-purpose, locally-deployable language model with an emphasis on:

- 🧠 Reasoning and structured problem solving
- 🔢 Mathematics and quantitative tasks
- 💻 Programming, debugging, and code execution
- 🛠️ Agentic tool use (Python execution, function calling)
- 💬 General instruction following
- 🏠 Private, self-hosted, and offline inference

> **Important:** Keefe-Discere is an independent model project and is **not** an official Qwen model.

---

## Quick Start

### 🤗 Transformers (Python)

```python
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
from peft import PeftModel

# 1. Load the merged base model
base_model_id = "KeefeBuild/Keefe-Discere"
tokenizer = AutoTokenizer.from_pretrained(base_model_id)
base_model = AutoModelForCausalLM.from_pretrained(
    base_model_id, torch_dtype=torch.bfloat16, device_map="auto"
)

# 2. Attach the v1.1 LoRA adapter for enhanced coding/tool use
model = PeftModel.from_pretrained(base_model, base_model_id)

messages = [
    {"role": "system", "content": "You are Keefe-Discere. Write clean Python code and use print() to output final answers."},
    {"role": "user", "content": "Calculate the sum of the first 15 prime numbers."}
]

inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
outputs = model.generate(inputs, max_new_tokens=512, temperature=0.1)
print(tokenizer.decode(outputs[0][inputs.shape[-1]:], skip_special_tokens=True))