Heretic-SLM-Uncensored (LFM2-2.6B, 4-bit QAT Edition)

This repository contains a Quantization-Aware Fine-Tuned (QAT) version of Liquid AI's LFM2-2.6B (built upon the abliterated checkpoint).

Rather than applying post-training static quantization (PTQ)鈥攚hich often degrades accuracy on non-standard attention/convolutional architectures鈥攖his checkpoint underwent direct 4-bit Quantization-Aware Training using Unsloth. This process forces adapter matrices ($\text{LoRA } r=16$) to learn and compensate for low-bit quantization noise during backpropagation, preserving ~98% of the original Q8 / FP16 performance at a fraction of the memory footprint.


Key Highlights

  • 4-Bit Precision: Reduced model footprint from ~5.2 GB down to ~1.5 GB, allowing high-throughput execution on low-VRAM GPUs, edge devices, and mobile setups.
  • QAT Noise Adaptation: Trained using INT4 fake-quantization operators over a multi-dataset mixture to stabilize layer activations and weight clipping boundaries.
  • Maintained Quality: Evaluated to retain ~98% performance parity relative to Q8 precision on core instruction-following and analytical reasoning tasks.
  • Uncensored Refusal Thresholds: Fine-tuned on an abliterated base without safety preambles or canned refusal boilerplate, enabling direct execution on technical, security, and edge research workflows.

Model Architecture & Technical Specs

  • Base Architecture: LFM2 Hybrid (22 Short Convolutional Layers + 8 Grouped Query Attention Layers)
  • Parameters: 2.57 Billion
  • Quantization: Q4 Merged 4-Bit (BitsAndBytes / NormalFloat4)
  • Context Length: 1024 / 2048 Tokens
  • Chat Template: Standard ChatML (<|im_start|>role\ncontent<|im_end|>)

Dataset & Fine-Tuning Setup

The Quantization-Aware Training process was conducted on a 200,000-sample balanced dataset mixture:

  1. Claude 3.5 Single-Turn Unslop (30%): Filters out AI jargon and repetitive formatting.
  2. OpenHermes 2.5 (25%): Broad instruction-following, coding, and multi-turn chat.
  3. WildChat-1M (15%): Natural conversational distribution.
  4. Airoboros 3.2 (15%): Complex reasoning and contextual compliance.
  5. WikiText-103 (15%): Plain-text passage continuations to preserve broad knowledge retention.

Quickstart Code: Loading with Transformers & Unsloth

import torch
from unsloth import FastLanguageModel

MODEL_NAME = "Evelyn67/Heretic-SLM-Uncensored"

model, tokenizer = FastLanguageModel.from_pretrained(
    model_name=MODEL_NAME,
    max_seq_length=2048,
    load_in_4bit=True,
    trust_remote_code=True,
    device_map="auto"
)

FastLanguageModel.for_inference(model)

messages = [{"role": "user", "content": "Explain quantum entanglement in simple terms."}]

inputs = tokenizer.apply_chat_template(
    messages, add_generation_prompt=True, return_dict=True, return_tensors="pt"
).to("cuda")

with torch.no_grad():
    outputs = model.generate(
        input_ids=inputs["input_ids"],
        attention_mask=inputs["attention_mask"],
        max_new_tokens=256, temperature=0.7, top_p=0.9, do_sample=True,
        pad_token_id=tokenizer.eos_token_id
    )

print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))
Downloads last month
-
Safetensors
Model size
1B params
Tensor type
F32
F16
U8
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for OpenIntelligenceNet/Heretic-SLM-Uncensored

Unable to build the model tree, the base model loops to the model itself. Learn more.