ULTRON-Core-512M-Instruct (Generation 3 Aligned)

ULTRON-Core-512M-Instruct is an aligned conversational and reasoning foundation model trained from scratch with custom Multi-Head Latent Attention (MLA) and Chain-of-Thought (<think>) supervision.

Architecture & Alignment

  • Base Parameters: 562,382,848 (~512.05M Non-Embedding)
  • Attention: Multi-Head Latent Attention (MLA: $d_c=512, d_R=64$) + QK-LayerNorm
  • Vocabulary: 24,576 (Pure-Python Byte-Level BPE, NFC Standard)
  • Alignment: Supervised Fine-Tuning (SFT) with dynamic prompt masking and multi-pattern <think> reasoning traces.
Downloads last month
11
Safetensors
Model size
0.6B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support