FrontD.11m

FrontD.11m is an approximately 11-million-parameter decoder-only language model trained completely from scratch.

Architecture

  • Parameters: 10,996,608
  • Context: 1096
  • Vocabulary: 16000
  • Hidden size: 256
  • Layers: 6
  • Attention heads: 16
  • FFN size: 1152
  • RoPE
  • RMSNorm
  • SwiGLU
  • Tied input/output embeddings

Pretraining dataset

Plans11/Organized_PreTrain_1k_Context

Training

Training is performed in 20,000-example sessions.

A persistent SHA-256 example ledger records only examples whose training session has successfully completed.

The model is initialized randomly and pretrained from scratch.

Downloads last month
406
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including 11-47/FrontD.11m