BDH-150M-EN β€” Byte-Level BDH Scaling Study Checkpoint

A 150M-parameter BDH (Fast Weight Layers) model trained to completion (30k steps) on a 970M-token byte-level research mix (80% EN fineweb-edu / 20% PL speakleash). This is the largest rung of the SlayerLab BDH scaling ladder.

Training summary

Metric Value
Params ~151M
Steps 30,000 (completed)
Seq len 2048 (byte-level)
Data research-mix-v1, 970M tokens, byte-level (vocab 256, no tokenizer)
Optimizer AdamW lr=1e-3 β†’ 1e-4 cosine, wd 0.1, warmup 1000
Grad clip ZClip
Final val loss 1.099
Hardware RunPod A4000, single GPU

Learning curve (val loss per 500 steps, from step 15500)

step   val_loss
15500  1.233
16000  1.325
16500  1.115
17000  1.273
17500  1.460
18000  1.169
18500  1.250
19000  1.258
19500  1.270
20000  1.260
20500  1.203
21000  1.175
21500  1.152
22000  1.241
22500  1.402
23000  1.199
23500  1.125
24000  1.079
24500  1.114
25000  1.215
25500  1.197
26000  1.286
26500  1.178
27000  1.390
27500  1.216
28000  1.531
28500  1.090
29000  1.115
29500  1.186
30000  1.099

Full metrics (incl. steps 0–15000 from the pre-resume backup) in bdh_150M_en.json.

Architecture

  • BDH from pathwaycom/bdh β€” looped latent recurrence / fast weights
  • n_layer=8, n_embd=256, n_head=4, mlp_internal_dim_multiplier=768
  • Vocab: byte-level (vocab 256) β€” no tokenizer, works directly on UTF-8 bytes
  • ~151M parameters Β· seq length 2048

Files

  • model.safetensors β€” weights (state dict, model.* keys, compatible with pathwaycom/bdh)
  • config.json β€” architecture + training config
  • bdh_150M_en.json β€” full training metrics (loss, grad_norm per step)
  • bdh_150M_en.meta.json β€” run metadata

Load

import sys, torch
sys.path.insert(0, "path/to/pathwaycom/bdh")
import bdh as bdh_mod
from safetensors.torch import load_file

cfg = bdh_mod.BDHConfig(
    n_layer=8, n_embd=256, n_head=4,
    mlp_internal_dim_multiplier=768, dropout=0.1,
    vocab_size=256,
)
model = bdh_mod.BDH(cfg).bfloat16()
sd = load_file("model.safetensors")
model.load_state_dict({k.replace("model.", ""): v for k, v in sd.items()}, strict=False)

Context

Part of the BDH scaling study (25M β†’ 50M β†’ 150M) β€” results & baselines comparison in SlayerLab/bdh-scaling. Smaller rung: SlayerLab/bdh-25m-pl.

Downloads last month
97
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support