OLMo-7B Full Fine-Tune — Chemistry SMILES CPT

Model Description

This model is a full-parameter fine-tuned version of HuggingFaceTB/SmolLM-135M trained on chemistry SMILES strings from the Codemaster67/Causal_lm_chemistry_1M_rows dataset.

The base model's tokenizer was pre-extended with ~300 SPE (SMILES Pair Encoding) chemistry tokens plus <|start_of_smiles|> / <|end_of_smiles|> special tokens, and its embedding & LM-head layers were resized with mean-initialised vectors for the new tokens.

Training Details

Parameter Value
Method Full Fine-Tune (all weights updated)
Parallelism FSDP (Fully Sharded Data Parallel)
Epochs 1
Learning Rate 5e-06
Batch Size (per device) 32
Gradient Accumulation 1
Max Sequence Length 128
Warmup Ratio 0.1
Weight Decay 0.01
Scheduler Cosine
Precision bf16
Augmentation OFF
Training Samples Full dataset
Eval Samples Full dataset (10%)

Evaluation Results

Metric Value
Final Eval Loss 3.204535961151123
Final Eval Perplexity 24.644061560759287
Training Loss 3.1543

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("Codemaster67/Test_run", trust_remote_code=True)
tokenizer = AutoTokenizer.from_pretrained("Codemaster67/Test_run", trust_remote_code=True)

smiles_input = "<|start_of_smiles|>CC(=O)Oc1ccccc1C(=O)O<|end_of_smiles|>"
inputs = tokenizer(smiles_input, return_tensors="pt")
outputs = model.generate(**inputs, max_new_tokens=128)
print(tokenizer.decode(outputs[0], skip_special_tokens=False))

Intended Use

Chemistry-domain language modelling, SMILES generation and completion, and downstream molecular property prediction via fine-tuning.

Limitations

  • Trained primarily on SMILES strings; natural-language instruction-following ability may degrade compared to the base OLMo checkpoint.
  • Augmentation was disabled for this run.
Downloads last month
42
Safetensors
Model size
0.2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Codemaster67/Test_run

Finetuned
(127)
this model

Dataset used to train Codemaster67/Test_run