SmoLLM — 109M Parameter Language Model from Scratch
try it — live inference
109M params · CPUAsk the instruct checkpoint something. It's a 109M-param model running on a free-tier CPU box, so answers can be fluent-sounding but wildly wrong, and generation is slow — even more so on the first request if the box was asleep. Each prompt is answered on its own — there's no conversation history, so it won't remember what you asked before.
# try:
# try:
# try:
Implemented a Llama-style decoder-only transformer from scratch: 109.5M parameters, 12 layers, 768 hidden dimensions, 12 attention heads, 512 token context, a custom BPE tokenizer (~32k vocab), RoPE positional encoding, RMSNorm, SwiGLU, and multi-head causal attention.
Pretrained on FineWeb-Edu (sample-10BT) at a Chinchilla-optimal token budget of roughly 2.2B tokens (~20 tokens per parameter). Diagnosed and fixed a training bug where EOS tokens were never appended at document boundaries, then ran a continued-pretraining phase to recover sequence-termination behavior.
Instruction-tuned on databricks-dolly-15k and released both base and instruct checkpoints publicly. The base model reliably completes high-frequency memorized sequences and produces fluent text, but shows limited in-context pattern learning and no factual reliability — it memorizes patterns rather than facts, which is expected at this scale.
Measured WikiText-2 perplexity of 74.57 on the base checkpoint (out-of-distribution evaluation, since training used FineWeb-Edu rather than WikiText).