garagelm.org
hf
the garage intelligence lab · real models, cheap hardware, honest numbers
home  ·  models  ·  notes  ·  competitions  ·  learn  ·  team

$ models

Weights that came out of the research program, plus the toy that powers Learn. Click a name for the full write-up.

{{ modelCount }} models · 2 on hugging face, 1 in your browser
{{ m.name }} {{ m.tag }} {{ m.meta }}
{{ m.excerpt }}
{{ m.specLine }}
{{ s.text }}
{{ pp.text }}
{{ lk.label }}{{ lk.sep }}

$ benchmarks

0-shot accuracy, n=300/task, every model re-run through the identical local lm-eval harness. Tokens seen in parentheses.

{{ benchChart }}

source: milestone 09 eval table · smollm2 (2T tokens) leads everywhere, as pre-registered · full table in the 600× note

milestone 11 tested four training-efficiency levers against this recipe; all four failed their pre-registered gates; the recipe stands. post-training has since run (milestone 12). milestone 13 then tested capacity at a fixed memory budget: 2.5× the parameters bought validation loss but no measurable downstream gain, and slower decode. distillation is the remaining priced lever.
home · models · notes · competitions · learn · team
real models · cheap hardware · honest numbers
hf  
© 2026 garagelm.org