the garage intelligence lab · real models, cheap hardware, honest numbers
$ models
Weights that came out of the research program, plus the toy that powers Learn. Click a name for the full write-up.
{{ modelCount }} models · 2 on hugging face, 1 in your browser
{{ m.excerpt }}
{{ m.specLine }}
$ benchmarks
0-shot accuracy, n=300/task, every model re-run through the identical local lm-eval harness. Tokens seen in parentheses.
{{ benchChart }}
source: milestone 09 eval table · smollm2 (2T tokens) leads everywhere, as pre-registered · full table in the 600× note
milestone 11 tested four training-efficiency levers against this recipe; all four failed their pre-registered gates; the recipe stands. post-training has since run (milestone 12). milestone 13 then tested capacity at a fixed memory budget: 2.5× the parameters bought validation loss but no measurable downstream gain, and slower decode. distillation is the remaining priced lever.
real models · cheap hardware · honest numbers
© 2026 garagelm.org