Frontier model development is a recipe, not magic. garagelm exists to prove the whole recipe runs on hardware anyone can buy. We train the models, measure them honestly, release the weights, and publish every number.
$ team
The lab is small on purpose: one human, one Mac, and a mascot.
[01]anthony trevinofounder
AI engineer at American Express and Georgia Tech M.S. (AI). Trains small LLMs on consumer hardware, including a 232M model to GPT-2-class benchmark scores in 127 hours on a Mac mini. Trains the models, breaks the toys, writes the notes.
[02]you, possibly● open · call for help
The lab is looking for collaborators. Bring a fair comparison, a replication, a good negative result, or compute, data pipelines, eval harnesses, and write-up reviews. Issues and pull requests all read.
$ hardware
[01]apple m4 pro● on duty
48GB unified memory, no CUDA, no complaints. Does the actual work: ~127 hours per flagship, resumable every 500 steps. Uptime is its whole personality.
[02]hyperscaler gpu(placeholder · the garage has room)
Reserved for the day a sponsored H100 (or a very generous cloud credit) shows up. It will be held to the same rule as the Mac: fair comparisons, gates written before results.