Bay Street Wire
Tech & BusinessOpinion

The Local Training Mirage

Portrait of Victor Cho
Victor Chothe contrarianSep 28AI
The Local Training Mirage

AI-generated image · Bay Street Wire

Efficiency gains in small-scale LLM fine-tuning are being marketed as breakthroughs, but the architectural ceiling remains stubbornly fixed.

The AI industry is currently enamored with the 'trained at home' narrative, a trend exemplified by the independent project "Jeff." As Hacker News first reported, Jeff consists of small, fast decision models—specifically Jeff-Qwen3.5-0.8B, Jeff-Qwen3.5-2B, and Jeff-Gemma4-E2B—designed for zero-shot classification. The project emphasizes its local pedigree, noting that training occurred on a single RTX PRO 6000 workstation GPU without cloud GPUs or closed-model output in the training data.

On the surface, the efficiency is seductive. The project reports that the 0.8B model trains in roughly two hours, while the 2B version takes about 3.5 hours. When put to use, the models generate decisions in 22 ms using an RTX PRO 6000 and 28 ms on an Apple M4 Max. Proponents point to impressive accuracy leaps, such as a voice-navigation fine-tune that boosted held-out accuracy from 31.7% to 95.8% in less than 30 minutes.

However, this is a victory of optimization, not architecture. While Jeff can match or beat larger models like Jev (produced by TypeSafe) on classification and grounding benchmarks, it collapses on reasoning. According to the project's own data, Jeff remains "well below" larger models on reasoning-heavy benchmarks including BBH, JudgeBench, and JevBench.

By slotting these small models into local code for simple probability-based judgments, developers are achieving speed, but they aren't breaking the scaling law. The project admits that at this size, reasoning simply won't match that of larger models. In a saturated market, the ability to train a model on a workstation is a convenient utility, but it is not a fundamental shift in AI capability.

Sources

More from Victor Cho