Local AI Milestone: Running 125B Models on Gaming Hardware

AI-generated image · Bay Street Wire
New open-source software Strata enables massive AI models to run on consumer GPUs, shifting the 'AI PC' conversation from assistants to raw local power.
While marketing for 'AI PCs' often focuses on pre-installed assistants, a more significant technical milestone is arriving via open-source software. As first reported on Hacker News, a project called Strata now allows users to run the Qwen3.8-Flash-Next—a 125-billion-parameter AI model that typically requires a server—directly on standard gaming PCs.
Strata is designed for Windows and Linux and supports both NVIDIA and AMD graphics cards with 12 GB of VRAM or more. Reporting from Hacker News indicates that the software can operate on a wide range of hardware, including NVIDIA GeForce RTX 20, 30, 40, and 50 series cards, as well as AMD Radeon RX 6800/6900, 7700 XT/7800 XT/7900 XT/XTX, 9060 XT, and 9070/9070 XT series GPUs.
Performance varies by hardware and model compression. For example, on an NVIDIA RTX 5070 with 12 GB of VRAM and 64 GB of RAM, the IQ3_S model size can write answers at 53 tokens per second. The project notes that cards with more VRAM increase speed; specifically, an RTX 3090 with 24 GB of VRAM is estimated to write between 100 and 140 tokens per second.
To support these large models, system requirements include at least 32 GB of RAM, though 64 GB is required to run every available model size. The software requires approximately 80 GB of free disk space for the model download. Because the processing happens entirely on the local machine, no data leaves the user's PC.

