FreeToken Lets Frontier Mixture-of-Experts Models Run on Consumer GPUs With Dynamic CPU-GPU Co-Execution
UC Berkeley and MIT researchers, with Databricks co-founders Matei Zaharia and Ion Stoica, released an open-source engine that serves up to 753-billion-parameter MoE models on a single workstation GPU.
Overview
Researchers from UC Berkeley and MIT have released FreeToken, an open-source inference engine built to run frontier-scale Mixture-of-Experts (MoE) language models on consumer hardware instead of datacenter GPUs, according to InfoQ. The project is co-authored by Databricks co-founders Matei Zaharia and Ion Stoica alongside Song Han, Kurt Keutzer and others, as InfoQ reports, and is released under the Apache License 2.0, according to the project’s GitHub repository.
In the paper describing the system, the authors write: “Frontier open-weight models are increasingly available, but serving them still largely assumes datacenter infrastructure.” FreeToken, they write, instead “treats a personal machine not as a small GPU, but as a unified, elastic inference platform.”
What We Know
- MoE architectures activate only a fraction of their total parameters per token, but decoding still requires routing across hundreds of billions of inactive weights that must be held somewhere accessible, according to InfoQ.
- On datacenter hardware, high-bandwidth interconnects like NVLink hide the cost of moving those inactive weights around. On consumer machines, PCIe throughput of roughly 16–64 GB/s and host RAM latency create what InfoQ describes as “severe decode bottlenecks,” and existing edge runtimes typically stream inactive weights from system RAM to the GPU synchronously, “completely stalling execution on cache misses.”
- FreeToken addresses this with a dynamic co-scheduling approach the authors call the q* policy: rather than pausing the GPU during a cache miss, the engine splits token computation between CPU cores and GPU tensor cores based on real-time interconnect throughput, per InfoQ.
- The system also uses a custom fast weight format (FTW) with full-layer double buffering, so that streaming weights over PCIe overlaps with active computation rather than blocking it, and an elastic memory manager that reallocates VRAM between KV cache entries and resident experts at runtime without reloading the model, InfoQ reports.
- FreeToken adds a feature aimed specifically at coding assistants and other agentic workloads, which InfoQ notes “introduce unique execution patterns: frequent prompt modifications, tool-call responses, and thinking blocks constantly alter the context window.” Standard inference engines discard their KV cache whenever a prompt prefix changes, forcing a full recomputation; FreeToken instead uses what InfoQ calls “semantic anchor checkpointing,” caching intermediate attention states at logical task boundaries so edited tool arguments or injected execution output don’t invalidate the whole cache.
- On Ollama and llama.cpp, InfoQ writes: “Optimised for GGUF quantisation and layer-wise offloading, but lack dynamic load splitting for sparse experts across host and device. FreeToken achieves 3–4x faster decode and 6–30x faster prefill on equivalent MoE models,” according to InfoQ.
- On the hardware side, InfoQ reports that FreeToken ran the Qwen3.6-35B model at roughly 39 tokens per second on an 8GB RTX 4060 laptop GPU, served the 284-billion-parameter DeepSeek-V4-Flash on an RTX 5090 desktop, and processed the 753-billion-parameter GLM-5.2 on a single workstation GPU.
- The paper independently corroborates the scale of that claim, stating that FreeToken “changes what these machines can practically serve, from a 35B model on a laptop to a 284B model on a gaming desktop and the 753B GLM-5.2 on a single workstation GPU,” and that the system “supports more than 20 MoE models and real coding and tool-using agents across hardware ranging from an 8GB laptop GPU to a single workstation GPU,” per the arXiv paper.
- The project’s GitHub repository lists native support for NVIDIA RTX 30, 40, and 50 series GPUs and describes the engine’s purpose as: “Run massive models locally, fast and efficiently,” per the FreeToken repository.
What We Don’t Know
- Neither InfoQ nor the paper’s abstract specifies exact memory or CPU requirements beyond the GPU models tested, so it’s unclear how the system performs on hardware configurations outside the RTX 30/40/50 series range described.
- The paper and InfoQ’s coverage do not detail how FreeToken’s benchmarked speedups were measured against Ollama and llama.cpp — for instance, whether comparisons used identically tuned configurations on both sides.
- AMD or Apple Silicon support is not mentioned in any of the cited sources, and the GitHub repository’s supported-hardware list is limited to NVIDIA RTX GPUs on Linux and Windows.
Analysis
The project lands in a period when open-weight MoE models with hundreds of billions of parameters — DeepSeek-V4-Flash and GLM-5.2 among them — are shipping faster than most developers’ access to datacenter-grade inference infrastructure. By reframing a gaming desktop or workstation as, in the paper’s words, “a unified, elastic inference platform” rather than an undersized datacenter node, FreeToken targets a specific pain point for developers building coding agents and other tool-using systems: running frontier-scale models locally without paying recurring cloud API costs or waiting on datacenter GPU allocations. The involvement of Databricks co-founders Matei Zaharia and Ion Stoica, both long associated with distributed-systems research, situates the project within a broader push to make large model inference tractable on hardware individual developers already own.