← Все репозитории

Новый проект

carloslfu/slotstream

Run Qwen3.8-Flash-Next (125B MoE, 104 GB at 4-bit) on Macs with a fraction of that RAM by streaming experts from SSD. MLX + Swift, Ollama-compatible API.

Открыть GitHub
Stars
89
Forks
4
Язык
Swift
Статус
В канале 1 янв. 1 г.

AI-разбор

Run Qwen3.8-Flash-Next (125B MoE, 104 GB at 4-bit) on Macs with a fraction of that RAM by streaming experts from SSD. MLX + Swift, Ollama-compatible API.