Tutorials Self-Host Kimi K3 Open Weights With vLLM
Run Moonshot's 2.8T open-weight model on your own GPUs with vLLM and MXFP4.
How-to content for builders, indie hackers, and AI engineers. Less theory, more shipped code.
Tutorials Run Moonshot's 2.8T open-weight model on your own GPUs with vLLM and MXFP4.
Tutorials Point the OpenAI SDK at Meta's agent model, add tools, let it self-manage a 1M-token context.
Machine Learning Run the first 27B-class model on a phone: MLX, llama.cpp, tool calls, and the memory math.
Tutorials Meituan's 1.6T open MoE topped OpenRouter as 'Owl Alpha.' Call it in Python with the OpenAI SDK.
Tutorials Turn a still image into a 720p video with native audio using xAI's Grok Imagine 1.5 in Python.
Tutorials Control thinking_level, media_resolution and thought signatures in the Gemini 3.1 Pro API.
Tutorials Compile llama.cpp with Vulkan in Termux and run a quantized LLM on your Android GPU, no root.
Tutorials Deploy Microsoft's new reasoning model and build a tool-calling triage agent.
Tutorials Call Microsoft's June 2 coding model via OpenRouter for cheap, fast refactors.
Tutorials Install grok-build-0.1, run plan mode, stream JSON in CI, and call the API from Python.
Tutorials Build a research-write-review multi-agent pipeline using Microsoft Agent Framework 1.0 in Python.
Tutorials Generate AI video in Python with Veo 3.1 — the model powering Google's Omni Flash launch.