~ / local
Running models on your own hardware
Open weights are free. The machine to run them is not. This section answers the five questions that decide whether self-hosting is worth it for you, with the numbers in tables rather than spread across a hundred forum replies.
- What configuration actually holds the model, card by card
- How much memory each model needs at each quantisation
- How fast it runs, and why capacity and speed are different problems
- How good the result is, expressed as the closed model it matches
- What it costs today, with the date the price was checked
Coding
Which GPU runs a coding model, what the ceiling is against Claude and GPT, why open weights often still need a server, and the cheapest hardware that actually works.
liveImage
Local image generation: model sizes, VRAM per resolution, throughput, and what a workstation costs against a subscription.
plannedVideo
Open-weight video models such as MiniMax H3, which quantisation fits a consumer card, and how long a clip takes to render.
plannedMusic and speech
Local audio generation and text to speech: model choices, memory, and real-time factors.
plannedWhy this section exists. The people selling API access have no reason to explain what you could run yourself, and the people who do run models locally mostly write it up in scattered threads where the hardware, the quantisation and the quality claims never appear in the same place. Every page here keeps those together, states the price with the date it was checked, and says plainly when a number has never been measured rather than filling the gap with a guess.