Loading...
Work out how much VRAM a model's context window consumes and the maximum context length your GPU can handle for a giv...
Based on shared tags
Check whether your GPU can run a given local LLM before downloading gigabytes. Compares your available VRAM against the model's requirements at each quantization level.
Calculate exactly how much GPU VRAM a local LLM needs at any quantization and context length. Enter a model and see its memory footprint before you download — no guesswork.
Pick the right GGUF quantization for your hardware. Compare Q4, Q5, Q6, Q8 and more on quality vs. VRAM so you fit the largest model your GPU can hold.