YOUR MEMORY BUDGET

Local AI models for 24GB VRAM.

Explore models by theoretical weight storage. This shortlist covers available downloadable language models with a known or name-derived parameter count.

22 GBleft for weights at 4-bit
24 GB budget − 2 GB reserve

The reserve is your assumption, not a measured overhead allowance. The default 2 GB may be too little for your runtime or context length. Results are not confirmed fits and have no verified speed rating. Weight estimates use decimal GB and do not include quantization overhead.

405 models within the weight budget

Largest parameter count first · not a quality ranking

NVIDIA · Downloadable

Qwen3.6-35B-A3B

Text generation

Generates text for conversations and writing.

Source listed
May 27, 2026
Parameters
35B
4-bit weights only
≈ 17.5 GB
License
apache-2.0
Compare this model →

Ornith · Downloadable

Ornith-1.0-35B

VisionCoding

Designed for coding and software tasks. Understands text and images and responds in text.

Source listed
Jun 21, 2026
Parameters
35B
4-bit weights only
≈ 17.5 GB
License
mit
Compare this model →

Ornith · Downloadable

Ornith-1.5-35B-A3B

VisionCoding

Built for coding agents and multi-step tasks. An open-weight model from the Ornith team.

Source listed
Aug 18, 2026
Parameters
35B
4-bit weights only
≈ 17.5 GB
License
mit
Compare this model →

Qwen · Alibaba · Downloadable

Qwen3.5-35B-A3B

Vision

Understands text and images and responds in text.

Source listed
Feb 24, 2026
Parameters
35B
4-bit weights only
≈ 17.5 GB
License
apache-2.0
Compare this model →

Qwen · Alibaba · Downloadable

Qwen3.6-35B-A3B

Vision

Understands text and images and responds in text.

Source listed
Apr 15, 2026
Parameters
35B
4-bit weights only
≈ 17.5 GB
License
apache-2.0
Compare this model →

Meta · Downloadable

CodeLlama-34b-hf

Text generationCoding

Designed for coding and software tasks. Generates text for conversations and writing.

Source listed
Mar 14, 2024
Parameters
34B
4-bit weights only
≈ 17 GB
License
llama2
Compare this model →

Meta · Downloadable

CodeLlama-34b-Instruct-hf

Text generationCoding

Designed for coding and software tasks. Generates text for conversations and writing.

Source listed
Mar 14, 2024
Parameters
34B
4-bit weights only
≈ 17 GB
License
llama2
Compare this model →

Meta · Downloadable

CodeLlama-34b-Python-hf

Text generationCoding

Designed for coding and software tasks. Generates text for conversations and writing.

Source listed
Mar 14, 2024
Parameters
34B
4-bit weights only
≈ 17 GB
License
llama2
Compare this model →

DeepSeek · Downloadable

deepseek-coder-33b-base

Text generationCoding

Designed for coding and software tasks. Generates text for conversations and writing.

Source listed
Oct 28, 2023
Parameters
33B
4-bit weights only
≈ 16.5 GB
License
other
Compare this model →

Microsoft · Downloadable

NextCoder-32B

Text generationCoding

Designed for coding and software tasks. Generates text for conversations and writing.

Source listed
May 3, 2025
Parameters
32B
4-bit weights only
≈ 16 GB
License
mit
Compare this model →

MiniMax · Downloadable

SynLogic-32B

Text generation

Generates text for conversations and writing.

Source listed
May 30, 2025
Parameters
32B
4-bit weights only
≈ 16 GB
License
mit
Compare this model →

NVIDIA · Downloadable

Cosmos-Reason2-32B

VisionReasoning

Understands text and images and responds in text. Supports a reasoning mode.

Source listed
Apr 29, 2026
Parameters
32B
4-bit weights only
≈ 16 GB
License
other
Compare this model →

Qwen · Alibaba · Downloadable

Qwen3-32B

Text generation

Generates text for conversations and writing.

Source listed
Jun 11, 2025
Parameters
32B
4-bit weights only
≈ 16 GB
License
apache-2.0
Compare this model →

Qwen · Alibaba · Downloadable

Qwen3-32B-MLX

Text generation

Generates text for conversations and writing.

Source listed
Jun 11, 2025
Parameters
32B
4-bit weights only
≈ 16 GB
License
apache-2.0
Compare this model →

Qwen · Alibaba · Downloadable

Qwen3-VL-32B-Thinking

VisionReasoning

Understands text and images and responds in text. Supports a reasoning mode.

Source listed
Oct 19, 2025
Parameters
32B
4-bit weights only
≈ 16 GB
License
apache-2.0
Compare this model →

Qwen · Alibaba · Downloadable

WebWorld-32B

Text generation

Generates text for conversations and writing.

Source listed
Feb 13, 2026
Parameters
32B
4-bit weights only
≈ 16 GB
License
apache-2.0
Compare this model →

What changes with a 24GB budget?

With 2 GB set aside, 22 GB remains for weights. At 4-bit that allows up to 44 billion parameters in the theoretical calculation. An 8-billion-parameter model alone needs approximately 4 GB at this precision, before its conversation cache and runtime.

Why might a listed model still fail to load?

Its actual weight format may need more space, the runtime may not support its architecture, or the conversation cache may exceed your reserve. Vision models can also need space for image processing. Check the original model card and measure your exact setup before relying on it.

Read the memory guide and technical sources · Compare two models