YOUR MEMORY BUDGET

Local AI models for 16GB VRAM.

Explore models by theoretical weight storage. This shortlist covers available downloadable language models with a known or name-derived parameter count.

14 GBleft for weights at 4-bit
16 GB budget − 2 GB reserve

The reserve is your assumption, not a measured overhead allowance. The default 2 GB may be too little for your runtime or context length. Results are not confirmed fits and have no verified speed rating. Weight estimates use decimal GB and do not include quantization overhead.

352 models within the weight budget

Largest parameter count first · not a quality ranking

Qwen · Alibaba · Downloadable

WebWorld-14B

Text generation

Generates text for conversations and writing.

Source listed
Feb 13, 2026
Parameters
14B
4-bit weights only
≈ 7 GB
License
apache-2.0
Compare this model →

Meta · Downloadable

CodeLlama-13b-hf

Text generationCoding

Designed for coding and software tasks. Generates text for conversations and writing.

Source listed
Mar 13, 2024
Parameters
13B
4-bit weights only
≈ 6.5 GB
License
llama2
Compare this model →

Meta · Downloadable

CodeLlama-13b-Instruct-hf

Text generationCoding

Designed for coding and software tasks. Generates text for conversations and writing.

Source listed
Mar 13, 2024
Parameters
13B
4-bit weights only
≈ 6.5 GB
License
llama2
Compare this model →

Meta · Downloadable

CodeLlama-13b-Python-hf

Text generationCoding

Designed for coding and software tasks. Generates text for conversations and writing.

Source listed
Mar 13, 2024
Parameters
13B
4-bit weights only
≈ 6.5 GB
License
llama2
Compare this model →

Meta · Downloadable

Llama-2-13b

Text generation

Generates text for conversations and writing.

Source listed
Jul 9, 2023
Parameters
13B
4-bit weights only
≈ 6.5 GB
License
llama2
Compare this model →

Stability AI · Downloadable

StableBeluga-13B

Text generation

Generates text for conversations and writing.

Source listed
Jul 27, 2023
Parameters
13B
4-bit weights only
≈ 6.5 GB
License
Not disclosed
Compare this model →

Z.ai · Downloadable

agentlm-13b

Text generation

Generates text for conversations and writing.

Source listed
Oct 8, 2023
Parameters
13B
4-bit weights only
≈ 6.5 GB
License
Not disclosed
Compare this model →

Z.ai · Downloadable

apar-13b

Text generation

Generates text for conversations and writing.

Source listed
Jul 22, 2024
Parameters
13B
4-bit weights only
≈ 6.5 GB
License
Not disclosed
Compare this model →

Stability AI · Downloadable

stablelm-2-12b

Text generation

Generates text for conversations and writing.

Source listed
Mar 21, 2024
Parameters
12B
4-bit weights only
≈ 6 GB
License
other
Compare this model →

What changes with a 16GB budget?

With 2 GB set aside, 14 GB remains for weights. At 4-bit that allows up to 28 billion parameters in the theoretical calculation. An 8-billion-parameter model alone needs approximately 4 GB at this precision, before its conversation cache and runtime.

Why might a listed model still fail to load?

Its actual weight format may need more space, the runtime may not support its architecture, or the conversation cache may exceed your reserve. Vision models can also need space for image processing. Check the original model card and measure your exact setup before relying on it.

Read the memory guide and technical sources · Compare two models