GeForce RTX 3090
The reference specs, practical tradeoffs, and local models worth trying. A starting guide—not a benchmark or a promise that every workload fits.
GPU specs reviewed 2026-10-05 · Model listings reviewed 2026-10-05 · No live price or stock feed
- VRAM
- 24 GBReference GPU [10]
- Memory
- GDDR6XReference GPU [10]
- Memory bus
- 384-bitReference GPU [10]
- Board power
- 350 WReference board, not PSU [10]
- Architecture
- AmpereReference GPU [10]
Gaming & local AI
Legacy high-end hardware; evaluate current game-specific results and inspect cooling and condition before considering a used board.
24 GB CUDA option to investigate on the legacy market. Inspect condition, thermals, warranty and seller carefully; no used-price or stock claims here. [1] [2]
Software first: Check your GPU architecture, NVIDIA driver and CUDA-capable runtime version. Blackwell cards need a runtime that supports them. [1] [2]
Legacy stock/used evaluation; this is not a claim of current availability. Check seller, warranty, condition, cooling, and exact board specifications. power_w is NVIDIA reference board power, 350 W (Founders Edition Graphics Card Power), not typical gaming consumption or a whole-system PSU recommendation; partner boards may differ.
Local models to try
These are capacity-screened candidates from our researched model library, not tested recommendations for this exact card. Published download size is not total VRAM required. Hardware/runtime support must be checked separately.
| Model / exact Ollama tag | Published download | Capacity screen | What to explore |
|---|---|---|---|
Qwen3 4Bqwen3:4b [27]4.02B parameters · Q4_K_M at review | 2.5 GBDownload, not VRAM required | First trialsCapacity screen only; not tested. | Editorial shortlist candidate for multilingual chat, writing and instruction-following experiments; evaluate outputs for your task.Text input and text output; the tag page does not advertise vision or audio. |
Gemma 3 4Bgemma3:4b [32]4.3B parameters · Q4_K_M at review | 3.3 GBDownload, not VRAM required | First trialsCapacity screen only; not tested. | Editorial shortlist candidate for text-and-image question answering and summarization experiments; verify visual interpretations.Text and image input, text output. Published 3.3 GB is the download listing, not a runtime VRAM requirement or usable-context guarantee. |
Qwen2.5 Coder 7Bqwen2.5-coder:7b [34]7.62B parameters · Q4_K_M at review | 4.7 GBDownload, not VRAM required | First trialsCapacity screen only; not tested. | Editorial shortlist candidate for code generation, explanation and repair experiments; review and test generated code.Text/code input and text/code output; the tag page does not advertise vision or audio. |
Qwen3 8Bqwen3:8b [28]8.19B parameters · Q4_K_M at review | 5.2 GBDownload, not VRAM required | First trialsCapacity screen only; not tested. | Editorial shortlist candidate for multilingual chat, writing and instruction-following experiments; evaluate outputs for your task.Text input and text output; the tag page does not advertise vision or audio. |
DeepSeek R1 0528 Qwen3 8Bdeepseek-r1:8b [33]8.19B parameters · Q4_K_M at review | 5.2 GBDownload, not VRAM required | First trialsCapacity screen only; not tested. | Editorial shortlist candidate for reasoning-oriented mathematics, programming and logic prompts; independently check conclusions.Text input and text output; the tag page does not advertise vision or audio. |
Qwen3 14Bqwen3:14b [29]14.8B parameters · Q4_K_M at review | 9.3 GBDownload, not VRAM required | First trialsCapacity screen only; not tested. | Editorial shortlist candidate for multilingual chat, writing and instruction-following experiments; evaluate outputs for your task.Text input and text output; the tag page does not advertise vision or audio. |
Qwen3 30Bqwen3:30b [30]30.5B parameters · Q4_K_M at review | 19 GBDownload, not VRAM required | Tighter memoryMore context/overhead risk; may offload. | Editorial shortlist candidate for multilingual chat, writing and instruction-following experiments; evaluate outputs for your task.Text input and text output; the tag page does not advertise vision or audio. |
Qwen3 32Bqwen3:32b [31]32.8B parameters · Q4_K_M at review | 20 GBDownload, not VRAM required | Tighter memoryMore context/overhead risk; may offload. | Editorial shortlist candidate for multilingual chat, writing and instruction-following experiments; evaluate outputs for your task.Text input and text output; the tag page does not advertise vision or audio. |
How these candidates are selected—and what we do not know
Editorial heuristic, not a memory calculator: downloads up to 65% of nominal VRAM are labeled “First trials”; above 65% and up to 85% are “Tighter memory.” Larger downloads are omitted from this single-GPU shortlist. These percentages are deliberately simple screening rules, not measured allocation budgets; the source sizes are rounded downloads and do not establish exact GB/GiB memory fit.
We have not calculated model-specific KV cache sizes, measured peak VRAM, tested this card/model/backend combination, or established a safe maximum context. Parameter count and weights quantization alone cannot answer those questions. Equal-VRAM GPUs receive the same capacity screen, not the same speed or software-compatibility claim.
A model's advertised maximum context is a capability limit, not a promise that this card can allocate it. A model can launch with CPU offloading and still be unsuitable for your desired speed. Inspect the actual allocation before deciding it “runs well.” [24] [25]
Try it, then check the allocation
After installing a supported Ollama runtime and GPU driver, use this first-trial command. This example is a cautious starting experiment—not an assertion that the configuration has been tested here.
ollama run qwen3:4bInside the interactive Ollama session, explicitly set the trial context before sending a prompt. Do not rely on a default context setting. [25]
/set parameter num_ctx 4096In another terminal, check the loaded model:
ollama ps- Inspect PROCESSOR and CONTEXT. CPU/GPU splitting means the trial is not fully GPU-resident. A “100% GPU” result is evidence for that running configuration, not for every future context or workload. [24]
- Test the workload you actually need. If allocation fails or unwanted offloading appears, try a smaller model or reduce context/concurrency. Avoid simultaneous games, other models, or GPU-heavy applications during the trial.
- Model weight quantization and KV-cache quantization are separate. Ollama documents
f16as its default cache type; cache quantization and supported Flash Attention can reduce memory usage, with model/backend support and potential quality tradeoffs. Do not assume they halve total model memory. [25] - Vision inputs can introduce additional allocations; a small text-only trial does not validate multimodal or long-context operation.
Sources & review notes
Manufacturer specs and exact model listing details are sourced below. Trial labels and workflow choices are GPUcrate editorial guidance. Model tags, downloads and runtime behavior can change; verify the listing and installed version before copying commands.