Yahya/Blog
Apr 05, 2026By Yahya

Running Open-Weights LLMs on Budget GPUs: Benchmarking RX 6500 XT, GTX 1660, Arc A750, and RX 6600

A practical comparative benchmark of budget consumer GPUs (RX 6500 XT, GTX 1660 Super, Arc A750, RTX 3050, RX 6600) running quantised LLMs (Llama-3, Qwen-2.5).

You do not need a $2,000 NVIDIA RTX 4090 or enterprise H100 to experiment with local LLMs. Budget consumer GPUs (under $200) can run 1B to 8B parameter open-weights models effectively if you understand VRAM limits, PCIe bus widths, and quantization setups.

Here is a practical comparative benchmark testing a handful of popular budget GPUs—AMD RX 6500 XT, NVIDIA GTX 1660 Super, Intel Arc A750, AMD RX 6600, and NVIDIA RTX 3050—running Qwen-2.5-1.5B, Llama-3.2-3B, and Llama-3-8B.


The Handful of Budget GPUs Tested

GPU ModelVRAM CapacityMemory Bus / PCIeBackend Driverapprox. price
AMD Radeon RX 6500 XT4GB GDDR664-bit (PCIe 4.0 x4)ROCm / HIP~$110
NVIDIA GTX 1660 Super6GB GDDR6192-bit (PCIe 3.0 x16)CUDA (cuBLAS)~$130
Intel Arc A7508GB GDDR6256-bit (PCIe 4.0 x16)oneAPI / SYCL~$180
NVIDIA RTX 3050 (8GB)8GB GDDR6128-bit (PCIe 4.0 x8)CUDA / Tensor Cores~$170
AMD Radeon RX 66008GB GDDR6128-bit (PCIe 4.0 x8)ROCm / HIP~$190

Benchmark Results (Tokens / Second)

All tests were conducted using llama.cpp with maximum VRAM layer offloading, context window capped at 2,048 tokens, and GGUF quantization.

1. Qwen-2.5 1.5B (Q8_0 Quantization, ~1.8GB VRAM)

  • AMD RX 6500 XT: 62.4 t/s
  • NVIDIA GTX 1660 Super: 74.1 t/s
  • Intel Arc A750: 81.5 t/s
  • NVIDIA RTX 3050: 88.2 t/s
  • AMD RX 6600: 95.6 t/s

Verdict: At 1.5B parameters, all budget GPUs easily hold the entire model in VRAM, delivering instant, highly responsive chat responses.


2. Llama-3.2 3B (Q4_K_M Quantization, ~2.4GB VRAM)

  • AMD RX 6500 XT: 38.1 t/s
  • NVIDIA GTX 1660 Super: 51.8 t/s
  • Intel Arc A750: 58.3 t/s
  • NVIDIA RTX 3050: 61.0 t/s
  • AMD RX 6600: 67.4 t/s

Verdict: The 3B sweet spot. Fits in under 2.5GB VRAM across all tested GPUs while maintaining excellent reasoning for general Q&A and coding tasks.


3. Llama-3 8B (Q3_K_M / Q4_K_M Quantization, ~3.8GB - 5.2GB VRAM)

  • AMD RX 6500 XT (4GB): 14.2 t/s (Q3_K_M, fits barely in VRAM)
  • NVIDIA GTX 1660 Super (6GB): 24.5 t/s (Q4_K_M)
  • Intel Arc A750 (8GB): 32.8 t/s (Q4_K_M)
  • NVIDIA RTX 3050 (8GB): 34.1 t/s (Q4_K_M)
  • AMD RX 6600 (8GB): 39.0 t/s (Q4_K_M)

Verdict: On 4GB cards like the RX 6500 XT, 8B models push memory limits to the absolute brink. 8GB cards (RX 6600, RTX 3050, Arc A750) allow standard Q4_K_M precision without VRAM spillover.


Key Hardware Takeaways

  1. The PCIe x4 Bottleneck (RX 6500 XT): If a model spills even 300MB into system RAM over PCIe x4, token generation tanks from 35 t/s down to 3 t/s. Keep VRAM utilization under 90%.
  2. The 8GB VRAM Threshold: Purchasing an 8GB budget GPU (e.g. used RX 6600 or RTX 3050) unlocks 8B models at full Q4_K_M precision with comfortable KV cache headroom.
  3. Driver Support: NVIDIA CUDA works out of the box; AMD ROCm 6.2+ is now stable on Navi 23/24; Intel SYCL/oneAPI via llama.cpp has improved dramatically on Arc GPUs.