Briefing

1080 Ti still viable for small LLMs but struggles with larger models

ai-dev
by /u/srodland01 · Llama

Keep using 1080 Ti for small models but plan an upgrade if context exceeds 4k tokens.

What to do now

Benchmark current VRAM usage and schedule an upgrade when context >4k tokens.

Summary

A 1080 Ti with 11 GB VRAM is still used daily, but it struggles with larger models. Qwen 2.5 7B and Llama 3.2 8B run at 8–9 tokens per second, while Mistral 7B can run fully on the card at Q5_K_S if the context window is kept short. Anything 13 B or larger requires heavy off‑loading, and the speed drops sharply. Context size is the real killer; past 4 k tokens the system crawls due to memory pressure. The card also lacks tensor cores, so it cannot use newer optimisations. For small workloads it remains fine, but for multi‑task or larger prompts it feels dated.

Key facts: 8–9 tps for Qwen 2.5 7B and Llama 3.2 8B, Mistral 7B runs fully on‑card at Q5_K_S, 11 GB VRAM limit, 13 B+ models need off‑loading, >4 k tokens slows performance, no tensor cores.

The post suggests keeping the 1080 Ti for small models but planning an upgrade when context exceeds 4 k tokens.

Key changes

  • Qwen 2.5 7B and Llama 3.2 8B run at 8–9 tps
  • Mistral 7B runs fully on‑card at Q5_K_S with short context
  • 11 GB VRAM limits larger models
  • 13 B+ models require heavy off‑loading
  • Context >4 k tokens causes system crawl
  • No tensor cores on 1080 Ti

Affects

internal

Customer impact

Analyzing matches…

Ask about this story

Impact on an agency? Which customers? Compare historically Risks of waiting