RTX 4090 vs RTX 5090 vs Mac for local LLMs
Two numbers decide this and the rest is noise. Memory size decides which models you can load at all. Memory bandwidth decides how fast they generate, because producing one token means reading the active weights out of memory once. Everything below comes from the same data as our hardware calculator.
By Evdokiia Petrovskaia, Director, Dubir GroupPrices and specs checked
Short version: buy the 5090 if you want the fastest answers on models up to about 35B. Buy a 128 GB Mac if you want to run the 120B class at all. Buy a used 3090 if you want most of the 4090 for a third of the money. The 4090 in 2026 is the awkward middle: less memory and bandwidth than a 5090, at more than twice the price of a 3090 that nearly matches it.
The specifications that matter
| RTX 4090 | RTX 5090 | MacBook Pro M5 Max | RTX 3090 (used) | |
|---|---|---|---|---|
| Memory | 24 GB | 32 GB | 128 GB unified | 24 GB |
| Usable by a model | 21.6 GB | 28.8 GB | 96 GB | 21.6 GB |
| Bandwidth | 1008 GB/s | 1792 GB/s | 614 GB/s | 936 GB/s |
| Board power | 450 W | 575 W | 90 W | 350 W |
| Street price | ~$2,900 | ~$5,000 | ~$6,700 | ~$1,350 |
| Models that fit (of 22) | 17 | 17 | 19 | 17 |
Prices are September 2026 street estimates, in a year when a GDDR7 shortage has pushed the 5090 a third above list and left a used 3090 costing more than it did new. All of it drifts. Check before you buy.
Speed, on real models
Estimated single-stream generation at Q4_K_M with 8k of context. These are estimates derived from memory bandwidth at 60% efficiency, not benchmarks, so read the ratios rather than the absolute numbers. The ratios hold.
| Model | Active params | 4090 | 5090 | M5 Max |
|---|---|---|---|---|
| Gemma 4 12B | 12B | 83 t/s | 148 t/s | 51 t/s |
| Qwen3.6 27B (dense) | 27B | 37 t/s | 66 t/s | 23 t/s |
| gpt-oss 20B (MoE) | 3.6B | 277 t/s | 493 t/s | 169 t/s |
| Qwen3.6 35B-A3B (MoE) | 3B | 333 t/s | 591 t/s | 203 t/s |
| gpt-oss 120B (MoE) | 5.1B | does not fit | does not fit | 119 t/s |
| Laguna S 118B-A8B (MoE) | 8.5B | does not fit | does not fit | 71 t/s |
Two things fall out of that table. The 5090 is a consistent 1.78x on anything both cards run, the exact ratio of their bandwidth. The second is stranger. A dense 27B model runs nine times slower than a 35B mixture-of-experts model on the same card, because the MoE reads 3B of weights per token instead of 27B. That is what people mean when they say a bigger model runs faster, and it is why the 2026 releases are nearly all MoE.
What the Mac buys and what it costs
A 128 GB Mac is the cheapest way to run a 120B model at home. Nothing in the consumer Nvidia line comes close. To match the capacity you are looking at an RTX PRO 6000, around $16,000.
The price is bandwidth. 614 GB/s against 1792 means the M5 Max generates at roughly a third of a 5090's rate on any model both can run. It also draws 90 W against 575 W. On a machine that is on all day, in a laptop, that is not a small difference.
Apple replaced the Mac Studio in September 2026, and the M5 Max version at 128 GB is the same chip and the same 614 GB/s as the laptop for $5,399, which makes it the value pick if you want capacity and do not need it to be portable. The M5 Ultra at 256 GB is the first box on our list that loads Qwen3.8 Flash-Next and DeepSeek V4 Flash at 4-bit, 21 of the 22 models, at 1.2 TB/s, for $9,499.
Recommendations
- The 5090 answers fastest on anything up to 35B. It costs about 70% more than a 4090 and returns 1.78x the speed, and the extra 8 GB buys real context headroom.
- A used 3090 is still the value pick: within 7% of the 4090 on generation speed, identical memory, less than half the price.
- Only the 128 GB Mac loads the 120B class, and at 71-119 tokens a second it is usable rather than a demo.
- For several people at once, none of them on their own. That is a batching question, and the runtime you pick matters more than the card. See Ollama vs LM Studio vs vLLM.
To check a specific machine against a specific model, use the hardware calculator. To decide whether to buy anything at all, the self-hosting breakeven guide is the more important read.
Questions, answered
Is the RTX 5090 worth it over a 4090 for local LLMs?
For speed, yes: 1792 GB/s of memory bandwidth against 1008 means roughly 1.8x the tokens per second on the same model, and the extra 8 GB gives real headroom for longer context. For which models you can run at all, barely: both fit 17 of the 22 models in our catalogue at Q4_K_M, and neither holds a 120B model. You are paying about $2,100 more for speed, not for capability.
Can a Mac run larger models than an RTX 5090?
Yes, and this is the whole argument for Apple hardware. A 128 GB Mac gives a model about 96 GB of usable memory against 28.8 GB on a 32 GB 5090. It runs gpt-oss 120B and Laguna S 118B, which no consumer Nvidia card can load. It runs them at roughly 71-119 tokens a second rather than the several hundred a 5090 manages on smaller models.
How many tokens per second does an RTX 4090 do?
It depends entirely on how many parameters are active per token. A dense 27B model at Q4_K_M runs about 37 tokens a second on a 4090. A 35B mixture-of-experts model with 3B active runs about 333, because only the active experts are read from memory for each token. The architecture matters more than the parameter count.
Is a used RTX 3090 still a good buy?
It is the best value on the list. 936 GB/s against the 4090's 1008 is a 7% difference in generation speed, the same 24 GB of memory, and roughly half to a third of the price. Its weakness is prompt processing and the age of the silicon, not token generation.
Do two GPUs double the speed?
No. Two 24 GB cards give you 48 GB of memory, which lets you run bigger models, but for single-stream generation the layers run in sequence and you get roughly the bandwidth of one card. Two cards buy capacity, not latency.
Private AI, from the spec to the running box
Dubir Group deploys open models on client hardware from Paphos, Cyprus, for the cases where the data is not allowed to reach a third-party API. We size the machine, install it and keep it answering.