Comparison

RTX 4090 vs RTX 5090 vs Mac for local LLMs

Two numbers decide this and the rest is noise. Memory size decides which models you can load at all. Memory bandwidth decides how fast they generate, because producing one token means reading the active weights out of memory once. Everything below comes from the same data as our hardware calculator.

By Evdokiia Petrovskaia, Director, Dubir GroupPrices and specs checked

Short version: buy the 5090 if you want the fastest answers on models up to about 35B. Buy a 128 GB Mac if you want to run the 120B class at all. Buy a used 3090 if you want most of the 4090 for a third of the money. The 4090 in 2026 is the awkward middle: less memory and bandwidth than a 5090, at more than twice the price of a 3090 that nearly matches it.

The specifications that matter

 RTX 4090RTX 5090MacBook Pro M5 MaxRTX 3090 (used)
Memory24 GB32 GB128 GB unified24 GB
Usable by a model21.6 GB28.8 GB96 GB21.6 GB
Bandwidth1008 GB/s1792 GB/s614 GB/s936 GB/s
Board power450 W575 W90 W350 W
Street price~$2,900~$5,000~$6,700~$1,350
Models that fit (of 22)17171917

Prices are September 2026 street estimates, in a year when a GDDR7 shortage has pushed the 5090 a third above list and left a used 3090 costing more than it did new. All of it drifts. Check before you buy.

Speed, on real models

Estimated single-stream generation at Q4_K_M with 8k of context. These are estimates derived from memory bandwidth at 60% efficiency, not benchmarks, so read the ratios rather than the absolute numbers. The ratios hold.

ModelActive params40905090M5 Max
Gemma 4 12B12B83 t/s148 t/s51 t/s
Qwen3.6 27B (dense)27B37 t/s66 t/s23 t/s
gpt-oss 20B (MoE)3.6B277 t/s493 t/s169 t/s
Qwen3.6 35B-A3B (MoE)3B333 t/s591 t/s203 t/s
gpt-oss 120B (MoE)5.1Bdoes not fitdoes not fit119 t/s
Laguna S 118B-A8B (MoE)8.5Bdoes not fitdoes not fit71 t/s

Two things fall out of that table. The 5090 is a consistent 1.78x on anything both cards run, the exact ratio of their bandwidth. The second is stranger. A dense 27B model runs nine times slower than a 35B mixture-of-experts model on the same card, because the MoE reads 3B of weights per token instead of 27B. That is what people mean when they say a bigger model runs faster, and it is why the 2026 releases are nearly all MoE.

What the Mac buys and what it costs

A 128 GB Mac is the cheapest way to run a 120B model at home. Nothing in the consumer Nvidia line comes close. To match the capacity you are looking at an RTX PRO 6000, around $16,000.

The price is bandwidth. 614 GB/s against 1792 means the M5 Max generates at roughly a third of a 5090's rate on any model both can run. It also draws 90 W against 575 W. On a machine that is on all day, in a laptop, that is not a small difference.

Apple replaced the Mac Studio in September 2026, and the M5 Max version at 128 GB is the same chip and the same 614 GB/s as the laptop for $5,399, which makes it the value pick if you want capacity and do not need it to be portable. The M5 Ultra at 256 GB is the first box on our list that loads Qwen3.8 Flash-Next and DeepSeek V4 Flash at 4-bit, 21 of the 22 models, at 1.2 TB/s, for $9,499.

Recommendations

  • The 5090 answers fastest on anything up to 35B. It costs about 70% more than a 4090 and returns 1.78x the speed, and the extra 8 GB buys real context headroom.
  • A used 3090 is still the value pick: within 7% of the 4090 on generation speed, identical memory, less than half the price.
  • Only the 128 GB Mac loads the 120B class, and at 71-119 tokens a second it is usable rather than a demo.
  • For several people at once, none of them on their own. That is a batching question, and the runtime you pick matters more than the card. See Ollama vs LM Studio vs vLLM.

To check a specific machine against a specific model, use the hardware calculator. To decide whether to buy anything at all, the self-hosting breakeven guide is the more important read.

Questions, answered

Is the RTX 5090 worth it over a 4090 for local LLMs?

For speed, yes: 1792 GB/s of memory bandwidth against 1008 means roughly 1.8x the tokens per second on the same model, and the extra 8 GB gives real headroom for longer context. For which models you can run at all, barely: both fit 17 of the 22 models in our catalogue at Q4_K_M, and neither holds a 120B model. You are paying about $2,100 more for speed, not for capability.

Can a Mac run larger models than an RTX 5090?

Yes, and this is the whole argument for Apple hardware. A 128 GB Mac gives a model about 96 GB of usable memory against 28.8 GB on a 32 GB 5090. It runs gpt-oss 120B and Laguna S 118B, which no consumer Nvidia card can load. It runs them at roughly 71-119 tokens a second rather than the several hundred a 5090 manages on smaller models.

How many tokens per second does an RTX 4090 do?

It depends entirely on how many parameters are active per token. A dense 27B model at Q4_K_M runs about 37 tokens a second on a 4090. A 35B mixture-of-experts model with 3B active runs about 333, because only the active experts are read from memory for each token. The architecture matters more than the parameter count.

Is a used RTX 3090 still a good buy?

It is the best value on the list. 936 GB/s against the 4090's 1008 is a 7% difference in generation speed, the same 24 GB of memory, and roughly half to a third of the price. Its weakness is prompt processing and the age of the silicon, not token generation.

Do two GPUs double the speed?

No. Two 24 GB cards give you 48 GB of memory, which lets you run bigger models, but for single-stream generation the layers run in sequence and you get roughly the bandwidth of one card. Two cards buy capacity, not latency.

Private AI, from the spec to the running box

Dubir Group deploys open models on client hardware from Paphos, Cyprus, for the cases where the data is not allowed to reach a third-party API. We size the machine, install it and keep it answering.