DeepSeek’s roughly 98% cache-hit discount, against an industry norm nearer to 90%, is the mechanism that has kept its measured cost per task at about 60% below Luna, even after Luna’s cost cut, he said.
“The schedule re-prices exactly that mechanism,” Gogia said. Flash’s edge over Luna decreases from roughly sevenfold to threefold off-peak, and 1.4 times at peak. “The cache is where the advantage genuinely erodes.”
Encouraging users to rethink their schedules
DeepSeek’s V4-Pro is now generally available, and V4-Flash is in beta. Both models have new flexible reasoning capabilities (low, high, max) and ‘thinking modes’ that use chain-of-thought (CoT) reasoning to improve answer accuracy. V4 Pro is now available on app, web, and via API, and users can try it using “Expert Mode.” V4 Flash is now in beta.