AI Industry
7 min read
How to Cut LLM Costs With Off-Peak API Pricing
DeepSeek's August 16 move to UTC-based peak/off-peak pricing for V4-Flash and V4-Pro pushes cache-hit costs up by more than 10x during busy hours, ending the era of near-free caching and forcing agent builders to actually schedule their workloads.