AI token prices hit new record lows as inference costs plunge 43% in ten weeks

1 hour ago 21

The price of processing AI tokens, the basic unit of work for large language models, has fallen to roughly $1.16 to $1.18 per million tokens in early August 2026. That’s the lowest average inference cost recorded this year, and it represents a 43% decline from $2.04 at the end of May.

To put that trajectory in perspective: frontier AI intelligence is now priced at approximately 12% of its March 2023 levels, according to BenchLM indices. Some benchmarks show comparable AI capabilities costing over 280 times less than they did in early 2023.

What’s driving the freefall

Two forces are crushing inference prices simultaneously, and neither shows signs of slowing down.

The first is OpenAI’s decision to slash pricing on its GPT-5.6 Luna models by 80% in late July 2026. After the cuts, input tokens cost $0.20 per million and output tokens run $1.20 per million.

The second force is the wave of cost-effective open-weight models emerging from China. Providers like DeepSeek have achieved price reductions of up to 99% earlier in 2026, making the American price war look almost restrained by comparison.

The result is a pricing environment where the average cost per million tokens dropped from $2.04 to $1.45 in late July, then kept falling to the current $1.16 to $1.18 range. ARK Invest tracked an even sharper decline in its own measurements, recording a move from $2.07 to $1.02 per million tokens in recent weeks.

The virtuous cycle that keeps feeding itself

ARK Invest describes the current dynamic as a “virtuous cycle” where lower prices drive broader adoption, which drives higher volumes, which justifies further infrastructure investment, which enables even lower prices.

Token processing volumes have reportedly surged up to tenfold in certain agent-driven applications. AI agents, which autonomously execute multi-step tasks and can chain together dozens or hundreds of model calls per interaction, are particularly sensitive to inference pricing. When the cost per call drops by half, developers don’t just save money. They build agents that make twice as many calls, tackling problems that were previously too expensive to automate.

The demand elasticity appears strongest in enterprise automation, autonomous systems, and applications where AI agents need to process large context windows repeatedly. These use cases were economically marginal at $2 per million tokens. At $1.16, they start looking like obvious investments.

What this means for the AI industry

OpenAI’s willingness to cut GPT-5.6 pricing by 80% signals a strategic bet that market share matters more than near-term margins. When the market leader prices at $0.20 per million input tokens, smaller competitors either match or differentiate.

The Chinese open-weight model ecosystem adds another layer of competitive intensity. DeepSeek and similar providers aren’t just competing on price. They’re releasing models with weights available for anyone to download and run, which puts a ceiling on how much any proprietary provider can charge.

Hardware improvements are compounding the software-side price war. Continual advancements in chip efficiency and processing architecture mean that even without competitive price cuts, the cost of running inference would be declining. The combination of better hardware and aggressive commercial pricing is what produces the 280x reduction in comparable intelligence costs since 2023.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.

Read Entire Article