Why Raising Model Token Efficiency Allows Price Increases

GPT5.5 is priced double compared to GPT5.4, and GPT5.4 is about 40% more expensive than GPT5.3codex. Both price increases share a fundamental logic: improved token efficiency. GPT5.2 is a painfully slow model, but the GPT5 series has high issue localization accuracy and few bugs, significantly reducing review…

GPT5.5 is priced double compared to GPT5.4, and GPT5.4 is about 40% more expensive than GPT5.3codex. Both price increases share a fundamental logic: improved token efficiency. GPT5.2 is a painfully slow model, but the GPT5 series has high issue localization accuracy and few bugs, significantly reducing the mental burden of review, so I have been reluctant to give it up. Starting from 5.3codex, each generation of GPT models has noticeably improved speed. According to official statements, the models enhance performance while reducing the token cost of thinking. That is, with unchanged hardware and improved performance, the model is more “to the point”—I interpret this as the model being more intelligent, with clearer logic, able to quickly reason out the correct conclusion rather than reflecting repeatedly. Those who have seen the reasoning chains of some open-source models will notice that most of their thinking process may be wrong, but by repeated reflection and repeatedly introducing detail attention, they may still arrive at the correct answer, only with an excessively long and messy reasoning chain. GPT5.5’s thinking efficiency is now arguably unique in the industry. Opus’s reasoning used to be virtually useless; since 4.7 it has had “Do you need…?” Based on my experience, the recommendation accuracy of GPT5.4 and later is already very high. I look forward to our models distilling GPT5 soon, and stop distilling opus.

Why Raising Model Token Efficiency Allows Price Increases Illustration 1

Why Raising Model Token Efficiency Allows Price Increases Illustration 2