deepseek price increase
Price doubled during working hours
Price doubled during working hours
By default it uses the stepfun model to judge the personal data obtainable from the GitHub API, quite fun https://githubroast.dev/ #vibecodingfestival[topic]#
Zhipu’s model performance is decent, but its engineering capability is limited, and subscription users can feel the service quality. The advantage of glm5.2 is rigorous thinking, at the cost of low token efficiency, leading to a higher overall price. Assuming most users do appreciate well-performing models, but their feet vote for cost-effective models, then cost-effective vendors will win the users. Assuming the actual profit model of AI model vendors resembles that of factories, valuing scale effects and cost control, then cost-effective vendors are more likely to gain profits…
Create multiple workspaces, each workspace can subscribe to one go. Workspaces can be created freely. No need to create multiple accounts, unless to take advantage of the $5 deal.
AI is surging in; let’s prepare for the future
From February to July last year, copilot was freely usable in vscode insider. During this period, I witnessed the gradual decline of copilot’s agent capabilities. Hard to imagine, the first month preview of copilot turned out to be the best in effect, and then the agent capability gradually declined. What exactly changed? That is, copi…
Over the past year and a half, some companies overestimated the role of AI and began layoffs. As time went on, companies gradually realized that AI is not as smart or cheap as expected, and began to reduce AI token consumption and slow down layoffs. A very important reason is that AI is not cheap and cannot take the blame. In the past, some industries joked that AI cannot go to jail, but based on my actual work…
Previously posted criticism that mimo-2.5-pro’s performance couldn’t justify its pricing. Its performance really can’t be compared to gpt5.x and opus4.x. It’s not that I wanted to compare them, but Xiaomi positioned them as competitors, so I pointed out the obvious gap. But now Xiaomi has shifted to benchmarking against deepseek, and I think it might still put up a fight. Selling LLMs…
The GPT-5 series has a high recall rate and can achieve very rigorous logic, making it excellent for debugging. At the same time, it can promptly detect contradictions in the context, offering a unique experience among current models. But recently, in my personal reverse engineering project, I found that GPT-5 has limited reverse engineering capabilities, clearly inferior to the opus model. It often explicitly replies that something won’t work, while switching to opus can quickly achieve functional reverse engineering…
AI writes code so fast that bugs per line are surely lower, but lines of code explode, and requirements we used to skip get cranked out, so bugs keep growing. Also, AI is unbeatable at finding others’ bugs, surfacing hidden ones. Major vendors and popular libraries never blew up with critical bugs every few days like now. Reverse engineering and attacks get easier, and people will gradually realize AI does more than create—future offense and defense may get interesting…
I actually have some aversion to the company anthropic, but it is currently at the height of its popularity, so I have to hold my nose and use it. I subscribed to a claudecode pro plan, but the quota is very small and can only be used for planning. Domestic models perform slightly worse, but the quota is more, offering clear cost-performance advantages. Domestic model providers currently do not offer a response API, so…
Someone shared very long and fancy prompts claiming to improve GPT5’s performance, which should be viewed with caution. Unlike the sonnet or opus4.6 and earlier models in claudecode, GPT5 is more noticeably negatively affected by randomly written custom prompts. GPT5’s system prompts should be carefully considered, otherwise it is easy to degrade the model’s performance.
First, the old friend openrouter. After depositing $10, you get 1000 free calls per day. Most of the time there are a few decent models in the free tier. The official policy reserves the right to set an expiry on the deposited amount, there may be a one-year validity period, but my $10 from over a year ago is still valid. Recently tried on a limited-time basis inclusion…
I have two hundred free accounts and one plus account, connected to cliproxyapi, reverse-proxied for use in various CLIs. These accounts are enough for heavy personal use. If you enable 1M context, fast mode, use subagent, high-res image parsing, the tokens of these few hundred accounts will still be stretched thin, and you still have to plan usage. F…
GPT5.5 is priced double compared to GPT5.4, and GPT5.4 is about 40% more expensive than GPT5.3codex. Both price increases share a fundamental logic: improved token efficiency. GPT5.2 is a painfully slow model, but the GPT5 series has high issue localization accuracy and few bugs, significantly reducing review…
Some people share that Xiaomi’s model can perform long-horizon tasks, executing hundreds of tool calls without errors. I don’t know who initiated this testing direction, but it’s like saying a student can hold their bladder for 100 minutes in an exam room—meaningless. Normal people understand that for a standardized exam, the only thing that matters is the score. To be more lenient, the score accounts for over 95% of the evaluation system…
And obtained the time-to-first-token and token rate at 4 PM Beijing time for some vendors. Used it for a few months, welcome to visit and share your feedback! Project link: #https://github.com/jqknono/coding-plans-for-copilot
Got a Xiaomi Pro Plan through my small project, trying it out this month. The project is a VSCode plugin that connects any chat-completion/responses/anthropic-compatible API to GitHub Copilot. Project address: #https://github.com/…
It may be a bit late to talk about escaping from Cursor now. Actually, over the past six months I have gradually reduced my use of Cursor, because my coding habits have changed, and Cursor’s strengths are no longer the reason I choose it. The feature with which Cursor still firmly ranks first is its unique completion model, supporting cross-line and cross-file completion in all directions. Whether in response speed or acceptance rate, no rival can match it…
Anthropic’s well-known advantage is common sense. Some development work relies on common sense, some on proprietary logic. Opus’s reasoning style is to guess many possibilities and then quickly verify. GPT5 tends to reason based on code implementation, repeatedly sorting out the business. The style is similar to the difference between breadth-first and depth-first. As…
The cost narrative does not fit the large model industry; everyone is losing money, even the winners. The losers feel this cannot go on, but lack the ability to persuade the winners to raise prices, so they decide to raise prices on their own outdated products first—this is reckless self-sabotage and digging their own graves. Domestic open-source model vendors make a name for themselves, and after their models gain some competitiveness they all go closed-source; qwen3.6/minimax2.7/glm5.1turbo, etc. have not been open-sourced. Other small vendors just tinker with the base models, distill the winners’ models, and then expect to charge money—they must be dreaming.
Half a year after the release of the GPT5 series, its logical rigor leads the field, but the most criticized aspect is its slow speed, especially GPT5.2, which is absurdly slow. For many ordinary requests, opus may start making changes after 20 seconds of context gathering, while GPT5.2 might analyze for 20 minutes before starting to make changes. The concurrent GPT5.2 codex model…
#https://jqknono.github.io/coding-plans-for-copilot/ Collected domestic and overseas open-source model plans, along with provider latency and token rates
Only curated domestic vendors with some reputation, not resellers, with a certain level of security assurance. In actual use, first-token latency, token rate, concurrency, and whether quantized are important but rarely mentioned by vendors. Comparatively, speed matters more, really. #jqknono.github.io/coding-plans-for-…
Those familiar with LLM benchmarks know that Opus often tops GPT in various leaderboards, but in actual use, GPT’s problem-solving ability is unmatched. As of posting, the latest major models in the GPT-5 series include gpt-5.2-mini, gpt-5.2, gpt-5.3-codex, gpt-5.3-codex-sp…
First, the task was merely to translate Chinese to English for a nodejs app that converts a mermaid into an image. This is a pure tool, and the task was only translation. So far, I have encountered the following issues while using the Zhipu model: 1. Extremely low concurrency limits, the more you need to use it, the lower the concurrency 2. Low token generation speed 3….
I was stunned for a moment when I saw it, then realized “plan” was translated as “package” GPT5.2high is ridiculously slow, GPT5.3 hasn’t been released yet, but 5.3-codex came out first. In actual testing, the speed improvement is significant, and the logic capability drop is not as obvious as the drop from GPT-5.2-Codex compared to GPT5.2. It is a specialized coding model worth switching to…
When some model developers first release a new version, the response speed is ideal and there is no intelligence degradation. But if too many people start using it, both speed and quality decline. This is the behavior of a leader—once performance falls behind, users churn and the stockpiled cards sit idle. To avoid this, one could consider decoupling heavy assets from intellectual property assets, selling model usage rights, or renting hardware resources.