I built Oginify — a free OG image generator. 3x daily, no signup. Try it →

Marketing & Growthpricing

Coding Plan Deep Dive: The Business Model of China's AI Subscription Pricing

From GLM to MiniMax, Alibaba to ByteDance — Coding Plan is China's unique AI subscription model, sitting between unlimited SaaS and pure token billing. This guide breaks down seven vendors, the business model logic, and three structural contradictions.

·Updated July 16, 2026·28 min

What Is a Coding Plan: The Third Pricing Model between SaaS and Token Billing

A Coding Plan is a subscription billing model designed specifically for China's AI developer market. Users pay a fixed monthly fee (¥29–¥699) and receive a dedicated API Key that works in Claude Code, Cursor, Cline, and other mainstream AI coding tools. Its essence: protect the user's psychological safety while protecting the vendor's cost baseline.

Why must this intermediate form exist? Because AI coding is fundamentally different from Notion or Figma — every API call consumes GPU compute. One heavy Claude Code user can consume as much in an hour as 50 light users in a week. Selling 'unlimited for a flat fee' means power users become profit black holes. Pure per-token billing makes users anxious checking bills daily — and Chinese developers especially aren't used to '¥X per million tokens' calculations.

So Coding Plan sits in the middle: more conservative than SaaS (three-tier rate limits), more user-friendly than token billing (fixed monthly fee, no mental math). On the business model spectrum, it occupies the gap between Notion (pure SaaS) and DeepSeek API (pure token) — precisely the space neither extreme can serve.

Fundamental Difference from Western AI Subscriptions {#}

Western AI coding subscriptions (Claude Code, GitHub Copilot, Cursor) are integrated products where you buy the tool and model together — a Claude Pro subscription bundles Claude models with Claude Code tools. Chinese Coding Plans are pure model subscriptions — the vendor gives you an API Key, and you inject it into any third-party tool supporting OpenAI or Anthropic protocols.

More critically, Chinese Coding Plans aggregate competitors' models. Alibaba Bailian's Coding Plan lets you switch between GLM, Kimi, and MiniMax. ByteDance Volcengine's offers Doubao, DeepSeek, and GLM. In Western markets this is nearly unthinkable — Claude's subscription will never let you call GPT-5.5. Chinese Coding Plan vendors are essentially API Key wholesalers — buying (or building) model access and bundling it into subscription packages for developers.

Seven Vendors: The Complete Landscape

As of July 2026, China's major AI vendors have formed a clear pricing landscape. But behind this neat-looking grid lies Baidu's four-month Coding Plan lifecycle — from launch in February 2026 to shutdown on June 25.

Zhipu GLM Coding Plan: The Market's Price Anchor {#}

Zhipu AI was the pioneer and pricing benchmark for China's Coding Plan market. Three tiers — Lite ¥49/Pro ¥149/Max ¥469 — defined the category's price range. GLM-5.2 consumes 3× quota during peak hours (14:00–18:00 UTC+8), 2× off-peak, dropping to 1× during a limited-time promotion. This time-based consumption multiplier is a patch that Coding Plan was forced to add because you can't adjust prices within a flat-fee model.

Third-party analysis by HyScaler notes that GLM Coding Plan's yearly-discount prices approach Claude Pro's $20/month, but 'at full price, Pro at $72/month costs more than three times Claude Pro.' This suggests Coding Plan's pricing competitiveness depends heavily on annual commitment discounts.

MiniMax Token Plan: From Coding Plan to Full-Modal Subscription {#}

MiniMax renamed Coding Plan to Token Plan in mid-2026, expanding coverage from pure text programming to full-modal (voice + video + image + music). The renaming itself signals the category trend — transcending the programming niche toward a unified AI resource consumption platform.

At Plus ¥49/Max ¥119/Ultra ¥469, the Ultra tier offers 7.1 billion tokens/month — the largest per-unit allocation across Chinese Coding Plans. But M3 launched without a highspeed variant (TPS 100 limited to M2.7-highspeed), reflecting the tradeoff between peak performance and cost efficiency.

Alibaba Bailian: Daily Limited-Release as Demand-Supply Control {#}

Alibaba Bailian's Coding Plan evolution path is the most instructive case study. From initial launch with Lite ¥40/Pro ¥200 plus ¥7.9 first-month discount, to Lite new-sale shutdown on March 20, Lite renewal shutdown on April 13, to current single Pro ¥200 tier with daily limited release at 9:30 AM, often selling out in seconds.

This trajectory reveals a deep demand-supply contradiction: Coding Plan relies on light users subsidizing heavy users — but light user profits can't sustain too many heavy users. The daily limited release is essentially filtering for 'light users willing to wake up early and compete for a slot,' rejecting indiscriminate inflow. If this pattern spreads, Coding Plan transforms from 'democratic subscription' to 'scarce resource' — a fundamentally different commercial narrative.

DeepSeek's Outlier Position: Firmly against Coding Plan {#}

DeepSeek researcher Chen Deli stated bluntly on Linux.do: 'This model has a fatal business logic defect — as user programming tasks increase, compute consumption rises sharply, causing vendors to lose more money the more users consume. This is textbook "losing money to buy buzz."'

DeepSeek insists on pure pay-per-use pricing: V4 Flash at $0.14/1M input tokens, V4 Pro at $0.435/1M, cache hit discounts exceeding 90%. Time-based pricing further regulates GPU load — off-peak 75% discounts on R1. Chen Deli's argument casts a shadow over the entire category's sustainability. In a red ocean where everyone is doing Coding Plans, the one refusing to participate gains a distinct brand position: cost-pricing advocate.

Three Structural Contradictions: Why Coding Plan Is an Unstable Intermediate Form

Baidu's four-month Coding Plan shutdown isn't an anomaly. It's the concentrated eruption of this model's inherent contradictions.

Contradiction 1: The Subsidy Flows Backward {#}

Traditional SaaS has heavy users subsidizing light users — Notion power users consume more but marginal cost approaches zero, so they effectively subsidize light users paying the same price. AI coding reverses this — light users (a few questions daily) are ¥180 profit on a ¥200 subscription, while heavy users (Claude Code agents running all day) can burn ¥500 in compute costs in a week.

Once the heavy-user ratio exceeds a threshold, the entire model collapses. That's why Alibaba killed the Lite tier and switched to daily limited releases — not to increase revenue, but to filter out heavy users.

Contradiction 2: The Aggregator's Profit Ceiling {#}

Alibaba Bailian's Coding Plan aggregates GLM, Kimi, and MiniMax — but GLM also sells its own Coding Plan, and Kimi has its own membership. Aggregator profits come from the spread between wholesale and retail prices. Whether this spread is sustainable depends on how quickly model vendors tighten their channel policies.

The deeper question: when the models you aggregate compete with your own business, how long can the 'competitor's model sold through my channel' business model survive?

Contradiction 3: Users Can't Rationally Compare Prices {#}

Zhipu uses prompt count as the quota unit, MiniMax uses request count, Alibaba uses request count — but the same request across different models can consume 10× differences in tokens. Users can't rationally compare 'which vendor is cheaper,' relying on trust and word-of-mouth rather than objective calculation.

This gives vendors significant pricing ambiguity room — but also seeds controversy. The moment one vendor publicly discloses a token-equivalent conversion standard (like GitHub Copilot's $0.01/credit), the entire category's pricing opacity shatters.

Historical Precedents: Coding Plan Didn't Emerge from a Vacuum

Flat-fee + usage-cap pricing has deep historical roots in international SaaS and infrastructure. Understanding these precedents helps assess Coding Plan's future trajectory.

The closest precedent is Heroku Eco Dyno — Salesforce's PaaS platform. $5/month provides 1,000 dyno hours; exhaust them and all apps are forced to sleep, with no option to purchase overage. Official documentation states explicitly: 'You can't purchase additional dyno hours. If you need your apps to be up and running, you can upgrade to the Basic dyno.' This is the exact prototype of Coding Plan's hard-upgrade signal.

Earlier still: Netflix DVD rental (2000–2005). Nominally $17.99/month 'unlimited rentals,' but the system identified heavy renters (18–22 DVDs/month) and artificially delayed their shipments, triggering a 2004 class-action lawsuit. Netflix's behavior shares the same logic as Coding Plan's consumption multiplier — unable to raise prices directly, regulate heavy users' service speed to protect margins.

On June 1, 2026, GitHub Copilot transitioned all plans from flat-rate to AI Credits usage-based billing — the largest pricing transformation in Western AI coding tools. It introduced a Base Credits + Flex Allotment dual-pool design — fixed plus variable, with $0.01/credit overage. This model is more flexible than Coding Plan's hard cap, but gives up Coding Plan's core selling point: the psychological safety of a completely fixed monthly fee.

International Hybrid Pricing Models Compared {#}

Notion AI's Custom Agents use a Credits model: $10/1,000 credits, auto-pausing on exhaustion. Intercom Fin uses outcome-based pricing: $0.99 per successful resolution. Salesforce Agentforce offers dual-track: $2.00 per conversation or Flex Credits at ~$0.10/action. These products represent the evolution from quota-based to finer billing granularity.

A telling signal: in June 2026, Salesforce signed an agreement to acquire Intercom Fin — the market reference price of $0.99 is now owned by a company simultaneously maintaining a $2.00/conversation product. As one industry observer put it: 'Independent AI-support pricing is disappearing into the suites.'

Evolution Direction: Where Coding Plan Goes from Here

Based on current trends, Coding Plan is likely to diverge into three paths. First: Token Plan-ification — MiniMax and Baidu are leading, abandoning prompt-count ambiguity for token-based quotas, while expanding from pure programming to full-modal. Second: limited-release normalization — Alibaba is the trailblazer, transforming Coding Plan from democratic subscription into scarce resource. Third: return to pure pay-per-use — if DeepSeek's path proves viable, declining compute costs could naturally close the window for flat-fee + cap models.

At the extreme, Coding Plan is a transitional form. It filled a market gap when compute costs were still too expensive, users weren't accustomed to token billing, and payment + proxy barriers existed. When these conditions change — compute costs drop low enough, users are educated into token billing, or barriers are reduced — this intermediate form may no longer be needed. Baidu's four-month Coding Plan lifecycle provides the most direct footnote to this assessment.

DeepSeek Time-of-Day Pricing vs GLM Consumption Multiplier: Two Solutions to the Same Problem {#}

GPU clusters have physical limits — everyone works during the day, massive idle capacity at night. Vendors must shift usage from peak to off-peak. DeepSeek uses price levers — 75% cheaper at night, users decide when to use. GLM uses quota levers — 3× consumption during peak hours, users feel 'afternoons waste quota.' The former is a natural extension of pure pay-per-use; the latter is a patch Coding Plan was forced to apply — because you can't change prices within a flat-fee model, you can only adjust consumption speed.

Conclusion

Coding Plan sits at an unstable intersection: between flat SaaS subscriptions and pure token billing, it trades predictable revenue for user goodwill — and the economics rarely survive scale. The structural contradictions are real: heavy users cap the vendor's margin, light users churn when unused credits expire, and every usage-multiplier patch (like GLM's consumption multiplier) is a workaround for a flat-fee model that cannot adjust price directly.

Treat Coding Plan as a launch and acquisition instrument, not a permanent pricing architecture. It works when a vendor needs to buy adoption, when costs are still falling fast, or when competition forces an accessible entry price. Plan the escape route from day one: a clear migration to usage-based tiers, meter data that carries over, and a story that explains why the flat plan graduates into metered pricing.

The direction of travel in China's market is toward token-level granularity — DeepSeek's time-of-day pricing and GLM's multipliers both point there. Expect Coding Plan to survive as a transitional product rather than a destination. If you are a buyer, read the quota and multiplier terms before committing; if you are a vendor, price the plan as customer acquisition with a defined path to recurring usage revenue.

References

  1. Zhipu GLM Coding Plan Official Documentation (Zhipu AI · 2026-07)Official documentation covering plan types, rate limits, consumption multipliers, and supported tools.
  2. MiniMax Token Plan Official Documentation (MiniMax · 2026-07)Complete Token Plan documentation including migration from Coding Plan, plan comparisons, and subscription key usage.
  3. Alibaba Cloud Bailian Coding Plan Documentation (Alibaba Cloud · 2026-07)Official documentation covering Pro plan details, supported model list, usage restrictions, and daily limited-release rules.
  4. Baidu Qianfan Coding Plan Upgrade Announcement (Baidu Cloud · 2026-06-25)Announcement of Coding Plan discontinuation, migration to Token Plan, user benefits, and transition timeline.
  5. China Coding API Roundup 2026 (CodePick · 2026-06)Cross-comparison of five Chinese coding APIs (Volcengine/Bailian/MiniMax/Zhipu/DeepSeek) covering pricing, quota mechanisms, and model support.
  6. GLM Coding Plan vs Claude/Copilot Review (HyScaler · 2026-06)Third-party pricing comparison of GLM Coding Plan against Claude Code and GitHub Copilot.
  7. A Guide to AI SaaS Pricing Frameworks (Stripe · 2026)Stripe's official guide: explains Hybrid Pricing model noting that 46% of SaaS companies use base+included allowance+overage.
  8. Heroku Eco Dyno Hours Official Documentation (Salesforce/Heroku · 2026)Official documentation for Heroku Eco plan: $5/month for 1,000 dyno hours, forced sleep on exhaustion, no overage option — identical to Coding Plan hard-cap logic.
  9. Intercom AI Agent Pricing Comparison (Intercom · 2026-07)Comparison of Intercom Fin $0.99/resolution with Salesforce Agentforce $2.00/conversation and Zendesk pricing models.

Your Product Deserves to Be Found.

Great AI products don't fail on quality — they fail on being invisible. We turn that around, one compounding system at a time.

Get started

This site uses cookies and similar technologies for analytics, personalized ads (via Google AdSense), and essential functions. By clicking “Accept All”, you consent to our use of cookies. You can reject non-essential cookies by clicking “Reject All”.

Privacy Policy