Anthropic Subscriptions Offer 5x+ More Value Than OpenAI Limit testing every AI subscription plan from Anthropic, OpenAI, Meta, SpaceXAI, MiniMax, Moonshot, Z.ai, Cursor, and Cognition Andrew Megalaa , Max Kan , and Dylan Patel Oct 05, 2026 ∙ Paid 76 7 Share Subscription plans are still the primary way consumers and small businesses pay for AI. These plans are highly subsidized—as we previously explained in June —but can still make economic sense as powerful customer acquisition and marketing tools. For example, the goodwill engendered by OpenAI’s generous resets is partially responsible for the recent surge in Codex adoption and has forced Anthropic to repeatedly walk back planned subscription nerfs to avoid getting clobbered in the court of public opinion. Furthermore, because subscription plans are so heavily subsidized, they can also have a large impact on margins and revenue per MW despite only being a small portion of total revenue. Consider the following rough numbers for Anthropic. Source: SemiAnalysis Tokenomics Model Despite being just 10% of overall revenue, subscriptions can take up over 40% of inference compute and lower blended revenue per MW by ~$36M. Subscriptions are even more important for OpenAI, as they make up a larger portion of their total revenue. For specific numbers, see our Tokenomics Model . In other words, if you want to accurately model AI lab financials, you need to understand their subscription limits. So how do limits actually work? The way to think about subscriptions is that your monthly payment grants you some numbers of “credits”. Each (model, token type) combo consumes a different amount of credits. Because credit cost ratios can differ dramatically from API price ratios, the “value” of the same plan changes depending on what model and workload you’re running. Put differently, it doesn’t make sense to say a plan is worth $X in isolation—you need to consider the full (plan, model, workload) tuple . Here’s how the API-equivalent value of the same $200/month Claude plan changes depending on the model and workload. Source: SemiAnalysis Tokenomics Model A single static snapshot isn’t enough either. Labs publicly change their limits all the time with promos and new model releases. They can also silently change limits whenever they want by tweaking credit costs. Ideally, you’d re-check the cost of each (plan, model, token type) daily so you can surface any changes in real-time. This is exactly what SemiAnalysis has done with our new Subscriptions Dashboard available exclusively to Tokenomics Model subscribers. Besides every OpenAI and Anthropic subscription, we also track Meta, SpaceXAI, Cursor, Cognition, Z.ai, MiniMax, and Moonshot. A screenshot of our dashboard showing a small subset of the available data. Source: SemiAnalysis Tokenomics Model New providers, plans, and models will be added as soon as they’re released. The rest of the article will give an overview of everyone’s current limits. All token and dollar amounts below assume agentic usage unless otherwise specified. Methodology Tokens are generally priced per MTok (million tokens) across the following types: Input: Fresh tokens added to the LLM’s context window that aren’t cached Cache write: Input tokens that are cached for multi-turn conversations. Generally slightly more expensive than regular input tokens. Cache read: Tokens from previous turns that are already cached in the conversation. Generally extremely cheap compared to uncached tokens. Output: Tokens generated by the model. The most expensive token type. Additionally, some models charge more for tokens above a certain context window. Those models typically compact before that expensive context window is reached, so we also keep our measurements below it. Source: OpenAI Subscription plans don’t expose this fine grained pricing. Instead, they provide a simple 0 to 100% usage meter across 5-hour and 7-day windows, and sometimes a separate meter for a model like Fable. Subscription tiers are then differentiated by usage multipliers relative only to the provider’s other plans. OpenAI for example used to advertise “Expanded Codex usage” in their $20 Plus plan, “5x more usage than Plus” in the $100 Pro plan, and “20x more usage” in the $200 Pro plan until they cut their $200 plan usage in half and removed all relative usage from their pricing page. Screenshot of Claude Code meters. Source: Claude Code CLI Setup We compute the subscription-rate of each (plan, model, token type) triple by running experiments that isolate one token type at a time and watching how far the model provider’s meter moves. An experiment is some number of repeated calls using a specific prompt that maximizes one token type while minimizing all the others. Source: SemiAnalysis Tokenomics Model Input, cache writes, and cache reads share the same prompt template. We use a portion of War and Peace since some models will refuse to respond if given large blocks of gibberish. For the input token experiments, we use a random tag on every call to ensure nothing gets cached. The cache write experiments run the same way but with the prompt marked for caching, so each new tag forces a new cache entry. The cache read experiments use a fixed tag, so the first call writes the cache and every repeat reads it. For the output token experiments, we used a technical essay to force long outputs because models will refuse mechanical prompts like “repeat SemiAnalysis 100,000 times”. Computing the Results For every call we record two things: how many tokens of each type the provider billed, and what the usage meter read. Turning those into a price takes three steps. 1. Measure the rate of each token type Providers generally report usage on a meter that moves in fixed amounts. This could be a whole percentage point, a credit, a cent etc... A single request often does not move the meter at all, so the cost of one request cannot be measured directly. Instead, we measure in steps. As requests run, we keep a running total of the tokens used. Each time the meter rises, we store that total. The tokens used between two meter moves make up one step, which is the cost of moving the meter by one unit. We drop two partial steps. When a run starts, the meter is already part of the way to its next move, so the tokens before the first move are less than a full step and would make the rate look higher than it is. The tokens after the last move never finish a step, so we drop those as well. To get a rate, we add up how far the meter moved over the complete steps and divide by the tokens used in them. Providers charge differently for input, output, and cached tokens, so we calculate a separate rate for each. Source: SemiAnalysis Tokenomics Model Caveats: Because we see the meter only once per request, a single step can be off by up to one request. These errors don’t add up with more consecutive steps, so we track a range instead of a single number, and the range gets narrower the more steps we count. We keep adding steps until the range is within ±5%, then report the rate across all of them. No request contains only one type of token. Some plans, for example, require a set of instructions on every request. We subtract that extra amount using the prices we measured for the other types. Reading from the cache costs almost nothing on some plans and may not move the meter at all. If the meter does not move after 500M cache read tokens, we assume that cache reads are free. 2. Convert rates into tokens per window and per month We read every meter a plan shows after every request. This could be a 5-hour limit, a weekly limit, and for models like Fable, its own weekly limit. So each token type gets a price on each meter. This is how much of that limit a million tokens uses. From each price we work out how many tokens fit e.g. if a million tokens use 5% of a limit, the whole limit holds 20 million. Source: SemiAnalysis Tokenomics Model The 5-hour limit resets many times within a week, so a month is capped by the weekly limits. 3. Pricing Lastly, we compute the API-equivalent value by converting the limit % consumed by a million tokens of each type into a dollar amount. For example, if you consume 1% of a $200 plan’s monthly limit, that’s $2. Then, we assume a workload shape, and compute how many total tokens you can get from each plan given this ratio. For the agentic workload, we use our own usage ratios in September as reported in our Tokenomics Model . Finally, we multiply by the blended price per MTok at API prices to get the API-equivalent value. Source: SemiAnalysis Tokenomics Model Catching a Provider A/B Test While refining our methodology, we ran into a very confusing situation where just 1 of 3 of the same subscription we tested for a particular provider had ~20% lower limits than the other 2. This unlucky account happened to be significantly older as well, and we were worried that the provider had some cursed setup where they changed subscription l
Anthropic Subscriptions Offer 5x+ More Value Than OpenAI Limit testing
Anthropic Subscriptions Offer 5x+ More Value Than OpenAI Limit testing every AI subscription plan from Anthropic, OpenAI, Meta, SpaceXAI, MiniMax, Moonshot, Z.ai, Cursor, and Cognition Andrew Megalaa , Max Kan , and Dylan Patel Oct 05, 2026 ∙ Paid 76 7 Share S
这条信息对 FDE 的直接价值在于提醒交付人员持续关注模型、智能体与企业流程之间的变化。面对类似项目,应先确认客户的真实业务目标、数据边界、权限条件和验收指标,再选择工具并用最小场景验证结果,避免只追逐功能更新。