Epoch AI:研究、数据与评测 · 2026/9/22

The plunging price of thought | Epoch AI Key Takeaways Over the past t

The plunging price of thought | Epoch AI Key Takeaways Over the past three years, the cost of a given level of AI performance has fallen an average of some 47% per quarter. That is a 13-fold drop every year – a faster rate than any other transformative technol

The plunging price of thought | Epoch AI Key Takeaways Over the past three years, the cost of a given level of AI performance has fallen an average of some 47% per quarter. That is a 13-fold drop every year – a faster rate than any other transformative technology in history. This rate is measured across the five benchmarks of AI capability for which we have the best data from the past three years — covering mathematics, hard sciences, and games of skill — as well as less sophisticated analysis with earlier data. We see somewhat slower cost drops on game-based puzzles, at 39–43% per quarter, and faster progress on math problems, at 50–52% per quarter. Costs have probably been falling this fast since the dawn of commercial LLM inference in November 2021, when OpenAI fully released GPT-3. The cost of a given level of performance often falls fastest right after that level is first achieved, that is, when it is state of the art (SOTA). Three of our five benchmarks exhibit this pattern. Averaging across all five, cost falls 66% per quarter (75x per year) for performance that has just debuted as SOTA. Two years later, prices fall half as fast, at 32% per quarter (4.7x per year). Despite falling prices, AI spending could remain high. If an important job for AI, like reviewing thousands of scientific papers for errors, demands as much cognition as running one of these benchmark tests a million times, then the spending would still add up even at a penny per run. Moreover, while prices fall, AI’s capability could keep rising. The price of passing a first-grade math test may now be trivial. The price of proving hard theorems is not. Data and code are on GitHub . An overlay page has many plots and tables to explore. Overview The “GPT” in “ChatGPT” stands for “Generative Pre-trained Transformer,” a technical description of how the AI inside it works. Surely, though, the creators of OpenAI’s GPT models were nodding to an older meaning of the initialism: general-purpose technology . They correctly foresaw that — like the steam engine, electricity, and the internet — large language models would someday touch every aspect of society. It is now widely understood that the AI boom is a macroeconomic force powerful enough to raise prices for the inputs it demands: chips, power, even the labor of electricians. Less well recognized is a paradoxical flip side: the price of the output from all those data centers is falling extraordinarily rapidly: The next chart shows some examples. On January 31, 2025, OpenAI released a new iteration in its series of “reasoning” models, called o3. We estimate that for an average cost of 30 cents per question, it could achieve a 75% score on GPQA Diamond , a multiple-choice exam covering PhD-level physics, chemistry, and biology. 1 Just under 18 months later, OpenAI released GPT-5.6 Luna. It scored just as well — for four hundredths of a penny per question ($0.0004). That is a 725-fold drop in the price of thought in under 18 months. It is like the sticker price on a new car falling from $50,000 to $69. No other general purpose technology in history appears to have gotten so cheap so fast. To measure these trends, we analyzed performance with a new and more comprehensive dataset that includes five AI performance benchmarks covering mathematics, hard sciences, and games of skill over the last three years. Across that time, we find that the price for a given level of performance has fallen about 47% per quarter, or 13x per year. We see slower drops on game-based puzzles, at about 39–43% per quarter (7–10x per year), and faster progress on math problems, at 50–52% per quarter (16–19x per year). The cost decline for a given level of performance does tend to slow over time, though this pattern is not universal. One possible explanation: when a performance level is first achieved, AI companies can briefly charge a premium for it, before competition and technological improvement quickly drive down the price. In time, that dynamic slows. Averaging across the five primary benchmarks, cost falls 66% per quarter (75x per year) at first. Two years later, it falls half as fast, at a “mere” 32% per quarter (4.7x per year). Our analysis comes with major caveats. AI companies may be expressly training their models for some benchmarks (“benchmaxxing”), so that improvement on the benchmarks outstrips improvement for real-world tasks. Even if they are not, doing well on a benchmark is not synonymous with useful work. Because we focus on the frontier — the absolute cheapest model capable of any given level of performance — we implicitly posit an AI user who relentlessly searches for the most cost-effective model for each task, when real users do not switch models so often, and therefore do not reap quite the same savings. Our data are incomplete and noisy: the timeframe is barely three years, and we do not include all combinations of AI model and benchmark. Prices drop differently for different models, benchmarks, time periods, and performance ranges, and there are many reasonable ways to average over this variegated experience. Overall, while we believe that our bottom-line numbers are reasonably representative of reality, they should not be read as exact. The rest of this report details our analysis. Parts of it are technical. Previous work We are not the first to quantify how fast the price of AI is falling. A 2024 post by Guido Appenzeller for Andreessen Horowitz documented how, in the three years following the general release of GPT-3, LLM costs fell by a factor of 1000, i.e., 10x per year. A few months later, in March 2025, an analysis by Epoch AI found 9–900x per year drops across six performance benchmarks. Both of those early analyses measured prices per token for models capable of achieving a given performance as distinct from actual cost to achieve given performance . With the advent of reasoning models — OpenAI released o1 in December 2024 — it has become more problematic to ignore this distinction. Reasoning models can productively consume far more tokens, but as a result extract good performance from an underlying LLM that is smaller and cheaper to run per token. In September 2025, Håvard Tveit Ihle shared an analysis on LessWrong that directly compared performance and cost on two suites of coding challenges. The analysis finds that costs halved every 1.4–2 months (64–380x per year). The most thorough analysis yet is the March 2026 paper by Gundlach et al. It, too, compares actual costs to performance. As in the present analysis, it estimates trends both in the full body of data and in the subset of models defining the cost frontier at any given time. It also disaggregates by level of performance — prices fall faster at the high end — and by model type (open or closed, dense or mixture of experts). Overall they find declines of 5–10x per year. Data The major novelty in the present analysis is to analyze trends in costs using a methodology that captures each model’s full expense-performance continuum. If one LLM can score 80% on a benchmark, and costs $1 to do so, and another peaks out at 60%, for a price of $0.50, it is not obvious which is more cost-effective. Perhaps the stronger model would, if run with fewer reasoning tokens, achieve 60% more cheaply than the weaker model. More generally, our interest is in mapping the “Pareto” frontier of cost-effectiveness — the cheapest way to attain each performance level from the models available at any given time. We would leave a lot of territory unmapped, and potentially a lot of frontier as well, if we only estimated cost and performance when LLMs are given an unlimited budget. Any given modern LLM can produce a range of performance levels, depending on whether its reasoning is set to low, medium, high, or max, and depending on whether a budget limit is imposed. One way to trace a model’s expense-performance curve is to run it many times against a benchmark, at various reasoning levels and token budgets. That process would, however, be expensive and slow. Instead, we follow a procedure developed by the federal Center for AI Standards and Innovation (CAISI) . It uses the transcript from a benchmark run with a high (or no) budget constraint to predict performance under tighter budgets. The core idea is that since benchmarks consist of many questions, one can estimate how many an LLM would answer before consuming a given budget. The transcript provides the needed information: how many tokens the model consumed in answering each question, and which it got right. 2 For any arbitrary per-question budget of X tokens, we sum up the number of questions that were answered correctly in fewer than X output tokens. Think of this as the maximum expense the model may incur before being forced to give up. When a model does not answer in time for a particular budget threshold, that question is scored as the probability of guessing correctly, which could be, for example, 0.25 on four-way multiple c

FDE 判断

这条信息对 FDE 的直接价值在于提醒交付人员持续关注模型、智能体与企业流程之间的变化。面对类似项目,应先确认客户的真实业务目标、数据边界、权限条件和验收指标,再选择工具并用最小场景验证结果,避免只追逐功能更新。