01Training costs tens of millions; inference fell 280-fold in two years
AI costs split into two layers: "training"—building a model—and "inference"—running it. The Stanford HAI AI Index 2024 estimated GPT-4's training cost at roughly $78 million and Google Gemini Ultra's at about $191 million. According to Epoch AI, the cost of frontier-model training has grown at 2.4× per year since 2016, and the largest training runs are projected to exceed $1 billion by 2027.
Inference prices, meanwhile, are moving in the opposite direction. Epoch AI's tracking shows that the price to achieve GPT-3.5-level performance fell from $20 per million tokens in November 2022 to $0.07 in October 2024—a roughly 280-fold decline. Training costs balloon; inference costs collapse. That asymmetry is the core of AI's cost structure.
02High fixed costs and falling marginal costs push toward concentration
Once a model is trained, the training expenditure becomes a sunk fixed cost. Inference is the variable cost, incurred each time a user runs the model—but its unit price is dropping fast. This cost curve resembles those of telecommunications and semiconductor fabrication: large upfront investment, declining per-unit cost as usage scales.
In economics, industries with this structure tend toward concentration among a few large operators. Only a handful of organisations can invest hundreds of millions of dollars in a single training run, so frontier-model development is converging on a small number of firms. Yet the steep fall in inference prices lowers the barrier to using models. Concentration on the supply side and diffusion on the demand side advance simultaneously. This dual movement is not a temporary phase; it is a structural feature of any industry where the ratio of fixed to variable cost is extreme and widening.
03Inference price declines vary sharply by task
The drop in inference prices is not uniform. Epoch AI tracked six benchmarks and found annual decline rates ranging from 9× to 900×. On GPQA Diamond, a PhD-level science benchmark, the price of reaching GPT-4 Turbo-level performance fell from $15 per million tokens in November 2023 to $0.12 in December 2024—about 40× per year. On the general-knowledge MMLU benchmark, prices fell from $60 in November 2021 to $0.07 in October 2024—roughly 1,000× in three years.
The implication: price disruption is fastest in the tasks AI handles best. For routine work such as summarisation, translation, and code generation, inference costs are approaching negligibility. For tasks requiring specialised judgement, high-capability model pricing remains elevated, and cost-effectiveness calculations still matter. The practical consequence is that organisations should evaluate AI adoption task by task, not as a blanket decision, because the economics differ dramatically across use cases.
04The cost breakdown is chips, people, and power—and power's share is rising
Epoch AI's analysis breaks down frontier-model development costs as follows: hardware (GPUs, servers, networking) accounts for 47–67%, R&D staff compensation for 29–49%, and energy for 2–6%. Energy's share is small today, but power demand is growing rapidly as training scales up.
According to the IEA's Energy and AI report, global data-centre electricity consumption stood at roughly 415 TWh in 2024, about 1.5% of world electricity use. The IEA projects this will reach 945 TWh by 2030—approximately 3%. That is an annual growth rate of about 15%. In the United States alone, data-centre power demand is expected to increase by 240 TWh between 2024 and 2030, accounting for roughly half of the country's total electricity demand growth. AI-accelerated servers are projected to grow at 30% per year, constituting nearly half of the net increase in global data-centre consumption.
| Cost element | Current share | Trend | Industry implication |
|---|---|---|---|
| Hardware (GPUs etc.) | 47–67% | Absolute spend rising; share stable | Continued dependence on semiconductor suppliers |
| Staff (researchers) | 29–49% | Upward pressure from talent competition | Geographic concentration of development hubs |
| Power | 2–6% | Share increasing | Becoming a constraint on site selection |
| Data (acquisition & curation) | Rarely disclosed | Scarcity rising | Proprietary data gains value |
05In pharma and healthcare, falling inference costs accelerate adoption
AI's cost structure affects pharma through two channels. The first is the computational side of drug discovery. AI models used for molecular design and protein-structure prediction require substantial compute for training, and that cost enters the R&D budget as a fixed expense. Falling inference prices make it cheaper to screen candidate compounds at scale, increasing the number of hypotheses a team can test computationally before moving to wet-lab validation.
The second channel is clinical and operational use. When AI assists with diagnostic support or documentation at the point of care, each inference call carries a cost. At the 280-fold reduction observed over the past two years, the computational cost per diagnostic assist becomes negligible. The binding constraint shifts from compute to institutional costs: verifying AI output, assigning liability for errors, and aligning with regulatory requirements. The same applies to promotional-material review: the cost of generating a draft with AI is falling fast, but the human cost of checking that draft against regulatory standards is unchanged.
06Training-data scarcity is becoming the next binding constraint
Training AI models requires data. Epoch AI estimates the total stock of high-quality, publicly available human-generated text at roughly 300 trillion tokens. GPT-4 was reportedly trained on 6–13 trillion tokens. As models grow larger and "overtraining" techniques consume more data per parameter, demand is accelerating. Epoch AI projects that training-data demand will exceed the supply of public text between 2026 and 2032.
The finiteness of data has two implications for AI's cost structure. First, organisations that hold proprietary data gain bargaining power. Medical records, clinical-trial datasets, and regulatory-submission archives—data that is not publicly available—become scarce inputs for AI training. Second, synthetic data is expanding to fill the gap. AI-generated data used to train subsequent models has shown results in narrow domains such as mathematics and code, but in areas where verification is difficult—drug discovery, clinical reasoning—the risk of compounding errors through successive training generations remains real. In pharma specifically, clinical-trial databases and adverse-event reports represent data assets whose training value will grow as public text becomes saturated. Managing data quality and quantity is a hidden cost embedded in AI's economics.
07Three lenses for reading the cost structure before deciding
First, separate training cost from inference cost. For an organisation adopting AI rather than building it, training cost is someone else's fixed expense; inference cost is the variable expense tied to usage. With inference prices falling at double-digit multiples per year, a use case dismissed as uneconomic a year or two ago may already be viable. Tasks that process large volumes of text daily—promotional-material review, clinical documentation—stand to benefit most from the decline.
Second, factor in power and data constraints. Data-centre power consumption will continue to grow, and regulatory and infrastructure constraints on siting will tighten. Organisations using AI services will increasingly need to know where their provider's data centres are located and what power sources they use. At the same time, this is the moment to reassess the value of proprietary specialist data. Clinical records and regulatory archives are appreciating as raw material for AI training. Organisations that inventory and curate these assets now will be better positioned as data scarcity intensifies.
Third, do not overlook the costs beyond compute. Falling inference prices are solving the compute-cost problem. What remains are the human costs of verifying AI output, the liability costs when errors occur, and the design costs of regulatory compliance. In pharma and healthcare, these institutional costs—not compute—are the rate-limiting factor for AI adoption. Getting the institutional design right is not a secondary concern; it is the primary determinant of whether falling compute prices translate into actual productivity gains.
- AI's cost structure is defined by the asymmetry between soaring training fixed costs (GPT-4 at ~$78 M, growing at 2.4× per year) and plummeting inference variable costs (down 280-fold in two years). This drives concentration on the supply side and diffusion on the demand side.
- Power and data are the emerging constraints. Data-centre electricity consumption is projected to more than double from 415 TWh in 2024 to 945 TWh by 2030, and high-quality training data may be exhausted between 2026 and 2032.
- In pharma and healthcare, the compute-cost barrier is disappearing. The binding constraints are now institutional: verifying AI output, assigning liability, and achieving regulatory compliance. Designing those institutional costs well will determine the pace of AI adoption.
Read the cost structure and you can read where the AI industry is heading. Rising training costs concentrate development among a few organisations; collapsing inference prices spread usage broadly. Growing power demand and data scarcity introduce new constraints that will reshape the cost curve. For pharma and healthcare, compute cost is no longer the primary barrier. Who verifies AI output, who bears liability for errors, and how regulatory compliance is achieved—the design of those institutional costs will determine how much value AI actually delivers.
- Stanford HAI. The AI Index 2024 Annual Report. Stanford University, 2024. https://hai.stanford.edu/ai-index/2025-ai-index-report
- Epoch AI. How Much Does It Cost to Train Frontier AI Models? 2024. https://epoch.ai/blog/how-much-does-it-cost-to-train-frontier-ai-models
- Epoch AI. LLM Inference Prices Have Fallen Rapidly but Unequally across Tasks. 2024. https://epoch.ai/data-insights/llm-inference-price-trends
- IEA. Energy and AI. International Energy Agency, 2025. https://www.iea.org/reports/energy-and-ai/energy-demand-from-ai
- IEA. Electricity 2026. International Energy Agency, 2026. https://www.iea.org/reports/electricity-2026
- Villalobos, P. et al. Will We Run Out of Data? Limits of LLM Scaling Based on Human-Generated Data. Epoch AI, 2024. https://epoch.ai/publications/will-we-run-out-of-data-limits-of-llm-scaling-based-on-human-generated-data
- Epoch AI. Inference Economics of Language Models. 2024. https://epoch.ai/blog/inference-economics-of-language-models
- a16z. LLMflation: LLM Inference Cost. Andreessen Horowitz, 2024. https://arxiv.org/html/2511.23455v2