Technology

Your cloud bill is rising while your AI bill falls

Cloud prices are climbing in 2026 while AI inference gets cheaper. Here is the memory shortage behind it, and what it does to your runway.

James Hitch
James Hitch· COO
Published Sep 11, 2026
21 min read
Your cloud bill is rising while your AI bill falls

Two major technology costs are moving in opposite directions. AI model inference is becoming cheaper, with some providers reducing prices or cancelling planned increases. At the same time, the cost of the infrastructure needed to run ordinary cloud workloads is starting to rise.

The reason is closely tied to the AI boom itself. Demand for memory and computing hardware has increased sharply, putting pressure on the supply chain and raising the cost of servers and storage. These increases are now beginning to reach cloud customers, although they may not always appear as a simple increase in hourly instance prices.

Recent data shows how quickly this shift is happening. US producer prices for computer storage devices rose by more than 28% between September 2025 and June 2026. Prices for finished electronic computers have also started to increase after remaining relatively stable for more than a year. At the same time, memory manufacturers such as Micron are reporting record revenues and exceptionally high margins.

Meanwhile, AI providers are moving in the other direction. Anthropic recently confirmed that a planned price increase for Claude Sonnet 5 would not go ahead, keeping its existing $2 per million input tokens and $10 per million output tokens pricing in place.

For founders, this creates an unusual budgeting problem. The cost of using AI may be falling, while the infrastructure supporting the rest of your product becomes more expensive. Understanding where those increases appear, whether through storage, hardware, networking, or changes to cloud pricing structures, will become increasingly important when planning technology costs.

In this article

  • What changed this month?

  • Why is memory the thing that got expensive?

  • Is AI getting cheaper or more expensive?

  • What does this do to your runway?

  • What should you do about it this month?

  • Where this leaves your engineering budget?

  • Conclusion

  • FAQ

What actually changed this month?

Four significant pricing changes have landed within the past twelve weeks. Three increase the cost of running infrastructure, while one reduces the cost of using an AI model. On their own, these changes can look like routine pricing updates. Together, they show that the cost of technology is shifting in different directions depending on what you are paying for.

OVHcloud is changing its pricing on 1 October 2026. The advertised hourly price of its instances will remain the same, but some services previously included in certain plans will become separate charges. Local storage and public IPv4 addresses, which were previously bundled into B3, C3 and R3 instances, will now appear as individual line items on customer invoices.

This distinction matters. A company could look at OVHcloud's advertised instance price and assume nothing has changed, while still paying significantly more for the same overall configuration. Depending on the instance size, OVHcloud says the total increase for an identical setup can range from 1.4% to 21.9%.

Hetzner made a more direct pricing change in June. Several cloud plans increased substantially, particularly memory-heavy configurations. Some plans more than doubled in price, while others saw smaller increases. Hetzner attributed the changes to rising procurement costs and confirmed that the increases affect new orders and server rescaling.

Google Cloud also increased prices in 2026, although for a different part of the infrastructure stack. On 1 May, the company raised peering egress prices across several regions. In North America, the price doubled from $0.04 to $0.08 per GiB. Google did not attribute the increase to hardware costs, instead saying the changes were intended to align pricing with the value and performance of its services.

Anthropic moved in the opposite direction. Claude Sonnet 5 was expected to increase from $2/$10 to $3/$15 per million input and output tokens on 1 September. That increase has now been cancelled, with the original introductory pricing becoming the standard price instead. For companies that had already updated their AI budgets based on the expected increase, that forecast now needs to be revised.

The public reaction to these changes also helps explain why some cost increases receive more attention than others. Memory pricing has become a growing topic of discussion, and Hetzner's price changes generated extensive debate among developers. Google Cloud's network pricing increase received comparatively little attention.

That difference is important. Developers tend to notice changes when they appear clearly on a price list. They are less likely to notice increases caused by services being unbundled and added as separate charges. The result is that some of the biggest changes to infrastructure spending can happen quietly, even when the advertised price of the core product stays exactly the same.

Why is memory the thing that got expensive?

Memory matters because it is one of the biggest factors in the cost of building a cloud instance. Over the past year, memory manufacturers have also shifted more of their production capacity toward products used in AI infrastructure. That combination has pushed up the cost of conventional memory and created pressure further down the infrastructure market.

What actually determines an instance price

Research into infrastructure-as-a-service pricing has found that RAM and CPU capacity are among the strongest factors influencing the price of a cloud instance. Storage and data transfer matter, but they have historically had less influence on the base price.

The specific price levels from older research are no longer relevant, but the underlying structure remains familiar. Cloud providers still organise their products around combinations of CPU and memory. A provider can change processor generations or storage configurations, but RAM remains one of the fundamental resources being sold.

That means a substantial increase in memory costs is difficult for providers to absorb indefinitely. If one of the main inputs behind an instance becomes significantly more expensive, the increase eventually has to appear somewhere in the customer's bill.

What happened upstream

The current pressure does not appear to come from a sudden disruption such as a factory closure. Instead, it comes from manufacturers allocating their limited production capacity toward more profitable products.

Demand for high-bandwidth memory, or HBM, has grown alongside the expansion of AI infrastructure. HBM is used alongside advanced AI accelerators and generates significantly more revenue than conventional memory. When manufacturers have limited advanced production capacity, it makes commercial sense to prioritise these higher-value products.

The consequence is less available capacity for conventional DRAM. TrendForce described this shift in 2025, noting that major DRAM suppliers were prioritising high-end server memory and HBM, which reduced capacity available for other markets.

The financial results of memory manufacturers also show how significant the shift has become. Micron's revenue and margins increased dramatically during this period. That kind of growth suggests that the industry is benefiting from much higher prices and demand, rather than simply shipping a larger volume of memory at unchanged prices.

What the official price indices actually show

There is also evidence outside of analyst forecasts.

The US Bureau of Labor Statistics tracks the prices manufacturers receive for computer storage devices. That index increased from 62.597 in September 2025 to 80.201 in June 2026. This represents a 28.1% increase in nine months.

The historical context is important. Between early 2023 and late 2024, the same index generally moved slowly downward. That makes the sharp increase in 2025 and 2026 particularly noticeable.

A broader index covering electronic computer manufacturing tells a similar story. Prices remained relatively stable for roughly eighteen months before rising sharply in July 2026. In other words, the cost pressure appears to have started with components before eventually reaching finished computing equipment.

European data points in the same direction. Eurostat's producer price index for computer, electronic and optical products across the EU increased by 6.1% between September 2025 and July 2026.

The useful comparison is the broader semiconductor market. If this were simply a shortage affecting all chips, semiconductor prices should be rising across the board. They were not. The relevant US producer price index for semiconductors actually declined slightly over the same period.

That makes the current situation more specific. This is not evidence of every type of chip becoming more expensive. The pressure is concentrated around memory and storage, with the effects now spreading into the products built around them.

The number you have probably seen, and why it is not the number

You may have seen reports claiming that DDR4 prices increased by 158% and DDR5 prices increased by 307%. Those numbers originated from TrendForce, but they are often repeated without the context that came with the original report.

They referred to movements in the spot market during late 2025. They were not a measure of the contract prices paid by large cloud providers buying memory in bulk.

TrendForce itself cautioned that spot prices are not the best indicator of broader pricing trends and suggested focusing on contract prices instead. This distinction matters because cloud providers typically purchase memory through long-term supply agreements rather than buying it on the spot market.

The figures are still useful, but they should be interpreted correctly. They demonstrate how extreme the market movement became during that period. They should not be presented as the percentage increase paid by every company buying memory.

There is also a gap in the public data. No freely available official index tracks DRAM prices directly. That is why paid analyst estimates often dominate reporting on memory prices, even though broader official manufacturing indices provide useful evidence of the same trend.

How a bundled item becomes an invisible increase

Once an input becomes more expensive, a cloud provider has a choice about how to recover the additional cost.

It can increase the advertised price of an instance. This is the obvious approach because every pricing page and comparison tool immediately reflects the change.

Alternatively, it can keep the headline instance price unchanged while separating services that were previously included. Storage, IP addresses or other resources can become separate charges. The advertised compute price remains technically unchanged, but the total cost of running the same configuration increases.

OVHcloud has chosen the second approach. Its published examples show that the percentage increase varies significantly depending on the size of the instance. Smaller instances experience much larger percentage increases than larger machines.

The reason is straightforward. The additional charges are largely fixed. A storage allocation or IP address represents a similar euro amount regardless of whether it is attached to a small or large instance. That fixed increase becomes a much larger percentage of the bill for a cheaper machine.

For a founder, this matters because early-stage infrastructure is often built from smaller instances. A pricing change that looks modest when viewed as a fixed monthly amount can have a disproportionately large effect on the services a startup is most likely to use.

The result is a regressive increase. The smallest machines take the largest percentage hit, while the largest instances absorb the same additional cost much more easily.

Is AI getting cheaper or more expensive?

AI is getting cheaper on published rate cards, at least for the models most businesses are likely to use. The exception is the newest flagship tier, where vendors continue to introduce more expensive models for workloads that need maximum capability.

The distinction matters because a new premium model is not automatically a price increase. Your costs only rise if you move to it.

Line itemBeforeNowEffective change
Claude Sonnet 5, per million tokens$3/$15 scheduled$2/$10 permanentIncrease cancelled
Claude Opus 5, per million tokensUnchanged$5/$25Current list price
Claude Haiku 4.5, per million tokensUnchanged$1/$5Current list price
GPT-5.6 Terra, per million tokensUnchanged$2/$12Current list price
GPT-6 Astra, per million tokensUnchanged$10/$50New premium tier
OVHcloud instance, identical configBundled services+1.4% to +21.9%Increase from October 2026
Hetzner CCX13, Germany and Finland€15.99/month€42.99/monthIncreased June 2026
Google Cloud peering egress, North America$0.04/GiB$0.08/GiBDoubled May 2026

Three points are worth stating directly.

First, the new premium OpenAI flagship is not a price increase for existing users. GPT-6 Astra sits above the existing model range rather than replacing it. GPT-5.6 Terra remains available at a much lower rate. The risk is behavioural rather than contractual: costs increase if someone changes a production default to the more expensive model without considering the impact on usage.

Second, published token prices are not the same as effective costs. Anthropic notes that newer generations use a tokenizer that can produce approximately 30% more tokens for the same text. That means a reduction in the price per token does not necessarily translate into an identical reduction in the cost of processing the same workload.

The direction can still be downward while the real saving is smaller than the rate card suggests.

Third, not every comparison deserves to be included simply because a number exists somewhere online. Alibaba's Model Studio documentation lists its current flagship model but does not provide enough official pricing information to make a reliable comparison. Third-party trackers also disagree on the price of its predecessor. Without a vendor-published rate card, the more honest choice is to leave the number out.

The broader picture is therefore more complicated than either "AI is getting cheaper" or "technology is getting more expensive." Model inference prices are falling or holding steady for mainstream tiers. At the same time, the physical infrastructure required to run software is becoming more expensive in several areas.

The final row is important because it complicates the argument. Data processing and hosting prices in the US have remained almost flat despite increases in several underlying infrastructure inputs.

That does not mean the individual price increases are imaginary. It means providers have not passed every increase through evenly across their entire product catalogue.

The clearest conclusion is that AI and infrastructure are moving in opposite directions. The cost of generating intelligence is falling for many workloads, while some of the physical resources needed to run conventional software are becoming more expensive. For founders, the important question is no longer simply whether technology costs are rising or falling. It is which part of the stack you are buying.

What does this do to your runway?

It changes the direction of one of your largest technology costs. For much of the last decade, founders could reasonably expect cloud infrastructure to become cheaper over time. That assumption is becoming less reliable, particularly for companies running conventional workloads that depend heavily on memory, storage and large numbers of small instances.

Research into historical AWS pricing found that cloud costs fell consistently between 2009 and 2016. Compute prices declined by an average of around 7% per year, while database and storage prices fell even faster. A generation of startup financial models was built around that pattern. If usage remained stable, infrastructure was expected to become more efficient and cheaper over time.

The recent data complicates that assumption.

There is an important counterpoint. The US Bureau of Labor Statistics producer price index for data processing, hosting and related services increased by only 0.2% between September 2025 and July 2026. At an industry-wide level, hosting prices have barely moved.

Both observations can be true because they measure different parts of the market. Broad hosting prices reflect large providers with long-term customer commitments and reserved contracts, which can delay the impact of rising input costs. Manufacturing data measures the cost of the underlying equipment, and those costs have already increased.

The most reasonable interpretation is that the pressure has appeared upstream without yet showing up consistently in the industry average. That may change over the coming months. If the hosting index remains flat, however, then this will turn out to be a more limited story about specific providers rather than a broad shift in infrastructure pricing.

The immediate impact also depends on how a provider passes costs through.

Fixed charges for storage, IP addresses and other previously bundled resources affect companies with many instances more than companies running a small number of large machines. OVHcloud's own example shows this clearly. A configuration consisting of one gateway and ten private instances increases from €758.10 to €860.70 per month.

That is not an unusually large architecture. It resembles the kind of infrastructure a growing company can accumulate as it adds services, environments and internal systems.

A company running four large machines may barely notice a fixed additional charge. A company running thirty smaller instances will feel it much more because the cost increases with the number of machines rather than their overall computing capacity.

There is also a cash-flow issue hiding behind changes to commitment terms. If shorter commitments disappear and the best discounts require a twelve-month commitment, the effective minimum commitment has increased. That may look like a simple change to a pricing catalogue, but it affects startup runway directly.

A large company can exchange flexibility for a discount because it has a relatively predictable infrastructure requirement. An early-stage company often cannot. Its architecture, customer base and resource requirements may change significantly within a few months.

For founders building a financial model today, the infrastructure line deserves more attention than it did a year ago. It should no longer be treated as a cost that will automatically become cheaper with time.

This matters when estimating the total cost of building a SaaS platform, because infrastructure costs continue throughout the life of the product rather than appearing as a one-off development expense. It also matters during MVP development, where founders are often deciding how to divide a limited budget between engineering talent and the infrastructure needed to support the product.

The practical implication is not that startups should stop using cloud services or assume a major cost crisis is coming. It is simpler than that: rebuild the infrastructure forecast using current prices and leave room for increases.

For years, falling cloud costs were an assumption you could make without much thought. That assumption now needs to be tested rather than inherited.

Hire the top 2%.

Vetted developers, part-time or full-time, remote and ready, from $9.99/hr.

What should you do about it this month?

The good news is that you do not need to migrate providers to respond to these changes. The biggest savings can come from understanding what you are actually being charged for and removing resources you no longer need.

There are five steps that you should take:

  • Start with your invoice, not your dashboard.

The easiest change to miss is an included service becoming a separate charge. Your instance dashboard can show the same compute price while your total bill increases. Compare your monthly invoice with the amount of work your infrastructure is handling. Looking at the cost per workload is more useful than watching the headline instance price.

  • Count your public IPv4 addresses.

Remove any that are no longer necessary. OVHcloud will charge around €1.97 per IP address each month from October. Its own example also shows that routing private instances through a gateway instead of assigning each one a public address can save roughly €2 per instance per month. Across a larger estate, those recurring charges add up quickly. Reducing unnecessary public addresses also limits the number of systems directly exposed to the internet.

  • Consolidate before you renegotiate.

The new costs are partly fixed per instance, while OVHcloud's discounts apply to compute rather than the separately billed storage and IPv4 charges. That means reducing the number of instances can have a larger impact than negotiating a slightly better compute rate. Where your workload allows it, moving from many small instances to fewer larger ones reduces the number of fixed charges you pay.

  • Review your commitments before you need to change them.

OVHcloud's new Savings Plans took effect on 1 September 2026. They offer discounts for longer commitments, while the previous 1, 6 and 24 month subscription terms are no longer available. If you previously relied on a six month commitment to maintain flexibility, you now need to decide whether a twelve month commitment makes sense.

Hetzner requires a different approach. Its June price changes apply to new orders and rescaling existing servers. An existing machine therefore does not automatically receive the new price simply because the provider changed its catalogue. Rescaling can trigger the new pricing, so capacity changes should be planned rather than made casually.

  • Reforecast your AI spending.

If your budget assumed Anthropic would increase Claude Sonnet 5 from $2/$10 to $3/$15 per million tokens, remove that increase. Anthropic has confirmed that the scheduled price rise will not happen.

There may also be opportunities to reduce the cost further. Anthropic offers a 50% discount on input and output tokens for batch processing. Work that does not need an immediate response, such as reporting, enrichment and many back-office tasks, can potentially be moved to batch processing.

The overall strategy is straightforward: audit what you are actually paying for, remove unnecessary infrastructure, consolidate where possible, make commitment decisions deliberately, and take advantage of lower AI pricing where your workload allows it.

You do not need a migration project to start saving money. You need a more accurate picture of what your current infrastructure is costing you.

Where this leaves your engineering budget

The uncomfortable part of this month's price increases is that most of them are outside your control. You cannot negotiate down the cost of memory. You cannot opt out of an IPv4 charge unless you remove the address. When providers are passing on genuine increases in their own costs, there is little room for customers to push back.

Engineering capacity is different. This is one area where you still have control over the numbers.

RocketDevs Associate engineers start at $9.99 per hour, with mid-senior engineers at $21.99 and senior engineers at $30.99. All have 100% EU timezone overlap. These rates are based on a screening process rather than a reduction in engineering quality. Candidates complete 6 to 8 hours of structured assessment, and only the top 2% are placed.

The useful comparison is what those costs mean alongside a rising infrastructure bill. An additional €100 per month for a small fleet becomes €1,200 over a year. One hundred hours of Associate engineering costs $999. For a founder managing a fixed budget, that is a real allocation decision.

The same calculation applies when you are working out the total cost of hiring a remote developer. Infrastructure is one recurring expense. Engineering is another. Understanding both lets you decide where additional budget creates the most value.

If you want to test the model before making a longer commitment, RocketDevs offers a 14-day risk-free trial with a money-back guarantee. You could use that period to give a vetted engineer the infrastructure audit outlined above and evaluate the results before deciding whether to continue.

Conclusion: The infrastructure bill is changing. Your strategy should too.

The important story is not that cloud computing suddenly became expensive, or that AI suddenly became cheap. Both statements are too simple.

What changed is the relationship between the different costs of building software. Memory and storage are under pressure. Cloud providers are passing some of that pressure through in ways that are not always obvious from the headline instance price. Network costs are also moving in some parts of the stack. At the same time, the cost of using capable AI models is continuing to fall. That creates a different kind of engineering budget.

The old assumption was that infrastructure would gradually become cheaper while engineering remained the larger investment. You could afford to let a collection of small cloud instances grow because the underlying unit economics were improving. That assumption is no longer safe enough to leave untested.

The answer is not to panic or start moving workloads between providers every time a price changes. It is to understand your actual cost per workload and know which parts of the bill are within your control.

Remove resources you do not need. Consolidate instances where it makes sense. Review commitments before they lock you in. Take advantage of cheaper model inference where the workload allows it. Most importantly, stop treating the cloud invoice as a fixed cost of doing business.

For an early stage company, every recurring cost competes with something else. A thousand euros spent keeping unnecessary infrastructure running is a thousand euros that cannot be spent on engineering, product development or acquiring customers. That is the real lesson from this month's price changes.

The cost of technology is not moving in one direction. Your budget should not assume that it is. The startups that handle this well will not necessarily be the ones with the cheapest cloud provider or the cheapest AI model. They will be the ones that know exactly what they are paying for, understand which costs are rising, and move their budget toward the work that creates the most value.

Infrastructure will keep changing. Your ability to respond to it is the part you can control.

Frequently asked questions

Why are cloud prices going up in 2026?

Memory and storage costs have increased, and some cloud providers are passing those costs on to customers. The US producer price index for computer storage device manufacturing rose 28.1% between September 2025 and June 2026 after a prolonged period of falling prices. Prices for finished electronic computers also began rising in July 2026 after remaining broadly stable for eighteen months.

The pressure comes partly from memory manufacturers shifting advanced production capacity toward high bandwidth memory used in AI accelerators. That has reduced the capacity available for conventional server memory. The broader semiconductor price index has not increased, which suggests this is primarily a memory and storage issue rather than a general chip shortage.

Is AI getting cheaper or more expensive?

AI is getting cheaper at several of the tiers most companies are likely to use. Anthropic has kept Claude Sonnet 5 at $2/$10 per million input and output tokens after cancelling its planned increase to $3/$15. OpenAI's new flagship sits at a higher price, but it was introduced as a premium tier rather than replacing the cheaper models already available.

There is one important caveat. Newer models can use more tokens to process the same text because of changes to their tokenizers. Anthropic says some newer models produce roughly 30% more tokens for the same text. The headline price can therefore fall by more than the effective cost of processing an equivalent workload.

Should I commit to reserved cloud capacity now?

Only if you are confident you will need that capacity for the full commitment period. OVHcloud's new pricing structure offers discounts for twelve and thirty six month commitments while removing the shorter subscription terms.

You also need to check exactly what the discount applies to. At OVHcloud, the discount covers compute but not the newly separated storage and IPv4 charges. If the components becoming more expensive are excluded from the commitment discount, a longer contract may provide less protection than the headline percentage suggests.

How much of a startup's budget should go to infrastructure?

There is no reliable public benchmark that gives every startup a sensible infrastructure percentage. The right amount depends on the product, architecture, usage and stage of the company.

The more useful question is whether your infrastructure costs are moving in the direction your financial model assumes. Cloud prices historically fell over time, but several providers are now increasing prices for specific infrastructure components.

For the next four quarters, it is safer to model infrastructure as a cost that could rise rather than one that will automatically become cheaper. Keep flexibility where you can, particularly in areas where your workload is still changing rapidly. Then use your available budget on the engineering work and infrastructure that directly supports growth.

James Hitch, COO at RocketDevs.LinkedIn

Sources

Serious devs. Serious value.

The top 2% of applicants, rigorously vetted, from $9.99/hr. Part-time or full-time, dedicated to your team.

  • Top 2% of applicants
  • 6–8 hours of human vetting
  • 14-day risk-free trial
James Hitch

Written by

James Hitch

COO

James Hitch is the COO of RocketDevs, where he runs sales, recruiting, and the vetting operation that accepts only the top 2–3% of developer applicants. He cares about putting accessible, elite engineering talent within reach of founders and startups worldwide, at a fair price. He writes about technical hiring, building AI-native engineering teams, and how startups can access elite developers affordably.

Share this article

Help others discover this content