Why Your Enterprise Customers Still Have Data on AWS After Signing GPUs With You

Enterprise AI teams move compute to NeoCloud but leave training data on AWS. The cause is not inertia, it is a product gap: no storage tier to consolidate onto means data defaults to where the tiers are. Every training run then generates cross-cloud egress charges the tenant associates with your platform. This post explains the trap, how the hyperscaler credit window created it, and why attaching flat-rate, zero-egress storage at GPU contract signature is the only fix that closes it without a migration.
Stefaan Vervaet
August 12, 2026

The fix for stranded tenant data is a storage tier in the GPU contract itself, agreed at signature. Attach S3-compatible, flat-rate, zero-egress storage to the compute deal and the tenant consolidates off AWS, the recurring egress bill disappears, and your platform stops absorbing blame for a hyperscaler invoice. Retrofitting it six months later means a migration nobody wants to run.

A single training job on your GPU cluster reads 15TB out of an AWS S3 Standard bucket and runs about $1,350 in egress the tenant never quoted for. On AWS US East, moving that data to another provider costs roughly $0.09/GB for the first 10TB and $0.085/GB after that. They run the job again next week with a tweaked hyperparameter. Another egress bill. Month after month, every training iteration is metered. You won the compute contract. AWS is still collecting a monthly check, and your platform is the line item that looks expensive.

NeoCloud operators are winning the compute contract. They are not winning the storage conversation that should come with it. The result is a tenant paying twice: once to you for GPUs, and once to the hyperscaler for the data those GPUs cannot stop reading. You can hand a tenant a flat-rate storage exit before the second bill lands, but the reason the data is stuck there is worth understanding first.

What is data gravity in cloud computing? Data gravity is the tendency of data to attract related applications, compute workloads, and services to its physical location. The larger the dataset, the stronger the pull. For enterprise AI teams, training data stored on AWS keeps pulling compute back toward AWS even after a NeoCloud contract is signed, because every training read generates cross-cloud egress charges.

Why the GPU Migration Completed and the Data Migration Did Not

The tenant moved compute to you because you were faster, cheaper per GPU-hour, or both. The data did not follow. Most people read that as inertia or switching cost. It is neither. It is a product gap.

Hyperscalers sell storage in tiers. Hot for active datasets, warm for infrequent access, cold for archive. A tenant can put every class of data with one provider and manage it in one place. Most NeoClouds sell compute first and treat storage as an afterthought,  and where a storage tier exists, it rarely spans the full hot/warm/cold range a tenant needs to fully consolidate off a hyperscaler. There is no complete tiering path to consolidate onto, so the tenant has no straightforward route to bring all their data with them even when they want to.

This is not a loyalty problem, and it is not a migration-effort problem. A customer cannot consolidate storage with a provider that only sells one thing. When your offering is compute-only, the data stays where the tiers are, which is the hyperscaler the tenant was trying to leave. The GPU contract closed. The storage question was never on the table, so it defaulted to AWS.

The Egress Bill That Grows With Every Training Run

Training is not a one-time read. Model iteration is the entire point of renting GPUs. Every epoch, every hyperparameter sweep, every fine-tuning pass pulls the dataset across the network again, and each cross-cloud read carries an egress charge from the hyperscaler.

The cost compounds with the work. A tenant running data-parallel training across a cluster reads the dataset repeatedly. A team iterating toward a production model runs the pipeline dozens of times before they ship. Every one of those reads leaves the hyperscaler bucket and lands on your GPUs, and every one of those crossings is metered by the provider the data sits with. On AWS US East that is $0.09/GB for the first 10TB and $0.085/GB thereafter, charges that accumulate on the tenant's invoice every time your GPUs touch their data.

The tenant sees one invoice from you and one from AWS, and the AWS number keeps climbing for reasons they associate with your platform, because your GPUs are what triggers the reads. You did not create the egress charge. You inherited the blame for it.

Training Scenario AWS Egress Cost (estimated per run) With Akave (zero egress) Estimated Annual Delta
15 TB dataset, monthly retrain ~$1,325 $0 ~$15,900 / yr
50 TB dataset, weekly retrain ~$4,300 $0 ~$223,600 / yr
100 TB dataset, continuous training ~$8,550 / mo $0 ~$102,600 / yr

AWS egress at $0.09/GB (first 10TB) and $0.085/GB thereafter, US East. Akave hot tier: $14.99/TB flat-rate, zero egress.

How Hyperscaler Credits Scattered Training Data Across Clouds

Follow the data back to how it got there. Most of this training data was ingested during a credit window. Hyperscalers often distribute large credit packages to fund the build phase (AWS Activate, GCP Startup Program, and Azure for Startups are the common vehicles), and those credits made storage effectively free at the moment teams were loading datasets and standing up pipelines. So the data went in, at scale, with no cost discipline attached, because the meter read zero.

The credits expire. The data does not move. What was a subsidized ingestion becomes a full-price storage footprint spread across whichever providers were cheapest to adopt at the time. There is no clean migration path out, because the tenant never architected one. They optimized for free ingestion, not for portability.

By the time the tenant signs a GPU contract with you, the credits are gone and the storage bill is permanent. The hyperscaler funded the trap and now charges rent on it. The tenant feels the squeeze first as an egress line item every time they train.

Solving the Data Gravity Problem as Part of the GPU Contract

There are two versions of this relationship six months in.

In the first, storage was never part of the deal. The tenant's data is still on AWS, every training run still generates egress, and you are on a call explaining why your platform looks more expensive than the quote. The data gravity problem is now a retention risk, because the cheapest thing the tenant can do is move compute back to where the data already lives.

In the second, you treated zero egress as a contract differentiator and put storage in the deal at signature. The tenant's data sits with you, on flat-rate storage, reads do not carry a per-gigabyte egress fee, and the training bill is the training bill. There is nothing to explain. The hyperscaler is out of the loop, and the tenant has no economic reason to look backward.

The difference is not the quality of the compute. It is whether you closed the data gravity problem at signature or left it open for AWS to keep billing. And the advantage compounds beyond price. Once the data moves off the hyperscaler, every object on Akave carries cryptographic provenance: the content-addressed ledger records every operation, every version is independently verifiable, and the data sits where you define the region, the sovereignty boundary, the compliance boundary. Hyperscalers win on scale. Akave wins on control. That distinction is the one regulated tenants are beginning to put in the contract language, and it is the one a generic discount-storage pitch cannot match. Including storage in the initial contract is the move that changes the economics permanently, and it is cheaper for the tenant to adopt at signing than to retrofit during a painful migration later. You can point a prospect at the S3 migration path at docs.akave.xyz as part of the same conversation.

S3-Compatible, Flat-Rate, Zero-Egress Storage That Migrates Without Re-Architecture

The reason storage gets left out of GPU contracts is that operators assume adding it means building and running a storage business. It does not. Akave gives you a storage tier to attach to the compute deal without becoming a storage vendor, and it turns the data footprint into storage revenue the operator never captured rather than a check the hyperscaler cashes.

Akave Cloud is S3-compatible, which means it is a drop-in replacement for the tenant's existing S3 workflows. Their pipelines already speak S3. Pointing them at Akave is an endpoint and credential change, not a rewrite of the data layer.

The pricing is the weapon. Akave hot tier is $14.99/TB flat-rate, zero egress. The tenant's training runs read the dataset as many times as iteration requires, and reads do not carry a per-gigabyte egress fee. The bill is predictable because it is tied to what is stored, not to how many times the GPUs touch it. Against hyperscaler storage, that is up to 80% lower cost, and the egress line disappears entirely.

For operators who want the storage to carry their own brand, Akave Cloud offers a white-label option: tenant-facing under your name, running on isolated endpoints, with the infrastructure managed by Akave. Your customers see your platform. You do not run the storage business underneath it.

Whichever path you take, the economics change at signature, not six months later. Put flat-rate storage in the GPU contract and there is no AWS egress bill to explain and no hyperscaler collecting rent on data your GPUs keep reading. Run the numbers for a tenant at akave.com/akave-cloud-pricing.

FAQ

Why does a tenant's AWS bill keep growing after they move compute to a NeoCloud?

Because training reads the dataset repeatedly, and each cross-cloud read out of a hyperscaler bucket carries an egress charge. The compute moved; the data did not. Every training run pulls the data across the network again and the hyperscaler meters every crossing.

Why can't the tenant just move their data to the NeoCloud?

Most NeoClouds sell compute only, with no storage tier to consolidate onto. Hyperscalers sell hot, warm, and cold tiers, so the data stays where the tiers are. Without a storage offering in the contract, the data defaults to the hyperscaler.

How does flat-rate, zero-egress storage change the math?

Akave's flat-rate, zero-egress pricing means the tenant's cost is tied to what is stored, not to how many times the GPUs read it, so training iteration does not compound the bill. Storage cost can run up to 80% lower than hyperscaler pricing.

Does moving storage require re-architecting the tenant's pipelines?

No. Akave is S3-compatible, so existing S3 workflows keep working. The change is at the endpoint and credential level, not a rewrite of the data layer.

Can a NeoCloud offer this under its own brand?

Yes. Akave Cloud offers a white-label option: tenant-facing under the operator's brand, on isolated endpoints, with infrastructure managed by Akave.

References

  1. AWS S3 Pricing,  S3 Standard egress $0.09/GB (first 10TB), $0.085/GB (next 40TB), $0.07/GB (next 100TB), $0.05/GB (>150TB); US East

Modern Infra. Verifiable By Design.

Whether you're scaling your AI infrastructure, handling sensitive records, or modernizing your cloud stack, Akave Cloud is ready to plug in. It feels familiar, but works fundamentally better.