Data platforms and modernisation

The old platform is still running, the new one is half built, and reporting is frozen

We decide whether to modernise at all, and say so when the answer is no. When it is yes, the platform uses open table formats in your own cloud account, old and new run in parallel with reconciliation between them, and the decommissioning date is in the plan from week one rather than the last. Freezes are short and planned.

Companies delivered for
20+
Median to first production-grade artefact
6 weeks
Decommissioning date agreed
Week 1, not month nine
Team on a platform move
Pod of three
Published case study for this pillar
None yet

Six reasons a platform conversation starts, and none of them is a technology preference

Nobody wakes up wanting a new data platform. Something else has gone wrong first, and it is rarely the technology itself.

  • The platform was right for the company that chose it

    It was chosen for the business you were then. The volumes, the number of source systems and the people depending on the output have all changed since. The platform has not, and neither has the decision that picked it.

  • The licence renewal has arrived and nobody can justify it

    Procurement wants a business case by a date. The only people who could write one are the people keeping the platform working next Tuesday, and they have no way of saying which workloads are driving the invoice.

  • The old system was never switched off

    The migration finished, or was declared finished. Both bills still arrive and nobody can say with confidence which reports still read from the old one, which is exactly why nobody will turn it off.

  • Every quote begins with a freeze on new reporting

    The business is asked to stop asking questions until the migration lands. It will not. A shadow reporting layer grows in spreadsheets during the freeze, and that layer is still there long after the cutover.

  • A previous attempt stalled halfway

    Half the workloads moved, half did not, and the sponsor who knew the plan has moved on. You now carry the operational burden of two platforms and the benefit of neither.

  • The bill grows with the data, not with what the data is worth

    Ingesting more raises the invoice whether or not anyone reads the result. A dataset nobody has opened costs the same to hold as the one the board reads before every meeting.

Most of these are commercial or operational rather than technical, which is why what they add up to is a platform decision rather than a reporting one.

The first question is whether to modernise at all

Four answers come out of the audit week, and only two of them involve a migration. The first one is a real answer here, not a polite one.

Answer 1Do nothing

The platform still fits and the pain is somewhere else

A slow dashboard and a late month end can both be modelling or query problems. A new platform carries those across intact, at considerable expense. If that is what the audit week finds, we say so, and you have spent a week rather than a year.

Answer 2Renegotiate

The problem is the contract, not the technology

Sometimes the platform is fine and the commercial terms are not. Knowing exactly which workloads drive the bill, and which of them could credibly run elsewhere, is what gives you something to say in the renewal conversation.

Answer 3Move part

One or two workloads leave and the rest stays

Sometimes only the heaviest workloads justify moving at all. The rest can stay where it is for years without anyone being worse off. Partial is a legitimate end state, not a migration that failed to finish.

Answer 4Move it

The platform has to change, and here is the order

When the answer is a full move, the plan names the first workload, the reconciliation that will prove it, and the date the old platform is switched off. All three are written in week one, or it is not a plan.

From the sentence you said on the call to the decision it actually is

Each complaint belongs to a different decision, and the order they get answered in matters more than any single answer.

  1. The platform was right when we chose itModernise or not, decided on evidence
  2. The licence renewal has arrivedCommercial and licence position
  3. Getting our data out means a rewriteStorage format and portability
  4. Every quote begins with a reporting freezeParallel run and cutover
  5. The old system was never switched offDecommissioning plan
  6. The bill grows with the dataCost model
Left, the sentence somebody says. Right, the decision it belongs to. They arrive tangled together, and a contract renewal date is what forces the order they get answered in.

What we have actually delivered

We have not published a named platform modernisation case yet, so these are firm-wide delivery figures, and we will say on the call which of them came from work like yours.

20+Companies delivered forAcross India, APAC and the US, from Pune, since 2023.
50+Projects deliveredAcross data engineering, analytics, applications and AI. Not all of them were platform moves.
6 weeksMedian to first production artefactNot a prototype and not a slide. Something running that your team can break and we fix.
Week 3Something works byOn a platform move that means the first workload running on the new platform beside the old one, with the comparison already visible.
ThreePeople in the podNo bench and no rotating juniors. The same three people from the audit week through to handover.

What the work is made of

Six parts. The first one can end the engagement, and the last one is the part that gets left until there is no budget left for it.

Decision01

Whether to modernise, argued from your own workloads

We profile what actually runs, what it costs and who reads the output, then say whether a move is justified. The evidence is your query history and your invoice, not a comparison table between two vendors.

Format02

Open table formats, so the next decision stays yours

Data lands in an open table format rather than a proprietary one, so the storage layer can be read by more than one engine. The point is not fashion. It is that leaving again later costs a configuration change rather than another migration.

Deployment03

Your cloud account, not a vendor's

Compute and storage run inside the customer's own cloud subscription wherever the workload allows it. Your security team already governs that account, the cost arrives on a bill your finance team can already decompose, and the data does not quietly become somebody else's asset.

Parallel run04

Old and new running side by side, reconciled

The new platform processes the same inputs as the old one for an agreed period and the outputs are compared row by row rather than sampled. Nobody is asked to trust the new numbers before the comparison has become boring.

Decommission05

Switching the old platform off, on a date

Named date, named owner, written down in week one. The old platform goes read-only, then archived, then off. Until it is off, the saving is a forecast rather than a saving.

Cost06

A bill that tracks value rather than volume

Usage and cost put side by side per dataset, and workloads priced by what they actually need rather than all held at the same service level. The aim is that holding more data does not automatically cost more money, and that switching something off is an ordinary decision.

Three ways a modernisation gets attempted

Only the third one has a date on which the old bill stops.

  1. Approach 1

    Big bang, with a freeze in front of it

    Everything moves on one weekend, and new reporting is frozen for the months before it. The business does not actually stop asking questions during a freeze, so a shadow layer of spreadsheets grows and then survives the cutover. The platform has been modernised and the reporting has moved into Excel.

  2. Approach 2

    Halfway, then stalled

    The straightforward workloads move, the awkward ones do not, and then the sponsor changes. Both platforms stay on, both bills arrive, and the team supports two of everything. Nothing was decommissioned, so nothing was saved.

  3. What we do

    Parallel, reconciled, with a decommissioning date

    One workload at a time runs on both platforms, the outputs are compared, and the old path is switched off once the comparison has stopped being interesting. The freeze is measured in the hours around a single cutover rather than the months around a programme.

How the parallel run actually works

Five stages per workload, repeated. Nothing is switched off until the stage before it has been uneventful for a while.

  1. 01

    Both platforms read the same sources

    The new platform ingests from the original source systems, not from the old platform's output. Chaining the new one behind the old inherits every quirk you are trying to leave and hides them until the day you disconnect.

  2. 02

    Outputs are compared, not sampled

    Counts, keys and money by period, old against new, on a schedule. Differences are listed individually. A percentage match is not a reconciliation, it is a way of not looking at the rows that disagree.

  3. 03

    Reads move first

    Reporting and downstream consumers are pointed at the new platform while the old one keeps running untouched. If something is wrong, pointing them back is a configuration change rather than an incident.

  4. 04

    Writes move, and the old path goes read-only

    The old platform stops receiving new data and stays readable. This is the freeze, and it is per workload rather than across the whole programme, which is the entire reason it can be measured in hours.

  5. 05

    The old path is removed, not left dormant

    Dormant pipelines get switched back on by somebody solving a problem at seven in the morning, and then two platforms are live again. The path is deleted and the schedule is removed with it.

Repeated per workload rather than run once across the estate. Only one workload is ever mid-cutover, which is what keeps the freeze short and the rollback cheap.

What the reconciliation has to prove before a workload switches

A workload has not moved when the data lands on the new platform. It has moved when the person answerable for a number is willing to quote the new one.

  • The comparison runs on every load, not once before the switch

    A reconciliation run once, the week before a switch, proves the state of the data that week and nothing else. Running it on every load is what turns the comparison into something nobody is nervous about.

  • Financial figures tie to the finance system, not to a shared extract

    If both platforms read the same intermediate file, they will agree with each other and can both be wrong. The check has to run against the system the finance team treats as the record.

  • Late and restated data behaves the same on both

    The awkward case is never the happy path. It is the record that arrives three days late and changes a period that was already closed. If the two platforms handle that differently, the reconciliation drifts quietly and you find out at quarter end.

  • Rounding, time zones and currency are pinned before the comparison starts

    Variance from those three looks like a data problem and is not. Settling them first keeps the argument on the records that genuinely disagree rather than on the fourth decimal place.

  • Remaining differences are explained and signed, not tuned away

    Some differences are correct, because the old platform was wrong. Those get written down and agreed by name, rather than adjusted until the two numbers match and nobody knows why.

  • Rollback is a configuration change

    Until the old path is removed, going back has to be a switch somebody can throw without a deployment. If rollback needs a release, nobody will agree to switch in the first place.

We write this list against your own workloads during the Week 1 audit, and you keep the one-pager whether or not the rest of the work goes ahead.

Decommissioning, the step that gets planned last and then dropped

Until the old platform is off, you have added a platform rather than replaced one. This is the part we insist on writing down in week one, because it is the hardest thing to fund once the new platform is already working and the programme is out of money.

  1. 01

    Name the date and the owner in week one

    Before any data moves, the plan says when the old platform goes off and who signs that it has. A decommissioning step with no name against it is a wish, and it is the first thing to fall out of the plan when the timeline tightens.

  2. 02

    Inventory what still reads from it

    Scheduled jobs, embedded connections, an integration somebody built for a customer, and a spreadsheet on a desktop refreshing over an old connection. The last category is the one that stops a shutdown, and it is only found by watching what actually connects rather than by asking.

  3. 03

    Make it read-only and see who complains

    A planned read-only period surfaces the consumers no inventory found. A complaint during that window is useful information. The same complaint after a deletion is an incident.

  4. 04

    Archive what has to be kept, then prove you can read it

    Retention obligations do not lapse with the licence. Archive into a format you can still open without the vendor, and restore something from the archive before the contract ends rather than after.

  5. 05

    Switch it off, then cancel the contract

    The technical shutdown and the commercial cancellation are two tasks with two different owners, and the second is the one that gets forgotten. A platform that is off but still invoicing has saved you nothing.

Why the bill grows with the data rather than with the value

The business case is built on cost, so it is worth being precise about where the money actually goes.

The pricing model is usually the design

Platform bills scale with what you ingest, what you store and what you scan. None of those three has any relationship to whether the output is read by anyone. The platform has no reason to tell you which datasets are worth their cost, so nobody finds out, and the invoice grows on its own.

Where the money actually goes

History retained forever because deleting it needed a decision nobody wanted to make. Non-production copies of production data that nobody ever switched off. Refresh schedules running faster than the source systems update. Full rebuilds where an incremental load would do. None of this is exotic, and all of it is measurable on your own workloads before anything is changed.

What a modernisation has to change about it

Put the usage of each dataset next to its cost, so switching one off becomes an ordinary decision rather than a project nobody sponsors. Separate the workloads that have to be fast from the ones that only have to finish before morning, and price them differently. If the new platform cannot tell you what a dataset costs, it has not fixed the thing you are paying it to fix.

What the first six weeks look like on a platform move

The same three phases as every engagement, and the interesting question is what has to be true by week three.

  1. Week 1

    The audit, fixed fee

    Two calls, your query history and your invoice, and one workload followed from the source system to the number somebody uses. You get a one-pager naming whether to move at all, what moves first, and when the old platform goes off. Yours whether or not we go further.

  2. Week 2

    The first workload is chosen

    A pod of three starts on the workload that carries real decisions and has the fewest dependencies, not the one that demonstrates best. The demonstration workload teaches you nothing about whether the reconciliation will hold.

  3. Week 3

    Both platforms are running it

    The first workload runs on the new platform beside the old one, with the comparison visible to your team. This is the week the awkward records surface, which is exactly why it is not week nine.

  4. Weeks 4 to 6

    Reads move, then writes

    Consumers are pointed at the new platform, the old path goes read-only, and the workload is done. Across our engagements the median time to a first production-grade artefact is six weeks.

  5. Week 7 onwards

    The rest of the estate, and the shutdown

    Quarterly reviews and on-call governance, the remaining workloads in the same pattern, and the decommissioning plan executed against the date agreed in week one.

Where week three falls in a six-week move

The point of week three is not speed. It is that the awkward records surface while changing the plan is still cheap.

Day 1, the fixed-fee auditDay 42, six weeks in
Six weeks drawn as 42 days. Six weeks is our median time to a first production-grade artefact, measured across all of our engagements. On a platform move, week three is the first parallel run rather than the cutover.

Questions platform owners and finance sponsors ask us

What buyers ask before anything is signed.

How long is the reporting freeze, honestly?

Per workload, hours rather than months, because only one workload is ever mid-cutover. What we will not promise is zero. There is a window where the old path goes read-only and the new one takes writes, and it is planned rather than discovered.

The long freezes are not a technical necessity. They come from having no reconciliation, so nobody will commit to a switch, and the only safe answer left is to stop everything.

What if the answer is that we should not modernise?

Then we write that in the week one one-pager, and you have spent a week rather than a year. It is a real outcome of the week, not a disclaimer. The platform is sometimes fine and the money is going somewhere else.

We already have engineers. Why bring in a pod?

Because your engineers know the exceptions and are already fully committed. A parallel run is mostly reconciliation work with a deadline attached, and it is the first thing to be dropped when production has a bad week. We take that part, alongside your team rather than around them.

Our last attempt stalled halfway and we now run both. Where do you start?

With the inventory of what still reads from the old platform, and a date. What is missing at that point is rarely engineering. It is the list of consumers nobody wants to own and a shutdown date somebody will sign.

Do we have to move everything?

No, and often you should not. Move the workloads where the cost or the constraint is real, leave the rest, and be explicit that leaving it is a decision that was made rather than a job that was never finished. Partial is a legitimate end state.

Why open table formats? Our vendor's own format is faster.

It may well be, today. The question a format answers is what leaving costs next time. An open table format keeps the storage layer readable by more than one engine, so the next decision gets made on merit rather than on exit cost.

Why run it in our own cloud account rather than a vendor's?

Because your security team already governs that account, and because the data stays something you control directly. It also means the spend arrives on the cloud bill your finance team can already decompose, instead of as a single line nobody can break down.

You have no published case study for this. Why should we believe you?

Because we told you, rather than showing you a result from another discipline and hoping you did not check. The published figures on this site belong to our analytics, data engineering and AI work. What we offer here instead is week one: two calls, a fixed fee, and a one-pager you keep either way.

Who owns what you build, and can our team run it afterwards?

You own it from the first commit, in your repositories and your cloud account. The reconciliation harness and the decommissioning plan are handed over with it. The test we hold ourselves to is whether your team can move the next workload without us in the room.

What happens after the first workload?

Week seven onwards is quarterly reviews and on-call governance, the remaining workloads in the same pattern, and the shutdown executed against the date agreed in week one. It is a service you can stop at any point.

Start with the Week 1 audit

Two calls, a fixed fee, and a one-pager saying whether to modernise at all, which workload moves first, what the reconciliation has to prove, and the date the old platform goes off. You keep it either way, including when the answer is to do nothing. NDA-friendly, fixed scope. Write to hello@woodfrog.tech.