Data Migrations

The Legacy Billing Job Ran One Last Time. The New Pipeline Ran the Same Day.

The Legacy Billing Job Ran One Last Time. The New Pipeline Ran the Same Day.

What a parallel run proves that a runtime number can't.

By Dean Cirielli, VP Engineering, Blue Orange Digital

On July 1 a healthcare claims company we work with ran its monthly billing on the old system for the last time. The same morning, the same billing cycle ran on the Databricks lakehouse we'd built, in UAT, against the same production data. Two jobs, one day, one set of numbers to compare. The client's finance team had the legacy output open in one window and ours in the other. Nobody on our side got to say "done" until the two agreed.

That morning is the part of the engagement I'd put in front of an operating partner. It's not the part that made it into the board deck.

the number everyone quotes

The figure that travels from this engagement is the pipeline runtime. The daily load used to take roughly 24 hours. On June 9 the engineer who owns the reload reported the first full production run at 1 hour 42 minutes. That is one run, not an average across a month, and it went into the client's board materials the following week with that caveat attached.

It's a good number. It's also the least interesting thing that happened.

Fast and wrong is still wrong.

A pipeline that finishes in under two hours and produces different figures from the one it replaced is a faster way to be wrong. So the question we ask before we quote a runtime is whether anyone has yet compared the new output to the old on a cycle that mattered. We've started calling that the parallel-run bar: a data migration is finished when the new system reproduces the old system's output on a real business cycle, on the same day, with the people who own that output watching. Until that date exists, the runtime is a promise.

what the client was actually buying

The company processes dental insurance claims. When a dental office submits a claim, this is the layer that validates it, routes it to the payer, and tracks it. Speed and accuracy at the transaction level are the product, so the data team's problems were felt directly by the finance team.

There were three. The daily pipeline ran for about 24 hours, so by the time an analyst opened a dashboard the numbers were already a day old. Every new report was a manual build, and the backlog of requests grew faster than anyone could clear it. And years of business logic lived inside SSIS jobs. SSIS is the older SQL Server tool for moving and transforming data, and its logic sits inside the job configuration rather than in code anyone can read, diff, or test. Nobody could say with confidence what some of those jobs did.

Finance wanted to close the flash revenue report in under five business days. Billing close was running eight. Neither target was reachable on the architecture they had, and both were the kind of number a sponsor's board asks about every quarter.

what we built, and what took longer than scoped

We built a silver layer in Databricks. Silver, in lakehouse terms, is the set of cleaned and joined tables analysts should query instead of raw source data. By late June more than 500 tables had been promoted into UAT, which is the environment the client's own team tests in before anything reaches production. The SSIS jobs were translated into Databricks notebooks, so the logic that used to live in a job configuration now lives in versioned code with tests around it.

Here is the honest part. In mid-May the silver tables were about 85 percent complete in development, and we told the client UAT promotion was two weeks out. It took longer than that. The backlog workstream started a week late because the client's engineers had a release of their own to ship first, and we moved the milestone dates rather than pretend otherwise. The engineer we brought in for the SSIS translation pushed back on the scope in his first week, and he was right to: some of the jobs were doing two things that should have been separated, and we re-cut the plan around that.

None of those are failures. They're what a migration on a live system looks like, and a status report that doesn't have items like this in it is hiding something.

the parallel run is the deliverable

A parallel run costs money in a way a runtime improvement doesn't. You pay to operate two systems for a full cycle. The engineers who know the reload can't rotate off in the middle of it, which constrains staffing on every other project they touch. In this case UAT cycles ran through most of July, and the production reload was scheduled for the week after the parallel run cleared.

We'd already had one smaller version of that moment in April. The client's VP of operations ran queries from the new environment during a live finance executive meeting and posted in the project channel that everything matched, except changes made in the last few minutes. That exception is the part I'd keep. A daily pipeline that finishes before the business day starts is not a real-time pipeline, and a finance team has to know which one it's looking at. We wrote that limit into the reporting layer rather than hope nobody noticed.

The steelman for leading with the runtime is real. It's the one figure a board can verify from a screenshot, and it's what got the engagement into the board deck. But the reason it matters is decision velocity, the measure we use instead of headcount eliminated: how many decisions get made faster, and how much throughput grows without the cost structure growing with it. The target for the billing close is five business days, down from eight. As of the July run that was a target rather than a measurement, and I'd rather report it that way than round it up. If it holds, the finance team makes its cash calls three business days earlier every month, and that shows up in the bridge in a way an hour-and-42-minute runtime never will on its own.

the number that lets you stop the old job

A migration status report that carries a runtime and no parallel-run date reports the promise and leaves out the proof. The runtime says the new system is fast. The date says the finance team trusted it enough to close the month on it, and that's the only version of done that survives an audit.

Speed is the number that fits on a slide. Agreement is the number that lets the finance team stop running the old job. If one of your companies is mid-migration and nobody can name the parallel-run date, that's the conversation to have: https://blueorange.digital/contact-us/?utm_source=linkedin&utm_medium=social&utm_campaign=field-note&utm_content=parallel-run-bar

Technology
Industry

BOD Newsletter

Stay ahead of the AI × Data × PE curve.

Practical field notes for operators and investors — join the BOD newsletter.

Ready to build?

Turn these insights into production systems.

Blue Orange builds data and AI systems that ship to production and tie back to EBITDA. Let's scope your opportunity.

Start a Conversation