← All insightsRemote sites

Why a vendor should not build site-to-cloud sync twice

Vendors with technology at customer sites often rebuild data collection for each deployment. Here is what the second and third builds really cost, and what a supported connection changes.

4 min read

If your product runs at customer sites, with its own database in each warehouse, plant or depot, sooner or later someone asks for the data in one place. A customer with six warehouses wants one picture instead of six.

The first answer is almost always a script. Someone on your team writes a job that exports a few tables, pushes them to the cloud and runs on a schedule. It works, and it ships quickly, which is exactly why it gets copied.

The first build is not the expensive one

The first collection job is cheap because it is built for one site, one link and one version of your product. It assumes the internet stays up during the export. It assumes nobody deletes rows.

None of those assumptions survive the second customer. The second site runs an older version of your database layout. Its link drops for an hour most afternoons. So the job is copied, patched and tested again, and now there are two jobs that look alike and behave differently.

The cost is not in moving the data once; it is in keeping it moving at every site, every day.

What the second and third builds really cost

By the third deployment the pattern is clear, and so is the bill, even if it never appears as a line item.

  • Engineering time. Collection jobs, code that retries failed sends, export monitoring and per-site fixes eat hours that were meant for your product.
  • Manual recovery. A link drops, an export fails halfway, and someone runs it again by hand, often after a customer notices a gap.
  • Stale views. A dashboard that is a day behind is one nobody trusts, so people go back to phoning the site.
  • Repeated testing. Every new site means the same collection job, rebuilt and tested again against a slightly different environment.
  • Risk at the site. Anything that competes with the operational database during a peak tends to get switched off by the site team, and then the data stops.

Together, these turn a reporting feature into a permanent maintenance job that nobody planned to staff.

Why the hard parts keep coming back

Moving rows from one database to another is not difficult. Moving them correctly, from many sites, over links you do not control, is. Each new script rediscovers the same hard cases.

Edited and deleted rows are the first trap. A job that only copies new rows slowly drifts from the truth, because the central copy still shows what the site has already changed or removed. Offline periods are the second. If a site is cut off for a day, the job has to know where each table stopped and resume from there without losing anything.

Silent failure is the third. A site that stops sending looks, from the centre, much like a site that has nothing to send. Fixing these once is a project. Fixing them in a slightly different way for every customer is where the time goes.

What a supported connection changes

The alternative is to treat site-to-cloud sync as one supported service rather than a script per deployment. This is the approach we take with redfly. A small service runs inside each customer site's network, beside the site database, and sends data outward only. Nothing reaches back into the site.

Your team chooses the tables and ranks them, so the important ones stay current even on a slow link. After a first full copy, only rows that were added, changed or deleted travel. For SQL Server, changes are read from its built-in change tracking feature on a short interval; for MongoDB and PostgreSQL, the database's own change features stream them.

Changes are queued on the site's disk, compressed and encrypted with a key unique to that site, then sent on. If the link drops, changes wait at the site and resume table by table from where they stopped. Delivery is at-least-once, which means nothing is lost, though a change may occasionally arrive more than once.

Dynamic throttling, which holds the service back when the site is busy, protects the performance of the operational database under load.

What lands in your cloud

Every site feeds one central store. Every row is stamped with the site it came from, and each site is kept in its own space. The store is low-cost document storage, far cheaper than a cloud relational database, and the central copy is current to within a couple of minutes.

Your application reads from that one place. Your team still builds the dashboards, support tools and analytics your customers see; we keep the data flowing to them. Because those reports run against the central copy, the onsite database carries no reporting load, which leaves headroom on the site servers so they stay resilient during peaks.

Being honest about the limits

A supported connection removes the rebuild, not the work of rolling out. Each new site still needs its own setup, driven by a person, and adding tables later is a setting someone changes, not a switch that flips itself. The queue at the site is a holding area, not a copy that site applications can read while cut off.

What changes is that the second and third sites become a rollout rather than a rebuild. The fixes for deleted rows, offline periods and busy sites are made once and carried to every deployment.

redfly is built for vendors whose technology runs at customer sites and who need data from every deployment in one place. We work with design partners, starting with one site.

Ready when you are

Stop reading. Start shipping.

Work with us as a design partner and see the difference on your own database.

redfly · Design-partner contract · Licensed directly from redfly