Adding a fourth warehouse: what it takes to bring a new site onto the central copy
Bringing a new warehouse onto a central copy of your data is real work done by people. Here is the honest setup list, and why each site after the first is easier.
Your company runs three warehouses, each with its own database, and head office reads one central copy of all three. Now a fourth building is opening, or one you acquired needs to join. The question from the finance side is usually simple: how much work is this, and who does it?
The honest answer is that a new site is a small project, not a switch. The good news is that most of the hard thinking was done when the first site went live.
What a central copy is, briefly
A central copy is one store, in the cloud, that holds the current state of the tables you chose from every warehouse. A small service runs inside each warehouse network, beside the warehouse database. It sends changes outward only; nothing reaches back in.
Every row that arrives is stamped with the site it came from, and each site keeps its own space in the store. Your reports and portal read from that one place instead of from each building. That keeps reporting load off the onsite servers, which leaves them headroom during peaks.
The setup list for a new site
Here is what bringing a new warehouse on involves, in the order it usually happens. We set it up together with a contact at the site, and a person drives every step.
- Name one technical contact at the new site. This person gets us access to a machine and the database.
- Find a home for the small service. It needs a server, a virtual machine or a PC inside the warehouse network that can reach the warehouse database and the outbound internet.
- Confirm the network rules. The service only makes outbound connections, so no port is opened to the internet.
- Give the service access to the database. It watches only the tables you choose.
- Agree the list of tables and their ranking. Important tables go first, so they stay current even on a slow link. Large backlogs trickle in behind them.
- Agree any filters that decide which rows matter most.
- Send the first full copy. Each table is sent once, table by table. After that, only rows that were added, changed or deleted travel.
- Watch it settle. The site reports its health, so a stalled site is noticed rather than discovered in a report.
None of this replaces anything at the warehouse. Its system and its staff work as before.
Why the first site is the hard one
The first site carries decisions that apply to every site after it. That is where most of the time goes, and it is time you only spend once.
- The central store and the service that receives data in the cloud are built and running.
- The architecture review, which is part of the design-partner contract, has already covered how your sites look and what the first deployment includes.
- Your team has already decided which tables matter and how to rank them.
- Your reports and portal already read the central copy and already understand the site stamp on each row.
When the fourth warehouse runs the same system as the other three, the table list and ranking can usually be reused as a starting point. Someone still checks that the new site matches, because sites drift. A building that was set up years apart from the others may run a different version, with tables that look the same but are not.
The first site is where you make decisions; every site after it is where you reuse them.
What still takes care at the fourth site
Easier is not the same as nothing to do. A few things are specific to each site and cannot be copied from the last one.
Its own key
Changes are queued at the site, compressed and encrypted with a key unique to that site. Every site has its own key, so the new warehouse gets a new one.
Its own link
Every warehouse has a different internet connection, and some are poor. If the link drops, changes wait on the site's disk and resume from where they stopped, table by table. Delivery is at-least-once, which means nothing is lost, though a change can occasionally be sent more than once. A poor link means the first copy takes longer to clear.
Its own database load
The service uses dynamic throttling: it eases off when the warehouse database is busy, so the operation comes first. If the warehouse runs SQL Server, changes are found using a built-in feature called change tracking, read on a short interval. That keeps reads light, but the new site's busy hours are still worth checking before the first full copy runs.
What changes for the people reading the reports
Once the new site is flowing, its rows sit in the central store beside the other three, each stamped with where it came from. The central copy is current to within a couple of minutes. Your team decides how the new warehouse shows up in reports and the portal; we supply the data and advise, and you build what your people see.
Adding tables later works the same way. It is a setting, but someone still has to choose the tables, rank them and check that the reports use them correctly.
redfly fits here as the part that keeps each warehouse's data flowing into one central copy, while your team owns the reports and the decisions about what to show.