Case study
What was at stake
Bold sells payment terminals. That means the company does not only move money. It moves physical objects: card readers, SIM cards, cards and point-of-purchase material, from warehouses and logistics operators into a merchant’s hands.
The logistics domain covers two worlds.
Inventory is about making the numbers reconcile and keeping traceability over every terminal, card and piece of POP material. It also covers the guardrails on how inventory is allowed to move, the integrations with systems that touch it, and the reporting on top.
Distribution is about getting the product to the customer. We integrate with several third-party logistics operators, and here too we have to guarantee traceability and reporting across the whole process.
I work across both, and I own the subdomain that cuts through them: logistics issues.
The constraint
The happy path is the easy half. What makes this domain hard is that the plan regularly stops being true, and the system has to keep working anyway.
A concrete case. A shipping order is being loaded with a logistics operator, and it turns out that operator does not have availability for the full set of inventory to be dispatched. The shipment cannot go as planned.
Something has to decide what happens next. Route it through a different operator, split it, hold it. That decision cannot live in an operator’s head or a chat thread. It has to be a resolution strategy implemented in code, with the normal case, the edge cases and the ugly in-between cases all accounted for.
The other constraint is human. Logistics analysts are the people who have to understand what the system just decided, while it is still actionable. A resolution that is correct but invisible is not much use to them.
What I built
The work spans three layers of the same codebase: the Python backend, the AWS infrastructure that runs it, and the dbt models the business reads.
An aggregate for things going wrong
Logistics issues had no home. They lived scattered across handlers, with no lifecycle, no traceability and no defined way to be resolved.
I built the aggregate from scratch: model, commands, domain events, repository, storage and the indexes to query issues by status and by error type. On top of it sits a resolution strategy registered per error type, so a new class of incident means writing one strategy rather than threading a new branch through existing code.
That chassis is what everything since has been built on.
Resolution strategies
Each class of incident gets an explicit, coded path to resolution instead of an ad-hoc intervention. Most of the work is in enumerating reality honestly: what actually goes wrong, how often, what the legitimate responses are, and which of them can be taken automatically versus which need a person.
That list does not come from the codebase. It comes from the logistics analysts who have been absorbing these incidents by hand. So the first part of the job is sitting with them, watching what they actually do when a dispatch fails, and separating the steps that are genuine judgement from the ones they perform only because nothing else will. Only the second kind should become code.
Two examples that used to be manual and now are not:
An order already reassigned or cancelled, which the operator then reports as delivered. The system now detects the forbidden transition instead of logging it silently, and the resolution pre-validates inventory availability before opening an instant transfer to the receiver. The non-obvious part is that the origin of that transfer determines whether billing fires, so getting it wrong is not a cosmetic error.
An operator with no stock for the order. Rules differ per operator, the issue is discarded automatically if the order was cancelled meanwhile, and it closes itself when the underlying order resolves.
Automating the call somebody used to make
This was the largest piece of work, and the one with the clearest before and after.
When an operator reports a delivery incident, a wrong address or an absent customer, somebody on the team used to call the merchant, correct the data and retry the order by hand. That whole loop is now automated: an event consumer evaluates configurable escalation rules and decides whether to escalate, an assistant contacts the customer, and a callback applies the correction to the real shipping order. When the order is finally delivered or cancelled, the issue closes itself through dedicated event consumers.
Three details worth naming, because they are where the work actually was:
The law, not the meeting. The requirement discussed verbally was “call between 8 and 5”. Colombian Ley 2300 de 2023 actually says Monday to Friday 7:00 to 19:00, Saturdays 8:00 to 15:00, and prohibits contact entirely on Sundays and holidays rather than reducing it. I implemented the statute and reused the holiday calendar already in the codebase.
Counting attempts honestly. The number of delivery attempts comes from the operator, with a fallback to our own count for when an operator reports nothing useful.
Every outcome is a different message. Contact failed, customer declined, customer confirmed receipt, data corrected. Each produces its own notification and follow-up note, so the analysts can filter by what actually happened rather than reading a wall of identical alerts.
The whole flow was exercised end to end with the logistics team before it was turned on in production.
Never rewriting inventory history
A closed transfer with an error left stock in the wrong place, and the only fix was editing the database by hand.
I replaced that with a compensating transaction. Instead of deleting or mutating the original transfer, the system generates a complete inverse transfer, validates stock at the destination, and executes it through the same closing path any normal transfer takes, then marks the original as void.
In a system with event sourcing and physical inventory, this is the only correct shape. You never rewrite history; you write the correction that cancels it.
One inventory view instead of several
Historical inventory data existed, but not in a shape anyone could ask business questions of. I migrated it into a dbt gold dimension, consolidating 446K+ terminals and 1.2M+ stock movements into a single view the business can actually read, and built the issue dimension as a full medallion slice from staging through to gold, including PII tagging for data governance.
This is the unglamorous half of traceability. Knowing where a terminal is only counts if someone outside the engineering team can find out without asking.
Opening Peru
Adding a country to a logistics domain is not a configuration change, because countries do not agree on what an address is. Colombia addresses resolve in two administrative levels; Peru needs three, since a district sits under a province which sits under a department. Any code that assumes a fixed depth does not throw when it meets the deeper hierarchy. It simply stops matching any distribution rule, which is the worst kind of failure: no error, no alert, and orders that quietly route nowhere.
I built an address-type module to model hierarchies of varying depth, made the distribution rules engine resolve operator, warehouse and delivery type per country, and parameterised the configuration lists by country code. The lakehouse and both operator adapters had to follow.
Talking to the outside
Two external systems now create and query logistics orders through machine-to-machine APIs, authorised by scope. Building those meant a combined authorizer so a single route serves both back-office users and M2M clients, a new index for lookups, and some deliberate contract decisions: the endpoint that takes a national ID neither caches nor logs it, unlike its sibling, because it handles personal data.
I also replaced a synchronous cross-domain HTTP call with an integration event consumed by a dedicated strategy. One less way for another team’s outage to become ours.
What came of it
Logistics issues stopped being interruptions handled by whoever noticed first. They are now a modelled part of the domain, with defined strategies, visibility while they are still actionable, and a trail afterwards. The contact loop that used to start with a person picking up the phone now starts, and often finishes, without one.
The general lesson is one I keep relearning. In any system that touches the physical world, the exception path is the product. The happy path is a small, well-behaved subset of what actually happens, and how gracefully you handle the rest is most of what people experience.