About Flight Recorder, and how it was built
Why this exists
When a sales system decides which accounts deserve attention, the final score is easy to see. It is much harder to explain why that score happened, especially after the account data and the scoring rules have changed. The record of what the rules could see, what they used and what they left out is usually gone by the time anyone asks.
Flight Recorder records the evidence and the rules behind an automated account-prioritization decision, so the person responsible for that workflow can inspect what happened and ask whether different rules would have produced a different result, without the present rewriting the past.
What I did
This was an AI-assisted build. I worked with AI tools for implementation and review. My contribution was directing the product: choosing the problem, setting boundaries, making product decisions, and deciding what evidence was needed before accepting the work.
One concrete example was how to handle a scoring revision. I approved introducing a new version while preserving the earlier rules and historical records. That lets decisions be compared without quietly changing the past.
I also kept the scope focused on account prioritization as the one fully replayable decision type. That gave the project a specific claim that could be built and checked.
How it works
- One collector. Every event, including the demo's own seed data, enters through one versioned, idempotent door. There is no other write path.
- An append-only ledger. Recorded decisions, their preserved context and their logic identity are never edited. Corrections, outcomes and new rules are appended beside them.
- A deterministic evaluator. Scoring rules are declarative data identified by a content hash plus an evaluator version, so a decision can be re-verified exactly. A version label alone is not identity.
- On-demand replay. A replay re-verifies the original decision, then evaluates the same preserved context under a selected rule version. The result is labelled a counterfactual and kept nowhere.
- Pages that only read. The public demo serves a fixed synthetic snapshot read-only: mutating requests are refused before routing, and the database itself is opened read-only.
Four decisions and their tradeoffs
- A new rule version instead of editing old rules
- Older decisions stay explainable under the rules that made them, and any two versions can be compared on the same evidence. The cost is that rules accumulate and each version must be registered and identified rather than patched in place.
- Replay is never stored
- A counterfactual can never be mistaken for something that happened, and no cache can go stale. The cost is that every replay is computed again on each request.
- One fully replayable decision type
- Account prioritization is proven end to end: record, preserve, replay, compare. The cost is breadth: no other decision class is replayable here.
- A read-only public demo from a fixed synthetic snapshot
- The demo is stable for every visitor, needs no accounts and cannot be changed by anything a visitor sends. The cost is that it shows one dataset, and only the local installation accepts new events.
Limitations
- All data is synthetic. RelayBridge and every prospect are fictional, and no real customer, vendor connection or result is behind any page.
- One decision type. Account prioritization is the only fully replayable decision class.
- No integrations. Events resemble vendor-like sources but nothing is connected to a live system.
- The Insights comparisons are descriptive: observed rates, sample sizes and windows on synthetic data. They are not causal, they do not measure lift and they validate no scoring weight for a real business.
Run it yourself
The source is public. Its README holds the local setup: a few commands seed the same synthetic dataset through the collector and serve the same pages, with the writable collector available locally.
Attribution
Created by Elias Skora · Consilience Operations House LLC