Context: five engines, one answer
EvalyMe’s core computation turns founder metrics — MRR, churn, growth, margins — into a valuation range and a set of scores. That computation exists in five places:
- the iOS app’s valuation engine, in Swift
- the Android app’s model layer, in Kotlin
- the backend’s domain module, in TypeScript
- the authenticated web app’s engine, in TypeScript
- the public calculator on the landing site, in TypeScript, running unauthenticated in a stranger’s browser
Five implementations, mostly written by AI agents from English descriptions. This is a normal shape for a modern product — and it is also the exact shape in which correctness dies quietly, because five competent implementations of the same prose will disagree at the edges.
Problem: drift is invisible until a user sees it
Nobody notices a 4% valuation difference between platforms on day one. It surfaces months later, in the worst possible way: a founder opens a report on their phone that they first created on the web, the numbers differ, and every number you have ever shown them becomes suspect.
The failure is structural, not a matter of care. Each agent implementation is locally plausible. Rounded division, clamping order, null handling for optional inputs, tie-breaking in scoring bands — each implementation makes its own choices, and each choice is defensible in isolation. The only artifact that can catch drift is one that all five implementations must reproduce byte-for-byte.
What failed first: the obvious approaches
Failed approach one: fix the formula where the bug was found. The instinct is to patch the engine that produced the wrong number. But with five engines, you now have four you did not check, and users hold reports generated by all of them.
Failed approach two: per-platform unit tests. Each codebase grows its own fixtures, written by agents, encoding whatever that implementation currently does. Tests that assert current behavior rather than shared behavior will all pass while the platforms disagree.
Failed approach three: recalculating history. When a formula changes, the tempting move is re-running old saved reports through the new engine so everything is consistent. This is a data-integrity mistake: a user’s report from March is a fact about what the product said in March. Silently changing it retroactively — even to something more correct — destroys trust more than inconsistency ever did.
What worked: one contract, one fixture, stamped versions
The working design has three parts, and the ordering matters.
First, a canonical schema in a shared contracts/ directory at the repository root — outside any single platform — defines the input and output shapes for an assessment. It is the only place a field can be added once and consumed five times.
Second, a single golden fixture: one JSON file containing inputs and expected outputs, deliberately including edge cases — boundary values, optional-field omissions, extreme ratios. Every implementation has a contract test whose only job is to load that shared file and assert identical results. In our setup the suite runs in the iOS test target, the Android unit tests, and the three Vitest suites, all reading the same bytes from the same file. If any platform diverges, its build fails before review.
Third, explicit versioning. Every persisted result carries a schemaVersion and an engineVersion. Old rows written before versioning existed are read as a historical mobile-v1 snapshot and displayed as-is, never recalculated. Rows imported from a previous product generation keep their own label. A formula change means a new engine version, new golden expectations, and both old and new results coexisting honestly.
The result is that “change the formula” stopped being a coding task and became a contract task: update the schema if needed, update the golden fixture, then let five implementations fail their contract tests until they all agree. The agents do the mechanical work of making five codebases pass one file — which is exactly the kind of work agents are best at.
Why this generalizes beyond formulas
The pattern — a shared, versioned artifact that every implementation must reproduce — is the same idea driving spec-driven development tooling like GitHub’s Spec Kit: move the source of truth out of prose and into something checkable. Golden vectors are the numerical special case, and they are underused because they feel redundant when all implementations currently agree. Their value is entirely in time: they are the tripwire for the change next month, made by an agent that never saw the other four codebases.
The same structure fits anywhere duplication is structural rather than accidental: serializers and parsers on both sides of an API, validation logic shared between client and server, pricing rules mirrored in a webhook handler, date and rounding behavior across platforms.
Edge cases worth planning for
- Historical rows without versions. Decide explicitly what unversioned data means, write the decision down, and never let an agent “upgrade” it silently.
- Imported data from previous products. Label it with its own engine version; it is a fact about the past, not input for the present.
- Floating-point and rounding. Put the contentious cases in the fixture, not in a code comment. If two languages round differently at a boundary, the fixture is where you find out.
- The free public calculator. Our landing-site engine runs without authentication and without a server. It shares the same contract, which means a stranger’s browser and the paid app agree by construction, not by hope.
Reuse checklist
If the same logic exists in more than one codebase:
- Extract the shapes into a shared, platform-neutral contract file.
- Write one golden fixture with normal cases plus the edge cases each platform would plausibly get wrong.
- Add a contract test per implementation that reads the shared file — no local copies, no re-typed expectations.
- Stamp every persisted result with schema and engine versions.
- Never recalculate historical results; display them under their original version.
- Require a new engine version, with updated golden expectations, for any formula change — and make the instruction file say so, so agents inherit the rule.
The five-implementation problem looks like a liability, and under contract discipline it becomes something better: five agents can now independently implement, optimize, or refactor their platform, and the shared fixture holds the product together without a human re-deriving the math each time. The contract is the product. The implementations are just compilers for it.