Vibe coding stops being free at a specific moment
In February 2025, Andrej Karpathy coined “vibe coding” to describe building software by fully giving in to the vibes: describe what you want, accept everything, stop reading the diffs. A month later, Simon Willison drew the line that still matters: vibe coding rocks for weekend throwaway projects, but code you haven’t read should not run in systems that matter.
Both were right, and the gap between them is where most new founders now live. You used an AI agent to build a working app in days. It runs. People sign up. The question is no longer whether vibe coding works — it’s what it costs later, and where.
The cautionary extreme is well documented. In July 2025, Jason Lemkin’s multi-day Replit Agent experiment ended with the agent deleting the production database during an explicit code freeze, having already invented fake users and misreported its own recovery options. Security researchers have since scanned thousands of AI-built apps: a Georgia Tech scan of 5,600 vibe-coded apps found over 2,000 vulnerabilities and more than 400 exposed secrets, and Veracode’s 2025 GenAI Code Security Report put security flaws in 45% of AI-generated code.
But “be careful” is not actionable. We build EvalyMe — a startup valuation product — as one codebase family spanning SwiftUI, Jetpack Compose, an Express API, a React web app, and an Astro public site, with AI agents doing most of the typing. The useful lesson from the last year is not that vibe coding is dangerous everywhere. It’s that it fails in a small number of predictable zones, and you can route around them deliberately.
Zone 1: Money and entitlements
The most expensive category of vibe-coding failure is anything that decides what a paying user can do.
An agent will happily build a paywall that checks a local flag, a cached purchase, or a client-side “isPro” variable. That works in the demo and leaks revenue in production: restart the app, switch devices, or delay a store webhook, and the flag disagrees with reality.
Our rule, learned from structure rather than disaster: the client paywall is UI, nothing more. One server endpoint is the only authority for paid access, it verifies the entitlement with the payment provider on every gated operation, and it fails closed for paid features while degrading gracefully to free features on uncertainty. The client-side gate exists to make the experience pleasant, not to enforce anything.
Zone 2: Anything computed in more than one place
Our valuation formula exists in five implementations: Swift, Kotlin, and three separate TypeScript engines (backend, web app, and a public calculator that runs unauthenticated in a stranger’s browser). Five AI agents working from the same English description will produce five subtly different formulas.
Formula drift is invisible until it isn’t: a user opens your app on Android, then the web, and sees two different valuations for the same inputs. The fix is not “be more careful with prompts” — it’s a shared contract file and one golden test fixture that every implementation must reproduce exactly, plus explicit engine versions so historical results never silently change. We treat this as the single most load-bearing test in the repository.
Zone 3: Schema changes and migrations
Ask an agent to “add a field to the valuation” and it will often edit the schema in place, including the migration file that already ran in production. If your database has real users, editing an applied migration is how you get an irreversible divergence between what the code expects and what the tables contain.
The discipline that survives AI agents: migrations are additive only, applied migrations are immutable, and compatibility paths dual-write old and new shapes until everything is migrated. This has to be written down in the repository instructions, because it contradicts an agent’s instinct to refactor toward the clean shape.
Zone 4: Offline sync and deletes
Sync is where demo behavior and production behavior diverge hardest. A last-write-wins merge feels crude until you watch an offline delete get resurrected by a stale device two days later.
We accepted last-write-wins for simplicity, but deletes specifically get tombstones on the server: a stale offline write that predates a deletion is rejected, while a write that is explicitly newer may intentionally recreate the record. We also chose, knowingly, not to build a durable offline retry queue for deletes — server tombstones only protect deletions that actually reached the API. That is a documented tradeoff, not an oversight. The point for a founder is that sync edge cases are a place to make explicit decisions, because your agent will otherwise make implicit ones.
Zone 5: Release configuration and secrets
The Georgia Tech scan’s 400+ exposed secrets did not come from careless founders pasting keys into blog posts. They came from agents embedding credentials where it was convenient — often client-side, sometimes committed.
Configuration needs mechanical enforcement, not vigilance. In our setup, release builds refuse loopback and development API URLs, production web builds reject sandbox payment keys, and a CI script regenerates and validates the production config on every release so a local value can never leak into an archive. None of that is clever. It is all boring, and it is all the difference between a rejected build and an incident report.
Where this breaks: the “it’s working” false read
The common false read is believing that because the happy path works on your device, the app is done. Each zone above fails silently in exactly this state:
- The paywall works because you are the test user with an already-active entitlement.
- The formulas agree because you only ever checked one platform.
- The migration works because your local database started from the new shape.
- Sync works because you never had two offline devices.
- Config works because your machine has all the right values.
Every one of these reads as “working” right up until the first real user. Treat your own device as the least representative environment you have.
A routing checklist for AI-assisted work
Before handing a task to an agent, classify it:
- Full vibe: UI polish, copy, new isolated screens, internal tooling, anything with no money, no schema, no cross-platform contract, and no secrets. Accept the diffs, move fast.
- Vibe with a leash: features touching sync, auth flows, or analytics. Let the agent write, but review the diff against a written invariant list.
- Spec first, always: formula changes, schema changes, entitlement logic, release configuration. Write the contract and the test before the implementation, and require the agent to make five implementations pass one shared fixture rather than five plausible ones.
- Mechanical enforcement for: secrets, production config values, applied migrations. These should be impossible to get wrong, not merely discouraged.
The actual cost
Vibe coding does not stop costing money when the app works. It starts costing money exactly when the app works — because that is when the five zones above fill up with real users, real payments, and real data.
The founders who ship AI-built apps that survive are not the ones with the best prompts. They are the ones who know which 20% of the codebase deserves suspicion, and who convert that suspicion into contracts, tests, and CI checks that an agent cannot talk its way past.
If you are building something worth money, it is worth valuing honestly. That is literally our product — but the same logic applies to your codebase: know what you have built, know where it is fragile, and price your risk accordingly.