The gap between “it works” and “it can ship” is a list, not a mystery
An AI agent will take you from empty folder to working app faster than any previous tool in software history. What it will not do — because no one asks it to — is produce the list of obligations that appear the moment your app has real users: the store guideline that rejects your binary, the privacy expectation that protects your users, and the release configuration that keeps your production keys out of your development loop.
None of this is intellectually hard. It is just invisible while building, because none of it affects the demo. Here is the list, drawn from shipping and maintaining a product across the App Store, Play Store, and web — with notes on how to make each item mechanical so an agent cannot regress it later.
Account deletion is a shipping requirement, not a feature
If your app has account creation, Apple’s App Store guidelines require offering in-app account deletion — a full deletion, not a “contact support” dead end. This is the single most common late-stage rejection for new apps, and it arrives after you thought you were done.
The engineering is more subtle than a delete button, because deletion must be sequenced across every system that holds user state. Our deletion endpoint removes app-owned data first, then the payment-provider customer record, then the auth provider’s user, and only then does the client sign out locally. Two details agents routinely get wrong:
- The client must sign out only after the server confirms success; a client-initiated “delete and logout” that fails halfway leaves an account that cannot be deleted again from a signed-out state.
- Deleting the account must not cancel a running App Store or Play Store subscription — the user keeps Apple’s and Google’s subscription-management screens, and your deletion flow must preserve their access to them. Deleting the subscription management surface along with the account is how you generate chargebacks you cannot see.
Release configuration must be mechanically incapable of being wrong
The classic production incident for AI-assisted apps is configuration drift: a development API URL, a sandbox key, or a local override riding into a release build because it was convenient during testing.
The defense is to make wrong configurations fail at build time, not review time. In our setup, the CI system writes the production configuration file from secrets at clone time — the real values never exist in the repository at all — and a build phase rejects loopback and development API URLs in release archives. The web production build validates its environment the same way: non-HTTPS API URLs, non-live auth keys, and sandbox payment keys all abort the build. Local development uses a separate ignored config file that release builds are forbidden from including.
This converts a category of silent failure (“it shipped pointing at localhost, or worse, pointing at the dev database”) into a red build. Agents then self-correct, because a failing check is feedback they act on, while a convention is a suggestion they eventually ignore.
Store metadata is testable, so test it
App Store metadata — descriptions, keywords, screenshots labels, version copy — is usually edited in a web form at two in the morning before a release, which is exactly when typos and broken links ship. Ours is tested in CI like code: metadata files live in the repository, and a shell test validates structure and content before every release. The same applies to the release-configuration scripts themselves; they are code, they have tests, and they run without building the app.
The general principle for founders: if a mistake in an artifact would cost you a release cycle, the artifact belongs in version control with a check. “We’ll be careful” does not survive contact with a deadline.
Analytics privacy is a design decision made once
Once instrumentation is in, an agent logging “helpful” context will happily include the most identifying data available: raw user inputs, URLs, emails, provider error messages. For a product handling financial inputs, that is not just sloppy — it is a privacy exposure you cannot walk back after the data is collected.
The workable rule is a short blocklist enforced by convention in one analytics module: no raw financial inputs, no identities, no URLs, no provider errors. Our public calculator emits only coarse bands — “MRR between $10k and $25k,” never the number — and the app honors Do Not Track and Global Privacy Control signals. Decide this at instrumentation time; anonymizing after the fact is not a real option.
Demo modes, migrations, and other sharp edges
Three smaller items that each cost someone a bad week:
- Demo and screenshot modes must be structurally debug-only. Marketing needs screenshots of the paid flow, so we have a screenshot mode that bypasses the paywall gate for deterministic demo data — gated by a build flag that exists only in debug builds. If your demo bypass is a runtime toggle, it will ship.
- Database migrations are append-only after deployment. Once a migration has run anywhere real, editing it forks reality. New schema changes get new additive migrations, and compatibility paths dual-write until old data is fully migrated. This must live in your instruction file, because refactoring an old migration is always what the agent’s “cleanup” instinct suggests.
- Deploy scripts should have a dry run. Ours supports a dry-run and a check-only mode per component, and the differences between what each deployment actually verifies are written down. Deploy tooling that does everything silently is indistinguishable from deploy tooling that does nothing, until the day it does.
Where this breaks: the silent agent
The failure mode across all of the above is not negligence — it is that agents optimize for the stated goal and stay silent about unstated obligations. An agent asked to “add a delete account button” will add the button. It will not mention store guidelines, subscription management, or sequencing. An agent asked to “wire up analytics” will instrument what is easiest to reach, including your users’ raw inputs, unless the blocklist already exists.
The pattern that works: every obligation in this article lives in the repository as either a test, a build validation, or an explicit written rule — and the instruction file points agents at all three. Anything enforced only by your memory will be violated the first time you delegate the area to an agent while tired.
The pre-launch checklist
- In-app account deletion that fully sequences across your data, payments, and auth providers — and preserves store subscription management.
- Release builds reject development URLs, sandbox keys, and wrong-platform credentials mechanically.
- Store metadata and release scripts are versioned and tested.
- Analytics has a written blocklist: no raw financial inputs, no identities, no URLs, no provider errors; DNT and GPC honored.
- Demo bypasses are debug-build-only by construction.
- Migrations are append-only once deployed; the rule is in the instruction file.
- Deploy tooling has a dry-run, and what each deploy step verifies is documented.
The demo was the easy 80%. This list is the hard 20% that decides whether the app survives its first contact with reviewers, auditors, and real users — and every item on it is exactly the kind of well-specified mechanical work AI agents are good at, once someone specifies it. That someone is you.