Backups are for data you cannot regenerate
PocketOS's backups died with the volume they protected. The fix is not better backups. Split your tables by derivability and give the recipe-class half a rebuild command instead of a copy.
In late April 2026 an AI coding agent deleted PocketOS's production database. Nine seconds, one GraphQL call. The detail that made it catastrophic is the one everybody quoted afterward: the volume-level backups were stored inside the volume, so a single deletion removed the data and its recovery at the same time. The founder, Jer Crane, put the architectural rule plainly in his post-mortem: real backups live in a different blast radius. They restored from a three-month-old backup, and the data gaps from everything since were still open as he wrote it.
Every serious follow-up concluded that the fix is better backup architecture: off-account storage, published recovery SLAs, restore tests, deletion protection, scoped tokens. All correct. None of it is interesting, because it accepts the premise that the answer to data loss is always another copy of the bytes.
There is a second answer, and it applies to most of the databases you run. Some data is worth copying. Some data is worth being able to rebuild. Treating the second kind like the first is why you have storage you do not need, protection you cannot verify, and credentials in places where a nine-second mistake is possible.
Two classes, and the question that separates them
The dividing question is not importance. It is derivability.
Where is the value: in the rows, or in the rule that produced them?
| Artifact data | Recipe data | |
|---|---|---|
| The value lives in | These exact bytes | The structure and the procedure |
| Losing it costs | Whatever is in the rows | The minutes it takes to rebuild |
| Recovery mechanism | Restore a copy you protected | Run the recipe against the schema |
| Recovery point objective | Age of the last copy | As current as the schema definition |
| Where it must live | Outside the blast radius | Anywhere, including inside it |
| Typical content | Customer records, transactions, user-generated content, invoices | Test fixtures, sandboxes, demo databases, preview environments, load-test targets, scratch staging |
A reservation for a car rental customer is an artifact. It encodes something that happened, and nothing can reproduce it. A thousand plausible reservations exercising the same code path are a recipe's output; any thousand that satisfy the foreign keys and the constraints will do, and the next thousand are just as good.
The reason this matters architecturally rather than philosophically is that the two classes fail differently. An artifact database fails when the copy is unreachable, stale, or destroyed alongside the original. A recipe database fails when the recipe is stale, which is a completely different problem with a completely different cost profile, and one that a backup policy will never catch.
Most estates contain both classes in the same instance, which is where the misclassification happens. The classification is per table, not per database.
Sort your tables, not your databases
Look at a typical application schema and ask of each table whether a row's value is the row itself or the rule that made it.
Artifact tables are the ones with a human or an external event behind every row: customers, orders, payments, invoices, support_tickets, documents, anything with a created_by that names a real person at a real moment. If a regulator, a customer, or your own accounting close needed those exact bytes, they are artifacts.
Recipe tables are the ones where any conforming row is equivalent to any other: sessions, job_queue and its dead-letter table, notification state, cart contents, rate-limit buckets, cache tables, materialized views, telemetry you have already aggregated away, and every database whose entire purpose is to have code run against it.
The interesting cases are the ones that feel artifact but are not. A daily_revenue_report table computed from payments is derivable, provided the computation is deterministic and the inputs survive. A products catalog is an artifact if a human wrote the descriptions, and a recipe if it is seeded from a supplier feed you can re-pull. audit_log is an artifact under compliance regimes that require the literal record, and disposable if you added it for developer convenience and nobody has queried it in two years.
That last category is usually larger than teams expect. The first useful outcome of this framing is a list of databases you are backing up, protecting, replicating, and paying storage for that do not need any of it, and would be strictly safer if you deleted the copies and committed the procedure instead.
The property of recipes that backups cannot have
A backup is tested on the day you need it. A recipe is tested every time somebody resets their development database.
This is the line that should change a decision, because it inverts the perceived risk. Teams assume restoring from a backup is the conservative option and regeneration is the cavalier one. In practice the artifact path is the one whose correctness is verified rarely, under duress, by whoever is awake, against a copy whose age was last questioned during an incident. The recipe path is executed daily by people who will file an issue within minutes if it produces something wrong.
The asymmetry compounds in three places:
- Staleness is visible. A three-month-old backup is silently, invisibly old until it is used. A recipe that no longer matches the schema breaks loudly on the next run, which is the only moment anyone will actually fix it.
- The recipe is code. It lives in version control, gets reviewed in a pull request, and carries the same history as the schema it describes. A backup is an opaque blob in a bucket whose provenance you reconstruct from timestamps.
- Nothing to leak. A committed procedure contains no customer rows. A snapshot of a production database contains all of them, and every environment that receives one inherits an obligation it probably cannot meet. This is the reason production dumps in sandboxes turn a data-loss incident into a disclosure incident, and it is why an agent's reach has to be measured against what it could copy, not only what it could delete.
The recipe, concretely
For a Postgres database whose content is derivable, the recipe is one committed text file and one command: the schema in version control, and a generator that turns schema plus intent into rows.
The schema comes out of the database it belongs to, done by a human or CI rather than by anything an agent can read:
pg_dump --schema-only -d app > schema.sqlThe rows come back from that text, into whatever target is reachable:
Foreign keys resolve in dependency order, parents before children, so every child row points at something that exists. Columns whose values repeat, like status or country, can inherit their observed frequencies from pg_stats when the recipe reads a live schema rather than a parsed one, which is what makes a rebuilt database behave like the original under a query plan instead of merely containing rows. Coherence across columns holds, so a derived email still matches the name it was built from and a total still equals its line items.
The result is that a wiped sandbox is not an outage. It is a build step.
- Give your coding agent a database it can break: the credential boundary, step by step
- Generate Postgres test data with no database: paste mode from
CREATE TABLEtext
Where recipes fail
Honest boundaries
A recipe is only as current as the schema it reads, which is its own failure mode and the reason drift must be handled deliberately: when a migration lands, the recipe has to be updated alongside it, or the rebuilt database quietly stops matching production. A recipe also cannot reproduce the specific bad row that exists only in production; if a bug depends on one customer's malformed address, no amount of regeneration finds it. Distributions you can inherit are the categorical ones already recorded in pg_stats, not a full statistical portrait of every column. Generated rows are constrained by the schema you declare, so check constraints and cross-column rules that live in application code rather than in DDL need to be expressed explicitly. And artifact data always needs the protected copy; nothing in this post applies to customer records, and pretending otherwise is how a company ends up with a three-month-old snapshot as its worst case.
There is a subtler trap worth naming, because it converts a recipe database into an artifact database without anybody deciding to. Copy real rows into your sandbox, whether by weavori sync or a restored dump, and the sandbox now holds customer data. sync moves real values between databases in foreign-key order and does not mask them; that is a feature in the right context and a compliance problem in the wrong one. The moment you use it, the target is artifact-class data, your backup policy has to cover it, and the credential that reaches it just became worth protecting.
That is a legitimate need. When you specifically need production's literal values rather than a reproducible stand-in, copying them is the right call. Do it deliberately, into a target you then treat as an artifact, and not as a side effect of convenience.
The rule that follows
- Classify every table as artifact or recipe, and let the answer be per table rather than per environment.
- For artifact tables: protected copies outside the blast radius, restore tests on a schedule, deletion protection switched on, and an honest recovery point objective written down before you need it.
- For recipe tables: commit the schema as text, commit the generation command, and stop backing them up. Deleting the copies removes the exposure.
- Keep the recipe current with the migrations. If the recipe only works against last quarter's schema, it is a backup with worse guarantees.
- Put nothing inside an agent's reach that is artifact-class and does not have to be. Reachable credentials are the boundary, and the credential's authority is the whole of it.
- When you genuinely need real values in a derived environment, mark that environment as artifact-class at the moment you create it, and protect it like the thing it has become.
Summary
- A nine-second deletion took PocketOS's database and its backups together because the backups lived inside the volume. The correct criticism of that design is real, and it is not the only available answer to data loss.
- Split data by derivability rather than importance: rows whose value is their own bytes, and rows whose value is the rule that produced them.
- Artifact recovery needs copies that survive the original. Recipe recovery needs a procedure that is current, and it gets tested daily instead of on the worst day of your life.
- Classify per table. Most estates contain more recipe data than teams believe, which means more storage, more exposure surface, and more reachable credentials than the architecture intended.
- You cannot delete data that regenerates. For the half of your databases this is true of, that is the cheaper and safer recovery model, and it is the one that lets an agent work against your schema without you holding your breath.
Start by naming one database you are protecting that does not need protecting, and give it a recipe instead:
- Postgres query plans without production data: what a rebuilt database has to get right for a plan to be trustworthy
- The Postgres migration that passed every test: the constraint class that only realistic rows expose
- How to generate realistic test data for PostgreSQL: the full generation walkthrough
- Cookbook: reproducible recipes against real Postgres