Firestore disaster backups
Production cross-project managed exports use the recovery-project custom role
myRecipesFirestoreExportWriter on only the approved recovery bucket. It grants
storage.buckets.get, storage.objects.create, storage.objects.list, and
storage.objects.delete to the production Firestore service agent. Do not
replace it with Storage Admin or add object-read, bucket-update, IAM, restore,
or retention-change permissions.
The scheduled freshness monitor uses the separate recovery-project custom role
myRecipesBucketMetadataReader, containing only storage.buckets.get, on that
same bucket. It has no object access and no bucket-list permission.
Issue #377 provides a 24-hour-or-better Firestore recovery point. Development
implementation and all failure testing run only in myrecipes-test.
Production promotion is a separately approved operation under #391.
Operators must follow the
development operations runbook
or production operations runbook
for environment targeting, IAM, deployment, rollback, and monitoring.
Selected mechanism
The first release uses two complementary Firestore artifacts:
| Artifact | Consistent | Project-isolated | Primary purpose |
|---|---|---|---|
| Native Firestore backup | Yes | No | Preferred normal database recovery |
| Cross-project managed export | Not guaranteed | Yes | Project-loss contingency |
| Protected image copy | Selected-generation relationship | Yes | Recipe-image recovery |
Native daily and weekly Firestore backups are the authoritative consistent database recovery points. They contain all Firestore documents and index configuration and restore into a new database, but remain in the application project. Prefer them whenever that project and its backups remain available and trusted.
The independently administered fallback is a full managed Firestore export to
private Cloud Storage. The exporter omits collectionIds, so every collection
and subcollection is included. It deliberately omits snapshotTime because
the development Firestore API rejected every explicit timestamp while PITR
was disabled. Google therefore provides no point-in-time consistency guarantee
for this artifact; documents may reflect different moments during the export.
Only a completed long-running operation without an error and a top-level export
metadata file are recoverable.
Daily and weekly exports are independent:
gs://BUCKET/firestore/daily/{exportStartTimestamp}
gs://BUCKET/firestore/weekly/{exportStartTimestamp}
Daily artifacts expire after eight days and weekly artifacts after 35 days. Those buffers preserve at least seven daily and four weekly generations across normal schedule jitter. Partial output from a failed or cancelled operation is not a backup and must never be imported.
PITR remains a future defense-in-depth enhancement for minute-level recovery within the preceding seven days, not a first-release requirement.
Scheduling and status
- native daily and Sunday weekly schedules, managed by Firestore;
- cross-project daily full export: 01:30 UTC;
- cross-project weekly full export: Sunday 04:00 UTC; and
- independent native-backup and cross-project-export watchdogs: every six hours.
The dedicated runtime starts and polls exports. The Firestore service agent,
not the application runtime, creates objects in the destination bucket. A
successful cross-project daily run writes content-free status to
backup_operations/firestore-backup-latest. Status and logs contain the source
database, cadence, start/completion timestamps, operation ID, destination
prefix, duration, document count, and byte count. They do not contain document
IDs, fields, collection contents, or recipe data.
The native watchdog records only backup resource name, source database, consistent snapshot time, state, and age. It never logs database contents.
IAM and privacy
myrecipes-firestore-backup is the development runtime. Its custom role can
start/inspect exports and write operational status; it cannot delete artifacts.
The Cloud Scheduler service agent may mint OIDC tokens only for this runtime
identity so scheduled HTTP delivery remains authenticated.
The private recovery bucket uses uniform bucket-level access and public-access
prevention and is absent from client Firebase configuration. Firestore rules
deny normal application users access to backup_operations through the final
deny rule. No service-account keys are created.
Production must use a separately administered backup project/bucket. Its application project service agent receives only the minimum object-creation access, and recovery operators receive separate audited read/import access.
Monitoring and response
The log metric myrecipes_firestore_backup_failures counts export failures,
inaccessible destination failures, and daily recovery points older than 30
hours. myrecipes_native_firestore_backup_failures independently counts a
missing, stale, or unavailable ready native backup. Cloud Monitoring routes
both policies to the private owner/operator channel. The combined recovery
watchdog allows 30 hours for the daily native snapshot to complete its bounded
restore and image-manifest qualification. Its alert metric uses a 12-hour
alignment window so a genuine failure remains open across scheduled checks
instead of appearing to recover when a single log entry leaves a five-minute
window.
Acknowledge within four normal-coverage hours. The condition is release
blocking until remediated or explicitly accepted by the owner.
Cost model
Native backup storage is billed independently from managed exports and restore
operations. Managed exports charge one document read per exported document. With D
documents, a normal month performs approximately 30D reads for daily exports
plus 4D reads for weekly exports. Storage is approximately the sum of eight
daily full exports plus four weekly full exports, subject to export encoding
and lifecycle timing. Add Cloud Storage Class A operations, function/runtime,
logging/monitoring, and import reads/writes when recovery is exercised.
At the August 20 development inventory of 5,752 documents and 2,412,390 export bytes, the schedule is approximately 195,568 export document reads per 30-day/four-week month and about 28.9 MB of retained steady-state export bytes before encoding variance. This is expected to remain below $0.01/month in development, but production must recalculate from its current document count, export bytes, regional storage price, and Firebase billing plan before #391.
The August 21 hybrid run exported 5,758 documents and 2,416,545 operation bytes, which does not materially change that development estimate. Native backup storage and restore charges must be measured after the first scheduled backup; they are tracked separately from export reads and bucket storage.
Recovery selection and validation
Restore each path independently during #382:
- restore a ready native backup into a newly named isolated database and validate document counts, indexes, ownership, relationships, and references;
- import a completed cross-project export into a different isolated database and run the same validation plus dangling-reference and relationship checks;
- report and quarantine any inconsistent fallback import instead of promoting it silently; and
- select protected images by the chosen Firestore artifact's completion or snapshot time, recording any temporal gap.
The ordinary database-loss RPO is the age of the latest consistent native backup and must not exceed 24 hours. The cross-project fallback has a completed artifact-age objective of 24 hours but does not independently promise a consistent 24-hour RPO. Images require a recoverable generation/copy no older than 24 hours.
Restore boundary
Managed exports contain Firestore documents only. They do not restore Firebase Auth users, Storage objects, Security Rules, indexes, TTL policies, Functions, Scheduler jobs, Eventarc triggers, App Check, IAM, Secret Manager values, or Apple configuration. Restore validation belongs to #382 and must target a separate named database. Recipe images use #376; user export/import uses