Skip to content

Firestore disaster backups

Production cross-project managed exports use the recovery-project custom role myRecipesFirestoreExportWriter on only the approved recovery bucket. It grants storage.buckets.get, storage.objects.create, storage.objects.list, and storage.objects.delete to the production Firestore service agent. Do not replace it with Storage Admin or add object-read, bucket-update, IAM, restore, or retention-change permissions.

The scheduled freshness monitor uses the separate recovery-project custom role myRecipesBucketMetadataReader, containing only storage.buckets.get, on that same bucket. It has no object access and no bucket-list permission.

Issue #377 provides a 24-hour-or-better Firestore recovery point. Development implementation and all failure testing run only in myrecipes-test. Production promotion is a separately approved operation under #391. Operators must follow the development operations runbook or production operations runbook for environment targeting, IAM, deployment, rollback, and monitoring.

Selected mechanism

The first release uses two complementary Firestore artifacts:

Artifact Consistent Project-isolated Primary purpose
Native Firestore backup Yes No Preferred normal database recovery
Cross-project managed export Not guaranteed Yes Project-loss contingency
Protected image copy Selected-generation relationship Yes Recipe-image recovery

Native daily and weekly Firestore backups are the authoritative consistent database recovery points. They contain all Firestore documents and index configuration and restore into a new database, but remain in the application project. Prefer them whenever that project and its backups remain available and trusted.

The independently administered fallback is a full managed Firestore export to private Cloud Storage. The exporter omits collectionIds, so every collection and subcollection is included. It deliberately omits snapshotTime because the development Firestore API rejected every explicit timestamp while PITR was disabled. Google therefore provides no point-in-time consistency guarantee for this artifact; documents may reflect different moments during the export. Only a completed long-running operation without an error and a top-level export metadata file are recoverable.

Daily and weekly exports are independent:

gs://BUCKET/firestore/daily/{exportStartTimestamp}
gs://BUCKET/firestore/weekly/{exportStartTimestamp}

Daily artifacts expire after eight days and weekly artifacts after 35 days. Those buffers preserve at least seven daily and four weekly generations across normal schedule jitter. Partial output from a failed or cancelled operation is not a backup and must never be imported.

PITR remains a future defense-in-depth enhancement for minute-level recovery within the preceding seven days, not a first-release requirement.

Scheduling and status

  • native daily and Sunday weekly schedules, managed by Firestore;
  • cross-project daily full export: 01:30 UTC;
  • cross-project weekly full export: Sunday 04:00 UTC; and
  • independent native-backup and cross-project-export watchdogs: every six hours.

The dedicated runtime starts and polls exports. The Firestore service agent, not the application runtime, creates objects in the destination bucket. A successful cross-project daily run writes content-free status to backup_operations/firestore-backup-latest. Status and logs contain the source database, cadence, start/completion timestamps, operation ID, destination prefix, duration, document count, and byte count. They do not contain document IDs, fields, collection contents, or recipe data.

The native watchdog records only backup resource name, source database, consistent snapshot time, state, and age. It never logs database contents.

IAM and privacy

myrecipes-firestore-backup is the development runtime. Its custom role can start/inspect exports and write operational status; it cannot delete artifacts. The Cloud Scheduler service agent may mint OIDC tokens only for this runtime identity so scheduled HTTP delivery remains authenticated. The private recovery bucket uses uniform bucket-level access and public-access prevention and is absent from client Firebase configuration. Firestore rules deny normal application users access to backup_operations through the final deny rule. No service-account keys are created.

Production must use a separately administered backup project/bucket. Its application project service agent receives only the minimum object-creation access, and recovery operators receive separate audited read/import access.

Monitoring and response

The log metric myrecipes_firestore_backup_failures counts export failures, inaccessible destination failures, and daily recovery points older than 30 hours. myrecipes_native_firestore_backup_failures independently counts a missing, stale, or unavailable ready native backup. Cloud Monitoring routes both policies to the private owner/operator channel. The combined recovery watchdog allows 30 hours for the daily native snapshot to complete its bounded restore and image-manifest qualification. Its alert metric uses a 12-hour alignment window so a genuine failure remains open across scheduled checks instead of appearing to recover when a single log entry leaves a five-minute window. Acknowledge within four normal-coverage hours. The condition is release blocking until remediated or explicitly accepted by the owner.

Cost model

Native backup storage is billed independently from managed exports and restore operations. Managed exports charge one document read per exported document. With D documents, a normal month performs approximately 30D reads for daily exports plus 4D reads for weekly exports. Storage is approximately the sum of eight daily full exports plus four weekly full exports, subject to export encoding and lifecycle timing. Add Cloud Storage Class A operations, function/runtime, logging/monitoring, and import reads/writes when recovery is exercised.

At the August 20 development inventory of 5,752 documents and 2,412,390 export bytes, the schedule is approximately 195,568 export document reads per 30-day/four-week month and about 28.9 MB of retained steady-state export bytes before encoding variance. This is expected to remain below $0.01/month in development, but production must recalculate from its current document count, export bytes, regional storage price, and Firebase billing plan before #391.

The August 21 hybrid run exported 5,758 documents and 2,416,545 operation bytes, which does not materially change that development estimate. Native backup storage and restore charges must be measured after the first scheduled backup; they are tracked separately from export reads and bucket storage.

Recovery selection and validation

Restore each path independently during #382:

  1. restore a ready native backup into a newly named isolated database and validate document counts, indexes, ownership, relationships, and references;
  2. import a completed cross-project export into a different isolated database and run the same validation plus dangling-reference and relationship checks;
  3. report and quarantine any inconsistent fallback import instead of promoting it silently; and
  4. select protected images by the chosen Firestore artifact's completion or snapshot time, recording any temporal gap.

The ordinary database-loss RPO is the age of the latest consistent native backup and must not exceed 24 hours. The cross-project fallback has a completed artifact-age objective of 24 hours but does not independently promise a consistent 24-hour RPO. Images require a recoverable generation/copy no older than 24 hours.

Restore boundary

Managed exports contain Firestore documents only. They do not restore Firebase Auth users, Storage objects, Security Rules, indexes, TTL policies, Functions, Scheduler jobs, Eventarc triggers, App Check, IAM, Secret Manager values, or Apple configuration. Restore validation belongs to #382 and must target a separate named database. Recipe images use #376; user export/import uses

375/#378/#379. Never import an exercise into (default).