← All articles

My Favorite Firebase Abstractions

For nearly a decade I have been contracting on Firebase backends, first and foremost for StaffTraveler and alongside that for Baymard. As these codebases grew I kept hitting practical problems the official SDKs did not have good answers for, and the libraries below are what I wrote to close those gaps. Some of these abstractions felt essential to stay sane.

The common thread is type safety at scale. Firestore returns DocumentData, Pub/Sub hands you Record<string, unknown>, Cloud Tasks wants you to serialize payloads by hand. You end up sprinkling satisfies and casts everywhere, which works until the day someone renames a field and nothing complains until production data starts to drift. These libraries buy back most of that safety with schemas declared once and inferred everywhere.

typed-firestore

This one is essential to me. Not just for the type safety, but for how much boilerplate it strips out of every Firestore call site. I wrote a long article about it a year ago, explaining the reasoning behind the abstractions, so I will not repeat the motivation here. What I want to do instead is show how it actually looks in a real codebase.

The pattern starts with a single typed reference map. Every collection in the database is listed once, with its document type attached to the reference:

import type { CollectionReference } from "firebase-admin/firestore";
import type { Airline, LoadsReport, User, UserConnectingFlight } from "@stafftraveler/common";
import { db } from "./firebase-admin-client";

export const refs = {
  airlines: db.collection("airlines") as CollectionReference<Airline>,
  loadsReports: db.collection("loadsReports") as CollectionReference<LoadsReport>,
  users: db.collection("users") as CollectionReference<User>,
  userConnectingFlights: (userId: string) =>
    db.collection(`users/${userId}/connectingFlights`) as CollectionReference<UserConnectingFlight>,
};

Every read and write in the codebase goes through these refs, and because the reference carries the type, everything downstream is typed automatically. A typical write looks like this, straight from the loads-request flow at StaffTraveler:

import { setDocument } from "@typed-firestore/server";
import { FieldValue, Timestamp } from "firebase-admin/firestore";

await setDocument(refs.userConnectingFlights(userId), groupId, {
  flightIds,
  createdAt: FieldValue.serverTimestamp(),
  expireAt: Timestamp.fromMillis(departureTimeMs),
});

The compiler enforces the full shape of UserConnectingFlight here. If a required field is missing, if a name is misspelled, or if FieldValue.serverTimestamp() is passed into a field that is not actually a Timestamp, TypeScript tells you. With the raw SDK this write is an UpdateData<T> free-for-all and you need a trailing satisfies to have any hope of catching typos.

The part I added more recently, and that I think is underappreciated, is a small set of helpers for Firestore triggers. Background-function events in the official SDK are typed as QueryDocumentSnapshot<DocumentData>, which means every handler starts with a cast. The helpers let you pass a typed ref purely for inference:

import { getBeforeAndAfterOnWritten } from "@typed-firestore/server/functions";
import { onDocumentWritten } from "firebase-functions/firestore";

export const syncUserData = onDocumentWritten(
  { document: makeDocumentHandlerPath(refs.users, "userId") },
  async (event) => {
    const { userId } = event.params;
    const [before, after] = getBeforeAndAfterOnWritten(refs.users, event);
    // before and after are typed as User | undefined
  },
);

No extra fetch happens here. The ref is only there to give the compiler something to latch onto, and before and after come out strongly typed against the collection schema. makeDocumentHandlerPath turns a typed ref and a parameter name into the users/{userId} shape Firestore wants, while preserving the parameter name in the type system so event.params.userId is typed as string. I never have to write the collection path twice, and I never lose types on the trigger parameters. Full API details and the React and React Native companions live on typed-firestore.codecompose.dev.

typed-pubsub and typed-tasks

These two libraries share a mental model, so I will cover them together. Both start with a Zod schema declared once, and from there you get a typed publisher or scheduler on one side and a typed handler on the other. Rename a field in the schema and the compiler walks you through every site that needs to change.

On the pain side, @google-cloud/pubsub delivers payloads as Record<string, unknown> and every handler has to parse and cast. @google-cloud/tasks is worse, because on top of untyped payloads you have to create queues manually, serialize JSON yourself, and roll your own deduplication if you want it. Neither SDK validates at the boundary, which means a publisher shipping a malformed payload can keep quietly poisoning a queue until somebody notices.

Here is how topic schemas are declared in StaffTraveler. Zod’s discriminated unions turn out to be a great fit for event payloads:

import { createTypedPubsub } from "@codecompose/typed-pubsub";
import { PubSub } from "@google-cloud/pubsub";
import { z } from "zod";

const schemas = {
  record_user_event: z.discriminatedUnion("type", [
    z.object({
      userId: z.string(),
      type: z.literal("loadsRequest"),
      metadata: z.object({
        flightId: z.string(),
        isPriorityRequest: z.boolean(),
      }),
    }),
    z.object({
      userId: z.string(),
      type: z.literal("loadsReportFlag"),
      metadata: z.object({ loadsReportId: z.string() }),
    }),
    // ...many more variants
  ]),
  replace_name_in_conversations: z.object({
    userId: z.string(),
    newName: z.string(),
  }),
} as const;

export const pubsub = createTypedPubsub({
  client: new PubSub(),
  schemas,
  region: "us-central1",
});

Publishing an event is then just a typed function call, and the compiler will not let you mix type and metadata from different variants:

await pubsub.createPublisher("record_user_event")({
  userId,
  type: "loadsReportFlag",
  metadata: { loadsReportId },
});

The corresponding handler is declared with the same topic name, and its payload is narrowed automatically:

export const replace_name_in_conversations = pubsub.createHandler({
  topic: "replace_name_in_conversations",
  handler: async ({ userId, newName }) => {
    // userId and newName are fully typed, already validated
  },
});

The last point is what I mean by “near end-to-end type safety”. The schema gives you compile-time inference on both sides, and Zod validates the payload at runtime inside the handler before your code runs. A publisher that accidentally drifts from the schema fails on arrival, loudly, rather than silently corrupting downstream state.

typed-tasks works the same way, with one extra wrinkle for Cloud Tasks: deduplication. Two flags on a task definition cover the common cases. useDeduplication: true generates a task name from the payload hash so identical payloads collapse, and deduplicationWindowSeconds debounces within a rolling window. Both are things you otherwise write by hand:

import { CloudTasksClient } from "@google-cloud/tasks";
import { createTypedTasks } from "typed-tasks";
import { z } from "zod";

const taskDefinitions = {
  syncDeviceTokens: {
    schema: z.object({ userId: z.string() }),
    options: { deduplicationWindowSeconds: 10 },
  },
  syncWithTypesense: {
    schema: z.object({
      collection: z.enum(["airlines", "airports", "cities", "tips", "trackedFlights", "users"]),
      documentId: z.string(),
      operation: z.enum(["upsert", "delete"]),
      typesenseData: z.record(z.string(), z.unknown()).optional(),
    }),
  },
} as const;

export const tasks = createTypedTasks({
  client: new CloudTasksClient(),
  definitions: taskDefinitions,
  projectId: "my-gcp-project",
  region: "us-central1",
});

Enqueue:

await tasks.createScheduler("syncWithTypesense")({
  collection,
  documentId,
  operation: "upsert",
  typesenseData,
});

Handle:

export const syncWithTypesense = tasks.createHandler({
  queueName: "syncWithTypesense",
  handler: async ({ collection, documentId, operation, typesenseData }) => {
    // ...
  },
});

Both libraries also set saner Firebase v2 defaults than the platform itself does. 3000 max instances on a handler is how you surprise-bill yourself on a weekend. Documentation for both lives on typed-pubsub.codecompose.dev and typed-tasks.codecompose.dev.

gcp-job-runner

I originally wrote this for database migrations. StaffTraveler has close to a million users, and because we store a document for every user event, some of our Firestore collection group queries need to walk billions of documents. That is not something you want to stream through your laptop when the database lives on the other side of the world. The moment your collections get that big, you end up on Cloud Run anyway, and that means Docker images, Secret Manager wiring, CLI flags for arguments, and some way of keeping parity between the version you run locally during development and the one that actually runs in production. I got tired of re-doing that plumbing for every migration.

What I realized over time is that the same plumbing is exactly what I want for a lot of other things too: occasional one-shot scripts, data analyses, exports that dump a CSV or JSON to a bucket. Anything that does not belong in a function or a service, but still needs to hit the production environment with the right credentials. gcp-job-runner treats all of that uniformly.

A job is just a file that declares its arguments and a handler:

import { defineJob, getFileWriter } from "gcp-job-runner";
import { z } from "zod";

export default defineJob({
  description: "Export product catalog snapshots to JSON",
  schema: z.object({
    apiKey: z.string().describe("API key for the upstream service"),
    dir: z.string().default("snapshots"),
  }),
  handler: async ({ apiKey, dir }) => {
    const writer = getFileWriter();

    const [products, offerings, packages] = await Promise.all([
      fetchProducts(apiKey),
      fetchOfferings(apiKey),
      fetchPackages(apiKey),
    ]);

    await Promise.all([
      writer.writeJson(`${dir}/products.json`, products),
      writer.writeJson(`${dir}/offerings.json`, offerings),
      writer.writeJson(`${dir}/packages.json`, packages),
    ]);
  },
});

The same file runs locally via job local run and on Cloud Run via job cloud run. Arguments are Zod-validated on both ends. Secrets come from GCP Secret Manager when they are not in the local env file. Docker images are cached by content hash so subsequent cloud runs skip the build. And getFileWriter() is the part I find genuinely delightful: locally it writes to a directory you configured, in the cloud it writes to a GCS bucket, and the handler code does not change.

Environments are declared once, at the root of the app:

import { defineRunnerConfig, defineRunnerEnv } from "gcp-job-runner";

export default defineRunnerConfig({
  localOutputFilesPath: "../../exports",
  environments: {
    stag: defineRunnerEnv({
      project: "stafftraveler",
      envFile: ".env.stag",
      outputFilesPath: "gs://stafftraveler-stag-exports",
    }),
    prod: defineRunnerEnv({
      project: "stafftraveler-prod",
      envFile: ".env.prod",
      outputFilesPath: "gs://stafftraveler-prod-exports",
    }),
  },
  cloud: {
    name: "cli-jobs",
    network: { name: "default", subnet: "default" },
  },
});

A typical migration example is initializing a new settings field on every existing user: a single defineJob that walks the users collection and updates in chunks using processInChunks. The same binary you would run against staging is the one that runs on Cloud Run in production. The gap between “I have a one-off script” and “I need to run this safely against a million documents” stops being a meaningful distinction. Full documentation is on gcp-job-runner.codecompose.dev.

firebase-tools-with-isolate

This is the shortest story. firebase deploy does not understand pnpm or npm workspaces, so in a monorepo it ships a package manifest that references workspace dependencies the cloud pipeline cannot resolve. I wrote isolate-package a few years ago to solve that, and later a fork of the Firebase CLI that runs the isolation step automatically at deploy time.

firebase-tools-with-isolate is a drop-in replacement for firebase-tools. Version numbers track upstream one-to-one, and a GitHub Action re-publishes daily whenever a new upstream release lands. You swap one dependency and keep running firebase deploy exactly as before. If the CLI detects that your functions source sits inside a workspace, it isolates. If not, it gets out of the way.

A nice side effect is that the Firebase “multiple codebases” feature actually becomes practical. You can split your functions into several packages, one per domain, list them all under functions in firebase.json, and deploy each one on its own with firebase deploy --only functions:<codebase>. Each package isolates independently, so they stay decoupled at deploy time, but the emulator starts them all up together so local development still sees the whole backend. I wrote a separate article covering the background, if you want the longer version.

Wrapping Up

Put together, these libraries let you treat Firebase as something much closer to a schema-first backend. Documents, triggers, topics, tasks, jobs: every edge that leaves or enters your service carries a schema, and the compiler keeps them all in sync. It is not a perfect story (Firestore where clauses remain untyped, for example), but it gets far closer than you would expect from the raw SDKs.

All of these are open source, and I am publishing them as I move on from Firebase myself. I am in the process of founding my own startup, and the requirements for that platform are taking me into new territory. These libraries have been used in production at scale for years now, and I hope other teams will enjoy using them too. Issues and PRs are very welcome on all of them.

Three smaller libraries I lean on constantly deserve a closing mention too:

  • process-in-chunks for batching and parallelizing server-side work with a sane default concurrency
  • get-or-throw for turning “get or undefined” lookups into “get or throw” ones, which pairs especially well with TypeScript’s noUncheckedIndexedAccess strict setting and removes a lot of defensive branching
  • typescript-config for a set of shared tsconfig presets I use in every project

They are not Firebase-specific, but you will see them in almost any codebase I work on, including in the snippets above. None of them do very much on their own, but together they take the edges off a lot of everyday TypeScript work.

If you would like to leave feedback or have questions about this article, feel free to contact me on Bluesky . You can also follow me there to find out when new articles are published.