Skip to content

feat(improve-pr): the improvement radar measures first, and states it… #79

feat(improve-pr): the improvement radar measures first, and states it…

feat(improve-pr): the improvement radar measures first, and states it… #79

Workflow file for this run

name: deploy
# Redeploy the FlareDispatch workers when upstream (main) updates.
# Container images (infra/Dockerfile.sandbox for the dispatcher,
# infra/Dockerfile.substrate for the substrate) are (re)built by
# `wrangler deploy` on the runner, so this needs Docker — ubuntu-latest
# provides it. The path filter keeps doc-only commits from triggering an image
# rebuild.
#
# ORDER. substrate → canary → dispatcher. The substrate goes first because
# consumers' service bindings resolve against a Worker that must already exist;
# the canary sits between them because ADR-0011 requires the egress floor to be
# proven on the running build *before* consumer traffic is admitted, and the
# dispatcher is a consumer.
#
# `workflow_dispatch` with `target: substrate` stops after the canary, so a
# security patch to the substrate never queues behind a product release.
#
# CREDENTIALS. `CLOUDFLARE_API_TOKEN` — account-scoped to FractalBox
# (c91d52c2…), no zone resources: Workers Scripts:Edit, Workers KV:Edit,
# Workers R2:Edit, D1:Edit, Containers:Edit, Account Settings:Read.
# Created 2026-08-06 · dash.cloudflare.com/profile/api-tokens · no expiry.
# Scope is recorded here because GitHub secrets have no comment field and a
# Cloudflare token cannot report its own policies — reading them needs a
# *second* token with API Tokens:Read, so an undocumented token's grants are
# unrecoverable in practice. `Containers:Edit` is the one the "Edit Cloudflare
# Workers" template omits; without it the substrate's four container
# applications fail to deploy.
#
# Zone resources are deliberately none: both workers serve on workers.dev, so
# a token that could touch DNS would be strictly more blast radius than the
# job can use.
#
# The token is account-scoped only, so `GET /user/tokens/verify` returns
# "Invalid API Token" for it — that endpoint lives under /user/ and needs
# user-level access. Verify against `GET /accounts/{id}` instead; a 200 with
# `success: true` is the real signal.
on:
push:
branches: [main]
paths:
- "apps/**"
- "packages/**"
- "runs/**"
- "infra/**"
- "schemas/**"
- "wrangler.jsonc"
- "pnpm-lock.yaml"
- ".github/workflows/deploy.yml"
workflow_dispatch:
inputs:
target:
description: "What to deploy"
type: choice
default: all
options:
- all
- substrate
# Never let two deploys race onto the same Worker.
concurrency:
group: deploy-flare-dispatch
cancel-in-progress: false
jobs:
ci:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: pnpm/action-setup@v4
- uses: actions/setup-node@v4
with:
node-version: 22
cache: pnpm
- run: pnpm install --frozen-lockfile
- run: pnpm lint
- run: pnpm typecheck
- run: pnpm test
# The substrate deploys BEFORE the dispatcher, and its own job rather than a
# step in `deploy`, because the ordering is a real constraint and not a
# preference: consumers' service bindings resolve against a Worker that must
# already exist (apps/substrate/specs/platform.md). `needs:` is what encodes
# it — a step ordering inside one job would be lost the moment someone adds a
# matrix or reorders for speed.
#
# It also means a substrate deploy that fails stops the dispatcher's, which is
# the correct blast radius: half a topology is worse than none of it.
substrate:
needs: ci
runs-on: ubuntu-latest
outputs:
url: ${{ steps.deploy.outputs.url }}
steps:
- uses: actions/checkout@v4
- uses: pnpm/action-setup@v4
- uses: actions/setup-node@v4
with:
node-version: 22
cache: pnpm
- run: pnpm install --frozen-lockfile
# Schema before code: the deployed worker reads tables this applies, and
# a hand-run migration is one that the next org's deploy cannot reproduce.
# Already-applied migrations are a no-op.
- name: Apply D1 migrations
run: pnpm exec wrangler d1 migrations apply flare-dispatch-substrate --remote -c apps/substrate/wrangler.jsonc
env:
CLOUDFLARE_API_TOKEN: ${{ secrets.CLOUDFLARE_API_TOKEN }}
CLOUDFLARE_ACCOUNT_ID: ${{ secrets.CLOUDFLARE_ACCOUNT_ID }}
# The deployed hostname is read back out of wrangler's own output rather
# than configured: the workers.dev subdomain is per-account, and a BYOC
# org forking this workflow should not have to edit it. `SUBSTRATE_URL`
# overrides for an org serving the substrate on a custom domain.
- name: Deploy the substrate
id: deploy
shell: bash
run: |
set -euo pipefail
pnpm exec wrangler deploy -c apps/substrate/wrangler.jsonc | tee deploy.log
url="${SUBSTRATE_URL:-}"
if [ -z "$url" ]; then
url="$(grep -oE 'https://flare-dispatch-substrate\.[A-Za-z0-9-]+\.workers\.dev' deploy.log | head -1)"
fi
if [ -z "$url" ]; then
echo "could not determine the substrate URL — set the SUBSTRATE_URL repo variable" >&2
exit 1
fi
echo "url=${url}" >> "$GITHUB_OUTPUT"
env:
CLOUDFLARE_API_TOKEN: ${{ secrets.CLOUDFLARE_API_TOKEN }}
CLOUDFLARE_ACCOUNT_ID: ${{ secrets.CLOUDFLARE_ACCOUNT_ID }}
SUBSTRATE_URL: ${{ vars.SUBSTRATE_URL }}
# The gate ADR-0011 asks for: a container fetch to an unlisted host must die
# 520 on THIS build before any consumer talks to it. A red canary means the
# deny-all posture did not survive the deploy — most likely an SDK bump moving
# the interception semantics the class configuration depends on — and the
# dispatcher below never deploys against it.
canary:
needs: substrate
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: SDK-pin canary
run: ./apps/substrate/scripts/verify-deploy.sh "$SUBSTRATE_URL" canary
env:
SUBSTRATE_URL: ${{ needs.substrate.outputs.url }}
# /health only answers `ok` once the canary above has passed for this
# deployment, so this asserts the health surface agrees with the verdict.
- name: Health
run: ./apps/substrate/scripts/verify-deploy.sh "$SUBSTRATE_URL" health
env:
SUBSTRATE_URL: ${{ needs.substrate.outputs.url }}
# The scratch-consumer round trip against a real container and a real clone.
# Parallel with the dispatcher deploy rather than ahead of it: it exercises
# the substrate's own facade, and a GitHub outage during the clone should show
# up as a red smoke test, not as a blocked release.
dogfood:
needs: [substrate, canary]
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Facade round trip
run: ./apps/substrate/scripts/verify-deploy.sh "$SUBSTRATE_URL" dogfood
env:
SUBSTRATE_URL: ${{ needs.substrate.outputs.url }}
deploy:
needs: [substrate, canary]
if: github.event.inputs.target != 'substrate'
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: pnpm/action-setup@v4
- uses: actions/setup-node@v4
with:
node-version: 22
cache: pnpm
- run: pnpm install --frozen-lockfile
# Schema before code — the same rule the substrate job states, which the
# dispatcher's own database was never held to. Its five migrations were
# applied by hand in a single run on 2026-07-11 and nothing here has ever
# applied one since, so a PR that adds a migration ships code querying a
# column production does not have. Already-applied migrations are a no-op,
# which is what makes this safe to add after the fact.
- name: Apply D1 migrations
run: pnpm exec wrangler d1 migrations apply flare-dispatch --remote
env:
CLOUDFLARE_API_TOKEN: ${{ secrets.CLOUDFLARE_API_TOKEN }}
CLOUDFLARE_ACCOUNT_ID: ${{ secrets.CLOUDFLARE_ACCOUNT_ID }}
# Deploy from the repo root, which wrangler.jsonc points at
# (main: apps/dispatcher/src/index.ts). Uses the pnpm-pinned wrangler.
- run: pnpm exec wrangler deploy
env:
CLOUDFLARE_API_TOKEN: ${{ secrets.CLOUDFLARE_API_TOKEN }}
CLOUDFLARE_ACCOUNT_ID: ${{ secrets.CLOUDFLARE_ACCOUNT_ID }}