Repository navigation
feat: vendor the Open Computer Use source so a build needs no upstream #1010
Workflow file for this run
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| name: Rust | |
| # BR-70 — the repo's first `cargo` CI. Two orthogonal axes: | |
| # * cross-check — does the code we SHIP (windows-gnu / linux-gnu) COMPILE, | |
| # including #[cfg(test)] code, using the release's own docker | |
| # recipe (scripts/cross-env.sh)? This is the PR gate. | |
| # * test — does the platform BEHAVIOUR work on real kernels (macOS | |
| # Seatbelt, Linux Landlock, Windows PowerShell + taskkill)? | |
| # A nightly job runs the full cross BUILD — link errors, the glibc floor, and | |
| # the shipped packages' runtime-library contract — none of which a `cargo check` | |
| # can see, because a check never links and so never emits a NEEDED entry or a | |
| # symbol version. | |
| on: | |
| pull_request: | |
| push: | |
| branches: [main] | |
| schedule: | |
| - cron: "0 7 * * *" # nightly: the expensive full cross-BUILD + glibc floor | |
| # A schedule-only job is a job nobody can check. cross-build-nightly has run | |
| # 36 times since 2026-07-15 — its entire existence — for 33 failures and 3 | |
| # cancellations and not one green. Part of why that stood for a month is that | |
| # there was no way to ASK it a question: you could read the log the morning | |
| # after, change something, and then wait a day to learn whether the change | |
| # helped. Manual dispatch turns a 24-hour edit loop into a 20-minute one. It | |
| # runs every job in this file, the nightly included; the three PR-gate jobs | |
| # are cheap next to the cross build and their signal is never unwelcome. Read | |
| # cross-build-nightly's OWN conclusion rather than the run's, though — the | |
| # matrix carries whatever is red on main at the time (windows-latest was, the | |
| # day this was written), and that says nothing about the cross build. | |
| workflow_dispatch: | |
| # Keyed by event as well as ref, and the event half is load-bearing. `github.ref` | |
| # for the nightly is refs/heads/main — the SAME group a push to main lands in — | |
| # so with cancel-in-progress a morning merge silently killed the night's cross | |
| # build. Three scheduled runs died that way (2026-08-01, 08-07, 08-09) and were | |
| # indistinguishable in the run list from a real failure. Adding manual dispatch | |
| # would have made it worse: pressing the button on main would cancel whatever | |
| # push-triggered gate was in flight. Separating by event keeps the behaviour that | |
| # was actually wanted — a newer push supersedes an older push on the same ref — | |
| # without letting three unrelated kinds of run compete for one slot. | |
| concurrency: | |
| # A group holds at most ONE pending run by default: a newly queued run cancels | |
| # the pending one, whatever cancel-in-progress says (that governs only the | |
| # running one). Keyed by ref, a merge burst to main cancels its own queued | |
| # runs and leaves those commits with no status at all. So only a PR's runs | |
| # share a group, where the newest should win; any other run is keyed by | |
| # run_id, not sha, which repeat dispatches of one commit share. | |
| group: rust-${{ github.event_name }}-${{ github.event_name == 'pull_request' && github.ref || github.run_id }} | |
| cancel-in-progress: ${{ github.event_name == 'pull_request' }} | |
| env: | |
| CARGO_TERM_COLOR: always | |
| jobs: | |
| # ── 1. Cheap guards ──────────────────────────────────────────────────────── | |
| guards: | |
| # ⚠ A job with no timeout burns GitHub's SIX HOUR ceiling when it hangs, | |
| # and this workflow has done exactly that: `test (macos-latest)` finished | |
| # BUILDING in ~13 minutes and then sat silent for 5h46m on three | |
| # `subagent_handler` tests that await `run_complete_subagent_task` | |
| # unbounded, until the ceiling cancelled it. Same three names on a feature | |
| # branch and on clean main; not reproducible on a developer's own macOS. | |
| # | |
| # These budgets are deliberately well ABOVE the measured maxima (warm | |
| # cache, 2026-08-27: guards 0.1m, cross-check 8.2m, test 15.0m, serve | |
| # 9.5m) so a merely slow run is never failed -- they exist to make a HANG | |
| # cheap and legible, not to police duration. A deadlock now costs 40 | |
| # minutes of runner time instead of six hours, on every branch that | |
| # pushes. | |
| timeout-minutes: 10 | |
| if: github.event_name != 'schedule' | |
| runs-on: ubuntu-latest | |
| steps: | |
| - uses: actions/checkout@v4 | |
| - name: Cross recipes have not drifted / glibc floor pin intact | |
| run: ./scripts/check-no-cross-drift.sh | |
| - name: Version consistency (CLI = daemon = GUI) | |
| run: ./scripts/check-version-consistency.sh | |
| # The capability is "Computer Use" wherever a person reads it -- never the | |
| # removed "Computer Controller", never vendor-qualified. The internal | |
| # `computercontroller` identifier is deliberately exempt; see the script. | |
| - name: Computer Use naming consistency | |
| run: ./scripts/check-computer-use-naming.sh | |
| # The vendored upstream Computer Use source must match its manifest AND be | |
| # fully committable. Repository-wide .gitignore rules matched 15 of its | |
| # files when it was first vendored; a tree that verifies locally while | |
| # those files never reach a fresh clone is worse than one that fails here. | |
| - name: Vendored Computer Use source | |
| run: ./scripts/check-vendored-computer-use.sh | |
| # ── 2. Cross COMPILE gate — the BR-70 core. Docker, the SAME images/flags as | |
| # scripts/release.sh (both source scripts/cross-env.sh). ubuntu-latest is | |
| # just the docker HOST; the TARGET triple is what is being proven. | |
| cross-check: | |
| timeout-minutes: 30 | |
| if: github.event_name != 'schedule' | |
| runs-on: ubuntu-latest | |
| strategy: | |
| fail-fast: false | |
| matrix: | |
| target: [x86_64-unknown-linux-gnu, x86_64-pc-windows-gnu] | |
| steps: | |
| - uses: actions/checkout@v4 | |
| # Registry + per-triple target dir, persisted across runs. The per-triple | |
| # target dir is why this is minutes, not a quarter-hour — and it makes the | |
| # mac/linux build-script clobber (release.sh) structurally impossible here. | |
| - name: Restore cross cache | |
| uses: actions/cache@v4 | |
| with: | |
| path: | | |
| .cross-cache/registry | |
| .cross-cache/target/${{ matrix.target }} | |
| key: cross-${{ matrix.target }}-${{ hashFiles('Cargo.lock', 'rust-toolchain.toml', 'scripts/cross-env.sh') }} | |
| restore-keys: cross-${{ matrix.target }}- | |
| - name: cargo check --workspace --all-targets (${{ matrix.target }}) | |
| run: | | |
| set -euo pipefail | |
| mkdir -p .cross-cache/registry ".cross-cache/target/${{ matrix.target }}" | |
| export CROSS_REGISTRY_MOUNT="$PWD/.cross-cache/registry" | |
| export CROSS_TARGET_MOUNT="$PWD/.cross-cache/target/${{ matrix.target }}" | |
| . scripts/cross-env.sh | |
| case "${{ matrix.target }}" in | |
| x86_64-unknown-linux-gnu) | |
| cross_linux "cargo check --workspace --all-targets --locked" /cross-target ;; | |
| x86_64-pc-windows-gnu) | |
| cross_windows "cargo check --workspace --all-targets --locked" /cross-target ;; | |
| esac | |
| # ── 3. Native unit tests on REAL kernels. This is where platform BEHAVIOUR is | |
| # proven: PowerShell selection + taskkill (windows), Seatbelt (macos), | |
| # Landlock (ubuntu, once BR-69 slice 2 lands — this runner is its venue). | |
| # NOTE: windows-latest is MSVC; we SHIP mingw/gnu. Job 2 covers the gnu | |
| # ABI, this job covers the Windows behaviour. Both are load-bearing. | |
| test: | |
| timeout-minutes: 40 | |
| if: github.event_name != 'schedule' | |
| strategy: | |
| fail-fast: false | |
| matrix: | |
| os: [ubuntu-latest, macos-latest, windows-latest] | |
| runs-on: ${{ matrix.os }} | |
| steps: | |
| - uses: actions/checkout@v4 | |
| - uses: dtolnay/rust-toolchain@master | |
| with: | |
| toolchain: "1.92" | |
| components: clippy | |
| # MSVC's test PDBs dominate the Windows link and cache footprint (the | |
| # biorouter harness alone produced a ~1.1 GiB PDB). This keeps debug | |
| # assertions, overflow checks and panic locations while omitting debugger | |
| # symbols that CI never consumes. Set it before rust-cache so the action's | |
| # CARGO_* environment hash separates these artifacts from debug=2 builds. | |
| # | |
| # macOS is included for a different reason: it is the only runner in this | |
| # matrix with NO Rust cache (see the cache-on-failure note below), so it | |
| # pays a cold full-workspace build every run and its debug symbols are | |
| # what would push the repository over GitHub's 10 GiB cache ceiling once | |
| # it finally saves one. | |
| # | |
| # ⚠ Ubuntu was left out while it had a warm 1.7 GiB cache and no problem | |
| # for this to fix. It has one now: the integration-binary step below | |
| # links ~117 more executables, and on Linux every one of them EMBEDS the | |
| # DWARF of everything it links rather than leaving it in a side file. | |
| # Measured 2026-09-11 (run 34628149105) with debug=2: that build took the | |
| # runner from 70 GB free to 1.4 GB and died with `No space left on | |
| # device` before it finished. The variable changes rust-cache's key, so | |
| # ubuntu builds cold (measured: lib + bins 3.7 -> 7.5 min) until main's | |
| # first run with it saves the new cache. A one-time price. | |
| - name: Reduce CI test debuginfo | |
| shell: bash | |
| run: echo 'CARGO_PROFILE_TEST_DEBUG=0' >> "$GITHUB_ENV" | |
| # ⚠ `cache-on-failure` is what breaks a self-perpetuating deadlock, not a | |
| # convenience. rust-cache's post step is skipped whenever the job does not | |
| # succeed, so a job that cannot finish never saves a cache, is cold on the | |
| # next run, and therefore cannot finish. That is exactly what happened to | |
| # `test (macos-latest)`: it was cancelled on main at d0a177f8 after 360 | |
| # minutes with its post step skipped, and `gh cache list` subsequently | |
| # showed no `v0-rust-macos-latest-*` entry at all while ubuntu and windows | |
| # both had one. Every macOS run since has compiled the whole workspace | |
| # from scratch — 56+ minutes against a 20 minute budget, versus 17.8m on | |
| # the last run that had a warm cache (44ac420f). | |
| # | |
| # Compiled dependencies are valid whether or not the *tests* passed, so | |
| # saving them on a red job costs nothing and means a broken test can never | |
| # again strand the cache. | |
| # ⚠ `save-if` restricts SAVING to main, and that is what keeps a cache | |
| # available to anyone at all. GitHub scopes a cache to the branch that | |
| # wrote it: a feature branch can read its own caches and **main's**, and | |
| # nothing else. So every feature-branch run was writing its own private | |
| # 1-1.7 GiB copy per OS that no other branch could ever use, while the | |
| # repository sat at 9.74 GiB against a 10 GiB ceiling — evicting main's | |
| # copies, which are the only ones shared. Measured on windows-latest, by | |
| # rust-cache's own restore time: | |
| # | |
| # restore 0.7m -> `cargo test` 9.4m (hit, branch's own earlier cache) | |
| # restore 0.1m -> `cargo test` 19.5m (miss) | |
| # restore 0.1m -> `cargo test` 17.4m (miss) | |
| # | |
| # With saving confined to main, every branch restores one authoritative, | |
| # continuously refreshed cache and none of them compete for the ceiling. | |
| # A PR that changes Cargo.lock still builds cold and cannot save; that is | |
| # the accepted cost, and it is much smaller than three OSes' worth of | |
| # write-only caches evicting the shared ones on every push. | |
| - uses: Swatinem/rust-cache@v2 | |
| with: | |
| key: ${{ matrix.os }} | |
| cache-on-failure: "true" | |
| save-if: ${{ github.ref == 'refs/heads/main' }} | |
| - name: Install protoc (build-script dep) | |
| uses: arduino/setup-protoc@v3 | |
| with: | |
| repo-token: ${{ secrets.GITHUB_TOKEN }} | |
| - name: Install Linux native build dependencies | |
| if: matrix.os == 'ubuntu-latest' | |
| run: | | |
| # ⚠ Drop the runner image's Microsoft apt repositories FIRST. | |
| # `packages.microsoft.com` has repeatedly answered 403 for the | |
| # `azure-cli` and `prod` lists that the hosted image ships | |
| # preinstalled, and `apt-get update` exits 100 on any repository it | |
| # cannot refresh — so an outage on a third-party CDN we do not use | |
| # fails this job with `E: ... is no longer signed`. Nothing here | |
| # installs from them. Removing them makes the update depend only on | |
| # the archives the packages below actually come from. | |
| sudo rm -f /etc/apt/sources.list.d/microsoft-prod.list \ | |
| /etc/apt/sources.list.d/azure-cli.list \ | |
| /etc/apt/sources.list.d/microsoft-prod.sources \ | |
| /etc/apt/sources.list.d/azure-cli.sources | |
| sudo apt-get update | |
| sudo apt-get install --yes libdbus-1-dev libxcb1-dev pkg-config | |
| # Start narrow (lib + bins on every OS) so the matrix is green on arrival; | |
| # widen as it proves stable. This selector leaves out `tests/`; those | |
| # binaries run in their own step below, ubuntu only, with no network. | |
| # (They were long assumed to need cassettes and credentials the runners | |
| # lack. Measured, they do not: `mcp_integration_test` REPLAYS its | |
| # cassettes offline and only records under BIOROUTER_RECORD_MCP, and the | |
| # tests that need a real provider or vendor CLI are `#[ignore]`d.) | |
| # | |
| # ⚠ It does NOT follow that nothing in this step touches the network. The | |
| # selector filters by TARGET, not by behaviour, and a test that lives in | |
| # `src/` is a unit test as far as cargo is concerned however far it dials. | |
| # `biorouter-server/src/tunnel/lapstone_test.rs` did exactly that against a | |
| # third-party Cloudflare worker, and because `biorouter-server/src/main.rs` | |
| # re-declares `mod tunnel;` rather than using the library, both `--lib` and | |
| # `--bins` carried it: 2 tests x 2 targets x 3 platforms, twelve live | |
| # requests per run, red whenever that worker had a bad day. It is now | |
| # opt-in behind BIOROUTER_TEST_LAPSTONE_TUNNEL. Gate the test, not the | |
| # selector, if another one appears. | |
| # ⚠ `--no-fail-fast` is load-bearing on this matrix, not a nicety. Without | |
| # it cargo stops at the FIRST test binary that fails, and everything | |
| # ordered after it never runs. On windows-latest that was `biorouter_mcp`, | |
| # so `biorouter_sandbox`, `biorouter_server`, `biorouterd` and | |
| # `biorouter_test` were unreported for a hundred consecutive runs: each | |
| # batch of Windows fixes made the job fail one binary further along, which | |
| # read as "a new rotating set of failures" and was really one hidden queue | |
| # being drained one CI round-trip at a time. One run now names them all. | |
| # ⚠ **The OS credential store must be off, or macOS hangs until the job | |
| # is killed.** `Config::read` resolves a secret through `keyring`, and on | |
| # a runner with no unlocked login keychain `SecKeychainFindGenericPassword` | |
| # blocks on an authorization prompt nothing is there to answer. The | |
| # secrets tests in `routes::action_required` sit on that call, so the | |
| # whole binary stops: `test (macos-latest)` ran SIX HOURS on main | |
| # (run 32997530845) and was cancelled. Reproduced locally and sampled — | |
| # `Config::read` -> `keyring::Entry::get_password` -> | |
| # `SecKeychainFindGenericPassword` -> `CSSM_DecryptData`, blocked. With | |
| # this set the same binary finishes in seconds. | |
| # | |
| # Windows reaches the same code through the Credential Manager. Whether | |
| # that is what has made windows-latest slow is NOT measured, so this is | |
| # not offered as the explanation there — it is set on the whole matrix | |
| # because the store is unwanted on every runner, not as a Windows fix. | |
| # | |
| # It costs no coverage: nothing in the workspace asserts keyring | |
| # behaviour — the store is incidental to these tests, which are about | |
| # routes. `serve` below and `release-artifact-smoke.yml` already set it | |
| # for the same reason; this is the one job that was missing it. | |
| - name: cargo test (workspace, lib + bins) | |
| env: | |
| BIOROUTER_DISABLE_KEYRING: "true" | |
| run: cargo test --workspace --lib --bins --locked --no-fail-fast --timings | |
| - name: Upload Windows Cargo timings | |
| if: matrix.os == 'windows-latest' && always() | |
| uses: actions/upload-artifact@v4 | |
| with: | |
| name: windows-cargo-timings | |
| path: target/cargo-timings/cargo-timing.html | |
| if-no-files-found: warn | |
| retention-days: 14 | |
| # The cross-platform-sensitive suites BR-70 exists to prove, run on every | |
| # OS they apply to. The workspace step above already runs the sandbox, | |
| # developer shell/background and security unit tests through their library | |
| # targets. The sandbox's integration target is the one `tests/` binary | |
| # that runs on every OS (the rest follow in the next step, ubuntu only); | |
| # name it directly rather than executing those unit tests a second time. | |
| - name: cargo test (sandbox integration) | |
| run: cargo test -p biorouter-sandbox --test sandbox --locked --no-fail-fast | |
| # ── Every other `tests/*.rs` binary, bar a table of exceptions ───────── | |
| # | |
| # Until this step no workflow ran them. The lib + bins selector above | |
| # leaves `tests/` out, and `clippy --all-targets` below only COMPILES | |
| # them, which says nothing about whether they pass. So the guards | |
| # CLAUDE.md treats as load-bearing (the privacy master-switch binaries, | |
| # the keyless-daemon binaries, `privacy_ar15_is_retired`, | |
| # `every_test_binary_is_sandboxed`) ran only when someone remembered to. | |
| # The cost is on record: from eb594ded until PR #229, | |
| # `the_documented_closure_is_the_one_the_code_performs` read the wrong | |
| # function, and would have stayed green with the AR-15 gate deleted. | |
| # | |
| # ⚠ The table is an EXCLUSION list on purpose. An allowlist would rebuild | |
| # the gap one file at a time: a new binary would run only if its author | |
| # remembered to add it. This way it runs the day it lands, and one that | |
| # cannot run hermetically gets a line saying why. A line naming a target | |
| # that no longer exists fails the step, so the table cannot outlive what | |
| # it excuses. | |
| # | |
| # ⚠ Never add `--ignored` or `--include-ignored`. The live tests (real | |
| # providers, the vendor CLIs, llama.cpp downloads, GitHub fetches) are | |
| # `#[ignore]`d inside binaries whose other tests belong here, so the | |
| # binary stays in and its live tests stay out. | |
| # | |
| # ⚠ No network, enforced rather than assumed. The tests run in a network | |
| # namespace holding nothing but loopback, so a test that dials out fails | |
| # at once with a connection error instead of passing whenever the third | |
| # party is up (the lapstone note above is what that looks like). | |
| # `setpriv` drops back to the runner user once loopback is up: as root, | |
| # every "permission denied" assertion would pass for the wrong reason. | |
| # Servers on 127.0.0.1 and ::1 work as usual. | |
| # | |
| # The table's second entry is the instructive one. `ui_example_apps` | |
| # bundles each example with `npx --yes esbuild` when no local esbuild | |
| # exists, and this job installs none. Offline it does NOT fail: measured, | |
| # npx sits on the unreachable registry until the bundler's 60 s timeout, | |
| # then the test passes on the type-stripper fallback without esbuild | |
| # ever running. That is a minute spent proving nothing, and a pass that | |
| # turns into a failure the day npm's retries give up before the timeout. | |
| # | |
| # Measured 2026-09-11, one binary at a time on hosted runners (scratch | |
| # runs 34628149105 and 34629544896). ubuntu, inside the namespace: all | |
| # 119 binaries green, 687 tests passed, 0 failed, 49 ignored. Building | |
| # them took 135 s on top of the lib + bins build (14.5 GiB of | |
| # executables, with the debuginfo setting above) and the 117 this step | |
| # selects spend ~135 s in their tests. macOS was green but for one | |
| # log-capture test that failed once in ~70 runs. windows-latest fails | |
| # 31 tests in 9 binaries that had never run there, which is why this is | |
| # ubuntu only: widening means fixing those first, and `unshare` is | |
| # Linux-only, so another OS also needs its own way to stay offline. | |
| - name: cargo test (integration binaries, no network) | |
| if: matrix.os == 'ubuntu-latest' | |
| env: | |
| BIOROUTER_DISABLE_KEYRING: "true" | |
| CARGO_NET_OFFLINE: "true" | |
| run: | | |
| set -euo pipefail | |
| # One line per target, no wrapping: the first word is the name. | |
| # target why it does not run in this step | |
| excluded=' | |
| sandbox "cargo test (sandbox integration)" above runs it, on every OS | |
| ui_example_apps needs the npm registry (`npx --yes esbuild`); apps-smoke.yml runs it after `npm ci` | |
| ' | |
| all=$(cargo metadata --format-version 1 --no-deps --locked \ | |
| | jq -r '.packages[].targets[] | select(.kind == ["test"]) | .name' | sort) | |
| # `--test` selects by name across the workspace, so a name two | |
| # packages share would run (or be excluded) twice over. | |
| dup=$(uniq -d <<<"$all") | |
| if [ -n "$dup" ]; then | |
| echo "::error::two packages have an integration target named: $dup"; exit 1 | |
| fi | |
| skip=$(awk 'NF { print $1 }' <<<"$excluded" | sort) | |
| stale=$(comm -13 <(echo "$all") <(echo "$skip")) | |
| if [ -n "$stale" ]; then | |
| echo "::error::the exclusion table names targets that no longer exist; delete their lines: $stale"; exit 1 | |
| fi | |
| awk 'NF { t = $1; $1 = ""; printf "not run here: %s --%s\n", t, $0 }' <<<"$excluded" | |
| args=() | |
| for t in $(comm -23 <(echo "$all") <(echo "$skip")); do args+=(--test "$t"); done | |
| echo "running $(( ${#args[@]} / 2 )) integration binaries" | |
| sudo -E unshare --net -- sh -c \ | |
| 'ip link set lo up && exec setpriv --reuid="$SUDO_UID" --regid="$SUDO_GID" --init-groups -- "$@"' sh \ | |
| env HOME="$HOME" PATH="$PATH" \ | |
| cargo test --workspace --locked --no-fail-fast "${args[@]}" | |
| - name: clippy (host, informational until warnings are burned down) | |
| if: matrix.os == 'ubuntu-latest' | |
| run: cargo clippy --workspace --all-targets --locked | |
| # ── 4. Nightly: the FULL cross BUILD. `check` cannot catch link errors | |
| # (aws-lc-sys / winpthread) or glibc symbol-version regressions. | |
| cross-build-nightly: | |
| timeout-minutes: 120 | |
| # The three jobs above are gated `!= 'schedule'`, so they pick up | |
| # workflow_dispatch on their own; this one has to name it, or the button | |
| # would run everything EXCEPT the job it was added for. | |
| if: github.event_name == 'schedule' || github.event_name == 'workflow_dispatch' | |
| runs-on: ubuntu-latest | |
| steps: | |
| - uses: actions/checkout@v4 | |
| - name: Build shipped binaries for both cross targets (release command) | |
| run: | | |
| set -euo pipefail | |
| . scripts/cross-env.sh | |
| cross_linux "cargo build --release --bin biorouterd --bin biorouter" | |
| cross_windows "cargo build --release --bin biorouterd --bin biorouter" "" "$WIN_DLL_STAGE" | |
| - name: Assert the glibc floor (2.31 / bullseye) is not raised | |
| run: ./scripts/check-glibc-floor.sh | |
| # The glibc floor check covers the libraries the binaries need from libc. | |
| # This covers the ones they need from everywhere ELSE — today exactly one, | |
| # libxcb.so.1 — and asserts that the packages we ship name it, so apt/dnf | |
| # put it on the box before the binary ever runs. See the script's header; | |
| # the short version is that a new system dependency is invisible to every | |
| # other gate in this file and surfaces as exit 127 on a user's machine. | |
| - name: The shipped binaries' runtime deps are declared by the shipped packages | |
| id: runtime-deps | |
| run: | | |
| set -euo pipefail | |
| ./scripts/check-linux-runtime-deps.sh | |
| echo "deb_packages=$(./scripts/check-linux-runtime-deps.sh --print-deb-packages)" >> "$GITHUB_OUTPUT" | |
| # This step used to `docker run` the bare ELF out of target/ in a stock | |
| # debian:bullseye, and had therefore failed EVERY night since it was | |
| # written — ~40 consecutive red runs, never once green. The failure was | |
| # real and reproducible and meant nothing: `biorouter` links libxcb.so.1 | |
| # (arboard's clipboard, xcap's screen capture), the base image ships no | |
| # X11 libraries at all, so the loader aborted before `main`. It was | |
| # reporting a catastrophe that no user can reach, because no user runs the | |
| # naked binary — they install biorouter-cli_*.deb or the .rpm, and those | |
| # DECLARE libxcb1/libxcb, so the package manager installs it as a matter | |
| # of course. Measured on a clean debian:bullseye and a clean rockylinux:9: | |
| # both packages install and both binaries report their version. | |
| # | |
| # What the step proves now is the thing that was worth proving all along: | |
| # that the declared dependency set is SUFFICIENT. It installs exactly the | |
| # packages the shipped specs name — derived from the ELF by the guard | |
| # above rather than hardcoded here, so the list cannot go stale in one | |
| # place while moving in the other — and nothing else. If a future crate | |
| # adds a system library, the guard above fails first and names the fix; if | |
| # the declared set were merely incomplete, this boot fails and says so. | |
| # Both binaries are booted: they ship together and link the same set, and | |
| # only `biorouter` was ever being checked. | |
| # | |
| # bullseye is the floor pin (GLIBC_FLOOR / LINUX_RUST_IMG in | |
| # scripts/cross-env.sh). Its apt mirrors are what make the install step | |
| # work; when Debian 11 finally leaves them, this image moves forward | |
| # together with that pin, not on its own. | |
| - name: Boot both Linux binaries on the oldest supported distro | |
| env: | |
| DEB_PACKAGES: ${{ steps.runtime-deps.outputs.deb_packages }} | |
| run: | | |
| set -euo pipefail | |
| docker run --rm -e DEB_PACKAGES -v "$PWD":/w debian:bullseye sh -euxc ' | |
| apt-get update -q | |
| apt-get install -y --no-install-recommends $DEB_PACKAGES | |
| /w/target/x86_64-unknown-linux-gnu/release/biorouter --version | |
| /w/target/x86_64-unknown-linux-gnu/release/biorouterd --version | |
| ' | |
| # `biorouter serve` end to end: the daemon serving the interface, the token | |
| # exchange, and the endpoints that were previously unauthenticated. | |
| # | |
| # Before this job there was NO CI coverage of browser mode at all -- `grep -rn | |
| # headless .github/` returned nothing while a tarball of it was a shipped | |
| # release asset. The one smoke test that existed asserted the served page | |
| # contained "<!doctype html>", which is true of a completely non-functional | |
| # app. | |
| # | |
| # It cannot drive a real agent turn: there are no provider credentials here, | |
| # and a browser session may not set a provider by design (decision SD-1 in | |
| # docs/deployment/serve-decisions.md). What it does assert is everything | |
| # between the browser and that boundary. | |
| serve: | |
| # Was 30 against a measured 9.5m. The lifecycle step below builds | |
| # biorouter-cli's tests, which unify its dev-dependency features and so | |
| # recompile part of the graph a plain `cargo build` already built (4m56s on | |
| # a loaded M4 Max); the budget keeps the same margin over that. | |
| timeout-minutes: 45 | |
| if: github.event_name != 'schedule' | |
| runs-on: ubuntu-latest | |
| steps: | |
| - uses: actions/checkout@v4 | |
| - uses: dtolnay/rust-toolchain@master | |
| with: | |
| toolchain: "1.92" | |
| - uses: Swatinem/rust-cache@v2 | |
| with: | |
| key: serve | |
| - uses: actions/setup-node@v4 | |
| with: | |
| node-version: 24 | |
| - name: Install protoc (build-script dep) | |
| uses: arduino/setup-protoc@v3 | |
| with: | |
| repo-token: ${{ secrets.GITHUB_TOKEN }} | |
| - name: Install Linux native build dependencies | |
| run: | | |
| # ⚠ Drop the runner image's Microsoft apt repositories FIRST. | |
| # `packages.microsoft.com` has repeatedly answered 403 for the | |
| # `azure-cli` and `prod` lists that the hosted image ships | |
| # preinstalled, and `apt-get update` exits 100 on any repository it | |
| # cannot refresh — so an outage on a third-party CDN we do not use | |
| # fails this job with `E: ... is no longer signed`. Nothing here | |
| # installs from them. Removing them makes the update depend only on | |
| # the archives the packages below actually come from. | |
| sudo rm -f /etc/apt/sources.list.d/microsoft-prod.list \ | |
| /etc/apt/sources.list.d/azure-cli.list \ | |
| /etc/apt/sources.list.d/microsoft-prod.sources \ | |
| /etc/apt/sources.list.d/azure-cli.sources | |
| sudo apt-get update -q | |
| sudo apt-get install -y --no-install-recommends libdbus-1-dev libxdo-dev | |
| - name: Build the interface bundle | |
| working-directory: ui/desktop | |
| run: | | |
| npm ci | |
| npm run build:web | |
| # Root base, not Forge's relative base: a relative-base bundle served | |
| # at the root resolves its assets against the current path. | |
| grep -q 'src="/assets/' src/web/index.html | |
| ! grep -q 'src="./assets/' src/web/index.html | |
| - name: Build the command and the daemon | |
| run: cargo build -p biorouter-cli -p biorouter-server | |
| - name: Serve, and check the whole browser contract | |
| run: | | |
| set -euo pipefail | |
| export BIOROUTER_PATH_ROOT="$RUNNER_TEMP/serve-root" | |
| export BIOROUTER_DISABLE_KEYRING=true | |
| mkdir -p "$BIOROUTER_PATH_ROOT" | |
| ./target/debug/biorouter serve \ | |
| --port 18777 --token citoken \ | |
| --web-dir "$PWD/ui/desktop/src/web" > "$RUNNER_TEMP/serve.log" 2>&1 & | |
| serve_pid=$! | |
| for _ in $(seq 1 90); do | |
| curl -fsS -o /dev/null "http://127.0.0.1:18777/?t=citoken" && break | |
| sleep 1 | |
| done | |
| fail() { echo "FAILED: $1"; echo '--- serve log ---'; cat "$RUNNER_TEMP/serve.log"; exit 1; } | |
| # The document carries the daemon secret, so it is gated. | |
| [ "$(curl -s -o /dev/null -w '%{http_code}' http://127.0.0.1:18777/)" = 401 ] \ | |
| || fail 'the shell was served without a credential' | |
| [ "$(curl -s -o /dev/null -w '%{http_code}' 'http://127.0.0.1:18777/?t=wrong')" = 401 ] \ | |
| || fail 'a wrong token was accepted' | |
| # The token is exchanged for a cookie and leaves the address bar. It is | |
| # not consumed (SD-9 in docs/deployment/serve-decisions.md): the | |
| # readiness loop above has already redeemed it. | |
| curl -s -o /dev/null -D "$RUNNER_TEMP/h" "http://127.0.0.1:18777/?t=citoken" | |
| grep -qi '^HTTP/1.1 303' "$RUNNER_TEMP/h" || fail 'no redirect after the token exchange' | |
| grep -qi 'set-cookie: biorouter_session=' "$RUNNER_TEMP/h" || fail 'no session cookie' | |
| grep -qi 'HttpOnly' "$RUNNER_TEMP/h" || fail 'the session cookie is readable from script' | |
| grep -qi 'SameSite=Strict' "$RUNNER_TEMP/h" || fail 'the session cookie is sent cross-site' | |
| # With the cookie, the shell and its runtime configuration. | |
| curl -fsS -b 'biorouter_session=citoken' http://127.0.0.1:18777/ > "$RUNNER_TEMP/shell.html" | |
| grep -q '__BIOROUTER_HEADLESS_CONFIG__' "$RUNNER_TEMP/shell.html" || fail 'no runtime config' | |
| grep -qE '"secretKey":"[0-9a-f]{64}"' "$RUNNER_TEMP/shell.html" || fail 'no daemon secret' | |
| # ABSENT, not empty: empty is falsy in the renderer and sends the | |
| # browser to a hardcoded 127.0.0.1:3000. | |
| ! grep -q 'apiBaseUrl' "$RUNNER_TEMP/shell.html" \ | |
| || fail 'apiBaseUrl is present; the renderer will not use its own origin' | |
| grep -q 'no-store' <(curl -s -D - -o /dev/null -b 'biorouter_session=citoken' http://127.0.0.1:18777/) \ | |
| || fail 'the shell is cacheable and it carries the secret' | |
| # The bundle is really served. | |
| asset=$(grep -o '/assets/index-[A-Za-z0-9_-]*\.js' "$RUNNER_TEMP/shell.html" | head -1) | |
| [ -n "$asset" ] || fail 'no bundle reference in the shell' | |
| curl -fsS -o /dev/null "http://127.0.0.1:18777$asset" || fail 'the bundle is not served' | |
| # The interface's own endpoints are behind the daemon's auth now. They | |
| # had none at all before. | |
| secret=$(grep -o '"secretKey":"[0-9a-f]*"' "$RUNNER_TEMP/shell.html" | cut -d'"' -f4) | |
| [ "$(curl -s -o /dev/null -w '%{http_code}' http://127.0.0.1:18777/headless/health)" = 401 ] \ | |
| || fail 'an interface endpoint answered without the secret' | |
| curl -fsS -H "X-Secret-Key: $secret" http://127.0.0.1:18777/headless/health \ | |
| | grep -q '"status":"ok"' || fail 'the interface endpoint did not answer with the secret' | |
| # And a path traversal is refused, which the retired implementation | |
| # allowed outright: fs_read had no validation of any kind. | |
| [ "$(curl -s -o /dev/null -w '%{http_code}' -H "X-Secret-Key: $secret" \ | |
| 'http://127.0.0.1:18777/headless/fs/read?path=/etc/passwd')" = 403 ] \ | |
| || fail 'the filesystem endpoint read outside its allowed roots' | |
| # Stopping serve by pid stops its daemon. It used to leave the daemon | |
| # running, still holding this port and still honouring the token. | |
| # `serve` reaps the daemon before it exits, so the port is closed the | |
| # moment `wait` returns. | |
| kill "$serve_pid" || fail 'serve was no longer running' | |
| wait "$serve_pid" || true | |
| ! curl -s -o /dev/null --max-time 5 http://127.0.0.1:18777/status \ | |
| || fail 'the daemon outlived serve' | |
| echo 'the browser contract holds' | |
| # SIGTERM and SIGINT to `serve`, and SIGKILL, which leaves only the | |
| # daemon's own parent watch to stop it. Here rather than in the workspace | |
| # test job, which runs `--lib --bins` only: this needs a `biorouterd` from | |
| # the same tree beside `biorouter`, which the build step above provides. | |
| - name: Stopping serve stops its daemon | |
| env: | |
| BIOROUTER_DISABLE_KEYRING: "true" | |
| run: cargo test -p biorouter-cli --test serve_lifecycle |