Bazel monorepo with BuildBuddy cache
In a monorepo, every change raises the same question: which services does it affect, and which of them need a deploy? This guide answers it for a Bazel workspace that shares BuildBuddy's remote cache.
The example is pipemesh/demo-bazel:
five Java services on three shared libraries, built with Bazel 9 and
rules_java. Its pipelines are public at
pipemesh.io/github.com/pipemesh/demo-bazel.
libs/money ─┬─ orders payments catalog
libs/events ─┼─ orders inventory notifications
libs/http ─┴─ every service
What happens on a change
Measured on the demo:
| Change | Services that receive it | Deploys |
|---|---|---|
| A README edit | none | — |
A comment in MODULE.bazel |
none | — |
| A new test in one service | that service | none: the jar is identical |
A new method in libs/money |
orders, payments, catalog | those three |
With every build step coming from BuildBuddy, a service build's Bazel commands take about 15 s on a warm runner. The fingerprint takes about 7 s.
How it works
The repository declares two kinds of pipeline:
- A dispatch pipeline fingerprints each service and hands it the revision.
- A service pipeline, one copy per service, builds and tests, then
deploys to staging and production. Each copy has its own history,
board and URL, such as
/github.com/pipemesh/demo-bazel/-/pipeline/orders.
Three ideas decide what runs:
- Fingerprints. bazel-diff hashes every target with its sources, its rule and its whole dependency graph, external repositories included. A service's fingerprint digests the hashes of every target in its package, so it changes exactly when what the service builds or tests can change.
- Dispatch with memory. A service gets the revision when its fingerprint differs from the last one it received, not the previous commit's. Held back for ten commits, it still gets all ten when the next change reaches it.
- Deploy only a new jar. If the jar is byte-identical (a test, a sibling's code), the service builds, tests and stops.
BuildBuddy then replays any build or test step already run, on any service's pipeline, from any earlier revision.
Set it up
1. The workspace
Tag each deployable target deployable, which the fingerprint looks for
by default. It fingerprints each one's package, and names the file after
the package's last path segment: //services/orders writes
fingerprints/orders.
java_binary(
name = "server",
main_class = "dev.pipemesh.demo.orders.OrdersService",
tags = ["deployable"],
runtime_deps = [":lib"],
)
Bazel's deploy jars are byte-for-byte the same for the same inputs, on any machine, which is what lets a deploy skip. Dispatch relies on the fingerprint, so it works either way.
The job image needs Bazel (or bazelisk), a C++ toolchain for Bazel's
Java launcher, and the bash, git, curl and tar every hosted
runner image needs. Step
7 builds one.
2. The dispatch pipeline: pipemesh.yaml
service: !include .pipemesh/service.yaml
dispatch:
type: pipeline
stages:
- image
- graph
- dispatch
jobs:
# The CI image (step 7).
build_ci:
job_type: build
stage: image
checkout:
- ci
script: echo "ci/ changed - building the CI image"
publish:
- context: ci
key: ci_image
# Checks out the whole tree: bazel-diff hashes every target.
graph:
job_type: build
stage: graph
uses: bazel/fingerprint@1
image_from: build_ci/ci_image
# External repositories: downloaded once per MODULE.bazel.lock.
cache:
key: bazel-repo-${checksum:MODULE.bazel.lock}
restore_keys:
- bazel-repo-
paths:
- .cache/bazel-repo
with:
bazel: tools/bazelw
extra: .bazelversion .bazelrc .pipemesh/service.yaml deploy
produces:
orders: fingerprints/orders
payments: fingerprints/payments
# … one entry per service
# One child pipeline per service. It checks out nothing, so its
# fingerprint is its only input.
orders:
job_type: pipeline
stage: dispatch
consumes:
- graph/orders
body: !ref service
variables:
SERVICE: orders
# … one dispatch job per service
pipemesh:
pipelines:
pipeline: !ref dispatch
workflows:
checks:
body: !include .pipemesh/checks.yaml
on: pull_request
graph is a build, so it reuses an earlier run when nothing in the
repository changed. service is a type: pipeline body, the shape a
job_type: pipeline job takes. See The definition.
bazel/fingerprint@1 takes these parameters:
| Parameter | Default | What it is |
|---|---|---|
image |
none | The job's image: bash, git and Bazel. Leave it out to run in an image another job built, with image_from: on the job. |
targets |
attr(tags, "\bdeployable\b", //...) |
The query naming the deployable targets. Each one's package gets a fingerprint. It's written inside single quotes, so it can't contain one. |
extra |
.bazelversion .bazelrc |
Paths added to every fingerprint, hashed by content. A directory covers every file under it. Use it for what target hashes can't see: the Bazel release and flags, the service pipeline, the deploy scripts. |
out |
fingerprints |
Where the files go. |
bazel |
bazel |
The Bazel command: a wrapper script such as tools/bazelw, or bazelisk under another name. |
If $BAZEL_DIFF names a bazel-diff in the image, the component uses it.
Otherwise it downloads a pinned release and checks its sha256.
3. The service pipeline: .pipemesh/service.yaml
type: pipeline
stages:
- build
- staging
- production
jobs:
# Checks out the whole tree: Bazel reads MODULE.bazel, .bazelrc,
# tools/ and every package the service depends on.
build:
job_type: build
stage: build
timeout_seconds: 1200
image_from: ../build_ci/ci_image
# External repositories: downloaded once per MODULE.bazel.lock.
cache:
key: bazel-repo-${checksum:MODULE.bazel.lock}
restore_keys:
- bazel-repo-
paths:
- .cache/bazel-repo
secrets:
- BUILDBUDDY_API_KEY
setup:
- uses: bazel/remote-cache@1
script: |
tools/bazelw test //services/$SERVICE/...
tools/bazelw build //services/$SERVICE:server_deploy.jar
mkdir -p dist && cp bazel-bin/services/$SERVICE/server_deploy.jar dist/$SERVICE.jar
produces:
jar:
path: "dist/*.jar"
deploy_staging:
job_type: deploy
production: false
stage: staging
consumes:
- build/jar
checkout:
- deploy
script: deploy/deploy.sh $SERVICE staging
deploy_prod:
job_type: deploy
production: true
stage: production
needs:
- deploy_staging
consumes:
- build/jar
checkout:
- deploy
script: deploy/deploy.sh $SERVICE production
image_from: ../build_ci/ci_imageruns the build in the image the dispatch pipeline built.../names an entry of the parent pipeline.bazel/remote-cache@1points every later Bazel command at BuildBuddy (step 5).- The deploys check out only
deploy/. They run on a new jar or a changed deploy script, and skip otherwise.
4. Pull request checks: .pipemesh/checks.yaml
Pull requests test only the services the change can affect, against the merge base.
type: workflow
stages:
- image
- test
jobs:
build_ci:
job_type: build
stage: image
checkout:
- ci
script: echo "ci/ changed - building the CI image"
publish:
- context: ci
key: ci_image
# A task, so it runs on every pull request revision: what it tests
# depends on the merge base, not only on the tree.
affected_tests:
job_type: task
stage: test
checkout: true
image_from: build_ci/ci_image
cache:
key: bazel-repo-${checksum:MODULE.bazel.lock}
restore_keys:
- bazel-repo-
paths:
- .cache/bazel-repo
script: |
base=$(git merge-base "origin/$CI_MERGE_REQUEST_TARGET_BRANCH_NAME" HEAD)
services=$(tools/affected.sh "$base" HEAD)
if [ -z "$services" ]; then echo "no service affected"; exit 0; fi
echo "affected: $services"
tools/bazelw test $(printf '//services/%s/... ' $services)
- A pull request that changes
ci/builds and tests its own CI image. Any other reuses the default branch's. tools/affected.shis the classic Bazel query: everydeployabletarget that depends, transitively (rdeps), on a package the diff touched. A change Bazel can't place in a package (MODULE.bazel,.bazelrc) tests every service.- The check reports as
pipemesh/checks/affected_tests, so it can be a required check. It works with GitHub's merge queue (see GitHub integration).
5. Connect BuildBuddy
Only the default branch writes to the cache:
- In BuildBuddy, create an API key for CI with read and write access to the cache (your organization's Settings).
- In Pipemesh, store it as the secret
BUILDBUDDY_API_KEY(the repository's Settings). - Only the service
buildjob lists it insecrets:, and that job runs only for revisions on the default branch.
Pull-request runs never receive a secret unless the secret allows it.
bazel/remote-cache@1 writes the cache and the key to ~/.bazelrc,
readable only by the job's user, for every later bazel command. On
pull-request runs it never uploads (--remote_upload_local_results=false),
whatever key the run has. Its log lines:
[bazel/remote-cache] remote cache grpcs://remote.buildbuddy.io. With no key, it says so and builds without the remote cache instead of failing.Streaming build results to: https://app.buildbuddy.io/invocation/…, a link to the invocation's timings and cache hits.
Let pull requests read the cache (optional). Unlike a token-less Nx Cloud run, a Bazel run without a key gets nothing from BuildBuddy.
Create a second key with Read-only key (disable remote cache
uploads) checked. Store it as a secret such as
BUILDBUDDY_READONLY_KEY and tick its pull request context: proposed
code can read the cache with it, never write. Then point the checks job
at it:
secrets:
- BUILDBUDDY_READONLY_KEY
setup:
- uses: bazel/remote-cache@1
with:
secret_var: BUILDBUDDY_READONLY_KEY
bazel/remote-cache@1 takes these parameters:
| Parameter | Default | What it is |
|---|---|---|
url |
grpcs://remote.buildbuddy.io |
The remote cache. Any gRPC remote cache that takes a key in a header works. |
header |
x-buildbuddy-api-key |
The header the key goes in. |
secret_var |
BUILDBUDDY_API_KEY |
The variable holding the key. |
bes_backend |
grpcs://remote.buildbuddy.io |
Where build results stream. Empty turns it off. |
results_url |
https://app.buildbuddy.io/invocation/ |
The build results page the log links to. |
6. Enable the repository
Enable it as in Getting started. The first revision dispatches every service, since none has received anything yet. After that, only what changed.
7. A CI image (recommended)
A stock image downloads and extracts Bazel on every job. demo-bazel's
ci/Dockerfile
bakes it in instead:
debian:bookworm-slim, the runner's tools and gcc.bazelisk and bazel-diff, both pinned by sha256.
The Bazel release
ci/.bazelversionnames, installed and already extracted, through/etc/bazel.bazelrc:startup --install_base=/opt/bazel/install common --repository_cache=.cache/bazel-repo common --repo_contents_cache=/tmp/bazel-repo-contents
build_ci checks out ci/, the image's whole build context, and
publishes it without a repo:. Pipemesh keeps the image in its own
registry (no registry token to store) and reuses one already built from
the same content.
External repositories stay out of the image. They download into the
workspace's repository cache, kept in a Pipemesh cache keyed on
MODULE.bazel.lock.
Try it
On a fork or a copy of the demo, make a change from the table and watch the boards. Edit the README, say: the graph job runs, and no service receives the revision.