Build Infrastructure, Caches and CI

How downloads, shared state, mirrors, clean builds, and CI fit into a maintainable Yocto workflow

Once a Yocto project moves beyond one developer machine, build infrastructure starts to matter.

The aim is not to make every build painfully slow in the name of purity. The aim is to make development builds fast enough to use, CI builds trustworthy enough to gate changes, and release builds controlled enough that you know what shipped.

That balance matters. A perfectly clean build that nobody can afford to run is not very useful. A very fast build that silently depends on random machine state is also not very useful. We are looking for the boring middle, which is where most good engineering lives.

Downloads and Shared State

Earlier, we set up a shared directory because Yocto builds reuse a lot of data. The two important caches are:

DL_DIR

The downloads directory. This stores fetched source archives, Git mirror data, and other downloaded inputs.

SSTATE_DIR

The shared state cache. This stores reusable task output so BitBake can avoid rebuilding work that has already been done.

The usual pattern is to keep these outside the build directory:

DL_DIR = "${HOME}/yocto/shared/downloads"
SSTATE_DIR = "${HOME}/yocto/shared/sstate-cache/${MACHINE}"

The setup lesson introduced this idea for a developer workstation: Setting up the System. In a production project, the same idea becomes part of the build infrastructure.

Why Share Downloads?

Yocto fetches a lot of source code.

Sharing DL_DIR avoids repeatedly downloading the same upstream source for every build directory, every developer, and every CI job. It also reduces your exposure to temporary upstream outages.

On a real project, I normally want a controlled downloads cache or mirror that CI can use. If upstream disappears, the build should not fail merely because a tarball moved or a Git server had a bad morning.

Why Share Sstate?

The shared state cache is about build output, not source downloads.

BitBake breaks a build into tasks. If a task’s inputs and signature match something already in sstate, BitBake can reuse the previous output rather than running the task again.

That is the useful conceptual model:

same task inputs
  -> same task signature
  -> reusable sstate object
  -> less rebuilding

You do not need to understand the internal storage format to use sstate well. You do need to understand that it is a performance tool, not a substitute for proper release control.

Mirrors

Mirrors let you control where BitBake looks for source.

Two variables are commonly involved:

PREMIRRORS

Locations BitBake tries before the original upstream URL.

MIRRORS

Fallback locations BitBake tries if the original fetch fails.

A project mirror can make builds faster and more reliable:

PREMIRRORS:prepend = " \
    git://.*/.* https://mirror.example.com/git/PATH;protocol=https \n \
    http://.*/.* https://mirror.example.com/sources/PATH \n \
    https://.*/.* https://mirror.example.com/sources/PATH \n \
"

The exact mirror layout is project-specific, so do not copy that fragment blindly. The important idea is that production builds should not rely entirely on the continued good behaviour of every upstream server.

Do Not Start by Deleting tmp/

When a build fails, one of the least useful instincts is:

rm -rf tmp

It sometimes makes the next build pass. It also destroys the evidence of what failed, throws away useful work, and teaches you very little. That is not debugging. That is making the crime scene tidier.

Start with the failing task, logs, generated run script, and recipe work directory. The debugging lesson covers that workflow: Debugging Yocto Builds.

There are times when removing build state is appropriate, but it should be a decision, not a reflex.

Cleaning Tasks

BitBake provides cleaning tasks with different levels of force.

clean

Removes most build output for the recipe so tasks can be rebuilt.

cleansstate

Removes build output and shared state for the recipe, forcing BitBake to recreate sstate for that recipe.

cleanall

Removes build output, shared state, and downloaded source for the recipe.

The commands look like this:

bitbake -c clean myapp
bitbake -c cleansstate myapp
bitbake -c cleanall myapp

Use clean when you want to force a recipe’s local work to be rebuilt.

Use cleansstate when you have a good reason to distrust the sstate for that recipe, for example while debugging a bad cached result or validating that a change rebuilds correctly from tasks.

Use cleanall when you also need to remove the fetched source, for example when testing fetch behaviour or dealing with a corrupted download. It is rarely the right first tool.

CI Build Environments

CI should make the project harder to break.

That usually means:

  • building from committed metadata
  • using pinned layer revisions or a manifest
  • starting from a known host environment
  • using controlled shared downloads and sstate
  • retaining logs and artifacts
  • running checks that match the product risk

The build host matters because Yocto relies on host tools. You can control that with dedicated build machines, virtual machines, or containers. The exact choice depends on the organisation. The goal is the same: the CI build environment should be boring and repeatable.

Containers are useful, but they are not magic. You still need to manage storage, permissions, supported host distributions, network access, and cache policy.

Clean Release Builds

A release build should be controlled, but that does not always mean it must avoid all caches.

There are two separate questions:

  1. Did the release build use the correct inputs?
  2. Did the release build avoid hidden local state?

Shared downloads are usually fine, provided they are controlled and match the expected source checks. Shared sstate can be fine too, but some projects choose to do release builds with an empty build directory and carefully controlled caches so the result is fast enough while still being reproducible.

For high-confidence release work, record:

  • CI job identifier
  • build host or container version
  • layer manifest
  • image and machine
  • key configuration
  • artifact checksums
  • logs
  • license and SBOM output

That gives you evidence if you ever need to explain or reproduce the release.

Retaining Logs and Artifacts

Do not only keep the final .wic or .tar image.

Useful artifacts include:

  • image files
  • kernel, device tree, bootloader, and firmware outputs
  • package feeds, if used
  • SDK installers, if released
  • build logs
  • task failure logs from CI
  • build configuration or manifest
  • license manifests
  • SPDX/SBOM output
  • CVE reports

Retention policy depends on product needs, storage cost, and regulatory requirements. The engineering principle is simple: keep enough to investigate an old release without rebuilding the entire world before you can even ask the first question.

Useful CI Checks

A production Yocto CI pipeline often grows over time.

Useful checks include:

  • parse the metadata
  • build key images
  • build the SDK
  • run recipe or package QA
  • verify no unexpected dirty workspace changes were used
  • check that release builds use fixed layer and source revisions
  • generate license manifests and SBOMs
  • run CVE checks
  • boot smoke tests in QEMU where practical
  • collect image artifact checksums

Do not try to build the world’s most elaborate CI system on day one. Start with checks that catch real failures, then add more as the project matures.

Determinism Versus Performance

There is always a trade-off.

Disabling all shared state can prove that a build starts from nothing, but it is slow. Reusing every cache aggressively is fast, but it can hide problems if the cache is poorly controlled.

In practice, I like a tiered approach:

  • developer builds use shared downloads and sstate heavily
  • CI builds use controlled caches and retain logs
  • release builds use pinned inputs, controlled hosts, and archived artifacts
  • periodic clean builds prove the project can still rebuild without relying on stale local state

That gives the team performance during normal development without giving up the ability to trust the release process.

Summary

Build infrastructure is part of project maintenance.

The main ideas from this lesson are:

  • DL_DIR stores downloaded source inputs
  • SSTATE_DIR stores reusable task output
  • shared downloads and sstate improve build speed and reliability
  • mirrors reduce dependence on upstream availability
  • deleting tmp/ is usually a poor first debugging step
  • clean, cleansstate, and cleanall have different levels of force
  • CI should build from controlled inputs in a controlled environment
  • release builds should retain logs, artifacts, and build metadata
  • deterministic builds and build performance need to be balanced deliberately

Quick quiz: build infrastructure

Check how caches and cleaning commands should be used in a production workflow.

Question 1What does `SSTATE_DIR` primarily store?
Question 2Why is deleting `tmp/` usually a poor first response to a build problem?
Question 3Which command removes fetched downloads as well as build output for a recipe?