Stay in the loop
How job fingerprints help us scale down faster
We've been rolling out our new platform, Depot Metal, for a few months now. Depot Metal scales machines up to support traffic and scales them down when we no longer need them.
We recently wrote about the evolution of the Depot Metal autoscaler and the different signals it uses to make scaling decisions. One of the signals is the estimated time required for a host to drain. To get those estimates, we fingerprint jobs so we can group historical runtimes from comparable executions.
Choosing a host to drain
Once we've decided we can remove a machine, we still need to choose which one. Scaling down first requires draining the host. We stop assigning new sandboxes to it, let its existing work finish, and then terminate it. Because we spread work across compute hosts to avoid hotspots, excess capacity can be scattered across machines that still have running jobs. Having enough spare capacity to remove a host doesn't mean there's an empty one waiting for us.
We started with the simplest approach of choosing a host with low utilization. We quickly ran into a problem, though. A machine might have only one job left, but that job could run for hours. And while it drains, we're paying for the entire machine, but we're no longer assigning new work to its unused capacity.
Consider a simplified example:
| Host | Running jobs | Estimated time until empty |
|---|---|---|
| A | 1 | 90m |
| B | 8 | 3m |
Assuming both hosts are otherwise eligible, B is the better candidate to drain. It has more jobs running, but all of them are expected to finish sooner.
Utilization told us how much work was on a machine. We also needed some idea of how long that work would take.
Giving slow drains a timeout
Our first change was a timeout. If a host took too long to drain, we could start draining another while the original continued. This was progress, but it didn't tell us whether the next host would empty any sooner. We needed to understand more about the work running on each host.
Finding the same job across runs
We already had historical job timings. Many CI jobs run the same build, test suite, or other task across pushes and pull requests. Their work can change over time, but previous runs give us starting points for estimating the next one.
The question was which previous runs to use. We could average all jobs for an organization or repository, but that would mix together different work with very different runtimes. A quick lint job and a long integration test suite might belong to the same workflow. Even a job name alone isn't enough since multiple workflows can have a job called test.
We needed a reusable identity that would let us find comparable executions. We already have a fingerprinting technique for test failures; stable identities let us connect individual test results to their history. For jobs, we want to connect an incoming job to runtimes of previous executions.
The runtime fingerprint includes:
- Whether the job comes from Depot CI or a Depot GitHub Actions runner
- The organization and repository
- The workflow identity
- The job key and display name
- The runner's architecture, CPU count, and memory
These fields distinguish jobs that share a name but run in different workflows or on different sized machines. Display names can also distinguish matrix variants.
We don't include job commands in the fingerprint because even a small change to a step would give the job a new identity, potentially separating it from useful timing history. Instead, we keep that history connected and let the rolling estimates adapt as the work changes.
Turning history into an estimate
Once we have a fingerprint, we can collect timings from successful, positive-duration runs with that identity. We use a rolling multi-day history, refreshed hourly. From that history, we use the 90th percentile duration as the estimate, meaning that roughly 90% of these historical runs finished within that amount of time.
Using the p90 gives the estimate room for slower executions, though it doesn't guarantee the next run will finish within that time. The rolling history lets estimates adapt as test suites grow or builds get faster, although sudden and drastic changes can leave them inaccurate for a while.
Before launching a sandbox, we generate the job's fingerprint and use it to look up a runtime estimate. We attach that estimate to the sandbox so the autoscaler can use it while the job is running. If we're missing information needed for the fingerprint, or the job has no timing history yet, the sandbox starts without an estimate.
Using estimates to choose a host
With an estimated duration attached to a sandbox, the autoscaler can subtract its elapsed runtime to estimate how much longer it'll likely run. A host has to wait for its last sandbox to finish, so we use the longest remaining estimate as the predicted drain time.
Not every sandbox has a usable estimate, and some aren't CI jobs at all. We currently require estimates for at least 80% of a host's active sandboxes before considering its predicted drain time. Among eligible candidates, we prefer the shortest predicted drain, subject to other host-selection rules.
What about jobs that have already exceeded their estimate? Treating those as having zero time remaining would make an overdue job look like it was about to finish. Instead, we treat the remaining time as unknown, which reduces the host's estimate coverage. If no eligible candidate has enough coverage, the autoscaler falls back to its existing utilization-based selection.
The 80% threshold lets us use incomplete information, but an unknown sandbox can still be the last one running and keep the host around.
How's it going
Looking at launch logs over a recent 24-hour period, we estimate that about 99% of job sandbox launches across GitHub Actions and Depot CI had enough information to generate a fingerprint, and about 98% had matching timing history available to estimate their runtime.
In a recent sample, the p99 for completed host drains was about an hour, compared with just over 3 hours before we introduced runtime estimates. More than 75% finished draining in under five minutes. Other autoscaler changes contributed too, but we're seeing fewer completed drains stretch on for hours.
We're continuing to update the autoscaler as we learn more about traffic and the workloads running on these machines.
Fingerprinting gives us a way to bring job history into scaling decisions. When we find a host with only one job left, we can now use that job's historical runs to help decide whether it's a good machine to wait on.
FAQ
How does Depot Metal estimate a CI job's runtime?
Depot Metal fingerprints the job, finds successful historical runs with positive durations, and uses the 90th percentile from a rolling multi-day history. We refresh that history hourly so estimates can adapt as workloads change.
How do runtime estimates help the autoscaler choose a host to drain?
The autoscaler subtracts each sandbox's elapsed runtime from its estimate. Because the last running sandbox determines when a host becomes empty, the longest remaining estimate becomes the predicted drain time. Among eligible hosts, the autoscaler prefers the shortest predicted drain.
Why doesn't the runtime fingerprint include job commands?
A small command change would create a new identity and discard useful timing history. Keeping the same fingerprint lets the rolling estimate adapt to changes without starting from zero.
What happens when a job has no runtime history or runs longer than its estimate?
A job without enough fingerprint data or timing history starts without an estimate. If a job exceeds its estimate, its remaining time becomes unknown instead of zero. Both cases reduce the host's estimate coverage, and the autoscaler falls back to utilization-based selection when no candidate reaches the 80% threshold.




