CI is changing banner
10 min read
perspectives

Stay in the loop

Get notified when we ship new posts.

CI is changing

Over the past month, companies like Anthropic, Linear, and Lindy have all reached the same conclusion: CI is now the bottleneck in software engineering.

It feels like the rest of the ecosystem has finally caught up to what we've been saying for a while: CI is changing. It's no longer an area where you can underinvest, because it's the last mile to getting this newfound code volume out the door. If you don't give it the attention it deserves, all of that velocity ends in a logjam that leaves everyone waiting for CI to finish.

This isn't just another argument for faster CI. We need to refocus on the purpose of CI and reconsider when and how agents validate code. I want to share how we think about CI at Depot.

Everyone is hitting the CI bottleneck

The focus on CI has recently started showing up in a few different areas. First there was the blog post Anthropic wrote about how agentic coding is straining CI. Then there was the post from Linear about how CI was a massive bottleneck that they had to rework to keep up with their new velocity. Then over the weekend, there was this tweet from the founder at Lindy about how CI is the bottleneck and their CI spend is becoming stratospheric.

All of these posts are talking about the same thing: the bottleneck is now CI. We went from a world where the act of writing software was the bottleneck, to a world where code can be generated, committed, and put into the top of the delivery pipeline in seconds. Effectively, a five-person engineering team can generate the code of a 50-, 100-, or even 500-person team.

But what we know, from the posts above and from our own experience helping engineers using Depot, is that folks often haven't invested in their CI to keep up with that new velocity. Either because it wasn't a priority, or because they didn't have the right tools to make it easy to invest in CI. Sometimes it's just a lack of understanding about how big of a bottleneck CI can become.

Thousands of papercuts

Even in cases where people do invest in CI, we can see that small hiccups come with massive ramifications when you're talking about CI volume growing 25x in six months.

Non-performant CI, in my experience, is generally death by a thousand papercuts. It's not one big thing that makes CI slow, it's a thousand little things that add up. Often over weeks and months. A test that takes 10 seconds longer than it should, a flaky test that randomly fails once in a while, a suite of tests that runs on every commit even though it doesn't need to, and so on.

Anthropic: test selection became its own bottleneck

Anthropic went way beyond the traditional CI optimizations. They thought far enough ahead to implement test selection, which is a non-trivial optimization that most folks don't think about until they reach a moderate scale. But even with that, they still found themselves running into CI bottlenecks.

The papercut problem for Anthropic was that the test impact analysis system they built wasn't fast enough to keep up with the new volume of CI jobs. Meaning they were often having to rerun tests that didn't necessarily need to be rerun. They implemented exactly the right optimization, and still discovered that the architecture behind it couldn't keep up. The bottleneck was still CI.

Linear: CI optimization is a stack

In the Linear post, they do a fantastic job outlining all of the optimizations they built up layer by layer to get their PR wait times down to about 5 minutes, even while they quadrupled the number of tests they run in CI.

For Linear, the papercuts were even more subtle:

  • Replaced tsc with tsgo
  • Rewrote custom rules for cheaper linting in their jobs and adopted oxlint (something we have done as well)
  • Reduced checkout size using sparse and blobless checkouts
  • Made checkout more resilient to network failures from GitHub
  • Moved nonblocking jobs to run outside of merge-blocking checks so the PR isn't waiting for non-essential jobs
  • Preinstalled common dependencies in the CI image to avoid downloading them on every run
  • Updated installs to target specific workspace packages and their dependencies
  • Removed caching that was slowing down jobs in favor of filtered installation
  • Loaded schema snapshots instead of replaying unchanged migrations
  • Implemented test splitting and balanced shards to keep one shard from holding up the rest of the test suite
  • Introduced more parallelism in their test shards after reducing setup overhead
  • Allowed eligible tests to share module state so they don't have to start from scratch every time

The length of this list gives you a sense of how many small optimizations it takes to keep CI performant at scale.

No single fix for slow CI

It's not about one big optimization, like sizing up a runner or adding more parallelism. Those things help, for sure, but they are actually part of a larger stack of optimizations that need to be implemented.

Oftentimes, in my experience, people recognize that CI is a bottleneck. But they don't have the right tools or experience to know how to optimize it. They start at the first or second layer, the big optimizations, and then they stop. They don't go deeper.

But the reality of machine driven code generation is that every single team is hitting these bottlenecks now. It's not just the FANG companies of the world, it's every team.

The future of trusting code written by agents

The blog posts and tweets above are all about CI being the bottleneck. It's kind of always been the bottleneck at scale. It's just that now, with agents at our side, everyone is capable of reaching that scale and that trend appears to only be accelerating.

I think we are at a point where we need to start reframing this problem. Taking a step back and thinking about what the purpose of CI is, and how we can make it better for this new world we live in.

CI exists to build trust

The purpose of CI is to give us confidence that the code we are shipping is correct. That it does what we expect it to do, and that it doesn't break anything else in the system. But CI is not keeping up with the new volume of code being generated. Therefore, in the best case, our mean time to trust is increasing, and in the worst case, we are shipping code that is not correct and breaking things in production.

The future of CI is not just about making it faster. It's about making it more intelligent and more capable of understanding the code that is being generated and the context in which it is being generated.

In other words, the paradigm of CI as a place to validate that code is correct is changing. That validation is shifting to where code generation happens.

I think we need to let go of the preconceived ideas we have about CI and start thinking about what it means to validate code quickly and efficiently, in terms of wall time and dollars, while maintaining trust.

Move validation to where code is generated

Today, we validate code after it is written and committed.

In the future, we will validate code as it is being written and generated. We will validate code before it is committed and before it is merged. We will validate code in a way that is aware of the context in which it is being generated and the impact it will have on the system as a whole.

It's about software delivery as a whole, not just CI.

I believe that you have to integrate the concept of validating code into agents themselves. This is one of the many reasons that Depot CI can be entirely driven via our API and CLI. It's so that agents have a real CI system with pre-defined workflows they can use to validate their changes. Additionally, they can write their own workflows on the fly and the CI system will run them. They can even choose to run just specific steps that they have determined must be run, instead of the entire workflow from the top.

This gives agents the ability to verify that their changes integrate and are green in real CI before they even commit.

Beyond pre-commit validation

But pre-commit validation only scratches the surface of what's possible. If you forget CI and just work the problem: how do I validate and build trust around this code? You can come to a lot of interesting ideas (some of which we are working on):

  1. If we know the code that was generated already passed CI via the agent running it, we can skip CI entirely after the code is committed. Or perhaps there is a specific workflow that runs after code is committed and the agent already validated everything else.
  2. If we have source control and we know what code was generated, we can dynamically determine exactly what tests need to be run rather than you having to implement that yourself.
  3. We should be able to know the results of our validation in production. If the code took down prod, we should be able to tie that back into the trust loop so that the system automatically avoids that scenario again.

Conclusion

CI is the bottleneck. But it's not a new one. We're just collectively reaching that bottleneck faster and faster, at smaller and smaller companies. Large companies have faced these problems for the past 10 years and built their own internal solutions to solve them.

But perhaps building internal solutions isn't the answer. Perhaps we need to rethink the entire concept altogether, to integrate CI into one holistic system, rather than treat it as something that happens after commit.

The future isn't just faster CI. It's giving agents a way to validate code and maintain trust as they work.

FAQ

Why is CI becoming the bottleneck for teams using coding agents?

Agents can generate and commit code much faster than traditional CI can validate it. That increase in code volume exposes every slow test, flaky check, unnecessary job, and repeated setup step in the delivery pipeline.

What does mean time to trust mean in CI?

Mean time to trust is the time between generating a change and getting enough evidence that the change is safe to ship. When CI cannot keep up with code generation, that time increases and teams either wait longer or ship with less confidence.

How can agents validate code before it is committed?

Agents can use Depot CI through the API and CLI to run predefined workflows, create workflows on the fly, or run only the specific steps they determine are necessary. That gives them a real CI system for validating changes before they commit.

Why is faster CI alone not enough?

The purpose of CI is to give us confidence that the code we ship is correct. As agents generate more code, validation must also become more intelligent, understand the context of each change, and move closer to where code generation happens.

Start building with Depot
in minutes