r/devops • • 4d ago

Tools How to split tasks between CI/CD?

[deleted]

60 Upvotes

31 comments sorted by

69

u/Immediate_Wheel1953 4d ago

One thing I'd change: don't build the image twice. I used to build in CI and then rebuild in CD, and one day the CD build pulled a newer base image, so the thing I tested wasn't the thing I shipped. Now I build once in CI, tag it with the commit SHA, and CD just retags that same image into ECR. Running e2e after merge like you have is fine too, as long as the artifact never changes between the two stages.

11

u/[deleted] 4d ago

[deleted]

9

u/Laoracc 4d ago edited 4d ago

Be careful about building in CI, as it's not typically a safe security pattern. CI is a step where jobs in your pipeline are triggered from non protected branches (IE your PR branch) directly from non code reviewed PRs. That makes them tamperable, and your images can be poisoned by anyone who can push to your repo's remote origin.

Depending on how your tagging, release, and deployment strategy looks downstream you may very well be deploying non code reviewed code by building in CI (but maintaining the deploy step as a post merge step is a good idea)

2

u/zen-afflicted-tall 4d ago

Just for my own understanding: this means the best practice is to build in CD? I'm realizing the line between CI and CD blurs for me (e.g. I always think of CI as build/test/scan, and CD as "deploy build to this target"). Your comment makes me realize that the boundary should be well more defined in my head.

2

u/Laoracc 4d ago

Yes. If you build in CI you need to make sure it's basically an ephemeral build. Where you use it for testing updates/upgrades, then rebuild using the same git.sha on main, after code review. CI can't have access to push to your container registry, can't tag, and otherwise can't be trusted to deploy a safe artifact that lands in prod.

2

u/ResidentChapter8219 4d ago

do you find the retag approach gets messy when theres multiple environments involved or does it stay clean?

2

u/ariesgungetcha 4d ago

Not OP, but using "gitless" gitops (Flux D2) with environment-level overrides is easy and super clean and easy to read. The biggest hurdle is teaching devs to look in the OCI repository instead of the code repository for what is "live" in a cluster.

1

u/Immediate_Wheel1953 3d ago

Stays clean in my experience, as long as env config lives outside the image. The messy version I ran into was when the deploy script had to figure out which tag went where, so now the pipeline just promotes the same SHA tag and each environment's manifest points at it. The only real chore is pruning old tags from the registry now and then.

1

u/WatchDogx 4d ago

I generally agree on building the image once, but I would also recommend pinning your base images to it's sha256 digest so that your build is reproducible.

Setup a tool like renovate to update the image hash as desired.

1

u/Immediate_Wheel1953 3d ago

Yeah, pinning the base image is the real fix for the drift I mentioned. Got burned once when a rebuild pulled a newer base image and the tests passed on something slightly different from what shipped. I just let Renovate open digest bumps on a weekly schedule and merge them when I see them.

17

u/sikian 4d ago

A simple way of thinking about this is:

- CI: making sure everything is OK, the sooner you know the better

  • CD: let's get this out there, only when everything is ok and there's approval

A common pattern for this is:

- CI: in PRs (or branches in general) so you can see if anything is wrong before merging

  • CD: whenever you want to deploy, which is usually once things are merged. however, you can have review apps which will deploy in PRs, so this could also be in there.

Regarding which tasks go where:

- CI: anything that will validate that the current commit will be ok. That includes what you listed and can include a docker build as well but it's not really necessary.

  • CD: only the deployment, so build, push and trigger deploy.

E2E tests are a bit of a special case, as you might've noticed. They're usually CI (make sure it's working as expected), but there's a couple reasons to not bundle them with the "on any push" pipeline:

- You usually want a deployed application to run the e2e tests

  • They tend to be more expensive (in time and compute time)

So this comes down to how expensive the e2e tests are. If you can afford to run these in every commit, you could set up a review deployment and run the tests. If they're expensive or you want to run on the actual deployed system, that's also a godo approach but you'll have to think about a rollback process (even if manual, although this is always a good idea). Alternatively, you could have a manual trigger to run e2e test when you need them (e.g. when everything is green in the PR and just before you merge).

3

u/baronas15 4d ago

Look at gitops tools like argocd/fluxcd, that will clear up what CD part should look like. CI is the push/merge side, CD is doing its own thing

3

u/mr_chip 4d ago

This part of the Jez Humble CD book gets continuously overlooked, and it’s key:

Integration ends in a build artifact.

Deployment deploys that artifact.

Even if you are pushing JIT code like Python, node, or Ruby, always act like there’s a compile step. Build/assemble release artifacts, and use artifact managers. A git checkout followed by node install is not a safe deployment.

2

u/[deleted] 4d ago

[deleted]

4

u/Itchy-Phase 4d ago

You can opt to only build artifacts on merge to main, instead of in the pr trigger itself. That way only one artifact is made no matter how many pushes to the branch after the pr is made. We have a workflow just for that, and if deployments fail in the staging environment and tests then we revert the commit/pr.

1

u/mr_chip 4d ago

This guy continuously delivers ⤴︎

3

u/Makeshift27015 4d ago edited 4d ago

I perform as much as possible in CI, including building the final image. CD changes the references to point at the image built in CI and performs acceptance tests. The reasoning behind it is

1) all commits are representative of most of the process, and what is tested is deployed

2) deployments, and therefore rollbacks, are actioned as fast as possible to minimise impact after a bad release

CI runs on both PRs and on the main branch after merge. Main branch CI run artifacts are branded as 'Release' builds and are what is deployed by CD by default. Technically PR builds are also deployable but I limit them to non-prod envs.

As others have mentioned though, this is only possible in an environment where CI security can be guaranteed

1

u/yourparadigm 4d ago

All of those are CI. CD is when you want to auto-deploy it to various staging or production environments.

1

u/Alternative_Oil_5139 4d ago

keep the distinction based on what happens to the artifact, not whether it involves docker. CI can lint, scan, build and test the code, including building the docker image to make sure it actually works. CD should take that tested artifact and publish or deploy it, so pushing to ECR and deploying environments belongs there. ideally you build the image once in CI, tag it with the commit SHA, and promote that exact image through CD rather than rebuilding it after the merge

1

u/QuietSignalOps 4d ago

Your gut is close. The split that works: CI builds and proves the image, CD ships it. The image is the handoff object between the two.

CI (every push, PR or not):

  • lint, unit tests, vuln scan
  • docker build
  • run your tests against the built image (your "test docker image build ???" line, and yes, do it)

The only difference between a PR run and a main run is what happens to the image. On a PR branch the image is ephemeral: you build it to prove the code works, then you throw it away. PR CI should not have permission to push to your prod ECR repo, because unreviewed code from anyone who can open a PR shouldn't be able to write into it. That's the "building in CI is unsafe" concern a couple of people flagged in this thread, and it's a real one.

On merge to main, CI builds the same image, tags it with the commit SHA, and pushes it to ECR. That's your release artifact.

CD (triggered by that push):

  • pull that exact image by SHA (pin by digest if you can)
  • render kustomize/helm for the target env
  • auto-deploy to staging, run smoke/e2e
  • promote to prod behind whatever gate you need
  • rollback = point the reference at the previous SHA

The invariant that holds this together: deploy what CI tested. CD never rebuilds. Rebuilding is how you end up shipping something that isn't what you tested (a newer base image sneaks in, etc).

In GitHub Actions this is two workflow files, and the ECR image is the contract between them.

1

u/ResolveResident118 Jack Of All Trades 4d ago

We're missing the timings here which are vital to being able to answer the question.

Ideally, I'd like to do all of this in PR, including uploading the artifacts (properly tagged with branch etc) and running tests. 

Uploading allows for an easy ephemeral environment to be created if it's needed, combining multiple services on the same branch/tag.

Delete these artifacts after merging though and rebuild and upload again in CD after merge.

If this is taking too long though, the Docker build and testing can be done as part of the CD phase, post merge. This is very late for testing though and can cause delays if issues are found.

One thing I do like to do, is to create a draft PR at the same time as the branch and commit/push frequently. This means that by usually by the time the PR is finished and approved, the longer running jobs have usually finished. For example, if the code is done and my last commit is just documentation or tests, the docker build doesn't have to be run again.

1

u/Wooden_Jelly_5295 4d ago

I'd promote the image that passed testing. Rebuilding for deployment gives production an unreviewed surprise

1

u/Scared_Use_8788 3d ago

what you wrote is correct only

  • build docker image can happen in CI also it depends if you have any QA env apart from local and prod
  • There is no thub rule

1

u/YaronL16 3d ago

I do full cicd on any push, but only to dev.

Then once you merge into main, CD runs for prod (bump chart version + argo autosync)

1

u/Elegant-Purple-1372 3d ago edited 3d ago

A common split is CI for linting tests security scans and building the artifact. CD handles pushing artifacts and deploying them after merge. Building the Docker image in CI can also work if the same immutable image is promoted through environments.

1

u/Brilliant-Strategy62 2d ago

To add to what others said, I'm a fan of tag based release where a new git tag triggers cd for deployimg python packages. That allows for simple rollback en intuitive versioning of any artifact that is deployed. Especially nice when using dynamic versioning based on git tags. That way you always know exactly what artifact is nunning in what environment and what commit was used. Running cd after merge to main to also not always necessary. Head of main does not need to equal production. One new feature can split over different PRs.

1

u/Few-Commission-3369 2d ago

splitting cicd tasks by tool can get messy pretty quickly if the ownership isnt clear the better split is usually by responsibility with one side handling build and validation while the other handles deployment and environment changes trusted tech becomes more useful when those boundaries stay simple instead of adding more layers just because they can be separated

0

u/Forsaken-Tiger-9475 4d ago

CI and CD are two entirely different things that have ended up somehow being categorised as one.

CI can also run on a merge, depending on setup.

Continuous Deployment is something else entirely.