Sam Austin on October 3, 2026

GitOps for Machine Learning: Version-Controlled Deployments

GitOps for Machine Learning: Version-Controlled Deployments
Contents
GitOps version-controlled machine learning deployments on Kubernetes
GitOps version-controlled machine learning deployments on Kubernetes

Figure 1: Every model deployment is a commit, every rollback is a git revert — the cluster state always matches what's in the repository

A model's accuracy dips right after a deployment. The usual scramble follows: manually rolling back, patching something, hoping it sticks. GitOps replaces that chaos with something deterministic. Every deployment is a Git commit, every rollback is git revert, and your cluster's actual state is always traceable back to exactly what's in your repository. Let's build this for ML specifically, where "deployment" means more than just application code — it's the exact pattern the MLOps beginners guide argues is the difference between a model that ships and one that stays stuck in a notebook.

The Core Principle, Adapted for ML

Standard GitOps is simple: your Git repository is the single source of truth, and an operator running in your cluster continuously reconciles actual state to match it. For ML, this needs one important adaptation, since model weights and training datasets can run into gigabytes or more, nobody wants that in a Git repository.

The practical split that's emerged as standard practice: Git holds code, pipeline definitions, Kubernetes manifests, and pointers to large artifacts; object storage holds the actual datasets and model binaries themselves. Your commit history becomes a complete audit trail of desired state, while S3, Azure Blob, or GCS handle the bytes. This isn't a workaround, it's the entire design, and it's what makes GitOps genuinely scale for ML rather than becoming a repository full of binary blobs nobody can diff sensibly.

ArgoCD vs Flux: Picking Your Operator

Both tools remain the standard GitOps operators for Kubernetes-based platforms as of 2026, and the choice between them comes down to how opinionated you want your automation to be.

ArgoCD gives you a rich UI for visualizing application state, sync waves for controlling deployment order when resources depend on each other, and an Image Updater component that can watch a container registry and commit new image tags back to your GitOps repo automatically, with a full audit trail, while ArgoCD itself handles the actual deployment.

Flux takes a more lightweight, CLI and automation-first approach. Its image automation feature is tightly integrated with the core Flux model: you mark image tags in your deployment manifests with a comment like # {"$imagepolicy": "flux-system:myapp"}, and Flux updates them automatically as new images land in your registry, eliminating the manual step of updating YAML by hand for every new build.

Neither is objectively better for ML specifically. Teams already invested in a rich visual dashboard for debugging deployment state tend to prefer ArgoCD; teams wanting a leaner, more automation-heavy pipeline often prefer Flux. Pick based on your team's existing Kubernetes tooling preferences rather than ML-specific feature gaps, since both handle the core ML deployment patterns well.

What Gets GitOps-Managed in an ML Platform

GitOps for ML covers the infrastructure and deployment lifecycle broadly, not just the model serving endpoint. That means ArgoCD or Flux ends up managing your model serving deployments and InferenceServices, scheduled training pipelines defined as CronWorkflows, feature store configurations, and the surrounding Kubernetes resources those all depend on, while a separate model registry handles the actual artifacts those resources point to. This is the same "automate the whole lifecycle, not just the training step" territory the CI/CD for ML article covers — GitOps just moves the source of truth from pipeline YAML scattered across tools into one reviewable repository.

A typical repository structure separates these concerns cleanly, with a top-level path for each model, environment-specific overlays underneath (staging, production), and the GitOps operator's own configuration kept distinct from the application manifests it's watching.

The Model Promotion Workflow

Here's where GitOps principles translate into a genuinely useful ML-specific pattern: promoting a trained model from staging to production becomes a Git operation, not a manual deployment step. A typical promotion workflow, triggered through a GitHub Actions workflow_dispatch, looks like this in practice:

name: Promote Model
on:
  workflow_dispatch:
    inputs:
      model_name:
        description: "Model name"
      model_version:
        description: "Model version"

jobs:
  promote:
    runs-on: ubuntu-latest
    steps:
      - name: Update model version
        run: |
          MODEL=${{ inputs.model_name }}
          VERSION=${{ inputs.model_version }}
          cd models/${MODEL}/overlays/production
          sed -i "s|storageUri:.*|storageUri: \"s3://models/${MODEL}/${VERSION}\"|" deployment.yaml
          git commit -am "Promote ${MODEL} ${VERSION} to production"
          git push

Once that commit lands, ArgoCD picks up the change and syncs the new storageUri to your production InferenceService, pulling the actual model weights from object storage. Nothing happens outside Git. Someone manually SSHing into a pod and swapping a model file is exactly the failure mode this pattern eliminates, since the cluster's actual running state can always be traced back to a specific, reviewable commit.

Data Versioning With DVC

Model artifacts aren't the only large thing you need to track reproducibly, training data needs the same treatment. DVC (Data Version Control) fills this role by keeping lightweight pointer files in Git while the actual datasets live in object storage — the full mechanism is covered in the data versioning with DVC article, so here's how it plugs into the GitOps loop:

dvc pull  # downloads data from the remote, using .dvc pointer files tracked in Git

A .dvc file committed to your repository records a hash pointing at the actual dataset version in your configured remote, Azure Blob, S3, or similar. When your training pipeline runs, dvc repro checks whether the referenced data is current, and pulls it only if needed. A single Git commit can capture the exact code, the exact data version, and the exact model configuration together, which is precisely what reproducibility requires: being able to check out any past commit and regenerate the identical model.

Tracing a Running Pod Back to Its Training Run

One of the more satisfying patterns in a mature GitOps-for-ML setup: being able to map a running pod all the way back to the training run that produced its model. A production endpoint exposing a /version route can report its model version, which maps to an image digest, which maps to a specific training run logged in your experiment tracker, like MLflow. Combined with MLflow's "champion" alias marking which model version is currently production-ready, your serving application can fetch the right model automatically without a human manually copying a file path anywhere.

This closes a loop that's genuinely hard to maintain manually: when someone asks "which exact model is serving traffic right now, and what data trained it," the answer is a few clicks or API calls away instead of a cross-team Slack archaeology session.

Rollbacks Become Trivial

This is where GitOps earns its keep operationally. Since every deployment is a Git commit, rolling back a bad model release is a git revert on the promotion commit, nothing more exotic than that:

git revert <promotion-commit-sha>
git push

ArgoCD or Flux detects the reverted commit and restores the previous image and model reference automatically. Compare that against the traditional failure mode: a model's accuracy dips, someone scrambles to remember what the previous version was, manually patches a deployment, and hopes it's actually the right rollback target. With GitOps, argocd app history myapp-production gives you the full deployment history directly, and reverting is deterministic rather than a stressful guess.

Handling Dependencies With Sync Waves

Real ML deployments often have ordering requirements: a feature store needs to be ready before a model server that depends on it, a model registry update needs to land before the serving deployment that references it. ArgoCD's sync waves handle this directly, by annotating resources with an explicit ordering:

metadata:
  annotations:
    argocd.argoproj.io/sync-wave: "-1"  # feature store, created first
metadata:
  annotations:
    argocd.argoproj.io/sync-wave: "2"  # model deployment, created last

Resources sync in ascending wave order, so your feature store and any PreSync hooks finish before the model deployment that depends on them even starts. Without this, you're relying on Kubernetes' own scheduling luck rather than an explicit, declared dependency order, a bad bet for anything where startup ordering genuinely matters.

Progressive Delivery for Model Rollouts

Shipping a new model version to 100% of traffic immediately is risky, exactly the kind of risk progressive delivery tools like Flagger exist to reduce — the same reason A/B testing and canary deployments matter for any risky change. Combined with GitOps, the pattern becomes: commit Flagger's canary CRDs to Git alongside your deployment manifest, and when ArgoCD or Flux applies that change, Flagger's controller takes over, deploying the new model version behind the scenes and routing a small percentage of live traffic to it before gradually increasing that share as metrics stay healthy.

This matters more for ML than typical application deployments, since a model's quality regression often doesn't show up as an obvious error, it shows up as subtly worse predictions that might only become apparent after enough real traffic has flowed through it. A canary rollout buys you a window to catch that with real production signal before fully committing.

Multi-Cluster and Multi-Environment Patterns

Larger ML platforms often span multiple clusters, a staging cluster, a production cluster, sometimes region-specific production clusters for latency or data residency reasons. Both ArgoCD and Flux support this natively: ArgoCD lets you register additional clusters directly (argocd cluster add prod-us-east), while Flux is typically installed independently on each cluster and configured to track the correct branch or path in your GitOps repository for that specific environment.

The practical pattern for promoting across environments mirrors the model promotion workflow above: a staging branch or path gets updated first, and once validated, a merge or explicit promotion step updates the production path, with each environment's GitOps operator picking up exactly the changes relevant to it.

A Practical Decision Framework

  1. Pick your operator first: ArgoCD when the team wants a visual sync dashboard and registry automation as a separate component, Flux when leaner CLI-native image automation matters more
  2. Split the repository from day one: code, manifests, and pipeline definitions in Git; model weights and datasets in object storage behind DVC pointers and storageUri references
  3. Structure paths per model and environment: models/{name}/overlays/staging and models/{name}/overlays/production, with the operator's own config in a separate path
  4. Make promotion a workflow, not a console session: a workflow_dispatch that edits storageUri, commits, and pushes — nothing touches the cluster directly
  5. Declare ordering with sync waves anywhere one resource genuinely must exist before another, instead of trusting scheduling luck
  6. Put every model rollout behind a canary (Flagger's CRDs committed to Git) before it ever sees full traffic
  7. Expose /version on the serving endpoint and map it to an image digest and an MLflow run, so any pod traces back to its training run
  8. Roll back with git revert only — if you find yourself kubectl apply-ing an emergency patch, stop and fix the Git path instead

The order matters more than the individual tools — get the Git-as-source-of-truth split right first, and the operator choice, promotion workflow, and canary step all become configuration rather than rearchitecture.

Common Mistakes People Make

Committing model binaries directly into Git defeats the entire design, repositories become unwieldy, cloning becomes painfully slow, and diffs become meaningless for binary blobs. Keep DVC pointers and registry references in Git; keep the actual bytes in object storage.

Skipping sync-wave ordering for genuinely dependent resources leads to flaky deployments that sometimes work and sometimes don't, depending on scheduling luck rather than a declared dependency graph.

Treating model promotion as a manual deployment step outside Git undoes the audit trail that's the entire point of this approach. If a model reaches production through any path other than a reviewed, merged commit, you've lost the deterministic rollback and traceability GitOps is supposed to provide.

Deploying new model versions to full production traffic without any canary or progressive delivery step means your first signal of a quality regression is a production incident, not a controlled, limited-blast-radius test.

CoverBookDescriptionGet it
Cover of “Designing Machine Learning Systems” Designing Machine Learning Systemsby Chip Huyen covers the deployment and monitoring lifecycle this whole pattern wraps around, including why reproducible releases matter more than raw model quality. View on Amazon
Cover of “Machine Learning Engineering” Machine Learning Engineeringby Andriy Burkov rigorous treatment of production ML pipelines where versioned promotion and traceability are first-class requirements, not afterthoughts. View on Amazon
Cover of “Accelerate” Accelerateby Nicole Forsgren, Jez Humble and Gene Kim the research behind why deploy frequency and fast, deterministic rollback (exactly what git revert gives you) actually predict team performance. View on Amazon
Cover of “Kubernetes in Action” Kubernetes in Actionby Marko Lukša the Kubernetes foundation you need before ArgoCD sync waves and multi-cluster setups make sense rather than feeling like magic. View on Amazon

Unlock AI That Actually Works

Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.

Click here to get GPTAstra Max now — one-time payment, lifetime access.

Frequently Asked Questions

What is GitOps for machine learning?

GitOps applies the standard single-source-of-truth model to ML platforms: a Git repository defines desired state for serving deployments, training pipelines, and infrastructure, and an operator like ArgoCD or Flux continuously reconciles the cluster to match it. For ML, the key adaptation is that large artifacts — datasets, model weights — live in object storage while Git holds code, manifests, and pointers.

Should model weights and datasets be stored in Git?

No. Model weights and training datasets can run into gigabytes or more, and committing binary blobs makes repositories unwieldy and diffs meaningless. Git stores code, Kubernetes manifests, pipeline definitions, and pointers — DVC files for datasets and storageUri references for models — while S3, Azure Blob, or GCS hold the actual bytes.

Should I use ArgoCD or Flux for an ML platform?

Both work equally well for ML. ArgoCD suits teams that want a rich visual dashboard, sync waves, and a separate Image Updater component; Flux suits teams that prefer a leaner CLI-first setup with image automation built into the core model. Choose based on your team's existing Kubernetes tooling preferences, not ML-specific gaps.

How do you roll back a bad model release with GitOps?

Run git revert on the promotion commit and push. ArgoCD or Flux detects the reverted commit and restores the previous image and model reference automatically, so the rollback is deterministic and traceable to a specific, reviewable commit instead of a manual deployment patch.

Wrapping This Up

GitOps for ML applies the same core principle as GitOps for any application — Git as the single source of truth, an operator reconciling cluster state automatically — with one essential adaptation: large artifacts live in object storage while Git holds code, manifests, and pointers. ArgoCD or Flux handle the Kubernetes side, DVC handles data versioning, and the combination gives you deterministic rollbacks, a full audit trail, and the ability to trace any running model straight back to the exact training run and data version that produced it.

Will this eliminate every production incident involving a bad model? No, a canary rollout still needs someone watching the right metrics, and GitOps won't catch a model that's technically healthy but subtly wrong. But the next time a deployment goes sideways, git revert and a few minutes beats a manual scramble through S3 buckets and deployment history trying to remember what "last known good" actually was.

What are You Looking For?

esc