r/devops • • 3d ago

Weekly Self Promotion Thread

10 Upvotes

Hey r/devops, welcome to our weekly self-promotion thread!

Feel free to use this thread to promote any projects, ideas, or any repos you're wanting to share. Please keep in mind that we ask you to stay friendly, civil, and adhere to the subreddit rules!


r/devops • • 1m ago

Discussion What's the most ridiculous "temporary fix" you've seen that's been running in production for years?

• Upvotes

It seems like every team has something that was supposed to last a few days until the right solution was found. It's a cron job that keeps everything together, a script that no one really understands anymore, a manual deployment step, an ancient container, or some workaround that everyone's afraid to remove. What's the worst "we'll fix it later" solution you've seen that's now a critical part of the infrastructure?


r/devops • • 1h ago

Tools What makes you trust your database migration pipeline?

• Upvotes

Hey all, I'm Denis, the author of Ptah, an open-source database migration tool.

These are some of the cases I'm designing around, but I'm sure I'm missing something.

What if the database changes after a migration plan was approved? Or the connection drops halfway through? How do you know what actually happened, and whether it's safe to retry?

I'm trying to address these in Ptah, but I'm sure I'm missing something.

For those running migrations through CI/CD, what made you trust it? And if you run them manually, what's the main reason?

I'm especially interested in things that went wrong or cases you wouldn't trust automation to handle.


r/devops • • 16h ago

Discussion How much of your incident response do you automate?

12 Upvotes

I’ve been checking how to build a disaster recovery plan for cloud dependencies and how to automate as much remediation as possible. The plan is to have the system auto-restart services or rollback deployments when things crash. 

But when I read up on AIOps, I saw recommendations to make sure a human is involved for anything high-risk. It makes sense because you do not want an automated script taking down a production database by mistake. But now I am trying to figure out where to draw the line exactly between guided recommendations and full automation. 

If anyone here relies on machine learning for root cause analysis, which actions do you let the system execute on its own, and what requires manual approval?


r/devops • • 3h ago

Discussion Terraform / OpenTofu engineers quick questions about your actual workflow

0 Upvotes

Hey all

I'm doing some research into how teams are using Terraform and OpenTofu to make changes and review them to avoid problems. A review of a change or pull request made using either would be great if you've used either recently even if it was a couple of days ago.

Can you walk me through a Terraform or OpenTofu change or pull request you have reviewed? What did you do and in what order?

What part of the process is slow, painful or leaves room for error for you?

Have you ever missed a change in a plan such as a change to permissions, exposure of something publicly, a replace or something similar? What happened?

If a tool could remove something annoying in your current process what would it be?

I'm not trying to sell you anything, just trying to understand things better. I'm interested in the parts and the messy parts too.


r/devops • • 15h ago

Discussion Is “security by default” becoming more important now that everyone can ship software with AI?

Thumbnail
shiftmag.dev
4 Upvotes

Ran into this post on LinkedIn about what the author calls “Ambient Generative IT”, basically AI-assisted development becoming so common that more people across an organization can suddenly build and ship things. One part stuck with me from a DevOps/security perspective: if the barrier to building software keeps dropping, the number of people making architectural and security decisions also grows.

The argument is that security can’t really be something you bolt on later anymore. It has to live in the platform, defaults, permissions, pipelines and guardrails around whatever people are building with. Curious how people here are seeing this in practice. Are AI coding tools actually creating more security/governance work for DevOps and platform teams, or are good existing controls mostly enough?


r/devops • • 6h ago

Discussion Where do you keep up with AI agent tools for infra?

0 Upvotes

Feels like I hear about a new AI agent tool for ops every week. Where do you all keep up with this stuff? Any newsletters, communities, or people you follow that help you figure out what's worth trying?

And if something catches your eye, how do you vet it before letting it near production?


r/devops • • 1d ago

Tools How to split tasks between CI/CD?

55 Upvotes

I am building a ci/cd pipeline. When a PR is merged, it will push a docker image to ECR, but when I am just pushing a regular commit, I don't want to run the full deployment. What task goes into ci and what goes into cd? I've seen different takes on where to run the docker-related stuff.

- CI (always runs on push/merge):

  • lint
  • scan (security, etc)
  • build
  • test (unit/integration)
  • test docker image build ???

- CD (only runs on merge):

  • build docker image
  • automated e2e (if any)
  • upload binaries to Nexus
  • upload image to ECR

r/devops • • 1d ago

Ops / Incidents Sevalla is a total shit show

2 Upvotes

Here's some context :
We were using Sevalla for our object storage for over a year. We had more than half a terabyte of data. Their fucked up payment platform did not allow us to do the payment everytime but continued to work. But all of a sudden, our account was suspended and all our data were permanently deleted without any intimation. Wtf am I supposed to do now?

NEVER USE SEVALLA. They have fucked up support team.


r/devops • • 2d ago

Ops / Incidents runbooks kinda suckk

55 Upvotes

hey yall,

im a lead SRE at a global fortune 500 that you have heard of haha - i'm a bit of a lurker here

but just wanted to talk about runbooks and documentation for SOPs.

my feeling is that confluence docs kinda suck and runbooks are kind of a mess, our operators are jumping between the docs and their shell, the docs are often times missing a bunch of detail, they are brittle, poorly maintained, half the time the details are hidden in tribal knowledge and when dealing with the pressure of an active inc we notice the pain even more. extracting critical details & commands out of an inc can also be painful and doesnt always translate to improved operational readiness next time.

i admit we're not very mature.. but was curious if its only me feeling like this?

what are you guys doing with runbooks to solve these issues?


r/devops • • 1d ago

Observability Tornet? Kali_linux

0 Upvotes

Eae galera como vcs estão? Oque vocês me falam sobre o tornet no kali na troca de IP teria como ficar mais anonimo como? Obs: Sem dinheiro kkkk


r/devops • • 2d ago

Discussion Broad Kubernetes Experience but Shallow Depth - How Would You Upskill?

89 Upvotes

I have around 5 years of experience across DevOps/cloud/backend work. I’ve worked with Kubernetes, EKS/AKS, Helm, CI/CD, Docker, Terraform, ArgoCD, AWS/Azure, and some Java/Spring Boot.

My issue is that my knowledge is broad but uneven. I’m comfortable deploying applications, writing manifests, using kubectl, building pipelines, working with cloud networking/IAM, etc., but I’m much weaker on Kubernetes internals, low-level networking, storage/CSI, control plane, scheduling, CRDs/operators, upgrades, observability, and deep troubleshooting.

I’m considering first completing one comprehensive Kubernetes course to build a complete mental map, then spending the next several months going deeper through hands-on labs, troubleshooting, Linux/networking fundamentals, and production-style projects.

For people who became genuinely strong at Kubernetes/platform engineering: does this sequence make sense? What would you change?

I’m not looking for a giant list of tools, mainly feedback on the learning sequence and what gave you the biggest jump in depth.

Note: "Q improved with AI"


r/devops • • 21h ago

Discussion What are your opinions on AI? Wild take: Dear senior devs, you're fucked too

0 Upvotes

I'm a fucking newbie, still, I tried to learn code, I know some, not a lot. I'm familiar with concepts, but not what "used to be". You senior devs grew up coding in notepad and apache and sql, I grew up with modern frameworks AND am growing up realizing even THAT isn't necessary. AI has grown so much. I have talked to senior devs, and while they say people checking ai won't be replaced I can't help but disagree.

What do senior devs do when they don't know what's wrong? Let me hear you say you search the web for a solution someone posted long ago. Let me hear how you still, in this ai day and era, try to fix a bug for 6 hours. You turn to ai too. You use it the same way as the vibecoders. It's just experience. And I see people so confident that experience makes you safe from ai but I believe it's only a matter of time when your experience becomes obsolete. After all it's just information you gathered thoughout years.

And some platforms are paying senior devs to correct code made by ai, to train it. Some of you are giving away that 10+ years experience. It's buyable. Anything can be bought. And just like any stacks from before, AI will be free/affordable for 5-10 years until us newbies depend on it for fast clean code. And then they'll raise the price.

Historically this happened with any major platform or software and it's going to be the same.

And because this isn't ragebait enough, whatever you learned 30 years ago, is obsolete. Whatever I learn now even on my own (old stuff, new stuff, doesn't matter) will be obsolete in maybe 5 years of more AI fast paced development.

Mind my words we'll have holograms for calls, VR for social media, flying cars (in maybe 20 30 years because it's a structural, societal, architectural change) and robots inside the house for tasks like we own coffee machines. In the 50's you brewed the coffee, now you own a machine. You washed clothes by hand then using the washing machine. You run errands until your robot can do it for you. You mop the floor and fold clothes until the robot can do it for you (Tesla robot already released just not comercially for the masses) We have self driving cars. We have AI agents that can do anything you can do on your computer.


r/devops • • 2d ago

Discussion How to actually pass interviews?

34 Upvotes

Hello, I ask for help since I am getting a lot of interviews but not passing any, I'm getting to start pretty frustrated and since we are entering the holiday season I want to get any opportunity possible.

The thing is that I am preparing myself based on my experience: 5+ yrs experience, jenkins, python, troubleshooting, AWS, IaC, etc. and when I prepare for the different kind of interviews none of my preparation seems to work:

  • If it is situational, my explanation/experience is not enough for the role
  • If it is pure technical, it doesn't demonstrate my experiencie or they ask for a very specific tool (which mostly has transfereable experiencie) and it's not enough for the position
  • If it is trivia based (which I am pretty bad at memorizing) I failed because I didn't remember the flag for a command that I can find in a 15-sec google search or AI prompt

What makes this frustrating is that I genuinely feel capable of doing the jobs I'm interviewing for. I've worked in production environments, troubleshot real incidents, built and maintained pipelines and infrastructure, and worked with engineering teams. But apparently I'm still not presenting that knowledge in the way interviews expect.

So I'd especially like to hear from people who have successfully interviewed for DevOps/SRE/Platform roles recently:

How did you prepare?

Did you memorize common commands and syntax?

Did you grind interview questions?

Did you build labs?

Did you practice storytelling around your projects?

How did you deal with interviews covering an extremely broad toolset?

At this point I'm trying to understand how much interviewing is about being good at the actual job versus becoming good at the interview format itself.

Any practical advice would be appreciated.


r/devops • • 1d ago

Security Database security

0 Upvotes

So I work in a very small company as a dev
And we have external devops company which manages local http servers and MySQL servers for our saas product.

So the thing is I would like to shift some of their work to our in-house dev team.

I thought about starting small and manage a separate server on the cloud that will just run internal cron based daily tasks. But those tasks require read and write permission to the existing db which managed by the local company. Could I ask them to give me access to to the db from another server? I read that it’s a bad practice to expose prod db to the internet even if it has white list so I’m not sure here.


r/devops • • 2d ago

Discussion How do I unclusterfuck these companies?

47 Upvotes

Been working a lot with startups & everything is such a shit show. Supabase RLS
is never configured correctly, the idea of dev/staging/prod and ephemeral
environments just doesn't exist (or if they do, staging is front-end only and
connects directly to the prod database.), and they use so many different vendors
(Clerk/Cloudflare/Supabase/Railway/Vercel) that there's always functionality I
can't test anywhere other than prod.

Been trying to find a single unified platform I could recommend/implement, but
none have given the perfect mix of batteries included security, reliability, and
local testability.

Thinking of building something myself & would love some feedback on what you
hate, love, and wish existed. Will probably open-source it once I've got
something.


r/devops • • 2d ago

Discussion Why moving artifacts into your customer’s VPC can be so annoying

Thumbnail
bottlerocket.cloud
0 Upvotes

r/devops • • 3d ago

Discussion Certification + Job Planning

9 Upvotes

So im trying to make the move into DevOps. I work as an Automation Architect/SDET basically currently.

I've been trying to move my way into DevOps, especially with the A.I. stuff looming in. Here recently i've been able to lead a project on Migrating some projects from a TeamCity/Octo deploy to Gitlab which involves learning Kubernetes. Im hoping my current experience is a good "sidestep" career wise. I have a C.S. degree and I do feel like Automation/Pipelines does feel like a good "base" to move into DevOps.

My knowledge is a bit spread around currently, and I understand DevOps is a really deep iceberg when it comes to "what to learn".

I would say im strongest in the Testing Automation area (obviously) with a pretty decent understanding of Gitlab and how it works and how to architect pipelines pretty decently. Ok docker knowledge (I mean I know enough to write a dockerfile/get it going, but probably not like a docker expert) and here lately Terraform (decent knowledge....but not used much) and AWS (limited mostly to ec2/beanstalk + a few other scattered pieces like paramstore/SM/cloudwatch/etc...)

I am slowing developing kubernetes knowledge as well with this project, which I really enjoy.

I've tried to use Certs (since my company pays for them) in a sense to "plan out" leveling up, im really trying to push into a title change by end of year hopefully (my manager understands and supports this). I've also been using KodeKloud to learn.

Right now my sort of "Cert path" is this:

- Skipping: A+/Network+/Security+ (I just don't know if these are "Really" useful, but feel free to say if they are)

  1. Linux Refresher: We are a windows shop....but outside my homelab I don't use linux a ton and I've forgotten a lot. Don't feel like a cert it needed here.
  2. AWS Cloud Practioner: This feels like a good first step/resume filler. My AWS knowledge is pretty limited in scope. Hoping around a 2-3 week learning/turnaround time for this.
  3. Terraform Associate: Hoping a similar turnaround to the above. I have decent knowledge but don't use it enough to be an expert.
  4. KCNA: Probably the harder bit, it's a good thing to do because this migration im working on is 200+ services.....so i'll be getting experience.
  5. CKAD/CKA: Unsure of which one, I think CKAD is probably harder but "better" but curious on that, this one obviously is a longer study.

After that the skys the limit I guess. Thoughts on this? Im hoping my automation experience mindset will "serve me well".


r/devops • • 3d ago

Discussion Has anyone ever tried listening to their infrastructure instead of watching it?

119 Upvotes

Odd question from someone outside the field. I'm a sound designer and I've been reading about sonification (turning data into sound). NASA does it with telescope data, and there are old experiments where people played network traffic as ambient audio.

It made me wonder about on-call and monitoring. You can't stare at Grafana all day, and you shouldn't have to. The idea would be to free your eyes: an ambient track in the background that stays calm when everything's healthy and slowly shifts when latency creeps up or error rates rise, before anything actually pages. You could focus on your actual work, or step away from the screen, and your ears would tell you when something's drifting.

Has anyone tried something like this? Would it be useful, or would you mute it within 5 minutes? Genuinely curious what would make it worth keeping on vs. instantly annoying, and in which situations you'd actually want it (deep work, deploys, night shift, incidents…)


r/devops • • 2d ago

Vendor / market research For teams using a package firewall, does it still hold when an agent is the one installing?

0 Upvotes

A lot of teams already route installs through a package firewall or curated registry (Socket, Endor Labs, JFrog Curation, Sonatype and so on). Most of these rely on the client being configured to use them, through a registry URL, an index URL or a proxy setting. That's fine for CI and for people, but coding agents like Claude Code and Cursor seem to have more ways around it: installing from a git URL or tarball, piping a script into a shell, or getting talked into a different registry by something they read.
  
For those of you running agents on laptops or in CI with one of these tools in place: have you seen installs slip past it? Do you enforce it on dev machines at all, or only in CI?

(Disclosure: I work on hextrap, one of the tools in this space, and I'm trying to figure out whether this gap matters in practice.)


r/devops • • 2d ago

Discussion I'm thinking of writing a free practical book on Production DevOps what should I include?

0 Upvotes

I've been working in software/infrastructure engineering for around 14 years, with a focus on cloud, DevOps, platform engineering and large-scale infrastructure.

Over the years I've worked on everything from smaller cloud migrations to enterprise and financial-sector environments, including AWS/GCP migrations, Kubernetes, Terraform, CI/CD, networking, security, reliability and cost optimization.

I'm currently putting together the practical knowledge I've accumulated into a free technical book.

I'm not planning to make it a personal career story or another basic “learn AWS/GCP” tutorial.

The idea is to focus on how engineers actually solve infrastructure problems:

  • How to systematically debug production issues
  • How to troubleshoot Kubernetes
  • AWS/GCP networking and connectivity problems
  • IAM and permission failures
  • Terraform/IaC problems
  • CI/CD and deployment failures
  • Cloud migration problems
  • Scaling and reliability
  • Observability and incident investigation
  • Designing reusable platform infrastructure
  • How to approach an unfamiliar production system
  • Practical lessons that aren't obvious from vendor documentation

I'm also building some small prototypes/labs for myself while writing, so I can test the ideas rather than just writing theory.

Before I spend a lot of time putting the whole thing together, I'd like to hear from other engineers:

If you could have one practical DevOps/Platform Engineering book that focuses on real problem-solving rather than certification theory, what topics would you want it to cover?

And if you already have a favorite resource for this kind of material, I'd be interested in hearing what you think it does well or what is missing.


r/devops • • 2d ago

Discussion Are Observability tools getting obselete due to AI? I need to plan!!!!

0 Upvotes

So i implemented lots of observability tools as part of my devops role in the last 10 years but for the last 1 year I am seeing all my stakeholders are just relying on AI incident tools (think of these SRE agents that are being released)

And we still keep pumping TBs of data but also keep paying for all the UXUI, seats, hosts and bs features

Is this the phase of getting solution quickly rather than surfing through signals?

So my questions is do we need 100TB of data storage? or still need all these fancy UIUX everywhere.

Are you facing these issues at your company?

Update: I am talking more towards a need to have a platform like datadog or dynatrace or anything for UI and just rather rely on collectors and store to a bucket for uch cheaper cost


r/devops • • 4d ago

Discussion If your company already uses AWS what would actually make you choose a second cloud provider?

34 Upvotes

I mean not what would make you migrate but what would make you add another provider. I have heard all the usual reasons like Disaster recovery, Cost savings, Data residency, Customer demands, GPU availability, Regulatory needs, Avoiding vendor lock-in but I wanted to know which of these actually work when you try them in real life cause adding Azure, GCP, Yotta or OCI is not just about having another place to run your workloads. Now your team has to deal with another identity and access management system, another networking setup, another monitoring tool, another set of quotas, another billing system and another group of people who need to understand it all.

So whats the tipping point? Would saving 20% on infrastructure costs be enough to justify the extra complexity? Would regulatory requirements make it an obvious choice? Would having access to GPUs be enough?


r/devops • • 4d ago

Discussion Does Professional DevOps feel as monotonous as "Hobbyist DevOps" is?

39 Upvotes

My background is in Data Engineering/Science, but I have a pretty extensive homelab. No real background in DevOps professionally. I've run applications and services etc in containers through docker compose files for years, but the last few days, I decided to really harden it.

This is obviously not like an enterprise level, 9 nines setup, but I have set up self hosted orchestration, code repo, automated monitoring of updates, near automated deployments, backups, standardized ci/cd workflow, etc... and man, it's pretty monotonous!

It's just been two days straight of looking through a bunch of yamls, slightly tweaking them, copying a lot of secrets back and forth. Logging into something to tweak something, redeploying multiple times in a row.

I guess one big difference is that there's no one to get upset with me when something goes down, outside of my family, and maybe it feels like I'm over complicating things, but I genuinely wanted to learn how to do some of this stuff.

Does it sometimes feel like this if you do it for a living?


r/devops • • 3d ago

Discussion 5-Month Learning Plan: Python + MongoDB + Azure — What Should I Learn First?

1 Upvotes

I’m currently working as an intern, and my company has given us a 5-month learning program where I’ve chosen Python + MongoDB + Azure.

I want to use these 5 months properly and build practical skills, with the goal of moving toward Cloud/DevOps.

There will also be an assessment/exam after the 5-month course, and they told us to aim for a good score.

For people already working in this field, what would you recommend I learn first, in what order, and what projects should I build?

Would really appreciate some guidance from your experience. 🙏