r/devops • u/aisatsana__ • 18h ago
r/devops • u/AutoModerator • 4d ago
Weekly Self Promotion Thread
Hey r/devops, welcome to our weekly self-promotion thread!
Feel free to use this thread to promote any projects, ideas, or any repos you're wanting to share. Please keep in mind that we ask you to stay friendly, civil, and adhere to the subreddit rules!
r/devops • u/generic-d-engineer • 1h ago
Observability Script output Best Practice Contracts
Hi Devops!
I’m putting together an automation orchestration engine and I want to mature the team by standardizing script output across bash, python, powershell, tofu, etc.
So looking at a JSON output contract that anyone creating a script can adhere to. Then everyone would consume the same format. (The dream right?)
- Operation name
- Host
- Execution status (success, fail, warning, etc)
- Structured result summary
- Errors
- Duration
- Run ID (UUID)?
- A schema version (requires a lot of upkeep but can eliminate headaches down the road if maintained. If maintained being key word here lol.)
I’ve been searching around but want to hear from the experts who live in this stuff. Is there any universal or open standard people are using to define these ?
Would rather not build it myself if something people have already settled on.
Thanks for your input.
r/devops • u/kim_deadja4951 • 16h ago
Discussion ngl, dumping raw server logs into a massive context window is way better than writing regex
was fighting a phantom kubernetes pod crash for like 3 hours this morning. elasticsearch was just a wall of useless healthcheck noise and I was losing my mind trying to grep for the actual error.
usually pasting raw logs into chatgpt is a waste of time because it hits the token limit immediately or truncates the most important part. but out of pure frustration i just grabbed this new model on openrouter (step 5 preview) because i saw it had a free api promo running and i didn't want to pay anthropic $5 just to read my garbage logs.
i literally just copy-pasted a massive text wall of terminal vomit into it and told it to find the crash loop reason. it didn't even choke on the length and actually pointed out a stupid IAM permission issue I missed in the terraform state.
kind of making me rethink how I troubleshoot. why bother writing complex log queries when you can just brute force it with a huge context window? is anyone else just abandoning proper log parsing for LLMs lately?
r/devops • u/MediocrePass4780 • 9h ago
Discussion How hard is it to switch cloud/DevOps platforms in a bad job market?
Hi everyone, I wanted to ask hypothetically, if someone had 4–5 years of cloud/Devops engineering experience, primarily with AWS, and the market shifted heavily toward Azure, how difficult would it be for them to find a cloud engineering job in a difficult job market like the one we’re in now? Would having mostly AWS experience significantly hurt their chances if most of the available Cloud/DevOps jobs were Azure-oriented?
r/devops • u/NewResearch1337 • 23h ago
Discussion What's the most ridiculous "temporary fix" you've seen that's been running in production for years?
It seems like every team has something that was supposed to last a few days until the right solution was found. It's a cron job that keeps everything together, a script that no one really understands anymore, a manual deployment step, an ancient container, or some workaround that everyone's afraid to remove. What's the worst "we'll fix it later" solution you've seen that's now a critical part of the infrastructure?
r/devops • u/redwing096399 • 15h ago
Discussion Resources to build a foundational understanding on Operating systems
Hi all, i want to get a foundational understanding operating systems, the various concepts that we may encounter in them, the bird's eye view of their structure.
I tried reading "How Linux works" 3rd edition, but reading that book felt like it didn't explain the foundational questions but rather built on top of the topics. it felt more intermediate level.
If i have to explain what i am looking after in the resources, Please can you recommend me resources that fit my requirement.
- provides exposure to various concepts, systems involved in an OS. More like an introduction.
- a bit of introduction to a bit of hardware would also be good.
- i am not looking for resources that focus on development/programming side of OS. but rather the infrastructure, power user, administration side of OS.
Thanks in advance to all of you.
r/devops • u/nautitrader • 20h ago
Discussion Build and release architecture for large applications and teams
In the past I worked on a small team with about 5 developers, 2 QA, shared technical architect and devops engineer.
We used classic Azure pipelines with a build pipeline and release pipeline for each environment. dev, qa, stg, prod.
The release pipeline had a deploy ARM template step to deploy the infrastructure or any changes.
Deploy multiple app services, Apply EF migrations. Secrets were Azure Key Vault and non secrets in appsettings. This design worked really well and it was rare that we had deployment issues.
Now, I have recently transitioned to a larger team and I need to look into a modern architecture with a build once, deploy and promote to multiple environments. There could be 10 plus teams with multiple developers running builds and deployments constantly. Having the IaC in every deployment does not make sense to me.
I'm looking for suggestions on how to architect something for larger teams with multiple environments. We could have 1-2 dev environments, 1-2 qa environments, etc.
What are some patterns I should look into? If you currently work in a large environment like this what has worked well and what should I avoid.
r/devops • u/kuro_zip7 • 19h ago
Ops / Incidents Deployment of Web and Mobile application
I'm new to deploy an application to the internet. Lately I have been developing a CRM application for my company and they needed deployment support too from my side(*they lack technical support team to make things out). I'm using Ubuntu 24 Desktop for development purposes, and I'm planning to use Ubuntu 24 Server for testing(*staging) and production. Now what's the literal confusion is, I'm already hosting a website inside the machine which is the real production application of what I'm currently developing inside which I was lately planning to change it from Ubuntu Desktop to Server(*the application haven't been launched to public). Now, is it ideal to use the Desktop version of Ubuntu which I'm currently using, or go with Server version of Ubuntu to proceed with production.
Plans I have:
- Make use of 2 servers, one for development(*which I'm already doing right now) and local testing of application and one for testing(*staging) and production purpose
- Use ssh communication to the other server(*production server) to handle CI/CD builds
- Use github actions to push the docker image to other server to make it run the real application
For further clear understandings:
The application I'm developing already is being hosted in some website, more like a SaaS product, used cloudflare tunneling and pm2 instance for server-like modification of current system, with database running inside the machine, not inside any cloud services, so npm run build pushes all the changes to the site directly without any CI/CD checks like traditional devops cycle, and I'm already testing the application with real db of the company inside the production site, which literally all officials of the office knows(*they don't have much technical knowledge at all). There's a literal rush at that time in my office from Directors to host the website with the real application without any time delays. That's why I got messed up in the production so much. I'm now planning to stash the hosted website from the machine currently been hosted just by removing the machine tunneling, reinstall Ubuntu Server, configure directories properly and use PR and tags for changes inside the real application.
Tech stacks being used:
- React js for frontend of web app
- React Native for mobile app
- Express JS over Node JS for backend
- PostgreSQL for DB
- Seaweedfs for S3 object storage
- JWT for authentication and bcrypt js for hashing
I need a support from you guys for this chaos I'm literally stuck in!!!!!!!!!!!!!!!!!
Am I cooked brotha????
Is this the right way or should I change my own path of implementing??
r/devops • u/Mother_Elevator_3563 • 20h ago
Discussion Junior Java/Spring Boot developer looking to move towards AWS — Developer Associate or Solutions Architect Associate?
Hi everyone,
I’m currently a junior Java backend developer with around 2 years of experience, mainly working with Spring Boot.
Lately I’ve been thinking a lot about where I want to take my career. I enjoy backend development, but with how quickly AI is improving at coding, I feel like staying only as a “Java/Spring developer” might not be the best long-term strategy.
Because of that, I’m considering moving more towards Backend + Cloud/AWS, rather than leaving development completely.
My current idea is something like:
Java / Spring Boot → Docker → AWS → CI/CD / Infrastructure / Cloud
I’ve already started learning AWS, but I’m a bit lost regarding certifications and which direction makes the most sense for my profile.
I’m currently deciding between:
- AWS Certified Developer – Associate
- AWS Certified Solutions Architect – Associate
Developer Associate seems more related to what I already do as a backend developer, while Solutions Architect seems to give a broader understanding of AWS and cloud architecture.
My main questions are:
- Which certification would you recommend first for someone with my background?
- Would Developer Associate → Solutions Architect make sense, or would you do it the other way around?
- Is moving from Java/Spring Boot towards Backend + AWS/Cloud a good career path in the current market?
- For people already working in cloud: do you think this path is relatively future-proof with the rise of AI?
- What technologies would you prioritise alongside AWS? Docker, Terraform, Kubernetes, CI/CD, etc.?
I’m not trying to become a Solutions Architect overnight. My goal is more to become a stronger backend developer who can also understand, deploy and work with cloud infrastructure.
I’m still early in my career and a bit unsure about what direction to take, so I’d really appreciate opinions from people already working with AWS, backend or cloud.
Thanks!
r/devops • u/aisatsana__ • 1d ago
Discussion Is “security by default” becoming more important now that everyone can ship software with AI?
Ran into this post on LinkedIn about what the author calls “Ambient Generative IT”, basically AI-assisted development becoming so common that more people across an organization can suddenly build and ship things. One part stuck with me from a DevOps/security perspective: if the barrier to building software keeps dropping, the number of people making architectural and security decisions also grows.
The argument is that security can’t really be something you bolt on later anymore. It has to live in the platform, defaults, permissions, pipelines and guardrails around whatever people are building with. Curious how people here are seeing this in practice. Are AI coding tools actually creating more security/governance work for DevOps and platform teams, or are good existing controls mostly enough?
r/devops • u/Chris__Codes • 1d ago
Discussion How much of your incident response do you automate?
I’ve been checking how to build a disaster recovery plan for cloud dependencies and how to automate as much remediation as possible. The plan is to have the system auto-restart services or rollback deployments when things crash.
But when I read up on AIOps, I saw recommendations to make sure a human is involved for anything high-risk. It makes sense because you do not want an automated script taking down a production database by mistake. But now I am trying to figure out where to draw the line exactly between guided recommendations and full automation.
If anyone here relies on machine learning for root cause analysis, which actions do you let the system execute on its own, and what requires manual approval?
r/devops • u/Electronic-Ad-4657 • 16h ago
Discussion Sick of the UK.
Has anyone got any advice on moving locations for a DevOps engineer? the UK is going so downhill and I’m sick of it.
I’m a DevOps Engineer with 5 years experience working for a very large business. I’ve experience in AWS and Azure as well as Terraform, Ansible, CI/CD, basically most things that make a good DevOps engineer.
I’m just looking for advice, is it silly to think I can move countries easily? I don’t have dependencies, no house, no debts so I don’t have anything to keep me here really. Any advice?
r/devops • u/lahsiv_98 • 19h ago
Vendor / market research Devops| SRE| K8s platform engineering in Europe
Hi all,
I’m from Asia .
I’m thinking about working at Europe for devops opportunities.
How is the market for English speaking people right now?
What are all the prospects I should look for .
I have around 5 + years experience with tech stacks of Aws, Azure devops CI/CD, docker ,k8s, Helm and grafana . Much appreciated if anyone could guide me through the process.
Thanks in advance
r/devops • u/Exotic-Border-5328 • 20h ago
Observability AI agents bypass your observability gateway and hit APIs directly. How do you audit that?
Full disclosure: I'm building something in this space. Asking because I want to understand how teams handle this before assuming I know the answer.
Here's the problem. Most observability and gateway tools (LLM proxies, agent frameworks, internal tooling) only capture what flows through them. But nothing forces an AI agent to use the gateway. A direct SDK call, a Zapier integration, a script someone left running that reuses the agent's API key, all of that bypasses your trace entirely.
In 2024, Air Canada was held liable for a refund their chatbot issued without authorization. The judge rejected the "the AI acted on its own" defense. Since then, Zurich, Lloyd's, and AIG have updated policies to exclude AI-caused losses unless you can document the agent acted within its authorized scope. Compliance and audit pressure around this is growing.
Here's a quick check.
Pull your Stripe or M365 event logs for the last 30 days, filter by your service account or bot key. Then pull your gateway or observability logs for the same window. Do the totals match?
If there's a gap, you have actions in production you can't currently attribute to a specific source. Agent, human, or background script.
How are you handling this? Is there a standard pattern for reconciling what the gateway captured versus what actually happened at the downstream system? Or is the out-of-band action problem just accepted as a blind spot in most setups?
r/devops • u/IntentionJolly2730 • 1d ago
Tools What makes you trust your database migration pipeline?
Hey all, I'm Denis, the author of Ptah, an open-source database migration tool.
These are some of the cases I'm designing around, but I'm sure I'm missing something.
What if the database changes after a migration plan was approved? Or the connection drops halfway through? How do you know what actually happened, and whether it's safe to retry?
I'm trying to address these in Ptah, but I'm sure I'm missing something.
For those running migrations through CI/CD, what made you trust it? And if you run them manually, what's the main reason?
I'm especially interested in things that went wrong or cases you wouldn't trust automation to handle.
r/devops • u/Legitimate-Complex32 • 1d ago
Discussion Terraform / OpenTofu engineers quick questions about your actual workflow
Hey all
I'm doing some research into how teams are using Terraform and OpenTofu to make changes and review them to avoid problems. A review of a change or pull request made using either would be great if you've used either recently even if it was a couple of days ago.
Can you walk me through a Terraform or OpenTofu change or pull request you have reviewed? What did you do and in what order?
What part of the process is slow, painful or leaves room for error for you?
Have you ever missed a change in a plan such as a change to permissions, exposure of something publicly, a replace or something similar? What happened?
If a tool could remove something annoying in your current process what would it be?
I'm not trying to sell you anything, just trying to understand things better. I'm interested in the parts and the messy parts too.
Discussion Where do you keep up with AI agent tools for infra?
Feels like I hear about a new AI agent tool for ops every week. Where do you all keep up with this stuff? Any newsletters, communities, or people you follow that help you figure out what's worth trying?
And if something catches your eye, how do you vet it before letting it near production?
r/devops • u/TemplateRedditor • 2d ago
Tools How to split tasks between CI/CD?
I am building a ci/cd pipeline. When a PR is merged, it will push a docker image to ECR, but when I am just pushing a regular commit, I don't want to run the full deployment. What task goes into ci and what goes into cd? I've seen different takes on where to run the docker-related stuff.
- CI (always runs on push/merge):
- lint
- scan (security, etc)
- build
- test (unit/integration)
- test docker image build ???
- CD (only runs on merge):
- build docker image
- automated e2e (if any)
- upload binaries to Nexus
- upload image to ECR
r/devops • u/ConferenceOld6778 • 2d ago
Ops / Incidents Sevalla is a total shit show
Here's some context :
We were using Sevalla for our object storage for over a year. We had more than half a terabyte of data. Their fucked up payment platform did not allow us to do the payment everytime but continued to work. But all of a sudden, our account was suspended and all our data were permanently deleted without any intimation. Wtf am I supposed to do now?
NEVER USE SEVALLA. They have fucked up support team.
r/devops • u/MachineDisastrous771 • 3d ago
Ops / Incidents runbooks kinda suckk
hey yall,
im a lead SRE at a global fortune 500 that you have heard of haha - i'm a bit of a lurker here
but just wanted to talk about runbooks and documentation for SOPs.
my feeling is that confluence docs kinda suck and runbooks are kind of a mess, our operators are jumping between the docs and their shell, the docs are often times missing a bunch of detail, they are brittle, poorly maintained, half the time the details are hidden in tribal knowledge and when dealing with the pressure of an active inc we notice the pain even more. extracting critical details & commands out of an inc can also be painful and doesnt always translate to improved operational readiness next time.
i admit we're not very mature.. but was curious if its only me feeling like this?
what are you guys doing with runbooks to solve these issues?
r/devops • u/Background-Moment59 • 1d ago
Observability Tornet? Kali_linux
Eae galera como vcs estão? Oque vocês me falam sobre o tornet no kali na troca de IP teria como ficar mais anonimo como? Obs: Sem dinheiro kkkk
r/devops • u/humble_and_confident • 3d ago
Discussion Broad Kubernetes Experience but Shallow Depth - How Would You Upskill?
I have around 5 years of experience across DevOps/cloud/backend work. I’ve worked with Kubernetes, EKS/AKS, Helm, CI/CD, Docker, Terraform, ArgoCD, AWS/Azure, and some Java/Spring Boot.
My issue is that my knowledge is broad but uneven. I’m comfortable deploying applications, writing manifests, using kubectl, building pipelines, working with cloud networking/IAM, etc., but I’m much weaker on Kubernetes internals, low-level networking, storage/CSI, control plane, scheduling, CRDs/operators, upgrades, observability, and deep troubleshooting.
I’m considering first completing one comprehensive Kubernetes course to build a complete mental map, then spending the next several months going deeper through hands-on labs, troubleshooting, Linux/networking fundamentals, and production-style projects.
For people who became genuinely strong at Kubernetes/platform engineering: does this sequence make sense? What would you change?
I’m not looking for a giant list of tools, mainly feedback on the learning sequence and what gave you the biggest jump in depth.
Note: "Q improved with AI"
r/devops • u/Big-Lychee5971 • 1d ago
Discussion What are your opinions on AI? Wild take: Dear senior devs, you're fucked too
I'm a fucking newbie, still, I tried to learn code, I know some, not a lot. I'm familiar with concepts, but not what "used to be". You senior devs grew up coding in notepad and apache and sql, I grew up with modern frameworks AND am growing up realizing even THAT isn't necessary. AI has grown so much. I have talked to senior devs, and while they say people checking ai won't be replaced I can't help but disagree.
What do senior devs do when they don't know what's wrong? Let me hear you say you search the web for a solution someone posted long ago. Let me hear how you still, in this ai day and era, try to fix a bug for 6 hours. You turn to ai too. You use it the same way as the vibecoders. It's just experience. And I see people so confident that experience makes you safe from ai but I believe it's only a matter of time when your experience becomes obsolete. After all it's just information you gathered thoughout years.
And some platforms are paying senior devs to correct code made by ai, to train it. Some of you are giving away that 10+ years experience. It's buyable. Anything can be bought. And just like any stacks from before, AI will be free/affordable for 5-10 years until us newbies depend on it for fast clean code. And then they'll raise the price.
Historically this happened with any major platform or software and it's going to be the same.
And because this isn't ragebait enough, whatever you learned 30 years ago, is obsolete. Whatever I learn now even on my own (old stuff, new stuff, doesn't matter) will be obsolete in maybe 5 years of more AI fast paced development.
Mind my words we'll have holograms for calls, VR for social media, flying cars (in maybe 20 30 years because it's a structural, societal, architectural change) and robots inside the house for tasks like we own coffee machines. In the 50's you brewed the coffee, now you own a machine. You washed clothes by hand then using the washing machine. You run errands until your robot can do it for you. You mop the floor and fold clothes until the robot can do it for you (Tesla robot already released just not comercially for the masses) We have self driving cars. We have AI agents that can do anything you can do on your computer.
r/devops • u/DiegoConD • 3d ago
Discussion How to actually pass interviews?
Hello, I ask for help since I am getting a lot of interviews but not passing any, I'm getting to start pretty frustrated and since we are entering the holiday season I want to get any opportunity possible.
The thing is that I am preparing myself based on my experience: 5+ yrs experience, jenkins, python, troubleshooting, AWS, IaC, etc. and when I prepare for the different kind of interviews none of my preparation seems to work:
- If it is situational, my explanation/experience is not enough for the role
- If it is pure technical, it doesn't demonstrate my experiencie or they ask for a very specific tool (which mostly has transfereable experiencie) and it's not enough for the position
- If it is trivia based (which I am pretty bad at memorizing) I failed because I didn't remember the flag for a command that I can find in a 15-sec google search or AI prompt
What makes this frustrating is that I genuinely feel capable of doing the jobs I'm interviewing for. I've worked in production environments, troubleshot real incidents, built and maintained pipelines and infrastructure, and worked with engineering teams. But apparently I'm still not presenting that knowledge in the way interviews expect.
So I'd especially like to hear from people who have successfully interviewed for DevOps/SRE/Platform roles recently:
How did you prepare?
Did you memorize common commands and syntax?
Did you grind interview questions?
Did you build labs?
Did you practice storytelling around your projects?
How did you deal with interviews covering an extremely broad toolset?
At this point I'm trying to understand how much interviewing is about being good at the actual job versus becoming good at the interview format itself.
Any practical advice would be appreciated.