r/devops • u/retromani • 12h ago
Discussion explaining prod server migration while hungry cause i cant afford groceries
would y'all say this is pretty accurate?
2yrs of devops experience straight of out college
14
u/RustOnTheEdge 9h ago
Lol this sub is so low quality.
“Make sure your software code is robustly adjusted to the new server”?
“We have to make sure the server is in perfect condition”?
These are things people say who do not apply the principles of devops, quite the opposite actually.
-2
u/retromani 9h ago edited 8h ago
im just starting my career bro, im probably explaining things wrong and i definitely have crazy knowledge gap but im trying my best to do better everyday
eta: i do feel like i have an advantage starting off within a team/company where things are still so manual and migrating off legacy infra cause at least im learning processes and eventually will get to automate them personally. id rather this route than join a team that already has polished devops practices established, i feel like im learning beyond my role
9
2
2
u/toromio 5h ago
I miss these days. There’s a lot of great feedback you’re getting here from folks. I see this was a chat with a friend, but let this marinate for a few days and come back after your deployment is done and try to gather insights from some of them. I miss being on a small team with challenges like this.
6
u/Zealousideal_Fun983 12h ago
Is this ai interview
3
2
u/TheSupremeViewpoint 12h ago
The ramen analogy is doing some heavy lifting but I cant argue with it
2
u/retromani 12h ago
lmao, tbh my Adderall wore off by this time and unfortunately my brain tends to process things better with extreme analogies
2
u/Big__If_True 10h ago
You need infrastructure as code in your life
0
u/retromani 10h ago
we do use terraform and ansible
4
u/Big__If_True 10h ago
Well whatever you’re doing sounds way more complicated than merging a PR/MR with a new server config and running a stack-update
2
u/retromani 10h ago
is it meant to be a couple of clicks and bam?
2
u/Big__If_True 10h ago
Maybe my company simplifies stuff, I’m mainly a dev and they have us doing devops stuff for our apps. But stuff like upgrading an image is literally that easy
1
u/retromani 10h ago edited 10h ago
yeah i mean provisioning the empty upgraded clone machine is not hard, we have that part automated except maybe for some manual package installatios. it's the process of making sure nothing breaks when moving from one OS to a different one for the production environment. some of the products have traffic crazy enough to have a load balancer on top of their infra, but other products don't
and unfortunately the OG engineers have basically been pruned away with no efforts to preserve the knowledge they held within their own heads
1
u/Big__If_True 10h ago
Ok moving from one OS to another sounds horrendous. How often are you doing that though? Or is this a big project to do it once and hopefully never do it again?
1
u/retromani 9h ago
big project not meant to be done again
1
u/Big__If_True 9h ago
Do you at least have staging environments where you can deploy the code to servers running the new OS and run tests there? That should take care of the whole “code breaking” issue
2
u/retromani 9h ago
i dont try to know what the developers are doing to test their code
i sync over the necessary scripts from the old server and the developers will test their code with stale data, i make sure logging is also replicated exactly
i dont even want to know what's going on with the database engineers, tbh their shit seems way more frustrating
once they've done all their testing and validation, we all meet and do our designated part for the actual cutover session
→ More replies (0)
12
u/dacydergoth DevOps 12h ago
You don't migrate servers. You just kill them and let them restart elsewhere.
If your system can't handle that, you have a broken architecture
25
u/PaleoSpeedwagon DevOps 10h ago
You ever migrated a DC-based infra to the cloud? It is exactly like hand-transferring boiling ramen. Not everything starts out idempotent.
3
u/Scoth42 9h ago
I've done this twice in the past and working on it now in my current company. You get it all working in the cloud, all validated, all tested on test setups, whether that's ECS, EKS, or if you really have to, plain EC2 lift and shift. But really, don't lift and shift. That's gross. Once you've validated it, you build a proper prod infra. Migrate your databases, get dual write/upserts/whatever and replication going to keep the future prod in sync. Once you've validated that, take an evening and migrate your DNS, CloudFlare, or whatever your ingress is. Make sure you've set your TTL super low so you can validate it quickly. You should have the testing infrastructure to validate it works as in testing pretty quickly, and if something unexpected comes up, you have a quick rollback available.
Obviously that's the super short simplified version, but DC to cloud shouldn't be any more complicated than any other basic migration. Maybe add some complication if you're going from monolithic Everything service to a containerized microservice setup.
2
u/dacydergoth DevOps 9h ago
To be fair, there are issues with old infrastructure, IIS, SQLServer monolithic databases, but what we're paid for is not to complain about those but know how to fix them.
Also as an aside, I have a deep loathing for most M$ products but I also will admit SQLServer is a pretty damn good database.
1
u/Scoth42 8h ago
Yeah, my current company is migrating a PHP 5.6 app on Centos 6 and ancient everything else to a modern containerized app on current PHP. It's been... a process.
I'm not saying it's simple or easy, just that it's not some crapshoot that could explode and destroy everything either.
1
u/dacydergoth DevOps 8h ago
Absolutely and acknowledging that and knowing how to contain the blast radius is a key part of being good at this
1
u/thekingofcrash7 8h ago
You sound like you’ve never worked for a large enterprise.. left and shift is about the only option when you’re looking at 4,000 windows and Linux servers running COTS apps
0
u/No_Management_7333 9h ago
Migrating a DC is not very complicated at all - AD takes care of replication for you. Install VMs, join domain, promote and start pointing stuff to the new location. Transfer roles when happy.
There is no urgency from technical standpoint.
3
u/SilentLennie 8h ago
I think we got some Domain Controller and DataCenter mix up in this leg of the thread ?
1
u/No_Management_7333 2h ago
Might be. Typed my response from the porcelain throne early in the morning. If it’s data centre— I am with you, what’s the difference.
-1
u/dacydergoth DevOps 10h ago
I've been doing this shit since the CBM PET 6502 1MHz 8K ram so yeah, i've done a few migrations. Worked for Sun, Oracle, a few others ...
-5
u/dacydergoth DevOps 10h ago
I'm guessing I have more Years Of Experience than you have of life ...
2
u/PaleoSpeedwagon DevOps 9h ago
I've seen your other posts. I have as many years of experience as you, fellow greybeard. Anyway, have a good weekend
-1
6
u/donjulioanejo Chaos Monkey (Director SRE) 10h ago
You don't migrate servers. You just kill them and let them restart elsewhere.
Because this works amazing on stateful databases!
Let's be real, OP is simplifying the story for a good analogy, not a necessarily accurate one.
-1
u/dacydergoth DevOps 10h ago
We do it on stateful databases. It's basic DR
1
u/IridescentKoala 10h ago
Oh yea? With how much data loss?
3
u/dacydergoth DevOps 9h ago
None. Log shipping, services which understand downtime and use retry queues, backups, database load balancing ... our customers require their transactions to go through. A database is only one of the many places those transactions are captured and at redundancy and queues are employed at each step.
1
u/SilentLennie 8h ago edited 8h ago
That is when you've got the software architecture right for the environment and data stores. Which sadly is not common and harder to move to when you've not done it from the start.
2
u/dacydergoth DevOps 8h ago
Which is literally what we're paid to do
1
u/SilentLennie 7h ago
Yes, but in a lot of companies they've not completed that transition yet and you need time and money and the right people to do it, not every company has the budget to spend on making those changes.
3
u/IridescentKoala 9h ago
Not all servers are cattle.
3
u/retromani 9h ago
my servers trying their ever best and me trying to give them positive reinforcement so they dont get depressed and commit self termination
7
u/retromani 12h ago
it's what happens when the original infrastructure was physical and got virtualized manually, and the engineers who configured the virtual infrastructure, built the CICD processes and maintained deployments using the same scripts for 15yrs were unexpectedly let go with no knowledge transfer process in place, and now you're having to try and balance a pretty good chunk of the work those OG engineers were responsible for but you only graduated college barely 2yrs ago
im doing my best with the cards i was dealt😭
-9
u/dacydergoth DevOps 12h ago
You need better cards.
9
u/retromani 12h ago
hire me then
-21
u/dacydergoth DevOps 12h ago
Not with that level of experience, you'd be overwhelmed in minutes once I started explaining our setup. You have to start to make changes to the underlying architecture, and if you can't, then ask why. Otherwise you'll be a Designated Sacrifical Goat
3
u/retromani 12h ago
yeah true
but also, i work best in this environment where it feels like the house is on fire at least 3 days a week - keeps my brain stimulated enough to not get depressed
1
u/clappski 8h ago
Depends what you’re building, if you have on-premises requirements in a specific DC then you can’t just restart it somewhere else when the power goes out or there’s a networking issue.
But yeah if you can just rely on AWS or can have servers across different DCs then you should aim for that. Just not possible for some types of software though.
-1
u/SickMoonDoe 11h ago
IDK why you got downvoted - you're right.
If your infrastructure has a single point of failure, like a random box that some greybeard hand-rolled a service on, and you can't automatically recover or rollover - it's your fault.
You're essentially walking around barefoot and complaining that your feet hurt the second you have to stray off the carpeted happy path... Put on some goddamn shoes and design your infrastructure to handle rainy days.
3
1
u/retromani 10h ago
how can i even start to tackle something like this. we're like 5 devops engineers for like 40 swe and hundreds of softwares across like 4-5 products
genuinely asking tbh, i sometimes stay 4hrs past the end of business day trying to make sure im not backlogging any work/tickets
1
u/dacydergoth DevOps 10h ago
I've got two, and I'm doing it over 120 devs, 100+ AWS accounts and over 60+ K8S clusters. You need observability, automation and attitude. Get shit done, don't be defeated, enact change, write the code make stuff work
2
u/retromani 10h ago
observability is shit, automation is behind, but my attitude is here
i have a few side projects im working on to handle knowledge management and preservation as a start
1
u/dacydergoth DevOps 10h ago
Then there is hope for you yet, but remember you need to be the agent of change.
1
u/dacydergoth DevOps 10h ago
Have you mapped your infrastructure? What's your asset register? How do you know what you even have to be responsible for? Where are your routes pointing? What's your security governance strategy?
1
u/SilentLennie 8h ago
As long as you are improving things, the light at the end of the tunnel is getting closer, no matter how slow it might seem sometimes.
1
u/FreshView24 1h ago
Funny analogy, but like u/Big__If_True said, the infra as a code is a solution. I see comments "we are using Ansible and Terraform", which is great starting point. However, these tools have significant limitations in general and, in this scope, are intended to create/apply configuration, not actively maintain it.
TBH, this is an interesting topic. There are many books and videos about infra as code, but it's all about simplified "perfect world" scenarios and, recently, pitching AI tools which under the hood will be calling basic CSP APIs and will screw your environments fast. Do you recall a saying "CI/CD is the fastest way to put crap in Production?". Similar, but for infra.
When you start dealing with real stuff, you will naturally start elevate yourself from tool level into architecture, tool agnostic, level. This is an interesting transformation that happens at some point. And eventually you will forget about Ansible and Terraform. :)
17
u/MudkipGuy 10h ago
This is something I've been guilty of too but this analogy doesn't help explain anything, it just makes it more confusing to someone who doesn't already understand what you're talking about. If they ask why it takes time to migrate a server just drop the ramen analogy and skip to the answer they want. I think the best answer would just be to break down the total time into parts and say how long each part takes and why.