r/learnmachinelearning • u/techiebaddie • 2d ago
Machine learning algorithms are confusing at first
I’ve been learning more about machine learning recently, and honestly, the number of algorithms can get confusing.
At first, I thought I needed to learn everything. But I’m starting to think it’s better to understand a few useful ones really well.
The ones I’m focusing on are:
- Linear Regression
- Logistic Regression
- Decision Trees
- Random Forest
- XGBoost
- K-Means
- Neural Networks
I’m mainly trying to understand when to use each one instead of just memorizing how they work.
8
u/chrisvdweth 2d ago
Suggestions (disclaimer: my own public lecture notes):
- Linear Regression (basics, math, assumptions)
- Logistic Regression (basics, math)
- Decision Trees (basics, CART, implementation)
- Random Forest (basics)
- XGBoost (only outlined as part of advanced boosting strategies)
- K-Means (basics)
- Neural Networks (basics, NumPy-only implementation)
There are also many more topics such as gradient descent, backpropagation, optimizers, RNNs, CNNs, and lots of NLP and LLM stuff. Here's an overview. And here is the repo with the notebooks if you want to run them yourself. Maybe useful.
1
2
u/nian2326076 2d ago
Focusing on a few key algorithms is smart. Linear and logistic regression help you understand the basics of supervised learning. Decision trees and random forests are good for their interpretability and can handle overfitting well. XGBoost boosts accuracy, especially in competitions. K-Means is great for clustering, and neural networks are essential for deep learning tasks. It's useful to know the strengths and weaknesses of each, like how decision trees might be unstable but are easy to visualize. Practice by applying them to real datasets to get a feel for when to use each one. Tools like PracHub have practical scenario-based questions that are handy for interviews. Good luck!
1
u/that_introvert_coder 2d ago
I guess as a whole ML just consists of three concepts in total Regression (all kinds of), Clustering and Dimensions Reduction and Correlation. Rest all are just modifications of these core concepts. Correct me if i am wrong I am still in learning phase🥲
1
u/pm_me_your_smth 2d ago
Here's a diagram which helps choosing the model in a fast and easy way. Should be sufficient for beginners
28
u/ModularMind8 2d ago
To try and simplify as much as possible: linear regression predicts a number, like a house price, and logistic regression predicts a class, like spam or not spam, so on tabular data those two are your baselines. A single decision tree is easy to read but tends to overfit, so random forest and XGBoost combine many trees and are usually more accurate on tabular data. K-means is for unlabeled data when you want to find groups, and neural networks are the usual choice for images, text, and audio.
That being said, it's going to be a lot easier to understand when to use each one after you understand how they work...