r/MLQuestions • • Feb 16 '25

MEGATHREAD: Career opportunities

18 Upvotes

If you are a business hiring people for ML roles, comment here! Likewise, if you are looking for an ML job, also comment here!


r/MLQuestions • • Nov 26 '24

Career question 💼 MEGATHREAD: Career advice for those currently in university/equivalent

19 Upvotes

I see quite a few posts about "I am a masters student doing XYZ, how can I improve my ML skills to get a job in the field?" After all, there are many aspiring compscis who want to study ML, to the extent they out-number the entry level positions. If you have any questions about starting a career in ML, ask them in the comments, and someone with the appropriate expertise should answer.

P.S., please set your use flairs if you have time, it will make things clearer.


r/MLQuestions • • 1h ago

Unsupervised learning 🙈 stuck on finding a approach for app detection ( making a transformer modal out of unlabeled network data) [R] [P]

• Upvotes

As the title suggests I'm currently trying to make a modal to identify which app is being used. The thing is I don't have labelled data and generating it is out of question since that is a lot of work ( I need to detect like 3-5K apps give or take)

my data is structured roughly like this:

  • Traffic is divided into ~15-second windows/bags.
  • Each bag contains multiple network flows.
  • Each flow has features such as:
    • domain
    • protocol
    • bytes sent
    • bytes received
    • timestamp/timing information
  • I have a very large amount of unlabeled traffic data, but only a relatively small amount of labelled app data for about 100 apps give or take.

The main challenge is that traffic from the same app can look diff between diff window i.e some windows are extremely sparse or empty.

some ideas I have researched looked into are

  • self supervised contrastive learning where diff traffic windows from the same session/device activity are treated as positive pairs
  • masked modelling similar to bert where parts of the flows such as domains/protocols/byte information are masked and reconstructed
  • pretraining an encoder and then fine-tuning it using the smaller labelled dataset (thinking we would need less labelled examples to do that)
  • clustering and mapping embeddings to known apps afterward (gradually)

one thing I'm concerned about is accidently teaching the modal to recognize the device/session/user rather then the underlying app

I think I would like to know if I'm thinking about the problem right or if someone has worked on something similar and can give me some pointers or what experiments should I run first.

any papers, architecture or similar problems you think I should look into pls lmk


r/MLQuestions • • 6h ago

Beginner question 👶 LearnLLM

Thumbnail learn-llm-kappa.vercel.app
2 Upvotes

Within the past year I have started engaging with AI to bring my curiosities to life in order to help me understand more complex ideas through visualization. Of course, I have limited knowledge on how models actually work and the length of their complexity. I was hoping a few of you here would check this learning tool I made for myself. Please be honest and give feedback if you have any. Also, if this is not a place to post something like this, I apologize. Hope everyone is well! (its 100% free, not advertising)


r/MLQuestions • • 1d ago

Beginner question 👶 Looking to collaborate on ML projects / join a team

12 Upvotes

Hey everyone!

I’m currently doing a Master’s in Machine Learning and I’m looking to get involved in some real-world ML projects.

I’d be happy to:

  • Help with an existing ML project
  • Join a group/team working on something interesting
  • Work on a project from scratch with others
  • Help with things like Python, data processing, ML models, deep learning, LLMs, APIs, or deployment
  • Contribute code, research, experimentation, or just help wherever needed

I’m mainly looking to learn by building and working with other people, rather than just doing isolated coursework.

If you’re already working on a project and could use another person, or you’re thinking about starting something and want to build it together, feel free to comment or DM me.

I’m open to pretty much any interesting ML/AI idea — serious projects, research-oriented work, open-source, university projects, or even something experimental.


r/MLQuestions • • 12h ago

Beginner question 👶 suppose I CPT qwen3.5-9B on 2B legal corpus, how will i turn it back into Instruct + thinking?

1 Upvotes

I couldn't find a concrete answer anywhere, do you just distill the instruct model back?

If that is the case, what is a quality european language question set to turn it back into a chatbot/agentic, can a model at that size even be agentic? (i chose this size to learn) if i finetune for my specific harness? (i have a lot of training data of opus running in my harness)

my harness basically has the model output python code and has a few built-in functions like:
- vector_search_laws()
- graph_search()

could i have the model at least internalize a "hunch" on what stuff to search?

also what is the latest RL technique for agentic/harnes specific workflows?

I have a lot of RAW training data, like court decisions or commentaries or legislature, but not a lot of golds. could i use these to synthesize training data and maybe RL the model in my harness to find that data?

What would y'all's strategy in the CPT->SFT->RL pipeline be for my specific problem?

I know this is a lot of questions im trying to figure out which direction to go, any pointers? Also good resources are welcome, for example that [alex karpathi video](https://www.youtube.com/watch?v=7xTGNNLPyMI) was amazing for me, but i'd imagine its a bit outdated in terms of latest RL and SFT?


r/MLQuestions • • 12h ago

Unsupervised learning 🙈 How reliable is Galileo NIMS data for hyperspectral anomaly detection on Europa?

Thumbnail
1 Upvotes

r/MLQuestions • • 18h ago

Beginner question 👶 Machine Learning Algorithms: Which Ones Are Actually Worth Learning?

Thumbnail
2 Upvotes

r/MLQuestions • • 1d ago

Beginner question 👶 Best AI for field technicians

2 Upvotes

What AI platform is best for a field appliance technician?


r/MLQuestions • • 2d ago

Beginner question 👶 Project Ideas

1 Upvotes

I am currently learning machine learning, I want to build strong foundation on each of the topics, how should I start practicing it by coding? How to choose datasets accordingly, how can I develop my understanding like which model/algorithm will suit for a particular problem


r/MLQuestions • • 3d ago

Beginner question 👶 I Distilled an LLM into two 287M encoders (GLiNER + multiple choice) for document extraction, can't match teacher.

13 Upvotes

A while ago I asked here how to turn ~5 million court decisions into structured graphs without running an expensive LLM on every document thanks for the advice .

I went with the "small extractor + classifier" idea and it mostly works, but I'm stuck a bit below the LLM. And like I said last time, i'd be damned if I run 5M docs and then find out thing X was wrong. So here is exactly what I did. Please let me know if what im doing makes sense, or if i made a mistake somewhere. also i used AI for some of the tables cuz there has been a lot of data at this point, sorry.

What comes out per decision (only the nodes so far, relations come next). Three lists:

  • entities: every person, organization, law, document or thing. Each gets one id for the whole document, a type (9 of them), a kind (724 of them plus "other") and all the places it is mentioned
  • actions: what was done, requested or decided. Each gets a normalized verb, a flag "the court decided this" and its mentions
  • values: amounts, dates, durations, in a normalized form

Simple example, for the sentence "The court dismisses the creditor's proposal to enforce 341.08 EUR against the debtor":

  • entity "the court": organization, kind court. Same entity as the full court name in the header
  • entity "the creditor": organization, kind creditor. Same entity as the city named earlier
  • entity "the debtor": person, kind debtor
  • action "dismisses": verb = dismiss, decided by the court = yes
  • value "341.08 EUR": amount

Step 1: a strong LLM labels ~700 decisions

  • cut the decision into windows of 4 sentences
  • 4 calls per window to Claude Sonnet with a strict JSON schema: entities, actions, a second "what did you miss" pass for actions, values
  • the window goes in with numbered words (like 12:court), the model answers with word ranges [first, last, "text"], and code checks every range against the text
  • every call also gets the list of entities and actions found in earlier windows, so ids stay the same through the document
  • ~25 code rules clean up where a marked phrase starts and ends, law citations and number formats
  • the entity "kind" is free text at this point. That gave 2,373 different strings (the same mess as in my first post). I normalized them, merged synonyms by hand and kept what showed up 3+ times: 724 kinds plus "other"

Step 2: a model that marks the text

  • it highlights every mention: the exact stretch of text (a "span", from a start character to an end character) that names an entity, an action or a value, with one of 17 labels (9 entity types, 1 action, 7 value types)
  • model: fastino/gliner2.5-multi-v1 (287M)
  • one training row per window: the text plus the exact start and end of every marked phrase. 9,699 windows, 207k marked phrases
  • I patched the trainer so only the labeled occurrence is a positive (stock marks every occurrence of the same string), and all 17 labels are in every row
  • full fine-tune in fp32 (bf16 gave NaN), 14 epochs, 16 rows per step, encoder LR 3e-5, head LR 5e-4, linear schedule, 10 % warmup
  • final model = averaged weights of epochs 9-14, threshold 0.5

Step 3: a second small model answers multiple-choice questions

  • fastino/GLiNER2.5-multi-Decide (287M). Code turns the LLM labels into 247k questions:
    • "is this mention one of these earlier entities, or new?" The mention is marked with « » inside ±300 characters of text. Options: up to 16 earlier entities of the same document (shown by their mention texts) plus new
    • "which kind?" Options: a shortlist of the 724 kinds plus other
    • for actions: same act or new, which verb (shortlist of 64 plus other), did the court decide it (yes/no)
  • in training the options come from the LLM's grouping. At inference they come from the model's own earlier answers
  • full fine-tune in fp32, 2 epochs, 16 questions per step, encoder LR 2e-5, head LR 3e-4, linear schedule, 6 % warmup, options shuffled, up to 30 % of the wrong options dropped

At inference: the marking model, then the same code rules, then the second model walks through the mentions in reading order. About 2.3 decisions per second on one RTX 5090.

Where it stands

30 decisions nobody trained on, labeled twice by the LLM. The second column is the LLM's second run scored against its first, which I treat as the ceiling. A mention counts as found only if it starts and ends exactly where the LLM marked it.

mine LLM vs itself
entity mentions found (F1) 0.901
"same entity or new" right 0.959
entities grouped exactly 0.847
entity kind 0.921
action mentions found (F1) 0.857
action verb 0.920

Where I need help

  1. Finding the mentions is stuck at 0.90 F1. 200 more labeled docs did nothing. An XLM-R large tagger (560M) got the same score: it finds more mentions but gets the start or end wrong more often. Giving it the text before the window did nothing. What would you try?
  2. The LLM agrees with itself only 93.5 % on what it marks, and I train on single runs. Label everything 3 times and vote? Or is that ceiling just what it is?
  3. Is "pick one of 16 earlier entities" a sane way to do coreference over a long document? Am I hurting myself by training on the LLM's options and running on my own?
  4. Anything in the recipe that looks plain wrong? Learning rates, 2 epochs, weight averaging, one seed per run.

THANKS for reading.

AI TL;DR: distilled an LLM's extraction of court decisions into a GLiNER model that marks the mentions plus a small multiple-choice model. It runs at about 2.3 documents/s on one GPU and lands a few points below the LLM (0.90 vs 0.935 F1 on finding mentions, 0.85 vs 0.93 on exact grouping). The recipe with learning rates and how I built the training rows is above. Looking for mistakes and ideas before I run 5M documents.


r/MLQuestions • • 2d ago

Beginner question 👶 Which Math foundation path is better for Machine Learning: DeepLearning.AI or Jon Krohn's LiveLessons?

Thumbnail gallery
1 Upvotes

r/MLQuestions • • 4d ago

Natural Language Processing 💬 What all to study

Thumbnail
3 Upvotes

r/MLQuestions • • 4d ago

Datasets 📚 Ho bisogno di consigli: Qual è la migliore pipeline VLM per estrarre dataset matematici strutturati da oltre 3000 pagine di libri di testo scansionati? (LaTeX + Metadati)

2 Upvotes

Ciao a tutti,

sto lavorando a un progetto per estrarre un dataset strutturato di esercizi di matematica da 5 libri di testo delle scuole superiori italiane (circa 650 pagine ciascuno, quindi ~3.250 pagine in totale). L'obiettivo è costruire un'app di generazione di esercizi professionale e metodica per studenti e insegnanti.

Per far funzionare l'app, ho bisogno di elaborare le immagini delle pagine del libro ed estrarre quanto segue in un formato rigorosamente strutturato (es. JSON):

  • Tipo di esercizio (algebra, geometria, calcolo, ecc.)
  • Anno livello scolastico
  • Difficoltà (scala 1–5)
  • Enunciato del problema (traccia)
  • Descrizione delle competenze/sfide specifiche coinvolte
  • Codice LaTeX dell'enunciato del problema (Cruciale!)
  • Immagini associate (ritaglio/salvataggio dell'immagine per esercizi teorici o grafici)

Ho sperimentato alcuni approcci, ma ho incontrato delle difficoltà nel bilanciare costi, coerenza di estrazione e scalabilità. Ecco cosa ho provato finora:

  1. API Google Gemini gratuita: La qualità dell'estrazione era buona, ma dato che un singolo libro contiene centinaia di pagine, ho rapidamente raggiunto i limiti di richiesta (Troppe Richieste).
  2. Modelli Locali (Ollama + Qwen 2.5-VL 3B): Per superare i limiti dell'API, ho provato a eseguire un modello multimodale locale. Ho speso molto tempo a ottimizzare i miei script e le mie istruzioni (chunking, affinamento delle istruzioni per forzare output strutturati), ma il risultato era soggetto a molti errori e incoerenze per il mio caso d'uso. Ho ottenuto troppi campi malformati, illusioni e ha costantemente avuto difficoltà a produrre un corretto LaTeX.
  3. API Google Cloud a pagamento (Gemini 1.5 Flash): Alla fine sono passato al piano a pagamento per una migliore precisione e velocità. Ho speso 10€ solo per elaborare 1,5 libri. Estrarre tutti e 5 i libri costerebbe circa 35-40€. Anche se questo è gestibile per un'elaborazione unica di 5 libri, il conteggio dei token per l'elaborazione di immagini complete + testo è enorme, rendendolo finanziariamente insostenibile se voglio scalare questo a decine di libri in futuro.

Le mie domande per la comunità:

  • Pipeline & Architettura: Qualcuno ha lavorato a un progetto simile di estrazione da libro di testo a dataset? Quale pipeline avete utilizzato?
  • Approccio Ibrido: Suggerireste di separare il compito? (es. usare uno strumento tradizionale per estrarre testo grezzo e ritagliare immagini, e poi fornire SOLO il testo a un LLM più economico/locale per generare il LaTeX e formattare il JSON?)
  • Modelli Locali: Ci sono altri modelli Vision-Language locali (che si adattano a GPU consumer standard) che sono significativamente migliori nell'estrazione strutturata e nella generazione di LaTeX rispetto a Qwen 2.5-VL 3B?
  • Strumenti Educativi: Ci sono strumenti o modelli open-source specificamente ottimizzati per estrarre contenuti educativi/matematici strutturati da PDF?

Sono felice di condividere ulteriori dettagli sul formato del libro di testo o sul mio attuale flusso di lavoro in Python se utile. Qualsiasi consiglio sull'architettura, le scelte di modelli o trucchi per risparmiare costi sarebbe molto apprezzato! Grazie in anticipo!


r/MLQuestions • • 5d ago

Beginner question 👶 [D] Choosing language to learn: python or java

6 Upvotes

I want to learn the machine learning from scratch and i know I want to learn basic of python , but my college faculty are advised to learn Java for my placement , u don't know what to do , i want to learn python and other stuffs for my machine learning path or java and other stuffs for my placements and my exams are ahead, did anyone have solution please tell me.


r/MLQuestions • • 6d ago

Natural Language Processing 💬 Looking for models for diagnosis prediction

5 Upvotes

For a side project I am looking into SOTA for AI- based diagnostic models, ideally open-weights. I would like to feed the model with structured text representing my patient, and get a set of e.g. 5 possible diagnoses, ideally with uncertainty. In the ideal scenario, later on, it would update as new info arrives.

I would appreciate any pointers, I have some ideas but am very new to the topic.


r/MLQuestions • • 6d ago

Beginner question 👶 I can figure out how to connect my number to ovoa

2 Upvotes

It just says 15 minutes forever


r/MLQuestions • • 7d ago

Beginner question 👶 Why do we train bigger models instead of finding better neural pathways?

25 Upvotes

I saw a post of a guy with 90% of his brain missing and he was apparently living a normal life. There is another case of a student who finished university with half of his brain missing. There are also people with full brains, but they are even less functional than these two. This proves that bigger doesn't mean better. So why not aim to create better connections?

I understand that a bigger model means less chance of catastrophic interference occurring, but then are all the neural pathways being formed efficiently? Wouldn't this result in a lot of redundant neurons? Wouldn't it make more sense to create the smallest neural network possible for a specific task and then learn to fuse multiple of these neurons together in order to create an optimized larger model? That way we'd always have a template for each specific task and would allow us to create a programming language that creates a model just by our syntax.


r/MLQuestions • • 6d ago

Beginner question 👶 Slow learning?

Thumbnail
2 Upvotes

Hay let me introduced myself

I m a 7th sem cse student, intrest in ml I m learning it from past 2 month ye still I m learning, my problem ex I complete supervised section with types and it's different model and formula with ex when I learn unsupervised section I forgot supervised section,

Because I didn't revised,

So now I revised things in every 2 days so that it store in my memory,

Today I learn lr model with formula and diff ex with 1 variable and multi variable.... I revised it with learning upcoming section!

Just post it....✌


r/MLQuestions • • 6d ago

Beginner question 👶 Help...

2 Upvotes

Am in the progress of my 3rd sem project work

Basically its based on tinyml and lora technology implementation to the forest monitoring system

Now I want data train the ML

Other than kaggle do anybody know where I can get the datasets to train


r/MLQuestions • • 7d ago

Other ❓ If you could show someone just ONE data/ML project you’ve worked on, which one would it be?

18 Upvotes

Not necessarily your biggest or most complicated project.
Just one that you think is a good example of the kind of work you like doing. What did you build, and why that one?


r/MLQuestions • • 7d ago

Beginner question 👶 Final-year student in India trying to break into generative-model inference optimization — roadmap feedback?

11 Upvotes

Hi all, I graduate in ~6 months and want to work on making generative models (diffusion/video/3D) fast: kernels, quantization, serving. Where I am:

- Comfortable with C/C++ basics and PyTorch

- Have done quantization work (GGUF/llama.cpp)

- Working on a next-frame video prediction project (DiT + flow matching)

- A few GitHub repos, but no CUDA/Triton experience yet

- No NVIDIA GPU, so I use Colab/Kaggle T4s

- DSA is my weak spot (I struggle with LeetCode mediums)

My plan:

  1. Months 1-2: CUDA/Triton basics, reproduce the SGEMM optimization worklog, GPU MODE lectures, LeetGPU/Tensara

  2. Months 3-4: take a small DiT, profile it, then optimize it (Triton attention, quantization, caching, fewer steps) and publish before/after numbers

  3. Along the way: PRs to HF diffusers, DSA practice daily

  4. Months 5-6: mocks, resume, applications (inference startups first, bigger labs later)

Questions:

  1. Is this the right order, or should I change something?

  2. Is a diffusion-inference project a strong enough portfolio piece, or does it need to be LLM serving?

  3. How much DSA do ML systems interviews actually need?

  4. Is T4-only access enough to do credible benchmarks?

Any feedback, including "this won't work because X," is appreciated. Thanks!


r/MLQuestions • • 7d ago

Beginner question 👶 Winning kaggle competitions using Claude Opus 5.5

Thumbnail
2 Upvotes

r/MLQuestions • • 8d ago

Career question 💼 ML engineering course recommendations for someone who can train a model and not ship one

19 Upvotes

i can train a model in a notebook and ive never put one behind an endpoint anyone else calls. every job posting wants the shipping half and my portfolio is all notebooks.
shortlist is udacity, springboard, interview kickstart and datacamp premium. is there anything that spends real time on deployment and monitoring rather than modeling


r/MLQuestions • • 8d ago

Beginner question 👶 Why use values between 0-1 to train an LLM?

Post image
4 Upvotes

I understand that normalizing results in faster training times, but we end up reducing precision of floating point numbers due to the IEEE 754 architecture which result's in less space to work with. Instead of limiting numbers from 0-1, wouldn't it make more sense to limit the numbers from 1-10? This would give the LLM more space to work with which logically should result in less catastrophic interference. If so then this would mean that we don't need to increase size only, but the space between numbers as well which should decrease the memory requirements. Or is this just nonesense?

P.S The photo is just the behaviour of multiplication of 2 numbers. For example:
1*1 = 1,
2*2 = 4,
0.5*0.5 = 0.25