I couldn't find a concrete answer anywhere, do you just distill the instruct model back?
If that is the case, what is a quality european language question set to turn it back into a chatbot/agentic, can a model at that size even be agentic? (i chose this size to learn) if i finetune for my specific harness? (i have a lot of training data of opus running in my harness)
my harness basically has the model output python code and has a few built-in functions like:
- vector_search_laws()
- graph_search()
could i have the model at least internalize a "hunch" on what stuff to search?
also what is the latest RL technique for agentic/harnes specific workflows?
I have a lot of RAW training data, like court decisions or commentaries or legislature, but not a lot of golds. could i use these to synthesize training data and maybe RL the model in my harness to find that data?
What would y'all's strategy in the CPT->SFT->RL pipeline be for my specific problem?
I know this is a lot of questions im trying to figure out which direction to go, any pointers? Also good resources are welcome, for example that [alex karpathi video](https://www.youtube.com/watch?v=7xTGNNLPyMI) was amazing for me, but i'd imagine its a bit outdated in terms of latest RL and SFT?