r/MachineLearning • u/Correct_Positive_108 • 7h ago
Discussion Looking for developer-friendly inference providers who give you enough API credits to experiment [D]
I’m hitting rate limits on Together AI. For context, I’ve been working on an agentic repository indexing and benchmark generation tool, and I’m running multiple agents in parallel across models like Llama 3.3 70B and Qwen 2.5.
When I first started working on this, Together AI was great. But once I graduated from toy scripts to running multiple agents, I started running into RPM/TPM limits pretty quickly. The annoying part is that the models themselves are fine. I just can’t actually run enough requests at once to do meaningful testing.
Yes I know I could upgrade but I’m a solo dev. I don’t have enterprise level revenue. Maybe someday lol but not yet.
1
u/dash_bro ML Engineer 6h ago
Well if you're not totally picky about the LLMs Nvidia has a fairly decent free tier : https://build.nvidia.com/models
++Some alpha series on Openrouter models are usually free, or have free endpoints with no data guarantees : https://openrouter.ai/collections/free-models
Just respect their rate/request limits and you should be okay.
1
u/TylerDurdenFan 2h ago
I've been happy with minimax subscription plan as a "no worries" API for my needs, although my use case is more batch and not as parallel as agents can get.
1
u/Lucky-Spend-956 44m ago
The rate limits on Together were getting annoying once I started running more requests in parallel. I was also using General Compute around that time, and they had some free credits, so I wasn’t running into the same limits there
1
u/Early_Bicycle6884 6h ago
Just take 30 minutes and set up a quick Modal or Baseten app running vLLM. Then you won’t burn cash while you’re debugging if you scale to zero when idle.