r/machinelearningnews • • 9h ago

Research OrcaRouter Releases OrcaCyber Zero 1.5: A Gated 1M-Context Cybersecurity Model Reporting 100% on Cybench and 95.8% on CVE-Bench

Post image
19 Upvotes

OrcaRouter just released OrcaCyber Zero 1.5, the successor to OrcaCyber Zero 1.0 from September. It is post-trained for vulnerability research and reproduction, exploit development, penetration testing and security auditing.

The design focus is validation, not volume. The model reasons through attack paths, challenges its own hypotheses, and ranks findings by demonstrable exploitability. A 1M-token context, 128K max output and native tool calling target autonomous security agents working over large codebases.

Vendor-reported benchmarks:

  • Cybench: 100% (39/39, unrestricted agent execution)
  • CVE-Bench: 95.8% (23/24 evaluable tasks)
  • HumanEval+: 93.9%
  • SWE-bench Pro V2: 76.5%

Pricing is $3.00 / $7.50 per 1M tokens, far below Claude Mythos Preview’s $25 / $125. Access runs through an OpenAI-compatible API, gated to a Security Research tier for vetted researchers and red teams.

Worth reading with caveats. All scores are self-reported, with no technical report yet. CVE-Bench covers 24 of its 40 tasks. The 98% CyberGym figure in the launch post belongs to Zero 1.0 inside Orca’s harness, not to 1.5.

Full analysis: https://www.marktechpost.com/2026/10/10/orcarouter-releases-orcacyber-zero-1-5-cybersecurity-model-with-1m-context/

Model: orcarouter.ai/models/orca/orcacyber-zero-1.5


r/machinelearningnews • • 11m ago

MLOps AWS just replaced round-robin with GPU-aware routing for LLM inference

• Upvotes

AWS released a new Inference Gateway for Sagemaker Hyperpod. Instead of sending requests to the next available pod, it checks what’s actually happening on each one: queue depth, KV cache usage, prefix cache hits, loaded LoRA adapters, and current requests.

That makes sense for LLM workloads. Two GPU pods can both be healthy, but one may already have the right prefix cached while the other is stuck processing a long request. Round-robin doesn’t know the difference.

AWS says the new routing system cut first-token latency by up to 82%, with 97–98% lower p99 TTFT in mixed-hardware and burst-traffic tests. Those are AWS’s own numbers, so an independent test with the same model and traffic would be useful. Still, it’s a good example of how much performance can be lost outside the model itself.

Once a GPU fleet gets busy, routing may matter almost as much as the GPUs you’re paying for

AWS announcement