The self-hosted reasoning stack: Qwen3-32B, NeMo, and £9K/month
- The problem
- What does a 24/7, in-region, self-improving reasoning layer actually cost to run?
- Why it matters
- Three model tiers sized to their jobs, an edge-filtering technique that cuts per-incident cost from $0.52 to $0.07, and monthly fine-tuning with sub-5-minute rollback.
- Who should read it
- ML platform engineers
- What you'll get
- A fully costed stack and its operating loop.
- Series
- The Autonomous NOC, Part 6 of 12
- Readers
- 9 views on Builder Center
The full article lives on Builder Center. This page is a short guide to what it covers.