Particle.news

Ramp Launches Router.com to Cut AI Inference Costs

The service routes each call to the lowest-cost model that meets a performance threshold, linking those routing decisions to Ramp's spend-visibility controls.

Overview

  • Ramp launched Router.com Wednesday, making a single API endpoint available in the U.S. with free routing through 2026 and a $26 credit for new users.
  • Router sends each request to the cheapest model that meets a developer-set performance bar and applies over 100 optimizations plus automatic fallbacks to protect latency and uptime.
  • The product supports OpenAI, Anthropic, SpaceXAI and a range of open models with more providers planned, while users continue to pay each model provider’s listed token prices.
  • Ramp says Router cut its own production inference costs about 30% and reports customers using the service have averaged roughly 40% lower inference spend.
  • Router ties continuous, production-grounded benchmarking (Ramp SWE-Bench), BYOK and U.S.-hosted zero-data-retention options to finance-facing dashboards so teams can track token spend, attribution, and governance.