Furiosa-LLM OpenAI-Compatible Server
The HTTP server started by `furiosa-llm serve `. It hosts a single model on RNGD NPUs and exposes an OpenAI-compatible surface - /v1/completions, /v1/chat/completions, /v1/responses (OpenResponses), /v1/embeddings, /v1/models, /v1/models/{model_id} - plus the vLLM-originated /score and /rerank pooling endpoints, a tokenizer API (/tokenize, /detokenize, /tokenizer_info), GET /version and a Prometheus GET /metrics endpoint. It is customer-hosted software, so the base URL below is templated on the operator's own host; the documented default is http://localhost:8000/v1. FuriosaAI publishes no OpenAPI for this surface - the parameter tables in the serving docs are the contract.
Documentation
Documentation
https://developer.furiosa.ai/latest/en/furiosa_llm/furiosa-llm-serve.html
APIReference
https://developer.furiosa.ai/latest/en/furiosa_llm/reference.html
Authentication
https://raw.githubusercontent.com/api-evangelist/furiosa/refs/heads/main/authentication/furiosa-authentication.yml
RateLimits
https://raw.githubusercontent.com/api-evangelist/furiosa/refs/heads/main/rate-limits/furiosa-rate-limits.yml