Furiosa-LLM OpenAI-Compatible Server
The HTTP server started by `furiosa-llm serve `. It hosts a single model on RNGD NPUs and exposes an OpenAI-compatible surface - /v1/completions, /v1/chat/completions, /v1/responses (OpenResponses), /v1/embeddings, /v1/models, /v1/models/{model_id} - plus the vLLM-originated /score and /rerank pooling endpoints, a tokenizer API (/tokenize, /detokenize, /tokenizer_info), GET /version and a Prometheus GET /metrics endpoint. It is customer-hosted software, so the base URL below is templated on the operator's own host; the documented default is http://localhost:8000/v1. FuriosaAI publishes no OpenAPI for this surface - the parameter tables in the serving docs are the contract.