Hypura Ollama-Compatible Inference API
Hypura is a storage-tier-aware LLM inference scheduler for Apple Silicon that places model tensors across GPU, RAM, and NVMe tiers so models larger than physical memory can run. Running `hypura serve ` exposes a local Ollama-compatible HTTP API, making it a drop-in replacement for tooling that talks to Ollama. The server is local-first; there is no hosted endpoint.