# MLflow LLM Evaluate

**Canonical:** https://apis.io/apis/evals/mlflow-llm-evaluate/  
**Provider:** Evals — https://apis.io/providers/evals/  
**Base URL:** https://mlflow.org  
**Documentation:** https://mlflow.org/

MLflow LLM Evaluate is one of 20 APIs that [Evals](https://apis.io/providers/evals/) publishes on the [APIs.io](https://apis.io/) network. Tagged areas include Open-Source, MLflow, Experiment Tracking, LLM Judges, and Apache. The published artifact set on APIs.io includes API documentation and a GitHub repository.

MLflow LLM evaluate extends MLflow's experiment tracking with mlflow.evaluate() support for LLM tasks. The API runs reference-based and reference-free metrics (toxicity, perplexity, BLEU, ROUGE, exact match, custom LLM judges) over a logged model or a function and persists results into MLflow's experiment store alongside traditional ML metrics. Sits inside the broader MLflow open-source project.

## Machine-readable artifacts (2)

- **Documentation** — https://mlflow.org/docs/latest/llms/llm-evaluate/index.html
- **GitHubRepository** — https://github.com/mlflow/mlflow

## Other Evals APIs (12)

- [OpenAI Evals](https://apis.io/apis/evals/openai-evals/)
- [Inspect AI](https://apis.io/apis/evals/inspect-ai/)
- [Braintrust](https://apis.io/apis/evals/braintrust/)
- [LangSmith Evaluation](https://apis.io/apis/evals/langsmith-evaluation/)
- [Promptfoo](https://apis.io/apis/evals/promptfoo/)
- [Helicone](https://apis.io/apis/evals/helicone/)
- [Patronus AI](https://apis.io/apis/evals/patronus-ai/)
- [DeepEval (Confident AI)](https://apis.io/apis/evals/deepeval-confident-ai/)
- [Arize AI (Phoenix)](https://apis.io/apis/evals/arize-ai-phoenix/)
- [Galileo](https://apis.io/apis/evals/galileo/)
- [Humanloop](https://apis.io/apis/evals/humanloop/)
- [TruLens](https://apis.io/apis/evals/trulens/)

## Tags

Open-Source, MLflow, Experiment Tracking, LLM Judges, Apache

---

Profiled by [API Evangelist](https://apievangelist.com) and published on [APIs.io](https://apis.io/apis/evals/mlflow-llm-evaluate/). The API's provider profile, Kin Score and agent-readiness rating are at https://apis.io/providers/evals/.
