Home | NVIDIA NeMo Evaluator

LLM evaluation framework: benchmark environments, pluggable solvers, multi-format reporting.

Installation

Quickstart

———
Get Started

Features

Everything is an Environment. Built-in benchmarks, NeMo Skills, Gym remotes, lm-eval tasks, and VLMEvalKit datasets all resolve through one registry.

@benchmark + @scorer. Define a complete benchmark in under 10 lines of Python.

Pluggable solvers.simple, harbor, tool_calling, gym_delegation, openclaw — swap inference strategy per benchmark via config.

Cluster backends. Run locally, in Docker, or on SLURM clusters with automatic model deployment.

Resilient suites. Per-benchmark checkpointing with failure isolation. Resume partially completed suites with --resume.

Statistical regression. Compare runs with McNemar’s exact test, paired flip analysis, and confidence intervals. Gate releases across benchmark suites with per-benchmark policy thresholds.

17 built-in benchmarks. MMLU, MMLU-Pro, MATH-500, GPQA, GSM8K, DROP, MGSM, TriviaQA, HumanEval, SimpleQA, HealthBench, PinchBench, XSTest, Terminal-Bench variants, and NMP Harbor.

Tutorials

Architecture & Deployment