Switch providers
Plune is provider-agnostic. The top-level provider block sets the default backend for the whole
suite; the API key is read from the environment based on provider.type and is never written to
disk.
The three providers
Section titled “The three providers”provider: type: anthropic # anthropic | openai | openrouter model: claude-sonnet-5-5provider.type | Environment variable | Example model |
|---|---|---|
anthropic | ANTHROPIC_API_KEY | claude-sonnet-5-5 |
openai | OPENAI_API_KEY | gpt-4o-mini |
openrouter | OPENROUTER_API_KEY | openai/gpt-4o-mini |
Switching backends is a one-line change — set type and model, and make sure the matching key is
in your environment (or .env — where the CLI reads it):
echo 'OPENAI_API_KEY=sk-...' >> .envTune the provider
Section titled “Tune the provider”The provider block accepts optional knobs (all validated):
provider: type: anthropic model: claude-sonnet-5-5 max_tokens: 1024 concurrency: 4 # parallel model calls timeout: 60000 # ms max_retries: 2temperature (0–2) is a knob too, and Plune sends it only when the config sets one. Leave it out for
Claude models released after Opus 4.6: they take none, and the API answers 400 to any value but the
default.
Override per eval
Section titled “Override per eval”Any eval can override part of the provider — handy for routing one hard eval to a stronger model while the rest of the suite stays cheap:
evals: - id: tricky-reasoning provider: model: claude-opus-5-5 # just the model; type is inherited prompt: "{{question}}" dataset: datasets/hard.jsonl assertions: - type: llm-judge criteria: "The reasoning is correct and shows its steps."Override the judge
Section titled “Override the judge”Model-graded assertions (llm-judge) can use their own provider — e.g. a cheaper judge than the
model under test:
assertions: - type: llm-judge criteria: "The answer is concise and on-topic." provider: type: openai model: gpt-4o-miniRun offline with the mock provider
Section titled “Run offline with the mock provider”Set PLUNE_MOCK_PROVIDER=1 to bypass real API calls entirely — deterministic, free, and key-less.
Ideal for validating that a plune.yaml is wired up, and for zero-cost CI smoke checks:
PLUNE_MOCK_PROVIDER=1 plune runPrice a run before you make it
Section titled “Price a run before you make it”plune run --dry-run estimates what the run would cost and makes none of it: no model is called and
nothing touches the network, so it needs no provider key. That holds from 0.16.0; before it, the
command stopped for want of the key, like a real run (Missing ANTHROPIC_API_KEY). Each row is priced
for the model its eval uses — the eval’s own provider.model, if it has one — and no assertion runs.
The price of a call comes from the pricing block of plune.yaml when it names the model, otherwise
from the cost the provider reports (OpenRouter does), otherwise from a built-in table of current
models. A model none of them knows costs 0, with a warning that names it, until you give its rates
per thousand tokens:
pricing: my-model: input_per_1k_usd: 0.002 output_per_1k_usd: 0.01