Skip to content

Plune CLI — Getting started

The Plune CLI (@plune-ai/cli) is AI-powered assertion testing for LLM apps — a test runner for model behaviour. You describe the checks in one plune.yaml; Plune calls the model, evaluates each assertion, caches results, and reports a pass/fail summary with token cost — locally, in CI, or as a regression diff between two runs. It is the Evaluate step in Plune’s Generate → Evaluate → Gate flow.

v0.17.0MITNode ≥ 20

Terminal window
npm install -g @plune-ai/cli # or: pnpm add -g @plune-ai/cli
plune --version

Or run it without installing:

Terminal window
npx -y @plune-ai/cli run
  1. Scaffold a plune.yaml, an example dataset, and .env.example:

    Terminal window
    plune init
  2. Add your provider key — read from the environment / .env, never written to disk:

    Terminal window
    echo 'ANTHROPIC_API_KEY=sk-ant-...' >> .env
  3. Run the assertions:

    Terminal window
    plune run
    # → 1/1 passed · 0 failed · 0 errored · $0.0008
  4. Re-render the last run, or diff two runs to catch regressions:

    Terminal window
    plune report --format markdown
    plune diff baseline.json current.json --fail-on-regression

Each run writes its full result to .plune/last-run.json.

Everything above generates checks, which is why it wants a provider key. If your tests already exist, there is nothing to generate — the results exist and only have to arrive, and that road needs no provider key at all:

Terminal window
plune run import ./junit.xml # Vitest, Jest, pytest, PHPUnit, Surefire, Cypress, Playwright…

Plune reads a single plune.yaml, discovered by walking up from the working directory (or passed with -c <path>). This is what plune init scaffolds:

version: 1
provider:
type: anthropic # anthropic | openai | openrouter
model: claude-sonnet-5-5
evals:
- id: example
prompt: "Answer concisely. {{question}}" # {{vars}} come from each dataset row
dataset: datasets/example.jsonl # a file path, or an inline `examples:` list
assertions:
- type: contains
value: "Paris"

Datasets are JSONL — one row per line, shaped { "vars": { ... }, "expected"?: "..." }.

model is any id the provider serves. plune init starts you on claude-sonnet-5-5 — its wizard offers OpenAI users gpt-4o-mini. Switch providers has the rest: a model per eval, temperature, and what a run costs.

The provider API key is read from the environment based on provider.type — never written to disk:

Providerprovider.typeEnvironment variable
AnthropicanthropicANTHROPIC_API_KEY
OpenAIopenaiOPENAI_API_KEY
OpenRouteropenrouterOPENROUTER_API_KEY

Every command reads its settings the same way: from the environment first — what your shell or CI job exported — then from a .env beside the config you named with -c, then from a .env in the current directory. The first place that sets a variable wins, so a .env never overrides what is already exported. That covers the keys above and every PLUNE_* variable (PLUNE_TOKEN, PLUNE_API_URL, PLUNE_RUN, …) for run import, sync, login and the rest, not only run and report — no dotenv wrapper in your scripts (from 0.16.0). The token plune login saved is the last resort, used only when PLUNE_TOKEN is set nowhere else.

Ten built-in assertion types cover plain text, JSON-schema, LLM-as-judge, and RAG metrics:

TypePasses when…
exact-matchoutput equals value (optional trim, ignore_case)
containsoutput contains value
contains-anyoutput contains at least one of values
contains-alloutput contains every one of values
json-schemaoutput validates against the JSON schema
llm-judgean LLM grades the output against criteria (≥ pass_threshold)
semantic-similarityembedding similarity to reference ≥ threshold
faithfulnessoutput is grounded in context (RAG)
answer-relevanceoutput actually answers the question (RAG)
context-precisioncontext is relevant to the question (RAG)
CommandSummary
plune runRun the suite. Flags: --dry-run (price the run from the model’s rates — it calls no model, so it needs no provider key; from 0.16.0), --only <id|tag> (repeatable), --bail, --no-cache, --concurrency <n>, --format console|json|markdown, -o, --output <file>.
plune reportRe-render the most recent run. Flags: --format, -o.
plune diff <baseline> <current>Compare two plune run --format json outputs and report pass→fail regressions. Flags: --fail-on-regression, --format, -o.
plune initScaffold plune.yaml, a sample dataset, and .env.example. Flags: --yes (non-interactive), --force.
plune run import <file>Turn a JUnit XML or Playwright JSON report into a run — no provider key, nothing generated. Flags: --format, --key <externalKey> to land several jobs in one run, --create to offer unmatched tests for review.

| plune login / plune logout | Save or remove a platform API token. The token is checked against the API before it is saved. See Platform sync. | | plune sync | Upload the latest local run to the platform. Flags: --file <path>. | | plune run start / finish / exec / report | Open, close, wrap and replay a platform run — the group that lets several jobs report into one run. See Import an existing suite. | | plune run delete <id> | Delete a run and everything it produced: its results and the review-queue entries it raised. Approved test cases and the audit log stay. Recoverable for six months (ask Plune to put it back), then gone for good. No prompt: it is your data, and this command belongs in scripts. | | plune ingest [dir] | Record a Cairn run in Plune. Omit [dir] for the newest run under ./runs. Generated cases arrive as review proposals — nothing is created until a person approves it. | | plune pull [file] | Write the project’s test cases as one Markdown document — Testomat’s classical format, plus Plune’s own columns — to plune/cases.md (or [file]). Refuses to overwrite a file git sees as modified; --force overrides, --suite <id> takes one suite or folder and what is under it. See Cases as Markdown. | | plune push [file] [--dry-run] | Send the document back. Cases are matched by id; a block without one becomes a draft; an unknown id is refused by line; nothing is deleted. Prints the report — created, updated, unchanged, refused, warnings — and with --dry-run writes nothing. | | plune plan grep <id> | Print a test plan as one --grep pattern — npx playwright test --grep "$(plune plan grep <id>)" or npx vitest run -t "$(…)" runs exactly what the plan collects. Only the pattern reaches stdout; an empty plan prints (?!), which matches nothing. |

Global flags: -c, --config <path> · -v, --verbose · --no-color.

The same engine that powers plune run is exported for use from your own code. Unlike the CLI, the library does not parse argv or auto-load .env — set the provider key in process.env yourself.

import { run } from "@plune-ai/cli";
import type { RunResult } from "@plune-ai/cli";
const result: RunResult = await run({ dryRun: false, configPath: "plune.yaml" });
console.log(result.summary); // { total, passed, failed, errored, ... }

Run Plune on every pull request and post a regression diff as a sticky comment with the companion GitHub Action, eval-action:

- uses: plune-ai/eval-action@v1
with:
config: plune.yaml
fail-on-regression: true

Runs stay on your machine unless you ask otherwise. plune login saves an API token and plune sync uploads a run to the Plune platform, where runs accumulate into a pass-rate and cost trend. Everything above works without an account, and continues to.

The token comes from the dashboard: the avatar at the top right → Account → API tokens, the project picked beside Generate a token. It is shown once — paste it into plune login.

Account, API tokens: the Project select, Generate a token, and the tokens listed by the project each writes to, with Revoke. Account, API tokens: the Project select, Generate a token, and the tokens listed by the project each writes to, with Revoke.

MIT © Plune Contributors.