Skip to content

Decision layer (optional)

Some of Cairn’s steps generate nothing: they pick from a finite list. Why did this test fail? Does this case cover that checklist item? Which element on the page did a broken locator mean? Since 0.8.0, a System One model can answer those picks. Such a model never writes text. It answers typed questions (yes/no, a choice, a score) with probabilities.

  • It never writes cases, code or repairs, and never gives the Pilot verdict. Those stay on the LLM.
  • Its answers are untrusted input. Below DECIDER_MIN_CONFIDENCE (0.75 by default), an answer is ignored and Cairn does what it always did.
  • It never sinks a run. A timeout, a server error, an input too long for the model or an invalid answer each falls back silently to today’s behaviour, and the trace records it.
  • Confidence is not the chance of being right. It describes the shape of an answer’s distribution. Each provider computes it differently, so a threshold tuned on one does not carry over to another.
UseWhereWhat a confident answer doesOn by default
repair-triageThe validate ⇄ repair loop (explore, automate --validate)“The app is broken” or “the environment is broken” keeps the failing test out of the repair hint, because rewriting test code cannot fix either. The run lists it under Not repaired.Yes, once DECIDER names a provider
coverageThe checklist_coverage score (explore, design --checklist)A yes/no answer per checklist item per case replaces the LLM judge’s number, and the score’s comment says soNo: name it in DECIDER_USES
locator-healThe repair loop, after triageFor a getByRole(…) that matched no element, or several, the decider picks the element the locator meant from the start page. A pick the browser matches exactly once joins the repair hint as a verified replacement, and the repair still writes the code.No: name it together with repair-triage; it needs the lib browser backend

The run reports the layer’s calls and fallbacks, the tests kept out of repair (notRepaired) and the healed locators (healed):

  • explore writes them to report.json and report.md.
  • design adds the summary to report.json.
  • automate prints them, and its MCP result carries the same keys.
jevlayacompat
WhatTypeSafe’s cloud APIAn open model on your machine (laya-serve)Any other server speaking Jev’s POST /v1/systemone
Data goes toTypeSafe (api.typesafe.ai)This machineWherever DECIDER_BASE_URL points
AddressTYPESAFE_BASE_URL (default https://api.typesafe.ai)DECIDER_BASE_URL (required)DECIDER_BASE_URL (required)
KeyTYPESAFE_API_KEY (required)LAYA_API_KEY (if the server wants one)DECIDER_API_KEY (optional)

Each provider reads only its own address and key, so switching between them never sends one provider’s key to another. If your TYPESAFE_BASE_URL already serves the TypeSafe SDK elsewhere, for example a local laya-serve, set CAIRN_TYPESAFE_BASE_URL=https://api.typesafe.ai and CAIRN_TYPESAFE_API_KEY together. Give the origin only, because Cairn appends /v1/systemone itself.

Laya ships its own Jev-compatible server. Cairn never starts it or downloads its weights.

Terminal window
pip install laya
LAYA_API_KEY=<a token you choose> LAYA_HOST=127.0.0.1 LAYA_MODELS=english,multilingual laya-serve
DECIDER=laya DECIDER_BASE_URL=http://127.0.0.1:8000 LAYA_API_KEY=<same token> cairn doctor

Laya picks a checkpoint by the language of the text. For an application whose texts are not English, set DECIDER_MODEL=multilingual. Laya’s context is short, so Cairn refuses an input that would not fit rather than let it be cut: such a call falls back.

Terminal window
cairn explore --url <url> --checklist checklist.md --decider laya --decider-shadow

In shadow mode the decider is asked at every enabled use point, which means all three unless DECIDER_USES names fewer. Nothing it says is acted on. The run’s prompts, tests and reports are what they would have been without it. Its answers land in runs/<id>/decider-shadow.json, next to what the run actually did.

A use point is worth turning on when a pilot shows the following:

  • agreement of at least 90 %;
  • for triage, a hand-checked precision of at least 85 %;
  • fewer than 10 % fallbacks.

Tune the confidence threshold per use point and per provider.

  • Variables: every variable and its default is under Configuration → Decision layer.
  • Flags: --decider <off|jev|laya|compat> on explore, design and automate overrides DECIDER for one run, and --decider-shadow turns shadow mode on for one run.
  • Setup check: cairn doctor shows the provider, where the data goes and the latency of one real call. A misconfiguration fails at start, with the reason.

The full guide is docs/decider.md. The decision and its measurements are in ADR 0022.