Decision layer (optional)
Some of Cairn’s steps generate nothing: they pick from a finite list. Why did this test fail? Does this case cover that checklist item? Which element on the page did a broken locator mean? Since 0.8.0, a System One model can answer those picks. Such a model never writes text. It answers typed questions (yes/no, a choice, a score) with probabilities.
What it does — and does not do
Section titled “What it does — and does not do”- It never writes cases, code or repairs, and never gives the Pilot verdict. Those stay on the LLM.
- Its answers are untrusted input. Below
DECIDER_MIN_CONFIDENCE(0.75 by default), an answer is ignored and Cairn does what it always did. - It never sinks a run. A timeout, a server error, an input too long for the model or an invalid answer each falls back silently to today’s behaviour, and the trace records it.
- Confidence is not the chance of being right. It describes the shape of an answer’s distribution. Each provider computes it differently, so a threshold tuned on one does not carry over to another.
Use points
Section titled “Use points”| Use | Where | What a confident answer does | On by default |
|---|---|---|---|
repair-triage | The validate ⇄ repair loop (explore, automate --validate) | “The app is broken” or “the environment is broken” keeps the failing test out of the repair hint, because rewriting test code cannot fix either. The run lists it under Not repaired. | Yes, once DECIDER names a provider |
coverage | The checklist_coverage score (explore, design --checklist) | A yes/no answer per checklist item per case replaces the LLM judge’s number, and the score’s comment says so | No: name it in DECIDER_USES |
locator-heal | The repair loop, after triage | For a getByRole(…) that matched no element, or several, the decider picks the element the locator meant from the start page. A pick the browser matches exactly once joins the repair hint as a verified replacement, and the repair still writes the code. | No: name it together with repair-triage; it needs the lib browser backend |
The run reports the layer’s calls and fallbacks, the tests kept out of repair (notRepaired) and the healed
locators (healed):
explorewrites them toreport.jsonandreport.md.designadds the summary toreport.json.automateprints them, and its MCP result carries the same keys.
Providers
Section titled “Providers”jev | laya | compat | |
|---|---|---|---|
| What | TypeSafe’s cloud API | An open model on your machine (laya-serve) | Any other server speaking Jev’s POST /v1/systemone |
| Data goes to | TypeSafe (api.typesafe.ai) | This machine | Wherever DECIDER_BASE_URL points |
| Address | TYPESAFE_BASE_URL (default https://api.typesafe.ai) | DECIDER_BASE_URL (required) | DECIDER_BASE_URL (required) |
| Key | TYPESAFE_API_KEY (required) | LAYA_API_KEY (if the server wants one) | DECIDER_API_KEY (optional) |
Each provider reads only its own address and key, so switching between them never sends one provider’s key to
another. If your TYPESAFE_BASE_URL already serves the TypeSafe SDK elsewhere, for example a local laya-serve,
set CAIRN_TYPESAFE_BASE_URL=https://api.typesafe.ai and CAIRN_TYPESAFE_API_KEY together. Give the origin
only, because Cairn appends /v1/systemone itself.
Running Laya locally
Section titled “Running Laya locally”Laya ships its own Jev-compatible server. Cairn never starts it or downloads its weights.
pip install layaLAYA_API_KEY=<a token you choose> LAYA_HOST=127.0.0.1 LAYA_MODELS=english,multilingual laya-serveDECIDER=laya DECIDER_BASE_URL=http://127.0.0.1:8000 LAYA_API_KEY=<same token> cairn doctorLaya picks a checkpoint by the language of the text. For an application whose texts are not English, set
DECIDER_MODEL=multilingual. Laya’s context is short, so Cairn refuses an input that would not fit rather than
let it be cut: such a call falls back.
Shadow mode: judge it before you trust it
Section titled “Shadow mode: judge it before you trust it”cairn explore --url <url> --checklist checklist.md --decider laya --decider-shadowIn shadow mode the decider is asked at every enabled use point, which means all three unless DECIDER_USES
names fewer. Nothing it says is acted on. The run’s
prompts, tests and reports are what they would have been without it. Its answers land in
runs/<id>/decider-shadow.json, next to what the run actually did.
A use point is worth turning on when a pilot shows the following:
- agreement of at least 90 %;
- for triage, a hand-checked precision of at least 85 %;
- fewer than 10 % fallbacks.
Tune the confidence threshold per use point and per provider.
Configuration and setup check
Section titled “Configuration and setup check”- Variables: every variable and its default is under Configuration → Decision layer.
- Flags:
--decider <off|jev|laya|compat>onexplore,designandautomateoverridesDECIDERfor one run, and--decider-shadowturns shadow mode on for one run. - Setup check:
cairn doctorshows the provider, where the data goes and the latency of one real call. A misconfiguration fails at start, with the reason.
The full guide is docs/decider.md. The decision
and its measurements are in ADR 0022.