LANCET command guard
A second opinion
on every command.
Before the agent runs a shell command, a small classifier on your own machine reads it and decides whether to let it through, ask you first, or stop it. No API key, no network, and no charge per check: a few milliseconds of CPU is the whole cost.
It ships off. Run /lancet-guard setup in Pi to download the model
once and switch it on. SpecPi never switches it on for you, and never changes your choice
when it updates.
What it is
01
specpi-lancet-guard is one of SpecPi's eight default packages. It watches the
commands the agent sends to bash, powershell and SpecPi's
background tool, and checks each one before it runs. The judge is
LANCET Nano, a code-trained
language model fine-tuned for one job: telling risky shell commands from routine ones, in Bash,
PowerShell and cmd.
It runs inside Pi, on your processor, through ONNX Runtime. Commands are read as text and
never executed by the guard, and they never leave the machine. There is no service to sign up
for, no key to manage, and nothing to go down. It replaced specpi-jev-guard, which
sent each command to the hosted Jev classifier through OpenRouter.
How it decides
02Fixed rules go first and the model second. The rules settle the obvious cases instantly, and the model cannot overrule them.
-
Local rules first. A hard-deny list blocks the clearly catastrophic:
deleting the filesystem root or a home directory, formatting a disk, writing to a raw
device, piping a download into a shell, fork bombs. Provably read-only commands pass
without a model call. Your own
safeCommands,allowedCommandsanddisallowedCommandslists apply here too. -
Three answers from the model.
riskyasks you before the command runs, or blocks it in block mode (/lancet-guard mode block).reviewis Nano's band for commands it is unsure about, and it always asks.not_flaggedruns. - What it cannot read or clear, it asks about. Nano reads Bash, PowerShell and cmd, so a PowerShell command it does not flag now runs. A command over 8,192 bytes asks. A command longer than one 512-token window is read in full, as overlapping windows, but never cleared: if Nano does not flag it, the guard still asks, because padding can hide a risky line (see Limits).
-
File writes are judged by where they go. Ordinary project files pass.
Writes to
.envfiles, keys,.ssh,.gitor anywhere outside the project ask first. -
It fails closed. A missing or damaged model blocks what the rules leave
open rather than waving it through. With no one to ask, as in a headless run, an ask
becomes a block unless you set
"uncertain": "allow".
What it costs to run
03The hosted guard it replaced paid for every check twice: in money, per request, and in time, for a network round trip. A local model pays for neither.
| Resource | LANCET guard |
|---|---|
| Cost per check | $0 |
| Network per check | none |
| API key or account | none |
| Time per command | ~10 ms (95th percentile under 30 ms) |
| Long commands | ~0.3 s per 512-token window |
| Memory, once loaded | ~180 MB |
| Loading, once per session | ~0.6 s |
| Disk | 116 MB |
| One-time download | ~109 MB, from GitHub |
| GPU | not used |
Times are from a Ryzen 9 3900X, one command at a time, on a desktop running other programs; they varied by about half between runs, and LANCET's own Python runtime took the same time, measured side by side. Nothing is loaded while the guard is off, and most commands never reach the model at all, because the rules settle them first. It adds nothing to the agent's prompt either: SpecPi's first request is 12,287 characters, the same as with the Jev guard it replaced.
How well it does
04These results come from LANCET's v0.4.3 release evaluation: 7,911 commands from three benchmarks, none of them used to train, tune or calibrate it. Each guard was scored once, at its shipped setting, without SpecPi's rules in front. The Triage Score gives a point for every risky command asked about or blocked, and shrinks in proportion when more than 10% of safe commands are stopped. Caught and stopped are averaged over the three benchmarks, weighted by size as the score is.
| Model | Triage Score | Risky caught | Safe wrongly stopped | Runs on |
|---|---|---|---|---|
| LANCET Nano v0.4.3 (this guard) | 68.3 | 77.0% | 9.0% | your CPU |
| LANCET Nano v0.4.2 | 54.4 | 77.1% | 13.7% | your CPU |
| LANCET Nano v0.4.0 | 47.1 | 55.7% | 10.5% | your CPU |
| LANCET Nano v0.4.1 | 44.9 | 67.9% | 14.6% | your CPU |
| LANCET Nano v0.3.0 | 44.7 | 55.3% | 11.6% | your CPU |
| Jev (hosted, the old guard) | 38.9 | 90.0% | 23.4% | hosted API |
| Laya (421M parameters) | 13.3 | 87.3% | 69.6% | GPU |
-
v0.4.3 interrupts less and decides more. On lancet-bench-2-next, LANCET's
release benchmark, it catches the same 78.0% of risky commands as v0.4.2 while stopping
12.4% of safe ones instead of 20.0%: about one in eight rather than one in five. Half of
the benchmark's risky commands come back
risky, against 17.7% for v0.4.2, so fewer land in the unsure middle. - Jev still catches more, and stops more. It catches 95.1% of lancet-bench-2-next's risky commands but stops 21.8% of its safe ones. It was also given the task the agent was working on; Nano sees only the command.
- PowerShell and cmd are new. v0.4.2 could only ask about them. The benchmark has 102 PowerShell and 36 cmd commands, so the evidence there is thinner than for Bash.
Each benchmark on its own, at each model's shipped thresholds:
| Benchmark | Model | Risky caught | Safe wrongly stopped |
|---|---|---|---|
| lancet-bench-2-next (3,204 commands) | v0.4.3 | 78.0% | 12.4% |
| lancet-bench-2-next (3,204 commands) | v0.4.2 | 78.0% | 20.0% |
| lancet-bench-2-next (3,204 commands) | Jev | 95.1% | 21.8% |
| ShellRisk-Bench test (4,194 commands) | v0.4.3 | 64.8% | 3.0% |
| ShellRisk-Bench test (4,194 commands) | v0.4.2 | 60.1% | 2.7% |
| ShellRisk-Bench test (4,194 commands) | Jev | 67.9% | 20.9% |
| Neutral set (513 commands) | v0.4.3 | 85.8% | 5.6% |
| Neutral set (513 commands) | v0.4.2 | 91.0% | 7.0% |
| Neutral set (513 commands) | Jev | 97.2% | 30.2% |
A harder benchmark. Earlier versions of this page used lancet-bench-1, 793 Bash commands on which the best guards were near the ceiling: v0.4.2 caught 92.4% there. lancet-bench-2-next has 3,204 commands built as 1,602 risky/safe twins from 289 scenarios in 40 areas, each repeated inside subshells, functions, pipelines and other wrappers, in Bash, PowerShell and cmd. Every guard scores lower on it. lancet-bench-1 was retired and is now part of LANCET's training data. lancet-bench-2-next is LANCET's own, kept private so it stays unseen; ShellRisk-Bench is labelled upstream, and the neutral set is 66 commands from an outside coding-agent benchmark and 447 from the Shell Safety v2 test split.
Per-area results and charts are on the LANCET-model page, and the full method is in its model card.
The model
05LANCET Nano v0.4.3 is Salesforce's CodeT5+ 220M encoder, fine-tuned to score Bash, PowerShell and cmd commands and quantized to 8-bit integers so it runs well on an ordinary CPU, with a small head that pools its output: 111 million parameters in all. A command longer than the encoder's 512-token window is read as overlapping windows, so nothing is cut off. v0.4.3 was trained afresh from that base on 55,827 labelled examples, including a new curriculum of 21,544 risky/safe twins across the three shells. Model: Apache-2.0, on a BSD-3-Clause base. Runtime: MIT.
- Labels from documentation, not from another model. Its training labels come from the wording of human-written examples in tldr-pages, from AWS's own API definitions (including which fields are marked sensitive), from the Azure, GitHub, kubectl and Docker command references, from ShellRisk-Bench's and Shell Safety v2's training splits as labelled upstream, and from LANCET's own authored risky/safe twins and contrast sets of risky commands with safe look-alikes. No language model or hosted API produced a label.
- The same answers as LANCET's own runtime. The guard runs a Node port of LANCET's Python runtime, on the same ONNX Runtime release. v0.4.3 brought a new runtime for windowed reading; the port follows it and matches it on the guard's 76-command parity set: the same tokens, the same windows and the same answers, with scores within 1e-15.
- Verified before it runs. Setup downloads the release ZIP from GitHub and refuses it unless its SHA-256 matches the digest built into the package. The six model files are then extracted and checked one by one, and checked again every time the model loads. A tampered or damaged model is refused, never scored.
Turning it on
06The guard is installed with SpecPi and stays off until you say otherwise. In Pi:
-
/lancet-guard setup: download and verify the model, try it on two sample commands, then choose on for this session, on for good, or off. -
/lancet-guard onoroff: switch it for this session. Add--globalto save that as your default. -
/lancet-guard mode askorblock: choose whether ariskyverdict asks you, the default, or is blocked outright, which also stops the agent. It is saved for future sessions;reviewasks either way. /lancet-guard check <command>: score a command without running it./lancet-guard: show whether it is on, and the model's state.
Settings live in ~/.pi/lancet-guard.json, and a trusted project can add its own
.pi/lancet-guard.json. The package README
lists every setting.
Limits
07-
It is a second opinion, not a sandbox. A command it lets through runs
with your full permissions.
not_flaggedmeans it did not flag the command, not that the command is safe. - Credentials, Windows and macOS administration, and deceptive previews are its weakest areas on lancet-bench-2-next. The permission system and your own review still apply.
- PowerShell and cmd are newer than Bash. They have less training data and fewer benchmark commands behind them. Other shells ask.
- Padding dilutes. Harmless lines in front of a risky command lower its score. Within one window most such commands still ask, but not all: one in nine came back not flagged in four of twelve padded runs. Past one window, anywhere from one to all nine did, so the guard asks about every command that long that Nano does not flag.
- Long commands take longer. About 0.3 s per 512-token window, 5 to 8 s at the 8,192-byte limit. Anything longer asks.
- Scores vary slightly by CPU. ONNX Runtime's 8-bit arithmetic differs between processor types, by up to about 0.03, so a command right at a threshold can land in a different band on another machine.
- No Intel Macs. ONNX Runtime ships no build for them.