LANCET command guard

A second opinion
on every command.

Before the agent runs a shell command, a small classifier on your own machine reads it and decides whether to let it through, ask you first, or stop it. No API key, no network, and no charge per check: a few milliseconds of CPU is the whole cost.

It ships off. Run /lancet-guard setup in Pi to download the model once and switch it on. SpecPi never switches it on for you, and never changes your choice when it updates.

~10 msper typical command, on the CPU
$0per check: no API key, no account, no network
78.0%of risky commands caught on LANCET's release benchmark
12.4%of safe commands wrongly stopped, on the same benchmark

What it is

01

specpi-lancet-guard is one of SpecPi's eight default packages. It watches the commands the agent sends to bash, powershell and SpecPi's background tool, and checks each one before it runs. The judge is LANCET Nano, a code-trained language model fine-tuned for one job: telling risky shell commands from routine ones, in Bash, PowerShell and cmd.

It runs inside Pi, on your processor, through ONNX Runtime. Commands are read as text and never executed by the guard, and they never leave the machine. There is no service to sign up for, no key to manage, and nothing to go down. It replaced specpi-jev-guard, which sent each command to the hosted Jev classifier through OpenRouter.

How it decides

02

Fixed rules go first and the model second. The rules settle the obvious cases instantly, and the model cannot overrule them.

Shell command bash, powershell Local rules instant, fixed ON YOUR CPU LANCET Nano ~10 ms rest Blocked rm -rf /, curl | sh Runs ls, git status risky asks you review unsure: asks not flagged runs Permission system checks every call too; either one can block
The rules come from the Jev guard and are unchanged. The model only sees what they leave open, which is where a fixed list runs out.
  • Local rules first. A hard-deny list blocks the clearly catastrophic: deleting the filesystem root or a home directory, formatting a disk, writing to a raw device, piping a download into a shell, fork bombs. Provably read-only commands pass without a model call. Your own safeCommands, allowedCommands and disallowedCommands lists apply here too.
  • Three answers from the model. risky asks you before the command runs, or blocks it in block mode (/lancet-guard mode block). review is Nano's band for commands it is unsure about, and it always asks. not_flagged runs.
  • What it cannot read or clear, it asks about. Nano reads Bash, PowerShell and cmd, so a PowerShell command it does not flag now runs. A command over 8,192 bytes asks. A command longer than one 512-token window is read in full, as overlapping windows, but never cleared: if Nano does not flag it, the guard still asks, because padding can hide a risky line (see Limits).
  • File writes are judged by where they go. Ordinary project files pass. Writes to .env files, keys, .ssh, .git or anywhere outside the project ask first.
  • It fails closed. A missing or damaged model blocks what the rules leave open rather than waving it through. With no one to ask, as in a headless run, an ask becomes a block unless you set "uncertain": "allow".

What it costs to run

03

The hosted guard it replaced paid for every check twice: in money, per request, and in time, for a network round trip. A local model pays for neither.

ResourceLANCET guard
Cost per check$0
Network per checknone
API key or accountnone
Time per command~10 ms (95th percentile under 30 ms)
Long commands~0.3 s per 512-token window
Memory, once loaded~180 MB
Loading, once per session~0.6 s
Disk116 MB
One-time download~109 MB, from GitHub
GPUnot used

Times are from a Ryzen 9 3900X, one command at a time, on a desktop running other programs; they varied by about half between runs, and LANCET's own Python runtime took the same time, measured side by side. Nothing is loaded while the guard is off, and most commands never reach the model at all, because the rules settle them first. It adds nothing to the agent's prompt either: SpecPi's first request is 12,287 characters, the same as with the Jev guard it replaced.

How well it does

04

These results come from LANCET's v0.4.3 release evaluation: 7,911 commands from three benchmarks, none of them used to train, tune or calibrate it. Each guard was scored once, at its shipped setting, without SpecPi's rules in front. The Triage Score gives a point for every risky command asked about or blocked, and shrinks in proportion when more than 10% of safe commands are stopped. Caught and stopped are averaged over the three benchmarks, weighted by size as the score is.

ModelTriage ScoreRisky caughtSafe wrongly stoppedRuns on
LANCET Nano v0.4.3 (this guard)68.377.0%9.0%your CPU
LANCET Nano v0.4.254.477.1%13.7%your CPU
LANCET Nano v0.4.047.155.7%10.5%your CPU
LANCET Nano v0.4.144.967.9%14.6%your CPU
LANCET Nano v0.3.044.755.3%11.6%your CPU
Jev (hosted, the old guard)38.990.0%23.4%hosted API
Laya (421M parameters)13.387.3%69.6%GPU
  • v0.4.3 interrupts less and decides more. On lancet-bench-2-next, LANCET's release benchmark, it catches the same 78.0% of risky commands as v0.4.2 while stopping 12.4% of safe ones instead of 20.0%: about one in eight rather than one in five. Half of the benchmark's risky commands come back risky, against 17.7% for v0.4.2, so fewer land in the unsure middle.
  • Jev still catches more, and stops more. It catches 95.1% of lancet-bench-2-next's risky commands but stops 21.8% of its safe ones. It was also given the task the agent was working on; Nano sees only the command.
  • PowerShell and cmd are new. v0.4.2 could only ask about them. The benchmark has 102 PowerShell and 36 cmd commands, so the evidence there is thinner than for Bash.

Each benchmark on its own, at each model's shipped thresholds:

BenchmarkModelRisky caughtSafe wrongly stopped
lancet-bench-2-next (3,204 commands)v0.4.378.0%12.4%
lancet-bench-2-next (3,204 commands)v0.4.278.0%20.0%
lancet-bench-2-next (3,204 commands)Jev95.1%21.8%
ShellRisk-Bench test (4,194 commands)v0.4.364.8%3.0%
ShellRisk-Bench test (4,194 commands)v0.4.260.1%2.7%
ShellRisk-Bench test (4,194 commands)Jev67.9%20.9%
Neutral set (513 commands)v0.4.385.8%5.6%
Neutral set (513 commands)v0.4.291.0%7.0%
Neutral set (513 commands)Jev97.2%30.2%

A harder benchmark. Earlier versions of this page used lancet-bench-1, 793 Bash commands on which the best guards were near the ceiling: v0.4.2 caught 92.4% there. lancet-bench-2-next has 3,204 commands built as 1,602 risky/safe twins from 289 scenarios in 40 areas, each repeated inside subshells, functions, pipelines and other wrappers, in Bash, PowerShell and cmd. Every guard scores lower on it. lancet-bench-1 was retired and is now part of LANCET's training data. lancet-bench-2-next is LANCET's own, kept private so it stays unseen; ShellRisk-Bench is labelled upstream, and the neutral set is 66 commands from an outside coding-agent benchmark and 447 from the Shell Safety v2 test split.

Per-area results and charts are on the LANCET-model page, and the full method is in its model card.

The model

05

LANCET Nano v0.4.3 is Salesforce's CodeT5+ 220M encoder, fine-tuned to score Bash, PowerShell and cmd commands and quantized to 8-bit integers so it runs well on an ordinary CPU, with a small head that pools its output: 111 million parameters in all. A command longer than the encoder's 512-token window is read as overlapping windows, so nothing is cut off. v0.4.3 was trained afresh from that base on 55,827 labelled examples, including a new curriculum of 21,544 risky/safe twins across the three shells. Model: Apache-2.0, on a BSD-3-Clause base. Runtime: MIT.

  • Labels from documentation, not from another model. Its training labels come from the wording of human-written examples in tldr-pages, from AWS's own API definitions (including which fields are marked sensitive), from the Azure, GitHub, kubectl and Docker command references, from ShellRisk-Bench's and Shell Safety v2's training splits as labelled upstream, and from LANCET's own authored risky/safe twins and contrast sets of risky commands with safe look-alikes. No language model or hosted API produced a label.
  • The same answers as LANCET's own runtime. The guard runs a Node port of LANCET's Python runtime, on the same ONNX Runtime release. v0.4.3 brought a new runtime for windowed reading; the port follows it and matches it on the guard's 76-command parity set: the same tokens, the same windows and the same answers, with scores within 1e-15.
  • Verified before it runs. Setup downloads the release ZIP from GitHub and refuses it unless its SHA-256 matches the digest built into the package. The six model files are then extracted and checked one by one, and checked again every time the model loads. A tampered or damaged model is refused, never scored.

Turning it on

06

The guard is installed with SpecPi and stays off until you say otherwise. In Pi:

  • /lancet-guard setup: download and verify the model, try it on two sample commands, then choose on for this session, on for good, or off.
  • /lancet-guard on or off: switch it for this session. Add --global to save that as your default.
  • /lancet-guard mode ask or block: choose whether a risky verdict asks you, the default, or is blocked outright, which also stops the agent. It is saved for future sessions; review asks either way.
  • /lancet-guard check <command>: score a command without running it.
  • /lancet-guard: show whether it is on, and the model's state.

Settings live in ~/.pi/lancet-guard.json, and a trusted project can add its own .pi/lancet-guard.json. The package README lists every setting.

Limits

07
  • It is a second opinion, not a sandbox. A command it lets through runs with your full permissions. not_flagged means it did not flag the command, not that the command is safe.
  • Credentials, Windows and macOS administration, and deceptive previews are its weakest areas on lancet-bench-2-next. The permission system and your own review still apply.
  • PowerShell and cmd are newer than Bash. They have less training data and fewer benchmark commands behind them. Other shells ask.
  • Padding dilutes. Harmless lines in front of a risky command lower its score. Within one window most such commands still ask, but not all: one in nine came back not flagged in four of twelve padded runs. Past one window, anywhere from one to all nine did, so the guard asks about every command that long that Nano does not flag.
  • Long commands take longer. About 0.3 s per 512-token window, 5 to 8 s at the 8,192-byte limit. Anything longer asks.
  • Scores vary slightly by CPU. ONNX Runtime's 8-bit arithmetic differs between processor types, by up to about 0.03, so a command right at a threshold can land in a different band on another machine.
  • No Intel Macs. ONNX Runtime ships no build for them.