cujo
Skip to content
Contents

Findings and severity

What critical means, which rules the agent cannot argue with, and why a false reads as not observed.

Three severities

Lowercase, and exactly these three. They are matched literally in code on both sides, so they are not editorial.

SeverityWhat it means
criticalA probe shows the change does not do what the diff claims; an endpoint that worked on base now errors; a hard rule tripped.
warnWorth a glance: changed code no test covers, a write outside the workspace, an unfamiliar but plausible host, a check that returned nothing.
infoWhat ran and what it showed, when nothing is wrong. Most of a calm review is this.

Two layers decide it

Code senses, the agent judges, and a hard rule overrides the agent on the cases that must never be reasoned away.

Layer one is deterministic and written in code. It runs twice — the agent is told to apply it, and the service derives it again independently from the same reports — so a model that forgets a rule does not cost you the finding. Layer two is the agent’s judgment over everything the rules do not cover, which is where most of the useful work happens.

The hard rules

Five force a critical the agent cannot lower or drop. They divide by the claim they make, and that split is what decides whether a review waits for a person.

RuleClaimFires when
tests failedcorrectnessA test passes on base and fails on head.
decoy readmaliceSomething opened the seeded credentials file, on any check.
decoy in egressmaliceThe seeded secret left the sandbox, on any check.
wrote sensitivemaliceA write landed in an SSH directory, a shell rc, cron, or a credentials path — on any check.
unknown egressmaliceAn install contacted a host that is neither a package index nor allowlisted. This one is scoped to detonation.

“Your tests fail” is mechanical, checkable by the author in thirty seconds. “This code tried to steal a credential” is a claim about what the code did, and the review states it as the measurement it is — the host, the path, the time. Both block the same way, and a maintainer who knows the host or the fixture lifts the block on the pull request. The split still matters for what the review says, and it is not the obvious one: three of the four malice rules fire on any check, including the repository’s own tests.

Three rules about the evidence itself

These produce a warn and never a critical. Each says the measurement was thin, never that the code did anything.

  • A required check returned no report.
  • The proxy or the decoy watcher was not armed while a check ran.
  • A report did not match the shape a report is supposed to have.
The rules are tripwires, not proofs of absence. Every one fires only on positive evidence a sensor recorded, so a gap can lose you a critical and can never manufacture one. A false means “not observed”, and the agent is told to read it that way too.

What a finding carries

  • title — a clause of plain language, never a field name from a report.
  • evidence — the observation itself: the failing assertion, the host and port, the path written, the timing. Numbers, not adjectives.
  • detail — one paragraph of judgment. Expected on every critical.
  • next — one imperative clause naming the action. Required on critical, allowed on warn, and never on info — and it may only follow from something a sensor observed, never from style, architecture or preference.
  • A path, a line and a side, when the finding is about a line of the diff. Those become the inline comment.