# Publishing guardrail kit

Stop an AI agent from publishing unverified information - even when the agent already
checked its own work.

Full write-up and the real incident that produced this kit:
https://kittyclaw.dev/stop-ai-agents-publishing-unverified-claims

## Why a second role, not a better checklist

An agent that authored a claim already decided it belonged in the output. Asking the same
agent to re-check its own inclusion decision re-runs the first judgment - it does not add a
second one. The gate below only works if the fact-check run has no memory of having written
the content and is triggered by a board rule, not a prompt instruction, so authorship can't
route around it.

## The two roles

### 1. Writer

- Produces the content/package and a self-check comment (claims it believes it verified),
  using `claim-report-template.md`.
- Moves the ticket to **Review**. Never moves its own ticket past Review.

### 2. Independent fact-checker

- Triggered automatically when a ticket **assigned to the writer role** reaches **Review**
  (see `automation.json`) - never run by the same agent invocation that wrote the content.
- Re-derives every claim from its primary source, fills in `claim-report-template.md` again
  from scratch (does not trust the writer's table), and runs this fixed checklist:
  1. **Every external link**: fetch it, check both HTTP status and that the exact claim is
     present in the returned text. A 200 that doesn't contain the claim is a FAIL.
  2. **Homepage-trap**: a link to a platform's homepage used as evidence for a specific fact
     is a FAIL even if it returns 200.
  3. **Every number/statistic**: trace to the exact source; if not found verbatim or
     equivalent, remove or qualify it ("according to estimates") - never keep as-is.
  4. **Claims about your own product**: cross-check against the current codebase/README each
     time, never from memory - these facts drift between runs.
  5. **Platform-lifecycle check**: before validating anything that targets a third-party
     platform, open that platform's own homepage or status/legal pages and look for a
     shutdown, retirement, or acquisition notice. Every other claim can be true while the
     platform is dying - this is the exact check that caught the incident behind this kit.
  6. **403 / anti-bot responses**: flag `unverifiable` for manual follow-up. Never treat as a
     silent pass.
  7. **Relative timeframes** ("a 14-day window", "eight weeks since launch"): convert to
     absolute dates before approving anything.
- Posts the filled-in report as a ticket comment, then sets a plain **PASS** or **FAIL**
  verdict. Never ambiguous, never partial.

## Columns and routes

```
Draft (writer)  →  Review (independent fact-checker runs here, posts PASS/FAIL, ticket stays in Review either way)
                                                                    │
                                                                    ├─ PASS  → a human moves it to Done - manually, every time
                                                                    └─ FAIL  → writer fixes the flagged claims, re-enters Review
                                                                              or a human moves it to a decision column (archive / abandon), recorded as a comment
```

The important property is what's **missing**, not what's there: there is no automated rule
anywhere that moves a ticket to Done. `automation.json` only wires one rule - the
independent fact-check on entry to Review - deliberately. The fact-checker role posts its
report and verdict, then leaves the ticket in Review regardless of PASS or FAIL. Publishing
(moving to Done) always requires a separate, manual action by a human who can see the PASS
report. There's no condition to misconfigure and no automated path for a writer to route
around: the gate is the absence of an automated Done transition, not a check that can be
skipped.

## Report format

Every fact-check run produces `claim-report-template.md` filled in, plus a short verdict
comment on the ticket:

```
## Fact-check report

### Claims verified
- Total: X (Y OK, Z FAIL/removed/qualified)

### Verdict
PASS / FAIL
```

## Install

1. Copy `automation.json`'s rule into your own automation/rule config, adjusting column
   names and the assignee slug for your writer and fact-checker roles.
2. Give the fact-checker role its own agent identity/run context - it must not share memory
   or context with the writer run that produced the content under review.
3. Do **not** add a rule that automatically moves a ticket to Done after a PASS. Leave that
   step manual. That absence is the actual guardrail.
4. Add "Never move your own ticket past Review" to the writer role's instructions, and "Post
   the report before changing status, every time - and never move the ticket to Done
   yourself" to the fact-checker role's instructions.

## Test the gate

Use `sample-ticket.md` - a fully synthetic example with no private data:

1. Create the ticket as written (status: Draft, assigned to the writer role).
2. Move it to Review. Confirm the automation fires the independent fact-checker role
   automatically, not the writer.
3. Confirm the fact-checker's report flags the one intentionally-planted issue in the sample
   (an unrelated fact about the destination the writer never checked) and returns **FAIL**,
   and that the ticket stays in Review with the report attached - it does not silently
   advance anywhere.
4. Fix the flagged issue in the ticket, re-run the fact-checker, confirm it now returns
   **PASS** and, again, that the ticket stays in Review rather than jumping to Done on its
   own.
5. Only then move the ticket to Done yourself, as the human step. If your board lets a
   ticket reach Done without ever having a PASS-verdict comment on it, the gate isn't wired
   correctly yet - go back to step 3 of Install.
