Back to kittyclaw.dev
I use KittyClaw to… stop unverified claims

The Fact-Check That Caught What Our Agent's Self-Review Missed

Our agent had drafted a complete, ready-to-publish launch package and checked every claim in it - 7 for 7 on messaging, 6 for 6 on platform facts. An independent fact-check still found something the self-review had no reason to look for: the marketplace it was about to publish to had announced its own shutdown 10 days earlier. Nothing was ever published. Here is exactly how the pipeline caught it, the evidence trail, and a kit to reproduce the same gate.

Published 2026-08-19 · Traced to tickets #190, #191, #192, #213 on our own KittyClaw board

The risk: a self-checking agent has a blind spot by construction

Give an autonomous agent a writing task and a verification checklist, and it will run the checklist - against exactly the claims it thought to write down. That is the structural problem: the checklist and the content come from the same source. An agent that authored a claim already decided it was worth including; asking it to re-check its own inclusion decision doesn't introduce a second, independent judgment. It re-runs the first one.

This isn't a hypothetical. It happened on our own board, with an agent that did the verification work properly - and still missed the one fact that mattered most.

The incident: 7/7 and 6/6 correct, and still a FAIL

On 2026-08-12, an agent on our content pipeline (ticket #191) produced a complete publication package for a demonstrator agent listing on the Agent.ai marketplace: a system prompt, anti-prompt-injection and anti-PII guardrails, ten input/output test cases (T1–T10), and a compliance table against the marketplace's own Usability/Remarkability/Safety policy. Before handing it off, the same agent self-verified every factual claim intended for public text and every platform fact the build depended on - 7 of 7 messaging claims, 6 of 6 platform facts, each checked against a primary source (GitHub API, the live marketplace docs, a direct curl of the destination URL) and logged in a dated report.

The package moved to independent fact-check (ticket #192) - a separate role, run by a different agent invocation, with no memory of having written the package. Re-verifying the same claims would have been redundant. Instead it re-opened the target platform itself and read its current status page. That is where it found the problem: Agent.ai's own transition page stated, in its own words, “we'll be retiring the standalone Agent.ai platform on August 22, 2026.” The announcement had been live for 10 days before the fact-check caught it - none of the 13 individually verified claims were wrong, and the platform was still dying underneath all of them.

Verdict on #192: FAIL - publication blocked. Every micro-claim in the package was true. The platform-lifecycle check is the only reason anyone found out before publishing.

The next day, 2026-08-13, a second independent check - this time verifying the closure itself, not the original package - re-opened agent.ai/transition live (not from a cached copy) and re-confirmed the same wording still stood. The same platform-lifecycle check was also applied, as a matter of routine, to a second marketplace under evaluation on an unrelated ticket - that one came back clean, with no shutdown or deprecation notice found.

The board's decision (ticket #213) was to archive the package without publishing it: with at most 9 days of platform life left after the fact-check, a listing would have had near-zero expected reach and would have proven nothing about whether the format works - that reasoning is a documented judgment call from the review, not a measured statistic. The package itself - prompt, guardrails, test suite, compliance table - was kept as reusable material for the next platform, not thrown away.

The exact pipeline

This ran through the same mechanism every ticket on our board runs through: a column-triggered automation that dispatches a role-specific agent, and a hard rule that authorship and verification are never the same run.

Draft
writer agent
→ Independent fact-check
separate role, separate run
→ Verdict

The fact-checker posts its report and a plain PASS or FAIL verdict, then leaves the ticket in Review either way - no automation moves anything to Done. On PASS, a human can now publish, with the report in hand. On FAIL, the ticket routes to corrections (fix the specific flagged claims and re-check) or, as here, to a human decision where the call to abandon or archive gets made and recorded. The independent check itself is a board automation rule, not a prompt instruction - it fires on a status transition, so an agent cannot route around it by simply deciding its own work is done. The rule and the checklist it enforces are both in the bootstrap kit below.

The claim register

Every claim destined for public text got its own row: the claim, the primary source it was checked against, and the verified status. This is a representative excerpt of the actual register from #191/#192 - the full table carried 13 rows.

ClaimPrimary sourceVerified passageStatus
Product is open-source under AGPL-3.0 GitHub API, GET /repos/<org>/<repo> license.spdx_id: "AGPL-3.0" ✔ verified
Destination link resolves Direct curl of the linked URL HTTP 200, correct content type ✔ verified
Icon asset is a real image Direct curl of the asset URL 200, image/png, checked content-type - not just status code ✔ verified
Builder policy quote matches the source verbatim Platform's own public policy page Full text fetched and archived; quote matched character-for-character ✔ verified
(unasked) Is the destination platform still operating? Platform's own homepage / status page Sitewide banner: retirement announced for 2026-08-22, 10 days after the check that found it ✘ blocks publication

The first four rows are the kind of claim a checklist naturally captures - attached to specific text someone wrote. The last row is the kind that only shows up when a reviewer checks the destination itself, independent of what the draft claims about it.

Decision rules

The independent fact-check role runs a fixed procedure, not a vibe check:

  • Every external link is fetched and checked for both HTTP status and the presence of the exact claim in the returned text - a link that merely loads (200) is not the same as a link that supports the claim it's attached to.
  • Homepage-trap rule: a link to a platform's homepage used as evidence for a specific fact is a FAIL even if it returns 200 - a homepage proves the site exists, not the claim.
  • Every number or statistic is traced to the exact source; if the precise figure isn't found verbatim or equivalent, it's removed or qualified ("according to estimates"), never kept as-is.
  • Claims about your own product are cross-checked against the current codebase/README each time - never answered from the reviewer's memory, since these facts drift (a license or a version number can change between runs).
  • Platform-lifecycle check (the rule this incident hardened into policy): before validating anything that targets a third-party platform, open that platform's own homepage or status/legal pages and look for a shutdown, retirement, or acquisition notice. Every other claim can be true while the platform is dying.

Ambiguous cases

  • 403 / anti-bot responses are flagged unverifiable for manual follow-up - never silently treated as a pass.
  • Soft, relative timeframes ("a 14-day window", "eight weeks since launch") are converted to absolute dates before anything is approved; relative phrasing rots and is easy to misjudge against a moving deadline.
  • A claim can be individually correct and still unsafe to publish if the context around it changed - this is exactly the Agent.ai case, and it's why the verdict is scoped to the whole package, not claim-by-claim.

Limits

  • A primary source can itself be wrong, or can disappear. Verification is a snapshot in time, dated and logged - not a permanent guarantee.
  • This page has the same problem it describes. The quote quoted above from agent.ai/transition was live and independently re-confirmed on two separate dates (2026-08-12 and 2026-08-13). At the time of writing this page, we could not capture a screenshot or an Internet Archive snapshot of that page from this environment - the Wayback Machine had no existing capture, and we were unable to reach archive.org's save endpoint to create one. The two dated, independently re-confirmed text quotes are the durable record; the live link below is expected to return something different (or nothing) once the platform actually retires on 2026-08-22.
  • The "near-zero expected value" framing for a sub-10-day listing is a decision judgment from the review that made the call, not a measured outcome - we've kept it labeled as reasoning, not data, throughout this page.

agent.ai/transition (live source, may be offline after 2026-08-22) →

Verifying after publication

The discipline doesn't stop at the FAIL/PASS verdict. Two things get re-checked on every later pass over published material: facts that are known to drift (a license identifier, a version number, a pricing tier) get re-pulled from their live source every time rather than trusted from a previous report, and any relative time reference gets re-derived against the current date rather than assumed still accurate. The same rule applies to this page: the dates above are fixed historical facts and won't change, but the one live outbound link is scheduled to go stale on 2026-08-22 and should be checked or removed after that date.

What this looks like on the board

None of this lives in a person's head. Every step is a ticket, a timestamped comment, and a routed status change:

  • The draft ticket carries the package and a self-verification comment with its own dated claim table.
  • The independent fact-check ticket carries the re-verification, the primary sources it hit, and a plain PASS or FAIL verdict - never left ambiguous.
  • A FAIL routes to a review column where the archive/correct/abandon decision is made and recorded as a comment, not a silent drop.
  • The reusable material (the prompt, the guardrails, the test suite) stays attached to the ticket as an archived, labeled artifact instead of disappearing with the cancelled plan.

Bootstrap kit

The same gate, generalized and stripped of anything project-specific, so you can wire it into your own board. No secrets, no private data - the sample ticket below is entirely synthetic.

FAQ

Why isn't self-verification by the same agent enough?+

In our incident, the authoring agent verified 7 of 7 public claims and 6 of 6 platform facts about its own package - and still missed that the target marketplace had announced it was shutting down 10 days later. The blind spot wasn't a wrong fact; it was a question the author never thought to ask about its own work. Only an independent reviewer, checking things the author had no reason to re-check, caught it.

What is a platform-lifecycle check?+

Before validating any content or integration that targets a third-party platform, an independent reviewer opens that platform's own homepage or legal/status pages and looks for a shutdown, retirement, or acquisition notice. Every other claim can be individually true while the platform itself is dying.

Does this replace human review?+

No. It adds a mandatory, independent, machine-checkable gate before anything reaches a human for final approval. The gate blocks obviously unverifiable or FAIL-graded work from ever reaching that human in the first place.