cesure

[ TRUST ]

how we measure trust

A review is only worth reading if you can check it. Every claim in a Cesure review links to its source passage, and every proposed change must carry a verbatim quote, mechanically verified. This page is the honest account of what that actually guarantees, and what it doesn't.

[ THE METHOD ]

01

Every proposed change needs a verbatim quote

No proposed change enters a review without at least one quote copied character-for-character out of a source paper — the reconciler cannot propose otherwise. Compiled claims that rest on abstract-level evidence are marked as such where they appear and counted openly in the SOURCED metric on /method.

02

Quotes are mechanically verified

Each quote is checked against the source text itself by character offset — not by asking a model to vouch for its own extraction. This is a code check, not a judgment call.

03

A firewall drops what fails

Evidence that doesn't verify is dropped before it ever reaches the document. A claim resting on a quote the code can't find in the source is never written.

04

An LLM judge audits groundedness

A sample of active claims — up to 40 per audit — is judged by a separate model for whether the claim is actually entailed by its cited quote, not merely whether the quote exists.

05

Contested labels get a second pass

Revisions labeled contradicts or supersedes — the ones that overturn something you already believed — go through a second verification pass before they're proposed.

[ LIMITATIONS ]

Audits are sampled, not exhaustive.

At most 40 claims are judged per audit run. A grounded% describes that sample, not a guarantee that every claim in the document has been individually checked.

The judge can be wrong.

Groundedness is an LLM's call on whether a quote entails a claim. It can misjudge, especially on hedged, comparative, or heavily caveated claims. Judge errors are dropped from the denominator, not counted as passes — but they are still a source of noise.

Verified is not the same as correct.

Mechanical verification proves the quote exists in the source at that offset. It says nothing about whether the claim's interpretation of that quote is right — that judgment still rests on the LLM audit and, ultimately, on you.

Legacy data is marked, not backfilled.

Reviews compiled before mechanical verification shipped show — ("not yet measured") for VERIFIED rather than a retrofitted number. We don't compute a score we can't stand behind.

The three numbers this produces — SOURCED, VERIFIED, GROUNDED — sit on every review's document page and on each card in the public gallery. None of them are shown unless they were actually measured. The current measured results, with the harnesses and run dates, live at /method. The full benchmarked account — discovery vs. a keyword baseline, back-test replays against pre-registered landmarks, and faithfulness audits, every number generated from committed run artifacts by a build that fails on unsourced figures — is the whitepaper (PDF). Every earlier edition stays published unedited beside it — first edition and second edition (PDF) — including the results they got wrong and we later corrected. A paper that revises its own findings should not be able to quietly retire the ones that did not survive.