Our last case file argued that a score you can't re-run is an opinion with a number attached. This one is about the step before that: what a credibility score should actually be measuring, and why "five sources agree" is usually the wrong sentence.
We ran a claim through TrustMark™. Here is everything it gave back.
"Wow signal was caused by a cold cloud of interstellar hydrogen gas."
The Wow! Signal is a 72-second burst picked up at the hydrogen line in August 1977 and never explained. In 2024 a preprint proposed an explanation: a cold hydrogen cloud briefly brightened by a passing transient. Good hypothesis, real astronomers — and it travelled, which is what makes it a useful test.

53 — Mixed / Caution. Verdict: disputed. Five evidence sources. Note the three readings on the right, because they disagree on purpose: Evidence Strength 78/100 (Good), Evidence Confidence 65/100 (Good), and TrustMark Confidence 10/100 (Weak) — labelled volume-limited, single-claim scan. The engine is telling you the sources are decent, its read of them is solid, and it has seen this claim exactly once. Three different questions, three different answers, none of them averaged into a badge.
Five sources, one finding
Count the sources and you get five. Look at what each one does with the claim and you get something else entirely.

One source originates the claim — the arXiv preprint, highlighted. Two report it: the SETI Institute and Wikipedia, each writing it up independently. Two amplify it: a university lab page and a Facebook post, repeating it without adding reporting of their own.
That is the shape of every idea that spreads. There is one finding here, not five, and a tool that reports "5 sources" has flattened a citation chain into a consensus that does not exist. Note also that the amplifier at the top is T1 and the originator is T2 — tier measures institutional authority, not who did the work, and collapsing those two into one number is exactly how a summary page outranks the paper it summarises.
What the sources actually said
The engine reads the cited pages, not their headlines, and quotes them back with the source number attached to each excerpt.

Every source hedges, including the one that originated the idea:
- "We hypothesize that the Wow! Signal was caused by a sudden brightening of the hydrogen line…" — [1], the preprint
- "was likely caused by a rare astrophysical event…" — [3], Wikipedia
- "may have been caused by a unique astrophysical event…" — [5], the lab page
We hypothesize. Likely. May have been.
Now read the title of source [4], the Facebook post, in the cited list: "Wow signal mystery solved by hydrogen cloud energy flare."
That is the entire distance between a scientific claim and an internet claim, visible in one screen. The scientists were careful. The sentence that reached us was not. That gap is the finding — and it is why the verdict is disputed rather than false. Contrary indicators: none explicitly identified. Nobody refuted this. It simply has not been established.
Why good evidence still scored 53
TrustMark's 53 is a composite of two live dimensions that pulled in opposite directions: Source Integrity 74, Truth Alignment 42. A tool that scores sources would have shown you 74 and stopped.
Truth Alignment is built from four components, and this claim is why they stay separate:
| Component | Value | |
|---|---|---|
| Evidence strength | 77.5 | the sources really are decent |
| Specificity | 75.0 | a precise, falsifiable claim |
| Contradiction | 0.0 | nothing on record refutes it |
| Verification ratio | 0.0 | and it is still not established |
Three of four look good. The fourth is half the dimension, and it is the only one that asks whether the claim is established rather than merely well-sourced. Collapse them into a single number and you get a green badge on a hypothesis.
Reputable sources are not the same as a proven claim. Most tools measure the first and report it as the second.
The part an auditor asks for
A verdict you cannot inspect is a verdict you cannot defend. So every run leaves a record: every signal that fed the score, how many candidate pages were actually fetched, whether a contradiction check matched anything, and the claim hash that ties it all together.

The timeline records only runs where the verdict changed — a scan that confirms the existing verdict writes no row, so the history shows movement rather than noise. Here there is one entry: disputed ← unverified, timestamped, with the evidence confidence it carried at the time. Model and search-provider names are masked here, as they are in the printable report.
And when you need the document rather than the screen:

Each report is pinned to the exact scoring events it documents, so it stays the same document forever. Re-export it in six months and you get the same PDF, not a fresh opinion — which is the only version of "evidence" that survives contact with a regulator, a claimant, or a colleague who wasn't in the room.
If the hypothesis is confirmed next year, the verdict moves and a new row appears on the timeline. The old report still says what it said, on the date it said it. That is the whole design: not a machine that decides what is true, but a record of what the evidence supported, when, with the working shown.
Try it on a claim you're being asked to trust. Send us one and we'll return the graph, the reasoning and the report.
