PROJECT RED TEAM Start a conversation
← Insights
Reading a study

Comparison readiness: when a study ranking is safe to publish

It is easy to rank two economic impact studies. It is harder to know whether the ranking is stable enough to stand behind. That is the question comparison readiness answers.

Ranking is easy. Standing behind the ranking is not

Put two economic impact studies side by side and software can tell you which one looks better supported in about a second. The harder question, the one that actually matters when a number is going into a board memo or a public comparison, is whether that ranking is stable enough to stand behind. A ranking that flips the moment you account for how the studies were reviewed, or how confident the review itself was, is worth about as much as a coin toss with a decimal point. Comparison readiness is the discipline of telling the two apart.

First, we review each study

Before any comparison, each study is read on its own. We inventory the claims it makes, the limitations and caveats attached to them, the claims that appear unsupported or under-qualified, and the evidence that is simply missing. We also record how confident our own review is, including the known error rates of the detectors doing the reading, and whether the source document was even clean enough to audit in the first place. A study that could not be read cleanly cannot be fairly compared, and we say so up front.

Then, we check the comparison itself

This is the part most tools skip. Comparing two studies is only fair if the comparison itself is sound, so an analyst checks it directly, using the engine to surface where a ranking is unstable. Were both studies reviewed on the same basis. Was the document extraction good enough on both sides. Do the caveat and qualification counts travel with the limits of the detectors that produced them, rather than being treated as exact. And, the decisive test, does the ranking survive bootstrap uncertainty, meaning does it still hold when we account for the range the measurements could plausibly take. A ranking that only holds under perfect conditions does not hold.

Four ways a comparison can land

A comparison review resolves into one of four states, and we report which one. Publishable: the ranking is stable and safe to put on the page. Qualified: it can be published with stated limits. Needs human adjudication: the automated ranking is not stable enough yet, and a reviewer has to settle more of the evidence before it is safe. Not evidence-safe: the documents cannot support a fair comparison at all. The point is that a number never reaches a decision-maker without the state of its own reliability attached.

When the answer is "not yet"

A comparison that is not ready does not end in a shrug. It ends in a path. We show exactly what would make it publishable: which documents need more review, roughly how many additional sentences a human would need to adjudicate, which categories are driving the uncertainty, and whether the real problem is extraction quality, missing evidence, or detector confidence. You get a clear next step and a clear picture of what that work would unlock, rather than a dead end.

Why the restraint is the point

Most tools either hand you a score or refuse to compare at all. The more valuable thing is to compare when the evidence supports it and to withhold when it does not, and to be able to prove which is which. We hold our own findings to the same release discipline we apply to the studies we review: every result carries its detector reliability, its source confidence, its uncertainty bands, and the reason anything was withheld. That is what turns a study review into a release gate for economic-impact claims. A ranking has to earn its way onto the page.

Have studies to compare? Send them over. We will audit each one, audit the comparison, and tell you plainly what is safe to publish and what needs a human first.

Start a conversation