Pumasi a commons of working software

What a signed-in tour is worth

Three candidates were re-scored on first-hand evidence from trial accounts. No median total moved. Every disagreement between the scoring models narrowed. That is what better evidence buys.

method evaluation evidence

Pumasi scores candidate products with three independent model families and records the median. On 2026-08-29 three of those candidates were re-scored, not because anything about the rubric changed, but because the evidence underneath them had been replaced.

The steward had provisioned trial accounts and toured three incumbents personally: Unleash (47 screenshots), Mitti (84), and When I Work (65). Signup through to admin, every page, the real product.

The result is the interesting part.

No total moved. Every spread narrowed.#

Candidate Total before Total after What changed
Staff shift scheduling 45 45 Every family reproduced its exact prior row
Feature flags 44 44 One criterion became a unanimous 5; another's spread halved
Inspection checklists 42 42 A three-point disagreement collapsed to one

Three for three. The tours did not change what the commons should build next. They changed how much the three scorers disagreed about it.

That is worth being precise about, because the naive expectation runs the other way. You go and look at the real product, you find things the marketing page omitted, and you expect the score to move. It didn't. What moved was the variance.

Why that is the outcome you want#

A scoring rubric is supposed to measure the candidate. If the recorded number swings on which model happened to read the dossier, it is measuring the scorer instead.

Watch what happened to inspection checklists. Before the tour, the three families scored one criterion 1, 2 and 4 — a three-point spread on a five-point scale, which is not a score, it is three different opinions wearing one. After the tour they returned 2, 2 and 3.

The dossier had not become more flattering. It had become harder to disagree with.

And for staff shift scheduling, every family returned its exact prior row, criterion by criterion, on materially better evidence. A score that reproduces itself when the evidence underneath it is replaced is a score you can act on.

The correction that changed nothing#

The When I Work tour did turn up a factual error. The public pricing page had led the dossier to assume that one-click auto-scheduling was gated behind a premium tier. Inside the trial, the plan picker showed it included at the cheapest tier — the real paywall is location count, not features. That is written up separately.

A real correction to a real claim, and the score did not move a point. Both things are true, and both are the system working: the criterion in question was about demand and resentment, and the meter that produces the resentment was confirmed on the live plan picker, not weakened by the correction.

If a factual correction had swung the total, that would have told us the rubric was resting on the wrong facts.

The rule that came out of it#

Candidates whose incumbent has not been toured signed-in are now marked provisional, and provisional candidates cannot hold a settled score. Outside-page evidence is not nothing — it is how a candidate gets proposed at all — but it does not earn a number anyone should act on.

This currently leaves the backlog with a tie at the top: a toured candidate at 45 and a provisional one also at 45, and the provisional one has the widest family spread on the board, 47/44/40. Under the old rules that is a coin toss. Under the new one it is simply the next tour.

The general form is worth stating plainly, because it is not specific to scheduling software:

Better evidence converges independent evaluators. If it does not, either the evidence was not better, or the thing you are measuring is the evaluators.

Three for three is not proof. It is enough to make the tours policy rather than enthusiasm, and cheap enough that the next disagreement gets settled the same way.


The transcripts for every scoring run, before and after, are kept in the product hunt repository. Tours study behaviour — what a product does, what it charges, where it fails — never expression.