---
title: "What a signed-in tour is worth"
description: "Three candidates were re-scored on first-hand evidence from trial accounts. No median total moved. Every disagreement between the scoring models narrowed. That is what better evidence buys."
url: https://pumasi.ai/blog/what-a-signed-in-tour-is-worth/
kind: post
published: 2026-08-29
tags: ["method", "evaluation", "evidence"]
licence: Apache-2.0
---

# What a signed-in tour is worth

Pumasi scores candidate products with three independent model families and
records the median. On 2026-08-29 three of those candidates were re-scored, not
because anything about the rubric changed, but because the evidence underneath
them had been replaced.

The steward had provisioned trial accounts and toured three incumbents
personally: **Unleash** (47 screenshots), **Mitti** (84), and **When I Work**
(65). Signup through to admin, every page, the real product.

The result is the interesting part.

## No total moved. Every spread narrowed.

| Candidate | Total before | Total after | What changed |
|---|---|---|---|
| Staff shift scheduling | 45 | 45 | Every family reproduced its exact prior row |
| Feature flags | 44 | 44 | One criterion became a unanimous 5; another's spread halved |
| Inspection checklists | 42 | 42 | A three-point disagreement collapsed to one |

Three for three. The tours did not change what the commons should build next.
They changed how much the three scorers disagreed about it.

That is worth being precise about, because the naive expectation runs the other
way. You go and look at the real product, you find things the marketing page
omitted, and you expect the score to move. It didn't. What moved was the
*variance*.

## Why that is the outcome you want

A scoring rubric is supposed to measure the candidate. If the recorded number
swings on which model happened to read the dossier, it is measuring the scorer
instead.

Watch what happened to inspection checklists. Before the tour, the three
families scored one criterion 1, 2 and 4 — a three-point spread on a five-point
scale, which is not a score, it is three different opinions wearing one. After
the tour they returned 2, 2 and 3.

The dossier had not become more flattering. It had become harder to disagree
with.

And for staff shift scheduling, every family returned its **exact prior row**,
criterion by criterion, on materially better evidence. A score that reproduces
itself when the evidence underneath it is replaced is a score you can act on.

## The correction that changed nothing

The When I Work tour did turn up a factual error. The public pricing page had
led the dossier to assume that one-click auto-scheduling was gated behind a
premium tier. Inside the trial, the plan picker showed it included at the
cheapest tier — the real paywall is **location count**, not features. That is
[written up separately](/blog/the-per-seat-tax/).

A real correction to a real claim, and the score did not move a point. Both
things are true, and both are the system working: the criterion in question was
about demand and resentment, and the meter that produces the resentment was
confirmed on the live plan picker, not weakened by the correction.

If a factual correction had swung the total, that would have told us the rubric
was resting on the wrong facts.

## The rule that came out of it

Candidates whose incumbent has not been toured signed-in are now marked
**provisional**, and provisional candidates cannot hold a settled score.
Outside-page evidence is not nothing — it is how a candidate gets proposed at all
— but it does not earn a number anyone should act on.

This currently leaves the backlog with a tie at the top: a toured candidate at 45
and a provisional one also at 45, and the provisional one has the widest family
spread on the board, 47/44/40. Under the old rules that is a coin toss. Under the
new one it is simply the next tour.

The general form is worth stating plainly, because it is not specific to
scheduling software:

> Better evidence converges independent evaluators. If it does not, either the
> evidence was not better, or the thing you are measuring is the evaluators.

Three for three is not proof. It is enough to make the tours policy rather than
enthusiasm, and cheap enough that the next disagreement gets settled the same
way.

---

*The transcripts for every scoring run, before and after, are kept in the
[product hunt repository](https://github.com/pumasi-ai/pumasi-product-hunt).
Tours study behaviour — what a product does, what it charges, where it fails —
never expression.*
