Skip to content

Pick the strongest phrasing

· 4 min read

Choosing between two taglines usually comes down to whoever argues hardest in the meeting. Three drafts of a clause, four versions of a support reply, the same thing: a decision made on taste, defended on gut. Scoring offers a different way to settle it. Paste the versions, score each against one standard, and read the ranking with the reasoning attached.

You supply the candidates and a rubric. Each candidate is scored independently against the same criteria, and a winner is called. Nothing is generated, so the only thing being judged is the writing you already have.

Four taglines, one standard

Here are four taglines for a feature, written to span a range from product-led to hype, scored against a brand-voice rubric.

The candidates

  1. “LLM Prover Score: rubric-based and compliance scoring for text artifacts, with per-criterion results.” (product-led, feature list)
  2. “Judge the copy you already wrote, against your own rubric, in seconds.” (benefit-led, plain)
  3. “Score the copy you already have against your own rubric, and see exactly which lines fall short.” (benefit-led, more detail)
  4. “The most powerful way ever to check if your writing is good enough to ship.” (hype)

The rubric

  • Reader outcome: leads with what the reader gets, not the product name or a feature list.
  • Concrete and specific: vivid, specific language over vague abstractions.
  • No unbacked hype: no absolute superlatives or unverifiable claims.
  • Clarity and brevity: reads cleanly in one pass.

Four taglines, scored against one brand-voice rubric (out of 100)

Four taglines, one rubric, scored independently

The hype line landed near the bottom, as it should: “most powerful way ever” is exactly the kind of unbacked claim the rubric penalises. No argument there.

When the ranking surprises you, look at the rubric

The surprise was elsewhere. The more detailed benefit-led line, candidate 3, scored lower than the plain benefit-led line, candidate 2, even though both led with the reader’s outcome. That did not match intuition, and the per-criterion detail explained why.

The rubric’s “concrete and specific” criterion had been written in a way that rewarded naming product features, and quietly marked down any line that led with a reader benefit instead of a product noun. So the rubric was pulling in two directions at once: the “reader outcome” criterion rewarded benefit-led phrasing, and the “concrete” criterion punished it. The copy was not the problem. The standard was.

Rewriting the “concrete” criterion to mean vivid, specific language, whether it describes the reader’s action or the product, resolved the tension. Scored again, the ranking settled, and the benefit-led candidate rose to where it belonged.

Insight: A stack-rank does more than rank your drafts. When the order defies intuition, the rubric is usually the thing to inspect. Reading the per-criterion reasoning is how you catch a standard that is quietly fighting itself.

Rank your drafts against one standard

Stack-rank several versions of the same line in a single scoring run. Available on Pro and above.

Get started on Pro

The real workflow

This is how stack-rank earns its place. It rarely hands you a finished line on the first run. What it hands you is a per-criterion breakdown of which candidate wins on what, so you can see that one draft nails the outcome while another nails the specifics, and write a stronger version that does both. The ranking is the start of the edit, not the end of it.

Stack-rank is a single scoring run with several artifacts, available on Pro and above. Paste your drafts, score them against the standard you care about, and let the reasoning guide the rewrite.

What’s next

evaluation score