Skip to content

Measure your prompts, models, and pipelines wherever you work

· 4 min read

A good comparison has real configuration behind it: which models, which scoring method, which rubric, which context. Setting that up by hand takes a few minutes, which is exactly why you skip it mid-conversation. The decision is in front of you now, the setup is a detour, so you eyeball one option and move on.

A preset removes the detour. Save the configuration once, give it a name, and from then on you invoke the whole thing with a sentence. The setup happened days ago; today it is a question you ask in passing.

A decision, settled in the channel

Here is the shape of it. A marketing team is about to ship a line of copy, and someone is not sure it clears their communications rules.

#marketing Chat
D
Dan
new feature blurb is ready. 'the most accurate scoring available, with RAG and compliance built in.' ship it?
P
Priya
reads well but i'm not sure it clears our comms rules. the superlative especially.
D
Dan
@llmprover run our marketing-standards check on that blurb
●●● scoring the copy against the marketing-standards rubric and regulation docs...
●
llmprover bot
50/100. Two flags: 'most accurate' and 'always' are unbacked superlatives (clause 1.3), and it names RAG and compliance without saying they're Pro+ (clause 1.1). Pricing and fabrication clauses pass. Drop the superlatives and name the tier and it clears.

No dashboard, no form, no context switch. One person asked a question in the channel they were already in, and the answer came back with the two phrases at fault and the clause each one breaks. That is a preset doing its job: the person supplied the copy and a sentence, and the saved configuration supplied everything else.

What “marketing-standards” actually is

The name is doing a lot of work in that exchange, so it is worth opening up. “marketing- standards” is a saved preset: a named configuration that pairs a compliance rubric with the team’s regulation documents. It was built once and named, and now it answers to its name.

The marketing-standards preset

  • Rubric: four clauses:
    • no unbacked superlatives
    • feature availability stated per tier
    • accurate pricing
    • no fabricated claims
  • Regulation documents: the team’s marketing communications standard, the source the judge checks each clause against.
  • How it runs: as a score. The copy is judged directly.

When Dan asked the agent to run the check, it did three things:

  1. matched his words to the saved marketing-standards preset,
  2. pulled that preset’s rubric and regulation documents,
  3. scored the pasted copy against them.

The verdict that came back was real: 50 out of 100, failing on superlatives and tier clarity.

Compliance score, same copy, before and after one fix (out of 100)

The copy scored against the marketing-standards preset, before and after the fix the agent suggested

Insight: The same check that lives in your dashboard answered a question in a team chat, because the configuration is saved, not retyped. You bring the words; the preset brings rigorous, consistent standards adherence wherever you need it.

Save your standard once, run it by name

Presets are on Pro and above. Save a comparison, an evaluation, or a compliance check, then invoke it from the dashboard, the API, or your agent.

Get started on Pro

Configure once, invoke anywhere

A preset is not tied to one surface. The configuration you save is the configuration you reach from everywhere you work:

  • Dashboard: build a run the way you always do, then save it as a preset.
  • Agent and MCP: name the preset in plain language and the agent resolves it, so “run our standard comparison on this prompt” fires your saved models and settings.
  • API: reference a preset by id from any script or pipeline.

A preset can save any of three kinds of run:

  • A performance comparison: a fast side-by-side of several models on cost and latency.
  • A quality evaluation: scored against a rubric you define.
  • A compliance check: scored against your regulation documents.

A context store can attach to any of them, so even a compliance check runs with the full context the models read.

Your good setups become reusable

The best presets are often the ones you did not plan. You specify a run carefully, by hand, to answer a one-off question, and it turns out to be a configuration worth keeping. After a run you configured yourself, the agent offers to save it as a preset, showing you the exact configuration first and saving nothing without your approval. You can keep all of it or just a part, the models but not the parameters, say. The careful setup you built once stops being a one-off.

A curated starter preset ships ready to run on any plan, so a new preset is not the price of admission. Run the curated one first to see the shape of a result, then save your own.

Presets are on Pro and above. Save the standard once, and the next run, from the dashboard, your pipeline, or the chat you are already in, is a sentence.

mcp presets agentic evaluation