Skip to content

The prompt you just tested? Monitor it daily

· 3 min read

You tested a prompt, got a result you trust, and moved on. A week later a provider updates a model, and the result you trusted quietly stops being true. The test you ran once told you the answer for one day. The question worth answering is whether it still holds.

A preset can answer that without any new setup. It already holds the models, the rubric, and the context, which is exactly what a scheduled benchmark needs, so you can promote it to an ongoing check in one step.

From a one-off run to a standing watch

A preset already holds the models and the rubric that decide pass from fail, so standing up a scheduled benchmark takes one sentence. You name the preset, give it a prompt and a cadence, and the benchmark runs that exact standard every day. There is nothing to rebuild.

#ai-quality Chat
R
Ravi
the support-triage prompt passed its check today. want to keep an eye on it in case a model update knocks it off.
R
Ravi
@llmprover run our support-triage preset on this prompt every day and alert me if it drops below 90
●●● seeding a daily benchmark from the support-triage preset...
●
llmprover bot
Done. Daily benchmark created from the support-triage preset, same models and rubric. First run is logged. I set a threshold alert at 90, so you get an email the day a run drops below it, and I saved a note so I can report the latest status next time we talk.

One message turned a prompt that passed once into a prompt that is watched every day, measured the same way each run, with an alert on the quality floor.

What the watch catches

Here is why the daily run earns its place. The benchmark holds steady for a week, then a provider update knocks the score down, and the scheduled run catches it the morning it happens instead of a user catching it later.

Daily benchmark score, seeded from a preset (out of 100)

Illustrative: a daily benchmark seeded from a preset, holding above the quality floor until a model update drops it on day 11

Because the benchmark reuses the preset’s configuration, the measurement itself did not move between runs. That rules your own setup out as the cause, so the drop on day 11 is a real change to investigate (a model or provider update, for instance), not an artifact of a rebuilt config.

Turn a tested prompt into a daily benchmark

Seed a scheduled benchmark from a preset in one step. Benchmarks are on Pro and above.

Get started on Pro

A one-off test tells you the answer on the day you asked. A scheduled benchmark tells you the day the answer changed, which is the one you actually need. Run the preset once to test, then seed a benchmark from it to watch, the setup you already did is the monitoring you did not have to build.

Scheduled benchmarks are on Pro and above, daily on Pro, more frequent on higher plans.

presets benchmarking mcp