Skip to content

Benchmark your own model against the frontier

· 4 min read

If you fine-tuned a model, host one yourself, or built a pipeline, the hard question is whether it actually holds up against the frontier. Running it once by hand and eyeballing the answer is not evidence. A benchmark that scores it against GPT, Claude, and the rest, on your task, with a run history, is.

That is what Bring Your Own Endpoint does. Register any OpenAI-compatible endpoint and it appears in the model picker alongside every hosted model, scored by the same rubric in the same run. The run below put a registered endpoint head-to-head with two frontier models.

The setup

Three models scored on one task: a registered endpoint and two hosted frontier models. The endpoint is a Grok 4.7 deployment registered through BYOE; from the pipeline’s point of view it is just another model in the picker. Everything else is held constant.

The task and rubric

Prompt: Summarise the three most important risks in this clause in under 50 words. (A contract indemnity clause with a deliberately misdirected liability cap.)

Rubric: Factual accuracy (40%), Completeness (35%), Conciseness (25%).

Models: your registered endpoint (Grok 4.7), Claude Sonnet 4.5, GPT-4.1. Temperature 0.

The result

Agent
> run_evaluation (clause task, rubric, 3 models: your endpoint + 2 frontier)
your endpoint (grok-4.7): quality 70 | cost $0.00299
claude-sonnet-4-5: quality 60 | cost $0.00154
gpt-4.1: quality 25 | cost $0.00053
winner by quality: your endpoint (70) | winner by efficiency: gpt-4.1

In this run, the registered endpoint scored highest.

ModelQualityCost per call
Your endpoint (Grok 4.7)70$0.00299
Claude Sonnet 4.560$0.00154
GPT-4.125$0.00053

The registered model won on quality and cost the most. Both facts matter.

The registered endpoint produced the most complete risk summary, catching that the liability cap was misdirected (it limits the Customer, not the Supplier) where GPT-4.1 missed it. It also cost roughly twice what Claude did and nearly six times what GPT-4.1 did for this call.

That is the whole point of putting your own model in the same run as the frontier. You do not get a verdict handed down from a leaderboard someone else ran on someone else’s prompts. You get your model, your task, your rubric, and the quality and cost sitting side by side, so the tradeoff is a decision you make on numbers rather than instinct.

Insight: A win on quality that costs twice as much is not automatically the right choice. It is a tradeoff, and the only way to make it deliberately is to see both numbers for your own model on your own task. The benchmark turns “we think our fine-tune is better” into “it scores 10 points higher at twice the cost, now decide.”

Put your own model in the ring

Register any OpenAI-compatible endpoint and benchmark it against the field. Pro and above.

Get started on Pro

It is not just models

Anything that speaks the OpenAI API format can be registered, not only hosted or fine-tuned models. A RAG pipeline, an agentic workflow, or a multi-step backend can be registered as an endpoint and benchmarked as a black box: the same prompt a real user would send goes in, the output is scored by the same rubric as any model.

That is the difference between unit-testing a model and measuring a system. A model that scores 95 in isolation can score far lower as the generation step of a pipeline with a weak knowledge base. Benchmarking the pipeline end to end catches what the isolated model test cannot. The store-content A/B shows how much the pieces around the model move the result.

Reproducible, not anecdotal

The value is not the single run above. It is that the run is repeatable. Register once, and the same benchmark reruns on a schedule or on demand: when you ship a new fine-tune, when a provider updates a frontier model, when your pipeline changes. The evidence stays current, with timestamps and exportable results a reviewer can audit.

BYOE is available on Pro and above. See everything you can benchmark your own endpoint against, including the compliance and regression-testing angles, on the Benchmark Your Own Model page. For the full framework this fits into, including how your model choice interacts with prompt, store content, and drift, see the practical framework for your LLM stack.

What’s next

byoe benchmarking evaluation cost-optimisation