Benchmark your own model against the frontier
If you fine-tuned a model, host one yourself, or built a pipeline, the hard question is whether it actually holds up against the frontier. Running it once by hand and eyeballing the answer is not evidence. A benchmark that scores it against GPT, Claude, and the rest, on your task, with a run history, is.
That is what Bring Your Own Endpoint does. Register any OpenAI-compatible endpoint and it appears in the model picker alongside every hosted model, scored by the same rubric in the same run. The run below put a registered endpoint head-to-head with two frontier models.
The setup
Three models scored on one task: a registered endpoint and two hosted frontier models. The endpoint is a Grok 4.7 deployment registered through BYOE; from the pipeline’s point of view it is just another model in the picker. Everything else is held constant.
The task and rubric
Prompt: Summarise the three most important risks in this clause in under 50 words. (A contract indemnity clause with a deliberately misdirected liability cap.)
Rubric: Factual accuracy (40%), Completeness (35%), Conciseness (25%).
Models: your registered endpoint (Grok 4.7), Claude Sonnet 4.5, GPT-4.1. Temperature 0.
The result
In this run, the registered endpoint scored highest.
| Model | Quality | Cost per call |
|---|---|---|
| Your endpoint (Grok 4.7) | 70 | $0.00299 |
| Claude Sonnet 4.5 | 60 | $0.00154 |
| GPT-4.1 | 25 | $0.00053 |
The registered model won on quality and cost the most. Both facts matter.
The registered endpoint produced the most complete risk summary, catching that the liability cap was misdirected (it limits the Customer, not the Supplier) where GPT-4.1 missed it. It also cost roughly twice what Claude did and nearly six times what GPT-4.1 did for this call.
That is the whole point of putting your own model in the same run as the frontier. You do not get a verdict handed down from a leaderboard someone else ran on someone else’s prompts. You get your model, your task, your rubric, and the quality and cost sitting side by side, so the tradeoff is a decision you make on numbers rather than instinct.
Put your own model in the ring
Register any OpenAI-compatible endpoint and benchmark it against the field. Pro and above.
It is not just models
Anything that speaks the OpenAI API format can be registered, not only hosted or fine-tuned models. A RAG pipeline, an agentic workflow, or a multi-step backend can be registered as an endpoint and benchmarked as a black box: the same prompt a real user would send goes in, the output is scored by the same rubric as any model.
That is the difference between unit-testing a model and measuring a system. A model that scores 95 in isolation can score far lower as the generation step of a pipeline with a weak knowledge base. Benchmarking the pipeline end to end catches what the isolated model test cannot. The store-content A/B shows how much the pieces around the model move the result.
Reproducible, not anecdotal
The value is not the single run above. It is that the run is repeatable. Register once, and the same benchmark reruns on a schedule or on demand: when you ship a new fine-tune, when a provider updates a frontier model, when your pipeline changes. The evidence stays current, with timestamps and exportable results a reviewer can audit.
BYOE is available on Pro and above. See everything you can benchmark your own endpoint against, including the compliance and regression-testing angles, on the Benchmark Your Own Model page. For the full framework this fits into, including how your model choice interacts with prompt, store content, and drift, see the practical framework for your LLM stack.
What’s next
BYOE Guide
Register any OpenAI-compatible endpoint and benchmark your own models against any model in the registry.
Benchmark Guide
Save a run configuration, schedule it on a cadence, and track quality and cost drift over time.
Comparisons Guide
Fire a prompt at multiple models simultaneously and compare cost, latency, and output side by side.