Skip to content

MCP Reference

MCP Examples

End-to-end examples for running comparisons, evaluations, and benchmarks via MCP.

All examples show the exact instructions to give your agent and the tool calls that result. The agent handles the async polling automatically when told to poll.


Example 1 – Comparison with custom temperature

What you say to your agent:

“Run a comparison with the prompt ‘In one sentence, explain why agent-first SaaS will be bigger than human-oriented SaaS.’ Use GPT-4.1 Nano, Grok 4.5, Claude Sonnet 4.5, DeepSeek V4 Flash, and GPT-4.1 Mini. Temperature 0.9. No system prompt, no context. Poll every 2 seconds until complete then show me the results.”

What the agent does:

  1. Calls list_models to confirm model IDs (optional but good practice).
  2. Calls run_comparison:
{
  "prompt": "In one sentence, explain why agent-first SaaS will be bigger than human-oriented SaaS.",
  "models": {
    "openai": "gpt-4.1-nano",
    "grok": "grok-4.5",
    "anthropic": "claude-sonnet-4-5-20250929",
    "together": "deepseek-ai/DeepSeek-V4-Flash-0731"
  },
  "params": {
    "temperature": 0.9
  }
}

Note: models takes one model per provider. To include a second OpenAI model (GPT-4.1 Mini), run a second comparison or use a benchmark. See Example 4 below.

  1. Receives job_id. Calls poll_job every 2 seconds.
  2. When status: complete, reads results from the poll response.

Results shape:

{
  "comparison_id": "cmp_...",
  "results": [
    {
      "provider": "openai",
      "model": "gpt-4.1-nano",
      "response_text": "...",
      "latency_ms": 1240,
      "cost_usd": 0.000004,
      "tokens_input": 22,
      "tokens_output": 28
    }
  ]
}

Example 2 – Evaluation with gold standard

Use this when you have a known correct answer and want to score how close each model gets.

What you say to your agent:

“Evaluate the prompt ‘What is the capital of France?’ against the gold standard answer ‘Paris’. Use GPT-4.1 Nano and Claude Haiku 4.5. Poll until complete.”

Tool call – run_evaluation:

{
  "prompt": "What is the capital of France?",
  "models": {
    "openai": "gpt-4.1-nano",
    "anthropic": "claude-haiku-4-5-20251001"
  },
  "expected_output": "Paris"
}

Results shape:

{
  "results": [
    {
      "provider": "openai",
      "model": "gpt-4.1-nano",
      "response_text": "The capital of France is Paris.",
      "quality_score": 92.5,
      "scores": [
        {
          "criterion": "coverage",
          "score": 100,
          "reasoning": "Response contains the correct answer."
        }
      ]
    }
  ]
}

Example 3 – Create and run a benchmark

Use this when you want to track a prompt over time or run it on a schedule.

What you say to your agent:

“Create a benchmark called ‘Agent SaaS thesis’ with the prompt ‘In one sentence, explain why agent-first SaaS will be bigger than human-oriented SaaS.’ Use GPT-4.1 Nano and Grok 4.5. Manual schedule. Then trigger a run and poll until complete.”

Step 1 – create_benchmark:

{
  "name": "Agent SaaS thesis",
  "prompt": "In one sentence, explain why agent-first SaaS will be bigger than human-oriented SaaS.",
  "models": {
    "openai": "gpt-4.1-nano",
    "grok": "grok-4.5"
  },
  "schedule": "manual"
}

Returns suite_id.

Step 2 – trigger_benchmark_run:

{
  "suite_id": "suite_..."
}

Returns job_id. Poll with poll_job until complete.

Results shape:

{
  "run_id": "run_...",
  "model_metrics": [
    {
      "provider": "openai",
      "model": "gpt-4.1-nano",
      "cost_usd": 0.000004,
      "latency_ms": 1240,
      "tokens_output": 28
    },
    {
      "provider": "grok",
      "model": "grok-4.5",
      "cost_usd": 0.000180,
      "latency_ms": 2100,
      "tokens_output": 31
    }
  ]
}

Example 4 – Comparison across more than one model per provider

The models field takes one model per provider. To compare multiple OpenAI models against each other, use providers as keys and pick the model you want per slot. If you need more than one model from the same provider in a single run, the current approach is to run separate comparisons and compare results manually.

For ongoing multi-model tracking within a provider, create separate benchmark suites – one per model – and compare their run histories in the Analytics tab.


Example 5 – Retrieve a past comparison

What you say to your agent:

“Get comparison cmp_abc123 and summarise the results.”

Tool call – get_comparison:

{
  "comparison_id": "cmp_abc123"
}

Returns the full comparison record including all model responses, cost, and latency.


Example 6 – Check benchmark run history

What you say to your agent:

“List the last 5 runs for benchmark suite_abc123 and tell me which model had the lowest cost on each run.”

Tool call – list_benchmark_runs:

{
  "suite_id": "suite_abc123",
  "limit": 5
}

Returns per-model metrics for each run. The agent can then summarise cost trends across runs.


Prompting tips for agents

  • Always tell the agent to poll until complete. Without this instruction some agents call run_comparison and stop without polling.
  • Say “poll every 2 seconds” – this matches the server’s intended polling rate and avoids rate limit waste.
  • Say “use model IDs from list_models” if you are unsure of exact IDs. The agent will call list_models first and pick the right IDs for your tier.
  • The models field is a map, not a list. If your agent tries to pass a list, correct it: “models should be a JSON object mapping provider name to model ID.”

What’s next