Skip to content

MCP Reference

MCP Tools Reference

Every LLM Prover MCP tool, its input fields, and what it returns.

How to read this reference

Each tool entry shows:

  • The tool name as it appears in your MCP client
  • Every input field with its type, whether it is required, and what it does
  • What the tool returns

Fields marked required must be provided. All others are optional.

The models field appears on several tools. It is always a map of provider name to model ID – e.g. {"openai": "gpt-4.1-nano", "grok": "grok-4.5"}. Use list_models to get current model IDs for your tier.


list_models

List all models available to your tier, grouped by provider.

No input fields required.

Returns: providers map, each with a list of models. Each model has id, display_name, min_tier, cost_per_1k_input, cost_per_1k_output, capabilities, and status. Use the id values in the models field of other tools.


run_comparison

Fire a prompt at multiple models in parallel. Returns a job_id – poll with poll_job.

FieldTypeRequiredDescription
promptstringyesThe prompt to send to all models
modelsobjectyesMap of provider to model ID. e.g. {"openai": "gpt-4.1-nano", "grok": "grok-4.5"}
system_prompt_idstringnoID of a saved system prompt from your dashboard
store_idstringnoRAG store ID – relevant chunks injected at run time (Pro+)
paramsobjectnoInference parameters – see below

params fields:

FieldTypeDescription
temperaturenumber0.0-2.0. Controls randomness. Default varies by model.
max_tokensintegerMaximum output tokens
seedintegerSet for reproducible results
top_pnumberNucleus sampling threshold

Returns: {"job_id": "...", "status": "pending"}. Pass job_id to poll_job.


poll_job

Poll an async job until complete or failed. Call every 2 seconds. Stop after 5 minutes.

FieldTypeRequiredDescription
job_idstringyesThe job_id returned by run_comparison, run_evaluation, or trigger_benchmark_run

Returns: {"job_id": "...", "status": "...", "result": {...}}. When status is complete, the full result is in result – no further call needed. When status is failed, the error is in error.


get_comparison

Retrieve a previously completed comparison by ID.

FieldTypeRequiredDescription
comparison_idstringyesThe comparison_id from a completed comparison result

Returns: Full comparison record including all model responses, cost, latency, and scores if scoring was attached.


list_comparisons

List your recent comparisons, newest first.

FieldTypeRequiredDescription
limitintegernoNumber of results to return. Default 20.

Returns: Array of comparison summaries (no full responses). Use get_comparison to retrieve a full result.


run_evaluation

Score a prompt response against a gold standard answer or a rubric. Returns a job_id – poll with poll_job.

FieldTypeRequiredDescription
promptstringyesThe prompt to evaluate
modelsobjectyesMap of provider to model ID
expected_outputstringno*Gold standard answer. Required if no rubric_id.
rubric_idstringno*Rubric ID from your dashboard. Required if no expected_output. Pro and above.
store_idstringnoRAG store ID for context (Pro+)
paramsobjectnoInference parameters – same fields as run_comparison

*At least one of expected_output or rubric_id is required.

Returns: {"job_id": "...", "status": "pending"}. Pass job_id to poll_job.


list_evaluations

List your recent evaluations, newest first.

FieldTypeRequiredDescription
limitintegernoNumber of results to return. Default 20.

Returns: Array of evaluation summaries.


create_benchmark

Create a named benchmark suite. The suite is created immediately – no async job.

FieldTypeRequiredDescription
namestringyesDisplay name for the suite
promptstringyesThe prompt to run on every execution
modelsobjectyesMap of provider to model ID
expected_outputstringnoGold standard for scoring
rubric_idstringnoRubric ID for scoring (Pro+)
store_idstringnoRAG store ID for context (Pro+)
schedulestringnomanual, daily, 4x_daily, hourly. Default manual. Schedule availability is tier-gated.
thresholdsobjectnoPass/fail thresholds – see below

thresholds fields:

{
  "quality_score": {"min": 80},
  "cost_usd": {"max": 0.05},
  "latency_ms": {"max": 5000}
}

Returns: Suite record including suite_id. Use suite_id with trigger_benchmark_run.


list_benchmarks

List all your benchmark suites.

No required input fields.

Returns: Array of suite summaries including suite_id, name, schedule, and last run date.


get_benchmark

Get the full config for a benchmark suite.

FieldTypeRequiredDescription
suite_idstringyesThe suite ID

Returns: Full suite config including prompt, models, scoring setup, thresholds, and schedule.


update_benchmark

Update models or thresholds on an existing suite. Changes take effect on the next run.

FieldTypeRequiredDescription
suite_idstringyesThe suite ID
modelsobjectnoNew provider-to-model map. Replaces the existing map.
thresholdsobjectnoNew thresholds. Replaces existing thresholds.

Cannot update prompt, expected_output, rubric_id, or schedule – delete and recreate to change those.


delete_benchmark

Delete a benchmark suite. Historical run records are not deleted.

FieldTypeRequiredDescription
suite_idstringyesThe suite ID

trigger_benchmark_run

Trigger an immediate run of a saved benchmark. Returns a job_id – poll with poll_job.

FieldTypeRequiredDescription
suite_idstringyesThe suite ID to run

Returns: {"job_id": "...", "status": "pending"}. Pass job_id to poll_job.


list_benchmark_runs

List historical runs for a benchmark, newest first.

FieldTypeRequiredDescription
suite_idstringyesThe suite ID
limitintegernoNumber of runs to return. Default 50, max 200.

Returns: Array of run summaries with per-model metrics (quality score, cost, latency, pass/fail).


get_benchmark_run

Get the full run record including individual model responses and scores.

FieldTypeRequiredDescription
run_idstringyesThe run ID

Returns: Full run record including per-model responses, scores, and metrics.


delete_benchmark_run

Delete a run record. The benchmark suite is not affected.

FieldTypeRequiredDescription
run_idstringyesThe run ID

list_rag_stores

List your RAG vector stores with file count, chunk count, and sync status.

No required input fields.

Returns: Array of stores. sync_status: ready means the store is queryable. A store must be ready before it can be used in a comparison, evaluation, or benchmark.


create_rag_store

Create a named vector store. The store is empty until you upload files to it.

FieldTypeRequiredDescription
namestringyesDisplay name for the store
store_typestringnogeneral (default) or compliance (Pro+)

Returns: Store record including store_id. Use store_id in comparisons, evaluations, and benchmarks.


get_rag_store

Get a store’s config and sync status.

FieldTypeRequiredDescription
store_idstringyesThe store ID

delete_rag_store

Permanently delete a store, all its files, and all indexed vectors. Cannot be undone.

FieldTypeRequiredDescription
store_idstringyesThe store ID
confirmed_namestringnoRequired if the store is used by active benchmarks. Must match the store name exactly.

list_files

List your uploaded files with ingest status.

No required input fields.

Returns: Array of files. ingest_status values: pending, processing, ready, failed.


delete_file

Delete a file and remove its vectors from the RAG store.

FieldTypeRequiredDescription
file_idstringyesThe file ID

What’s next