Skip to content

API Reference

API Quickstart

Make your first comparison via the LLM Prover API in under 5 minutes.

Before you start

You need an API key. Go to the Developer section in your dashboard, open the Keys tab, and create a key. Copy it – it is shown once only.

Your API base URL is https://api.llmprover.pysolvr.com. All requests require an Authorization header with your key.


Step 1 – Check which models are available

curl https://api.llmprover.pysolvr.com/models \
  -H "Authorization: YOUR_API_KEY"

The response lists every model available to your tier, grouped by provider. Note the id values – you will use these in the next step.


Step 2 – Run a comparison

curl -X POST https://api.llmprover.pysolvr.com/compare \
  -H "Authorization: YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "prompt": "Explain the difference between precision and recall in one sentence.",
    "models": {
      "openai": "gpt-4o",
      "anthropic": "claude-3-5-sonnet-20241022"
    }
  }'

The response returns a job_id immediately:

{
  "ok": true,
  "data": {
    "job_id": "a1b2c3d4-...",
    "status": "pending"
  }
}

Step 3 – Poll for the result

Wait 2 seconds, then poll:

curl https://api.llmprover.pysolvr.com/jobs/a1b2c3d4-... \
  -H "Authorization: YOUR_API_KEY"

Keep polling every 2 seconds until status is complete. Most comparisons finish within 10-30 seconds.

When complete, the full result is in the data.result field:

{
  "ok": true,
  "data": {
    "job_id": "a1b2c3d4-...",
    "status": "complete",
    "result": {
      "comparison_id": "...",
      "prompt": "Explain the difference between precision and recall in one sentence.",
      "results": [
        {
          "provider": "openai",
          "model": "gpt-4o",
          "response_text": "Precision measures how many of your positive predictions were correct; recall measures how many of the actual positives you found.",
          "latency_ms": 1240,
          "cost_usd": 0.000125,
          "tokens_input": 18,
          "tokens_output": 32
        },
        {
          "provider": "anthropic",
          "model": "claude-3-5-sonnet-20241022",
          "response_text": "Precision is the fraction of retrieved items that are relevant, while recall is the fraction of relevant items that were retrieved.",
          "latency_ms": 1890,
          "cost_usd": 0.000198,
          "tokens_input": 18,
          "tokens_output": 29
        }
      ],
      "total_cost_usd": 0.000323,
      "total_latency_ms": 1890
    }
  }
}

Python example

import time
import requests

API_KEY = "YOUR_API_KEY"
BASE_URL = "https://api.llmprover.pysolvr.com"
HEADERS = {"Authorization": API_KEY, "Content-Type": "application/json"}

# Submit
resp = requests.post(f"{BASE_URL}/compare", headers=HEADERS, json={
    "prompt": "Explain the difference between precision and recall in one sentence.",
    "models": {"openai": "gpt-4o", "anthropic": "claude-3-5-sonnet-20241022"}
})
job_id = resp.json()["data"]["job_id"]

# Poll
for _ in range(150):  # max 5 minutes
    time.sleep(2)
    poll = requests.get(f"{BASE_URL}/jobs/{job_id}", headers=HEADERS).json()
    if poll["data"]["status"] == "complete":
        print(poll["data"]["result"])
        break
    if poll["data"]["status"] == "failed":
        print("Failed:", poll["data"].get("error"))
        break

What’s next