Documentation

One key, two formats. OpenAI base URL https://helioroute.com/v1 · Anthropic base URL https://helioroute.com. Change two values in your existing SDK and you are done.

Quickstart

Create an account, copy a key from the console, and make your first request in under a minute.

shell
# 1. Get a key: https://helioroute.com/console/keys
# 2. Call any model (OpenAI-compatible)
curl https://helioroute.com/v1/chat/completions \
  -H "Authorization: Bearer $HELIOROUTE_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"gpt-4o-mini","messages":[{"role":"user","content":"ping"}]}'

Authentication

Every request carries your virtual key. Both surfaces accept two header styles, so any client just works:

  • Authorization: Bearer sk-… — OpenAI SDKs and Anthropic-style clients
  • x-api-key: sk-… — the Anthropic SDK and Claude Code

Keys are metered per token and spend is deducted from your prepaid credit in real time. Rotate or revoke any time from /console/keys.

OpenAI format

Drop-in for the official SDKs at https://helioroute.com/v1. Only base_url and the key change.

python
from openai import OpenAI

client = OpenAI(base_url="https://helioroute.com/v1", api_key="sk-hr-...")

resp = client.chat.completions.create(
    model="claude-sonnet-5",
    messages=[{"role": "user", "content": "Explain vector databases in 2 lines."}],
)
print(resp.choices[0].message.content)
node / typescript
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://helioroute.com/v1",
  apiKey: process.env.HELIOROUTE_KEY,
});

const res = await client.chat.completions.create({
  model: "claude-sonnet-5",
  messages: [{ role: "user", content: "hello" }],
});
console.log(res.choices[0].message.content);

Anthropic format

The Messages API at https://helioroute.com/v1/messages is served natively by HELIOROUTE — Claude SDK clients, Claude Code and Anthropic-shaped tools talk to it unchanged.

curl
curl https://helioroute.com/v1/messages \
  -H "x-api-key: $HELIOROUTE_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{"model":"claude-sonnet-5","max_tokens":256,"messages":[{"role":"user","content":"hi"}]}'
python (anthropic sdk)
import anthropic

client = anthropic.Anthropic(
    api_key="sk-hr-...",
    base_url="https://helioroute.com",
)

msg = client.messages.create(
    model="claude-sonnet-5",
    max_tokens=512,
    messages=[{"role": "user", "content": "hi"}],
)
print(msg.content[0].text)

Claude Code

Point Claude Code at HELIOROUTE and run it against any model on the gateway:

shell
export ANTHROPIC_BASE_URL="https://helioroute.com"
export ANTHROPIC_AUTH_TOKEN="sk-hr-..."
claude

Cursor

Use the OpenAI-compatible provider with a custom base URL:

cursor settings
# Cursor → Settings → Models → OpenAI API Key
# Override OpenAI Base URL:  https://helioroute.com/v1
# API Key:                   sk-hr-...
# Add a custom model, e.g.    claude-sonnet-5

LangChain

Any OpenAI-compatible LangChain integration works unchanged:

python
from langchain_openai import ChatOpenAI

llm = ChatOpenAI(
    model="gpt-4o",
    base_url="https://helioroute.com/v1",
    api_key="sk-hr-...",
)
print(llm.invoke("Summarise transformers in one sentence.").content)

Streaming

Server-sent events on both surfaces, identical to the native formats:

python
stream = client.chat.completions.create(
    model="claude-sonnet-5",
    messages=[{"role": "user", "content": "Write a haiku about latency."}],
    stream=True,
)
for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="")

Tool calling

Function/tool calling is supported for all tool-capable models:

python
tools = [{
  "type": "function",
  "function": {
    "name": "get_weather",
    "description": "Get current weather for a city",
    "parameters": {"type": "object", "properties": {"city": {"type": "string"}}, "required": ["city"]},
  },
}]

resp = client.chat.completions.create(
    model="gpt-4o",
    messages=[{"role": "user", "content": "weather in Mumbai?"}],
    tools=tools,
)

Vision

Send images as part of message content for multimodal models:

python
resp = client.chat.completions.create(
    model="gemini-3-pro",
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "What is in this image?"},
            {"type": "image_url", "image_url": {"url": "https://example.com/chart.png"}},
        ],
    }],
)

Image & audio

Image generation uses the standard OpenAI shape:

python
# Text-to-image (OpenAI-compatible)
client.images.generate(model="dall-e-3", prompt="isometric server room, neon")

# Speech-to-text
client.audio.transcriptions.create(model="whisper-1", file=open("clip.mp3", "rb"))

Errors

Errors follow the OpenAI envelope. Common codes:

  • 401 — invalid or revoked key
  • 402 — credit exhausted; top up in the console
  • 429 — rate limited; retry with backoff
  • 5xx — temporary server error; retries are automatic
json
{
  "error": {
    "message": "Budget has been exceeded",
    "type": "budget_exceeded",
    "code": "402"
  }
}

Rate limits

Limits scale with the credit pack you buy — from 30 requests/minute on Starter up to 1000 on Reseller. Limits are enforced per account.

Use exponential backoff with jitter on 429 and 5xx.

Docs · HELIOROUTE