Skip to content
Pearl Compute

Docs

Build with the Pearl Compute API

The API speaks the OpenAI Chat Completions format. If your code already uses an OpenAI SDK, change the base URL and the key; nothing else.

Quickstart

  1. Sign in with your email and add credit in the console.
  2. Create an API key under API keys and store it in the PEARL_COMPUTE_API_KEY environment variable.
  3. Point the SDK at the base URL below.
Base URL
https://api.ai.pearlsafe.xyz/v1

Send the key as a Bearer token in the Authorization header. Endpoints: POST /v1/chat/completions and POST /v1/completions.

# pip install openai
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.ai.pearlsafe.xyz/v1",
    api_key=os.environ["PEARL_COMPUTE_API_KEY"],
)

reply = client.chat.completions.create(
    model="llama-3.1-8b-instruct",
    messages=[{"role": "user", "content": "Say hello in five words."}],
)
print(reply.choices[0].message.content)

Streaming

Set stream: true to receive server-sent events as tokens are generated, in the same format as OpenAI. Add stream_options.include_usage to get token counts in the last chunk. The stream ends with data: [DONE]. Non-streaming responses always include usage.

stream = client.chat.completions.create(
    model="llama-3.1-8b-instruct",
    messages=[{"role": "user", "content": "Write a haiku about GPUs."}],
    stream=True,
    stream_options={"include_usage": True},
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)
    if chunk.usage:
        print("\n", chunk.usage)

Models and prices

Use the model id in the model field. Prices are Standard tier prices in US dollars per million tokens, excluding VAT, and follow the market: each model is priced 15% below the cheapest comparable provider on OpenRouter, and never below what we pay the miner plus our margin.

Models, availability and Standard prices per million tokens, excluding VAT
Model idContextStatusInputOutput
llama-3.1-8b-instruct
Llama 3.1 8B Instruct · 8B
32,768 tokens
max output 8,192
No node online$0.017$0.034
gemma-4-31b-it
Gemma 4 31B IT · 31B
32,768 tokens
max output 8,192
In validation$0.085$0.2805

A dash means the price feed hasn't published a price for that model yet. Every model's licence is on the licences page.

Privacy tiers

Each API key has a privacy tier, chosen when you create it.

Standard

We do not log or store prompts or completions, and we never train on them. We keep only request metadata (counts, times, costs) for billing. The same GPU work also mines PRL: our pool briefly checks small samples of the model's internal numbers. The privacy policy explains what that means. The prices above are Standard prices.

Private

The same zero retention, never used for mining, and only on vetted nodes. Priced from GPU cost — available once measured. No model is measured yet.

Confidential Coming soon

Requests inside GPUs with hardware isolation (confidential computing). Coming soon.

Rate limits and credits

  • Each key allows 60 requests per minute and 200,000 tokens per minute by default. The console shows each key's limits. Need more? Email us.
  • Over the limit you get 429 with a Retry-After header. Wait that many seconds, then retry.
  • Credits are prepaid. When a request starts we hold its worst case (your max_tokens at the model's price) and charge only what it used when it ends. If your available credit can't cover the hold, the request is refused with 402. Setting max_tokens keeps holds small.

Errors

Errors use the OpenAI shape: { "error": { "message": "…", "type": "…", "code": "…" } }. The OpenAI SDKs raise them as exceptions with the status code.

HTTP status codes
StatusMeaning and what to do
400The request is malformed (unknown field value, too many tokens for the model). Fix it; retrying won’t help.
401Missing, unknown or revoked API key.
402Not enough credit for this request. Add credit in the console.
404Unknown model id. See the models table above.
429Rate limit reached. Retry after the Retry-After header.
503No node can take the request right now. Retry with backoff (Retry-After when present).
504The request took longer than the deadline. Retry, or lower max_tokens.
500Something failed on our side. Retry with backoff.