Prerequisites
Before you start, make sure you have:
- A TokenRoc account. You can register in the console if you do not have one.
- Python 3 and the ability to install packages with
pip. - A terminal where you can set environment variables.
- A server-side place to run the code. Never put an API key in browser or mobile client code, where anyone can read it.
Step 1: Create an API key
Sign in to the console and open the API Keys page. Create a token there and copy it immediately into a secret store or an environment variable.
Create a separate key for each environment, such as local development, staging, and production. Separate keys can be revoked independently, so a leaked development key does not force you to rotate production.
Step 2: Point your client at the TokenRoc base URL
All API requests go to the TokenRoc gateway:
https://api.tokenroc.com/v1
Because the API is OpenAI-compatible, most existing clients need only this base URL and your TokenRoc key. The request and response shapes stay the same. The endpoints used in this guide are:
| Method | Endpoint | Purpose |
|---|---|---|
| GET | /v1/models | List models available to the current API key. |
| POST | /v1/chat/completions | Create a standard or streaming chat completion. |
A /v1/responses endpoint is also available where the model supports the Responses-compatible request format. The API quickstart lists the current endpoint reference.
Step 3: Choose a current model identifier
A request must name the model you want to run. Model availability changes over time, so do not copy an identifier out of a tutorial and hard-code it, including this one.
Instead, take the exact identifier from a live source:
- Open the live model catalog in the console and copy the identifier shown there, along with its current input and output rates.
- Or call
GET /v1/modelswith your key to list what that key can reach.
curl https://api.tokenroc.com/v1/models \
-H "Authorization: Bearer $TOKENROC_API_KEY"
The rest of this guide reads the identifier from an environment variable called TOKENROC_MODEL, so you can change models without editing code.
Step 4: Set your environment variables
Keep both the key and the model identifier out of your source files. Set them in your shell, your process manager, or your deployment secret store.
export TOKENROC_API_KEY="your-api-key"
export TOKENROC_MODEL="paste-an-identifier-from-the-catalog"
$env:TOKENROC_API_KEY = "your-api-key"
$env:TOKENROC_MODEL = "paste-an-identifier-from-the-catalog"
Step 5: Send the request from Python
Install the OpenAI Python SDK, which speaks the same protocol TokenRoc exposes:
pip install openai
Then create the client with your TokenRoc key and base URL, and send a chat completion. Only those two arguments are TokenRoc-specific; everything else is the standard OpenAI-compatible call.
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["TOKENROC_API_KEY"],
base_url="https://api.tokenroc.com/v1",
)
response = client.chat.completions.create(
model=os.environ["TOKENROC_MODEL"],
messages=[
{"role": "system", "content": "You are a concise assistant."},
{"role": "user", "content": "Say hello in one sentence."},
],
)
print(response.choices[0].message.content)
If you prefer to see the raw HTTP call, the same request with cURL looks like this:
curl https://api.tokenroc.com/v1/chat/completions \
-H "Authorization: Bearer $TOKENROC_API_KEY" \
-H "Content-Type: application/json" \
-d "{
\"model\": \"$TOKENROC_MODEL\",
\"messages\": [
{\"role\": \"user\", \"content\": \"Say hello in one sentence.\"}
]
}"
Authentication is a bearer token in the Authorization header, and the body is JSON, so the request needs Content-Type: application/json.
Step 6: Read the response
A successful chat completion returns a JSON object with the generated message and a token accounting summary:
{
"id": "...",
"object": "chat.completion",
"created": 0,
"model": "...",
"choices": [
{
"index": 0,
"message": {"role": "assistant", "content": "..."},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 0,
"completion_tokens": 0,
"total_tokens": 0
}
}
The fields you will use most often are:
choices[0].message.content— the generated text.choices[0].finish_reason— why generation stopped. Check this before treating output as complete.usage.prompt_tokens,usage.completion_tokens, andusage.total_tokens— the token counts for that call.
In Python, read them from the response object directly:
print(response.choices[0].message.content)
print(response.choices[0].finish_reason)
print(response.usage.total_tokens)
Step 7: Review usage
Every call reports its own token counts in usage, which is the fastest way to measure a single request while you are developing.
For the account-level view, use the console:
- Usage logs record individual requests, so you can confirm a call arrived and see what it consumed.
- The dashboard summarises activity across your account.
- The model catalog shows the current input and output rates for each model, which you need to turn token counts into cost.
Check usage logs early. It is much easier to notice a retry loop or an oversized prompt on the first day than after it has run for a week.
Common errors
Most first-request failures fall into a small number of HTTP responses:
| Status | Meaning | What to check |
|---|---|---|
| 401 | Authentication failed | Confirm the bearer token is present, active, and copied without extra spaces or a trailing newline. |
| 403 | Access or balance rejected | Check wallet balance, key permissions, and whether the key may use that model. |
| 429 | Rate or quota limit | Reduce concurrency, apply backoff, or review token limits. |
| 5xx | Gateway or upstream failure | Retry with exponential backoff and check the status page. |
GET /v1/models, rather than guessing a name. Identifiers are exact strings, not display names.If a request fails before it reaches the API at all, confirm that the base URL is https://api.tokenroc.com/v1 and that your SDK is not still pointing at another provider's default endpoint.
Next steps
Once a single request works end to end:
- Read the API quickstart for the full endpoint reference and cURL examples.
- Compare current rates in the model catalog before you commit to a model in production.
- Create environment-specific keys on the API Keys page and revoke anything unused.
- Add retries with exponential backoff for 429 and 5xx responses, and log
finish_reasonso truncated responses are visible. - Check the status page when behaviour changes unexpectedly.
Have a question about an integration? Contact TokenRoc and include the endpoint involved and any non-sensitive error details.
