Claude Opus 5.5 is Anthropic's newest agentic-coding and long-context reasoning model, and hiapi exposes it the same day it lands on the platform. Unlike hiapi's image, video, and audio models — which all go through the async POST /v1/tasks → poll pattern — Claude Opus 5.5 is a text-only chat model, so you call it with a single, synchronous request to /v1/chat/completions. This guide gives you a working curl call, the same call in Python, and the production details (reasoning effort, streaming, error handling) you need before you ship it.
What you need
- A hiapi account and an API key. Create one from API Keys — every hiapi model, including Claude Opus 5.5, shares the same key.
- No SDK install required. Claude Opus 5.5 uses hiapi's OpenAI-compatible Chat Completions format, so any HTTP client (curl,
requests,fetch) works out of the box.
The minimal working request (curl)
curl -X POST "https://api.hiapi.ai/v1/chat/completions" \
-H "Authorization: Bearer sk-YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-opus-5-5",
"messages": [
{ "role": "user", "content": "Give me three tips for reviewing a Python pull request." }
],
"stream": false
}'
The model id is passed bare — no route suffix, no provider prefix. A successful call returns a standard chat-completion object:
{
"id": "msg_...",
"object": "chat.completion",
"model": "claude-opus-5-5",
"choices": [
{
"index": 0,
"message": { "role": "assistant", "content": "1. ..." },
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 13,
"completion_tokens": 42,
"total_tokens": 55,
"usage_semantic": "openai",
"usage_source": "anthropic"
}
}
Read the answer from choices[0].message.content — there's no output[0].url to poll for, and no task id to track.
The same flow in Python
import requests
response = requests.post(
"https://api.hiapi.ai/v1/chat/completions",
headers={
"Authorization": "Bearer sk-YOUR_API_KEY",
"Content-Type": "application/json",
},
json={
"model": "claude-opus-5-5",
"messages": [
{"role": "user", "content": "Give me three tips for reviewing a Python pull request."}
],
"stream": False,
},
)
response.raise_for_status()
data = response.json()
print(data["choices"][0]["message"]["content"])
Because the endpoint speaks the OpenAI Chat Completions wire format, this same payload shape also works unmodified through any HTTP client in any language — there's nothing hiapi-specific to learn beyond the base URL and the bearer key.
Reasoning effort, thinking, and streaming
Claude Opus 5.5 keeps extended thinking always on — you can't disable it or hand it a manual token budget. What you control is how much of it the model does, via output_config.effort:
{
"model": "claude-opus-5-5",
"messages": [{ "role": "user", "content": "..." }],
"output_config": { "effort": "medium" }
}
effort accepts low, medium (the default if you omit the field), high, xhigh, or max — going from fast/cheap answers to deep, deliberate ones. The response's message includes a reasoning_details array holding the model's internal thinking trace; for a simple one-shot call you can safely ignore it and just read message.content.
For long answers, set "stream": true to get incremental output instead of waiting for the full response:
curl -N https://api.hiapi.ai/v1/chat/completions \
-H "Authorization: Bearer sk-YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"claude-opus-5-5","messages":[{"role":"user","content":"..."}],"stream":true}'
Each chunk is a data: {...} line shaped like chat.completion.chunk, with the incremental text in choices[0].delta.content. Keep reading chunks until one arrives with a non-null finish_reason (normally "stop") — that's the end of the response.
Claude Opus 5.5 also supports tool/function calling and has a 1M-token context window with up to 128K tokens of output, so it's built for large-codebase or long-document tasks, not just short Q&A.
Production notes
-
Auth errors are 401, not 400. A missing or invalid key returns:
{ "error": { "code": "permission_denied", "message": "This API key is invalid...", "type": "hiapi_error", "request_id": "..." } }Check for HTTP 401 specifically and surface the
request_idif you need to contact support — see the full invalid API key troubleshooting guide for other causes of this error. -
There's no callback/webhook option here. That pattern exists for hiapi's async task-based models (image, video, audio); Chat Completions is request/response only, so
stream: trueis your only alternative to a single blocking call. -
This call is idempotent by nature — there's no task id to accidentally double-create, so retries on timeout are safe as long as you're comfortable re-running the same prompt.
-
Watch your rate limits on high-concurrency workloads; see how to handle a 429 rate-limit error for backoff strategy.
-
Token pricing (input, output, and cached tokens are billed separately) is kept current on the Claude Opus 5.5 model page and the pricing page — don't hardcode numbers in your billing logic, read them from there.
Related reading
- Claude Opus 5.5 model page — full spec sheet and live pricing
- Claude Sonnet 4.6 API guide — a lighter, faster Claude option on the same Chat Completions format
- Building an OpenAI-compatible proxy in Python — useful if you're fronting multiple chat models behind one internal API
- Manage your API keys
FAQ
Does the Claude Opus 5.5 API support streaming?
Yes — set "stream": true and read the data: {...} chunks as described above.
Do I need an API key to call Claude Opus 5.5 in production?
Yes. Every request must carry Authorization: Bearer sk-<your key>; there's no keyless or anonymous access to the production API.
Can I turn off Claude Opus 5.5's thinking to save tokens?
No. Thinking is always on for this model; the only control you have is the output_config.effort level (low through max), which trades speed/cost for depth.
Is Claude Opus 5.5 called through /v1/tasks like hiapi's image models?
No. Only hiapi's image, video, and audio models use the async task-and-poll pattern. Claude Opus 5.5 is a text model on /v1/chat/completions and returns its answer in the same response.
What's the context window and max output length? 1,000,000 tokens of context and up to 128,000 tokens of output — large enough for whole-repo or long-document prompts.
Does Claude Opus 5.5 support tool/function calling? Yes, alongside streaming and the OpenAI-compatible message format.









