Skip to content
English

Claude Opus 5.5 API

POST Base URL: https://api.hiapi.ai /v1/chat/completions

This model uses an OpenAI-compatible Chat Completions endpoint. The same HiAPI API key can call enabled models in your account group; image, video, and audio models use /v1/tasks and a different request shape.

Model overview

Model nameClaude Opus 5.5
ProviderAnthropic
Context / max output1M / 128K tokens
EndpointPOST /v1/chat/completions
ThinkingAlways on · default medium effort
PricingLive HiAPI token rates

Anthropic's Claude Opus 5.5 is available through HiAPI Chat Completions for long-running agentic coding and knowledge work, with always-on adaptive thinking controlled by effort.

Production guidance

Request contract
  • Use model=claude-opus-5-5 with POST /v1/chat/completions. Media models use POST /v1/tasks with a different request shape.
  • Do not send thinking.type=disabled or a manual thinking budget. Omit thinking and use output_config.effort.
  • Allowed effort values are low, medium, high, xhigh, and max. Medium is the provider default.
  • Use tool_choice=auto. Forced any and named-tool selection return an upstream error.

Best suited for

Long-running agentic work

Complex coding, knowledge work, and multi-step tool loops where effort can trade cost and depth.

messagesoutput_config.efforttools

Request parameters

model string required

Use this exact public model ID.

example claude-opus-5-5 enum: claude-opus-5-5
messages array required

Ordered system, user, assistant, and tool messages. Preserve returned reasoning_details unchanged in tool continuations.

stream boolean optional

Set true for Server-Sent Events.

default false
output_config.effort enum optional

The only public thinking-depth control. Thinking is always on.

default medium enum: lowmediumhighxhighmax
max_tokens integer optional

Hard limit shared by thinking and visible output.

system content cache_control object optional

Optional five-minute prompt-cache write. One-hour cache writes are not in the initial HiAPI contract.

tools array optional

OpenAI-compatible function definitions.

tool_choice enum optional

Forced any or named-tool choices are rejected by this model.

default auto enum: auto

API examples

Request examples

Streaming

Read content and reasoning deltas until [DONE].

Request body
{
  "model": "claude-opus-5-5",
  "messages": [
    {
      "role": "user",
      "content": "Analyze the trade-offs between a monolith and microservices."
    }
  ],
  "stream": true,
  "output_config": {
    "effort": "medium"
  },
  "max_tokens": 4096
}
Function tool

Let the model choose a function with tool_choice=auto, then replay the tool result with the original reasoning metadata.

Request body
{
  "model": "claude-opus-5-5",
  "messages": [
    {
      "role": "user",
      "content": "Check task demo-123."
    }
  ],
  "stream": false,
  "output_config": {
    "effort": "low"
  },
  "max_tokens": 4096,
  "tools": [
    {
      "type": "function",
      "function": {
        "name": "get_task_status",
        "description": "Look up a task by ID.",
        "parameters": {
          "type": "object",
          "properties": {
            "task_id": {
              "type": "string"
            }
          },
          "required": [
            "task_id"
          ],
          "additionalProperties": false
        }
      }
    }
  ],
  "tool_choice": "auto"
}
Five-minute prompt cache

Add cache_control to a stable system text block and verify writes or hits in usage.

Request body
{
  "model": "claude-opus-5-5",
  "messages": [
    {
      "role": "system",
      "content": [
        {
          "type": "text",
          "text": "Answer as a concise architecture reviewer.",
          "cache_control": {
            "type": "ephemeral",
            "ttl": "5m"
          }
        }
      ]
    },
    {
      "role": "user",
      "content": "Review this migration plan."
    }
  ],
  "stream": false,
  "output_config": {
    "effort": "medium"
  },
  "max_tokens": 4096
}

Response schema

Read visible text from choices[0].message.content and readable thinking from reasoning_content when returned. Preserve reasoning_details unchanged for tool continuations. Usage is the billing source of truth.

{
  "id": "chatcmpl_example",
  "object": "chat.completion",
  "model": "claude-opus-5-5",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "Start with a migration boundary that can be rolled back.",
        "reasoning_content": "First compare deployment and data risks."
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 32,
    "completion_tokens": 48,
    "total_tokens": 80
  }
}
  1. Read choices[0].message.content for visible text.
  2. For streams, accumulate content and reasoning deltas until [DONE].
  3. Read input, output, and cache usage; thinking counts as output.

FAQ

What is Claude Opus 5.5?

It is Anthropic's model for long-running agentic coding and knowledge work, released with a 1M context window and 128K maximum output.

Which endpoint and model ID should I use?

Use POST /v1/chat/completions with model=claude-opus-5-5.

Can I disable thinking?

No. Thinking is always on. Requests that send disabled or a manual budget are rejected; use output_config.effort instead.

Which effort values are supported?

Low, medium, high, xhigh, and max are supported. Medium is the default.

How do tools work?

Declare OpenAI-compatible tools and use tool_choice=auto. Preserve returned reasoning_details and the assistant tool call when continuing with a tool result.

How are tokens billed?

Input, output, five-minute cache writes, and cache reads use the live rates on the model and pricing pages. Thinking tokens count as output. View live pricing.

Next steps