Model catalog

Browse every model Seldon exposes on the public router. Prices are USD per 1,000,000 tokens.

anthropic/claude-3-5-sonnet

Claude 3.5 Sonnet

Coming soon
$0.003 / $0.015Guaranteed 30%+ savings

anthropic/claude-haiku-4-5

Claude Haiku 4.5

Anthropic

Routable

Claude Haiku 4.5 is Anthropic's fast, cost-efficient Claude model for high-volume chat and tool-using workloads via AWS Bedrock.

$1.2 / $6Guaranteed 30%+ savingschatcodefast

anthropic/claude-sonnet-4-6

Claude Sonnet 4.6

Anthropic

Routable

Claude Sonnet 4.6 is Anthropic's balanced Claude model for coding, reasoning, and agentic workloads via AWS Bedrock Global CRIS.

$3.6 / $18Guaranteed 30%+ savingschatcodereasoning

deepseek/deepseek-v3.2

DeepSeek V3.2

DeepSeek

Routable

DeepSeek V3.2 is a large MoE chat model tuned for strong reasoning and coding at competitive cost, served via Azure AI Foundry.

$0.58 / $1.68Guaranteed 30%+ savingschatcodereasoningazure-foundry

microsoft/phi-4

Phi-4

Microsoft

Routable

Phi-4 is a compact open-weight model from Microsoft Research, efficient for text tasks. Note: latency can be high under load on shared Azure Foundry capacity.

$0.13 / $0.5Guaranteed 30%+ savingschatopen-weightazure-foundry

microsoft/phi-4-mini-instruct

Phi-4 Mini Instruct

Microsoft

Routable

Phi-4 Mini Instruct is a lightweight open-weight Microsoft model optimized for efficient instruction following and chat on Azure AI Foundry.

$0.075 / $0.3Guaranteed 30%+ savingschatopen-weightfastazure-foundry

openai/gpt-4.1

GPT-4.1

OpenAI

Routable

GPT-4.1 is OpenAI's flagship multimodal model with a very large context window, strong instruction following, and vision support.

$2 / $8Guaranteed 30%+ savingschatvisioncodeazure-foundry

openai/gpt-4.1-mini

GPT-4.1 Mini

OpenAI

Routable

GPT-4.1 Mini is a lower-cost GPT-4.1-series chat model with a large context window, strong instruction following, and structured-output support via Azure AI Foundry.

$0.4 / $1.6Guaranteed 30%+ savingschatvisioncodefast

openai/gpt-5.4

GPT-5.4

OpenAI

Routable

GPT-5.4 is OpenAI's latest flagship chat model with an extended context window and multimodal capabilities via Azure OpenAI.

$2.5 / $15Guaranteed 30%+ savingschatvisioncode

openai/gpt-5.4-nano

GPT-5.4 Nano

OpenAI

Routable

GPT-5.4 Nano is OpenAI's lightweight GPT-5.4 variant for low-latency, cost-efficient classification, extraction, and high-volume chat via Azure OpenAI.

$0.2 / $1.25Guaranteed 30%+ savingschatvisioncodefast

openai/gpt-5.5

GPT-5.5

OpenAI

Routable

GPT-5.5 is OpenAI's newer flagship chat model with extended context, reasoning, and multimodal capabilities via Azure OpenAI.

$5 / $30Guaranteed 30%+ savingschatvisioncodereasoning

openai/gpt-5.6-luna

GPT-5.6 Luna

OpenAI

Routable

GPT-5.6 Luna is the fastest and lowest-cost GPT-5.6 variant for high-volume and cost-sensitive routing via Azure OpenAI.

$1 / $6Guaranteed 30%+ savingschatvisionfast

openai/gpt-5.6-sol

GPT-5.6 Sol

OpenAI

Routable

GPT-5.6 Sol is OpenAI's flagship GPT-5.6 variant for complex professional reasoning, coding, and agentic workloads via Azure OpenAI.

$5 / $30Guaranteed 30%+ savingschatvisioncodereasoning

openai/gpt-5.6-terra

GPT-5.6 Terra

OpenAI

Routable

GPT-5.6 Terra balances GPT-5.6 intelligence, speed, and cost for general production workloads via Azure OpenAI.

$2.5 / $15Guaranteed 30%+ savingschatvisioncodereasoning

openai/o1-preview

o1 preview

Coming soon
/ Guaranteed 30%+ savings

openai/o3

o3

OpenAI

Routable

o3 is OpenAI's reasoning model with multimodal input. Uses ``max_completion_tokens`` instead of ``max_tokens`` on the chat API.

$2 / $8Guaranteed 30%+ savingschatvisionreasoningazure-foundry

xai/grok-4

Grok 4

xAI

Routable

Grok 4 is xAI's frontier model with multimodal input, long context, and strong general reasoning served through Azure AI Foundry.

$3 / $15Guaranteed 30%+ savingschatvisionreasoningazure-foundry

xai/grok-4-1-fast-reasoning

Grok 4.1 Fast Reasoning

xAI

Routable

Grok 4.1 Fast Reasoning is a latency-optimized variant of Grok 4 prioritizing speed while retaining strong chain-of-thought performance.

$0.2 / $0.5Guaranteed 30%+ savingschatvisionreasoningfast