Back to Models
Inception

Mercury 2.5 Preview

Available
80% offA 80% discount applies to input, output, and cached token pricing.

Description

Mercury 2.5 Preview is a fast diffusion-based reasoning model, delivering 1,000+ tokens/sec with a 260K context window. Best suited for coding, agents, and latency-sensitive workloads.

At a Glance

Key pricing and model details available for this model.

Input price

$0.20$0.04

per 1M tokens

Output price

$0.75$0.15

per 1M tokens

Context window

260K

tokens

Hallucination rate

0%

Token Pricing

Token pricing normalized to per-million-token rates.

Input / 1M tokens

$0.20$0.04

Output / 1M tokens

$0.75$0.15

Cache Read / 1M tokens

$0.02$0.0040

Token Pricing Details

Rates are shown per 1M tokens for easier comparison.

Input / 1M tokens$0.04
Input unit1M tokens
Output / 1M tokens$0.15
Output unit1M tokens
Cache Read / 1M tokens$0.02$0.0040
Cache Read unit1M tokens

Feature Availability

Capabilities explicitly listed in the current payload.

LLM

Available

Yes

Vision

Not listed

No

Function calling

Available

Yes

Reasoning

Available

Yes

Service tiers

Not listed

No

Reasoning Effort Levels

low
medium
high
xhigh

Supported Parameters

frequency_penalty
logit_bias
max_completion_tokens
presence_penalty
stop
temperature
tool_choice
tools
top_p

Code Samples

Quick start with the Routeway API

import OpenAI from 'openai';

const openai = new OpenAI({
  baseURL: "https://api.routeway.ai/v1",
  apiKey: "<YOUR_API_KEY>",
});

async function main() {
  const completion = await openai.chat.completions.create({
    model: "mercury-2.5-preview",
    messages: [
      {
        role: "user",
        content: "Explain quantum computing in simple terms"
      }
    ]
  });

  console.log(completion.choices[0].message);
}

main();