Qwen Flash

qwen/qwen-flash

A speed- and cost-oriented model with switchable thinking mode.

This model is available through the API. It is not currently offered as an agent's reasoning model.

Use through the API
Modalities
Text
In / out per 1M
$0.264 / $2.12
Context
998K

Overview

A speed- and cost-oriented model with switchable thinking mode. It is billed from the same balance as every other model and API here, and the price of each call comes back in the response.

Pricing

The default plan, in USD per million tokens.

MeterPer 1M tokens
Input$0.264
Output$2.12
Cache read$0.264
Cache write$0.264

Your balance and per-call records come from /v1/credits and /v1/generation.

Use it

Point your client at the gateway and pass this model id. Nothing else changes.

import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://gateway.agentsky.dev/v1",
  apiKey: process.env.AGENTSKY_API_KEY, // ast_…
});

const completion = await client.chat.completions.create({
  model: "qwen/qwen-flash",
  messages: [{ role: "user", content: "Summarise this ticket." }],
});
console.log(completion.usage.cost); // USD, in the response

Questions