Qwen Flash

by Qwen

qwen/qwen-flash

A speed- and cost-oriented model with switchable thinking mode.

Modalities
Text
In / out per 1M
$0.264 / $2.12
Context
998K

Overview

A speed- and cost-oriented model with switchable thinking mode. It is billed from the same balance as every other model and API here, and the price of each call comes back in the response.

Pricing

The default plan, in USD per million tokens.

MeterPer 1M tokens
Input$0.264
Output$2.12
Cache read$0.264
Cache write$0.264

Your balance and per-call records come from /v1/credits and /v1/generation.

Use it

Point your client at the gateway and pass this model id. Nothing else changes.

import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://gateway.agentsky.dev/v1",
  apiKey: process.env.AGENTSKY_API_KEY, // ast_…
});

const completion = await client.chat.completions.create({
  model: "qwen/qwen-flash",
  messages: [{ role: "user", content: "Summarise this ticket." }],
});
console.log(completion.usage.cost); // USD, in the response

Questions

Qwen Flash — AgentSky