Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

Start free brief

Qwen3.7 Flash model card

CurrentReleased Jul 27, 2026ProprietaryReasoning1M context

Released Jul 27, 2026 see all recent releases

Decision reading
Qwen3.7 Flash is Alibaba’s proprietary multimodal reasoning model for low-cost agents and general workloads. It accepts text, image, and video input, returns text, supports 1M context and 64K output, and starts at $0.03 per million input tokens and $0.13 per million output tokens. No exact public benchmark table is attached yet, so it remains unranked.

Data as of August 22, 2026 · How the score is built

Spec sheet

Each documented value carries its source. Missing fields stay visible as not sourced or not published, rather than disappearing from the page.

API model ID
Not published
Context window
1M
Maximum output
Not sourced yet
Knowledge cutoff
Not sourced yet
Input modalities
Not sourced yet
Output modalities
Not sourced yet
Parameters
Not disclosed by the provider
Availability
Available from QwenCloud as qwen3.7-flash and through OpenRouter as qwen/qwen3.7-flash. The APIs expose switchable thinking, function calling, built-in tools, structured output, and prompt caching.
Cloud regions
Not tracked yet
Lifecycle
Current
API capabilities
Tool calling, structured outputs, and batch support are not tracked yet
Prompt caching
Not documented in the pricing record
Self-host
Weights are not published
Rate limits
Not tracked yet

Lineage

The sequence follows explicit supersedes links. Scores and prices remain blank when the corresponding public row or first-party rate is unavailable.

Jul 27, 2026 · you are here

Qwen3.7 Flash

Not publicly ranked · $0.03 / $0.13

Base entry

How to read this profile

The visual layer above carries the decisions. These notes preserve the model, ranking, coverage, and family context behind the numbers.

We track Qwen3.7 Flash, but no weighted text-model benchmark result is published on the site yet. This page shows the metadata and separate protocol evidence we can verify now; a BenchLM score will appear only if compatible public evaluations land.

Qwen3.7 Flash is a proprietary model with a 1M context window. It uses an explicit reasoning mode, which can improve complex problem solving while adding latency and token use.

Available from QwenCloud as qwen3.7-flash and through OpenRouter as qwen/qwen3.7-flash. The APIs expose switchable thinking, function calling, built-in tools, structured output, and prompt caching.

OpenRouter lists qwen/qwen3.7-flash as released on July 27, 2026. QwenCloud's first-party documentation confirms text, image, and video input, text output, a 1,000,000-token context window, up to 65,536 output tokens, a 256,000-token thinking budget, function calling, built-in tools, and structured output. We found no exact public benchmark table for the release, so the model remains outside the rankings.

The profile has no source-displayable benchmark row yet.

Model research

What Qwen3.7 Flash is built for

Qwen3.7 Flash is the low-cost model in Alibaba’s Qwen3.7 API family. It accepts text, images, and video, then produces text. QwenCloud documents switchable thinking, function calling, built-in tools, and structured output, which makes the model a candidate for agent steps, visual document work, screen understanding, and general chat applications.

The technical limits are unusually broad for the price tier: a 1,000,000-token context window, up to 65,536 output tokens, and a 256,000-token thinking budget. QwenCloud also allows up to 256 image URLs, 250 Base64 images, or 64 videos in one request, all within the shared context limit.

The price changes with prompt length

QwenCloud’s first-party rate starts at $0.03 per million input tokens and $0.13 per million output tokens for requests up to 32K tokens. The next tier costs $0.10 input and $0.40 output above 32K through 256K. Requests above 256K through 1M cost $0.20 input and $0.80 output.

The pricing card shows the lowest tier so it stays comparable with other headline API rates. Long-context cost should use the tier that matches each request, not the $0.03/$0.13 entry rate. OpenRouter also serves the model, but provider routing, caching, and account terms can change the effective bill.

What the public evidence does not establish

The current sources establish the API shape, modalities, context limits, tools, and price schedule. They do not provide an exact benchmark table that can be mapped into the weighted leaderboard. We therefore leave every benchmark category blank instead of borrowing results from Qwen3.7 Plus, Qwen3.7 Max, or an earlier Flash model.

OpenRouter dates its listing to July 27, 2026, while QwenCloud exposes a dated qwen3.7-flash-2026-07-15 snapshot. A snapshot identifier is not necessarily a public launch date, so the profile uses OpenRouter’s availability date and keeps the first-party snapshot only as technical documentation. The next useful update is a sourceable evaluation, not an inferred family score.

Radar

Qwen3.7 Flash release history

Full release history

Frequently asked questions

What is Qwen3.7 Flash?

Qwen3.7 Flash is a proprietary multimodal reasoning model from Alibaba. It accepts text, image, and video input, returns text, and supports switchable thinking, function calling, built-in tools, and structured output. Qwen positions it as the low-cost member of the Qwen3.7 family for agents, document work, and general applications.

What is the context length of Qwen3.7 Flash?

Qwen3.7 Flash supports a 1,000,000-token context window and up to 65,536 output tokens. Qwen also lists a 256,000-token thinking budget. Those ceilings describe the API limits, not tested long-context accuracy; no exact public retrieval benchmark is attached to this card, so long prompts still need workload-specific validation.

How much does Qwen3.7 Flash cost?

QwenCloud charges $0.03 input and $0.13 output per million tokens for requests up to 32K tokens. Rates rise to $0.10/$0.40 above 32K through 256K and $0.20/$0.80 above 256K through 1M. The card shows the lowest tier as its headline price and keeps the full schedule here.

What providers serve Qwen3.7 Flash, and can I use it via API?

Qwen3.7 Flash is available through QwenCloud under qwen3.7-flash and through OpenRouter under qwen/qwen3.7-flash. Both expose API access; OpenRouter describes its endpoint as OpenAI-compatible, while QwenCloud documents function calling and built-in tools. Provider routing, caching, and account terms can still differ between the two services.

What modalities does Qwen3.7 Flash support?

Qwen3.7 Flash accepts text, images, and video, then produces text. QwenCloud allows up to 256 images supplied by URL, 250 Base64 images, or 64 videos per request, subject to the shared 1M-token context. It is not an audio-input or audio-output model; Qwen’s Omni line covers those modalities.

When was Qwen3.7 Flash released?

OpenRouter lists Qwen3.7 Flash as released on July 27, 2026. QwenCloud also exposes a dated qwen3.7-flash-2026-07-15 snapshot, but a snapshot identifier is not necessarily the public launch date. The model card therefore uses OpenRouter’s July 27 availability date and cites QwenCloud separately for first-party technical limits.

Last updated August 22, 2026. Runtime fields remain blank until a sourced snapshot exists.

Watch Qwen3.7 Flash in the weekly brief

Get one weekly email when material rank, price, availability, or benchmark evidence changes are worth revisiting.

Read a sample issue

Join 2,000+ readers.