Jul 27, 2026 · you are here
Qwen3.7 FlashNot publicly ranked · $0.03 / $0.13
Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.
Start free briefReleased Jul 27, 2026 — see all recent releases
Data as of August 22, 2026 · How the score is built
Each documented value carries its source. Missing fields stay visible as not sourced or not published, rather than disappearing from the page.
The sequence follows explicit supersedes links. Scores and prices remain blank when the corresponding public row or first-party rate is unavailable.
Jul 27, 2026 · you are here
Qwen3.7 FlashNot publicly ranked · $0.03 / $0.13
Base entry
The visual layer above carries the decisions. These notes preserve the model, ranking, coverage, and family context behind the numbers.
We track Qwen3.7 Flash, but no weighted text-model benchmark result is published on the site yet. This page shows the metadata and separate protocol evidence we can verify now; a BenchLM score will appear only if compatible public evaluations land.
Qwen3.7 Flash is a proprietary model with a 1M context window. It uses an explicit reasoning mode, which can improve complex problem solving while adding latency and token use.
Available from QwenCloud as qwen3.7-flash and through OpenRouter as qwen/qwen3.7-flash. The APIs expose switchable thinking, function calling, built-in tools, structured output, and prompt caching.
OpenRouter lists qwen/qwen3.7-flash as released on July 27, 2026. QwenCloud's first-party documentation confirms text, image, and video input, text output, a 1,000,000-token context window, up to 65,536 output tokens, a 256,000-token thinking budget, function calling, built-in tools, and structured output. We found no exact public benchmark table for the release, so the model remains outside the rankings.
The profile has no source-displayable benchmark row yet.
Qwen3.7 Flash is the low-cost model in Alibaba’s Qwen3.7 API family. It accepts text, images, and video, then produces text. QwenCloud documents switchable thinking, function calling, built-in tools, and structured output, which makes the model a candidate for agent steps, visual document work, screen understanding, and general chat applications.
The technical limits are unusually broad for the price tier: a 1,000,000-token context window, up to 65,536 output tokens, and a 256,000-token thinking budget. QwenCloud also allows up to 256 image URLs, 250 Base64 images, or 64 videos in one request, all within the shared context limit.
QwenCloud’s first-party rate starts at $0.03 per million input tokens and $0.13 per million output tokens for requests up to 32K tokens. The next tier costs $0.10 input and $0.40 output above 32K through 256K. Requests above 256K through 1M cost $0.20 input and $0.80 output.
The pricing card shows the lowest tier so it stays comparable with other headline API rates. Long-context cost should use the tier that matches each request, not the $0.03/$0.13 entry rate. OpenRouter also serves the model, but provider routing, caching, and account terms can change the effective bill.
The current sources establish the API shape, modalities, context limits, tools, and price schedule. They do not provide an exact benchmark table that can be mapped into the weighted leaderboard. We therefore leave every benchmark category blank instead of borrowing results from Qwen3.7 Plus, Qwen3.7 Max, or an earlier Flash model.
OpenRouter dates its listing to July 27, 2026, while QwenCloud exposes a dated qwen3.7-flash-2026-07-15 snapshot. A snapshot identifier is not necessarily a public launch date, so the profile uses OpenRouter’s availability date and keeps the first-party snapshot only as technical documentation. The next useful update is a sourceable evaluation, not an inferred family score.
Alibaba · Model release
Qwen3.7 Flash is a proprietary multimodal reasoning model from Alibaba. It accepts text, image, and video input, returns text, and supports switchable thinking, function calling, built-in tools, and structured output. Qwen positions it as the low-cost member of the Qwen3.7 family for agents, document work, and general applications.
Qwen3.7 Flash supports a 1,000,000-token context window and up to 65,536 output tokens. Qwen also lists a 256,000-token thinking budget. Those ceilings describe the API limits, not tested long-context accuracy; no exact public retrieval benchmark is attached to this card, so long prompts still need workload-specific validation.
QwenCloud charges $0.03 input and $0.13 output per million tokens for requests up to 32K tokens. Rates rise to $0.10/$0.40 above 32K through 256K and $0.20/$0.80 above 256K through 1M. The card shows the lowest tier as its headline price and keeps the full schedule here.
Qwen3.7 Flash is available through QwenCloud under qwen3.7-flash and through OpenRouter under qwen/qwen3.7-flash. Both expose API access; OpenRouter describes its endpoint as OpenAI-compatible, while QwenCloud documents function calling and built-in tools. Provider routing, caching, and account terms can still differ between the two services.
Qwen3.7 Flash accepts text, images, and video, then produces text. QwenCloud allows up to 256 images supplied by URL, 250 Base64 images, or 64 videos per request, subject to the shared 1M-token context. It is not an audio-input or audio-output model; Qwen’s Omni line covers those modalities.
OpenRouter lists Qwen3.7 Flash as released on July 27, 2026. QwenCloud also exposes a dated qwen3.7-flash-2026-07-15 snapshot, but a snapshot identifier is not necessarily the public launch date. The model card therefore uses OpenRouter’s July 27 availability date and cites QwenCloud separately for first-party technical limits.
Related resources
Last updated August 22, 2026. Runtime fields remain blank until a sourced snapshot exists.
Get one weekly email when material rank, price, availability, or benchmark evidence changes are worth revisiting.
Read a sample issueJoin 2,000+ readers.