Qwen
•
0
•
Released 
August 2026

Qwen
0

Frontier-class agentic coding in a dense 27B open-weight model you can self-host on a single GPU.

model ID
qwen/qwen3.8-27b:free

Output Speed *

0.00
tok/s

Intelligence Index *

0
/ 100

Context Window *

0
tokens

Input price

0
Anytoken

Output price

0
Anytoken
Qwen3.8 27B: Agentic Coding at Local Scale, Free to Deploy Qwen3.8 27B is a dense 27-billion-parameter, Apache 2.0 open-weight vision-language model from Alibaba's Qwen team. Positioned as the compact, single-machine member of the Qwen3.8 family alongside the 2.4T-parameter Qwen3.8 Max, it delivers agentic coding and computer-use performance that rivals far larger systems. What distinguishes it is the ratio: near-frontier software-engineering and tool-use results from a model that fits on a single high-end GPU. It suits teams building coding agents, IDE integrations, and long-horizon tool-calling workflows where self-hosting, privacy, and cost control matter. Start building with the Qwen3.8 27B API on AnyAPI.ai

Performance

Where a 27B Model Punches Above Its Parameter Count

Qwen3.8 27B's standout strength is agentic coding and computer use. On Artificial Analysis' Intelligence Index it scores 52 at xhigh reasoning effort—a 14-point jump over Qwen3.6 27B at identical parameter count, driven entirely by post-training. Vendor benchmarks report SWE-bench Pro 61.7 and OSWorld-Verified 84.3. This matters because it collapses the gap between self-hostable models and cloud APIs for software-engineering agents. In production, teams can run repository-level coding, tool-calling, and desktop-control workflows on their own infrastructure, though the model's high output speed and verbosity depend heavily on the reasoning-effort setting chosen.

Benchmarks

Independent and Vendor Benchmarks: What the Numbers Show

Artificial Analysis independently scored Qwen3.8 27B at 52 on its Intelligence Index at xhigh reasoning effort, above average for open-weight models of similar size, and 44 at medium, 35 with reasoning disabled. That flagship-tier result at xhigh matched much larger mixture-of-experts systems on the same index. However, Artificial Analysis measured output at roughly 47 tokens/second—notably slow—and flagged the model as highly verbose, generating far more tokens than the peer median. Vendor-reported coding scores (SWE-bench Pro 61.7, LiveCodeBench v6 90.3, Terminal-Bench 2.1 73.0) use the Claude Code harness and should be treated as approximate cross-harness comparisons.

Output Speed

*
0.00
tok/s

Intelligence Index

*
0
/ 100

MMLU *

Broad world knowledge and problem-solving
0
%

GPQA *

PhD-level scientific reasoning across physics, biology, chemistry.
0
%

HLE *

Adherence to multi-step structured instructions.
0
%

LiveCodeBench *

Tool-calling reliability in long agentic loops.
0
%

Technical Specifications

What the model supports

Qwen3.8 27B is a dense 27.78B-parameter vision-language model accepting text, image, and video and returning text. It has a native 262,144-token context window, extensible toward roughly 1M tokens with YaRN. Reasoning (thinking) mode is on by default and controllable via reasoning_effort. The most production-relevant detail: the default effort is xhigh, which produces excellent output but burns very large token counts and slows response. For coding-agent workloads, most practitioners drop effort to low or medium, which materially reduces latency and token spend with minimal quality loss.
Verified Specifications — 
0
*
Context window
0
 tokens
Maximum output tokens
0
Reasoning
No
Knowledge cutoff
August 2026
Pricing (standard)
0
 AnyTokens in
 / 
0
 AnyTokens out

Limitations & Trade-offs

Where 0 falls short

1
Overthinking by default. Qwen3.8 27B ships with reasoning_effort set to xhigh on every request. Community and independent testing show this burns roughly 60K reasoning tokens per turn even for trivial tasks, drives 40+ second time-to-first-token, and produces very verbose output. Artificial Analysis measured it generating far more tokens than the peer median. For IDE loops and high-frequency agent calls, you must explicitly lower effort to low or medium—otherwise latency and cost balloon. A model with saner defaults may be simpler for casual use.
2
Slow output speed. Independent measurement puts Qwen3.8 27B at roughly 47 tokens/second at xhigh on Alibaba's API—at the lower end for open-weight models of similar size (peer median near 90 t/s)—with time-to-first-token near 4 seconds. Turning thinking down or off improves throughput substantially in local tests, but for real-time chat or streaming UX where responsiveness is critical, a faster model or aggressively reduced reasoning effort is preferable.
3
Weaker on frontier knowledge reasoning. Vendor benchmarks show Qwen3.8 27B trailing larger frontier models on knowledge-heavy tasks such as Humanity's Last Exam (30.8 vs 40.0 for Claude Opus 4.6 Max), GPQA Diamond, and NL2Repo-Bench. Its strengths cluster in agentic coding and computer use, not the hardest multidisciplinary reasoning. For frontier math or deep scientific reasoning, a larger reasoning-focused model remains the better choice.
4
Vendor-reported benchmarks and undocumented training. Most launch coding scores are Alibaba-reported using the Claude Code harness, with several in-house or corrected benchmarks; the SWE-bench Pro comparison imports competitors' official scores rather than rerunning them. Alibaba has not published the training corpus, token count, or knowledge cutoff. Treat cross-harness comparisons as approximate and validate on your own workload before committing to production.

Best-Fit Workloads

Where this model earns its place

01

Repository-level coding agents

‍
Qwen3.8 27B's strongest evidence is in agentic software engineering: SWE-bench Pro 61.7, LiveCodeBench v6 90.3, and Terminal-Bench 2.1 73.0 (vendor-reported, Claude Code harness). It handles autonomous planning and environment feedback for end-to-end task completion, making it well suited to IDE integrations, bug-fixing loops, and multi-file changes. For agent loops, set reasoning_effort to low or medium to keep latency and token consumption controlled.

02

Computer use and desktop/mobile automation

‍
The model reports strong computer-use results—OSWorld-Verified 84.3 and AndroidWorld among its largest vendor-reported wins—reflecting reliable multi-step planning against interactive environments. This suits desktop-control agents, browser automation, and mobile-app testing workflows. Its native video and image input support screen understanding directly. Validate on your target environment, since these are vendor-reported figures under specific harnesses.

03

Self-hosted, privacy-sensitive deployments

‍
Under Apache 2.0, Qwen3.8 27B weights can be inspected, modified, and hosted behind your own controls, and 4-bit quantized builds run on a single 24GB GPU. This makes it viable to replace cloud API calls for meaningful classes of coding, document, and agent work in regulated or air-gapped environments. Full BF16 requires roughly 56GB, so plan hardware and KV-cache budget for your target context length.

04

Image-to-code and multimodal document tasks

‍
As a native vision-language model accepting text, image, and video, Qwen3.8 27B performs well on image-to-web and visual understanding tasks—it earned a top open-weight ranking on an image-to-WebDev leaderboard. This suits screenshot-to-frontend generation, UI reproduction, and document-to-structured-output extraction. Combine with structured outputs (JSON schema) for reliable downstream parsing.

Pricing in anytokens via AnyAPI
Input
0
₳
Output
0
₳
Cache write
0
₳
Cache read
0
₳

Integration

Access 0 via AnyAPI.ai

Access 0 through AnyAPI.ai using a unified API built for multi-model AI applications. Integrate 0 without maintaining a separate provider-specific connection, and keep the flexibility to test, switch, or combine models as your application requirements evolve.

01

One API integration

Access 0 and other AI models through the same API workflow instead of maintaining separate integrations for every provider.

02

Easy model switching

Test 0 against alternative models or switch models as your performance, capability, or cost requirements change without rebuilding your application around another provider API.

03

Flexible for production

Use 0 from experimentation through production while keeping your AI stack flexible as workloads, traffic, and model requirements evolve.

04

Multi-model applications

Use 0 for the workloads where it performs best and combine it with other models for tasks that require different capabilities, performance, or efficiency.

Frequently Asked Questions

Answers to common questions about integrating and using this AI model via AnyAPI.ai

Qwen3.8 27B is strongest at agentic coding and computer use. Vendor benchmarks report SWE-bench Pro 61.7, LiveCodeBench v6 90.3, and OSWorld-Verified 84.3, and Artificial Analysis independently scored it 52 on its Intelligence Index at xhigh reasoning. It's designed for repository-level coding, tool-calling agents, and multi-step automation on a single-GPU, self-hostable footprint.

Qwen3.8 27B has a native context window of 262,144 tokens, extensible toward roughly 1 million tokens using YaRN. For agentic workloads within 1M context, Qwen recommends allocating up to 262,144 tokens for internal reasoning and around 131,072 tokens for the final response.

Qwen3.8 27B defaults to xhigh reasoning_effort on every request, producing very long reasoning traces even for trivial tasks—community tests report roughly 60K reasoning tokens per turn and 40+ second time-to-first-token. Setting reasoning_effort to low or medium sharply reduces latency and token usage with minimal quality loss for most coding and agent tasks.

Yes. Qwen3.8 27B is a native vision-language model that accepts text, image, and video input and returns text. It supports function calling via tools and tool_choice, and structured outputs through a JSON schema in response_format, making it suitable for multimodal agents and pipelines requiring parseable output.

Qwen3.8 27B uses a near-identical architecture to Qwen3.6 27B—same layers, hidden size, and 262K context—but improves through post-training alone. Artificial Analysis measured a jump from 38 to 52 on its Intelligence Index at the same parameter count, with large vendor-reported gains in agentic coding, computer use, and instruction following.

* Benchmark data source: Artificial Analysis artificialanalysis.ai