Qwen3.8-27B vs Opus 4.7 Bakeoff
I used Claude Opus through Claude Code when it was the frontier model, and I remember it being far better than what I get now from Qwen3.8-27B running locally through OpenCode. Not on any single answer. Over a session, the local model seems to lose the thread: it forgets what it was doing, repeats work, stops for no reason. Opus did not.
Qwen’s model card says otherwise. The card for Qwen3.8-27B has a column labelled “Opus 4.6 Max”. On SWE-bench Pro the 27B scores 61.7 to Opus’s 53.4. On QwenSWEBench it scores 79.0 to 63.8. On Terminal-Bench 2.1 it trails, 73.0 to 78.2. The card’s own footnote says the Opus SWE-bench figure is Anthropic’s published number, while Qwen ran its model through the Claude Code agent at temperature 1.0 with 256k context. Not the same run, and every Qwen number is full precision on datacentre serving.

The text benchmark table from the Qwen3.8-27B model card. The right-hand column is the Opus 4.6 Max comparison.
What I run is the UD-Q4_K_M GGUF on one RTX 5090 through llama.cpp. The card cannot tell me what that copy of the model does inside a coding agent on my machine with a clock running, so I set that up: the quantised model, inside three different agents, on three scored tasks, against Opus 4.7 through Claude Code. It answers the capability half of my complaint for bounded tasks; the losing-the-thread half is a different test.
The scores matched. The clock did not.
The setup
The machine is an RTX 5090 with 32 GB of VRAM, an i9-13900K and 96 GB of DDR5. llama.cpp build 10524 serves the model behind llama-swap with full GPU offload, a single 262,144-token slot, native multi-token-prediction speculative decoding and a q8_0 KV cache.
| Model | Precision | Agent | Runs on |
|---|---|---|---|
| Claude Opus 4.7 | not disclosed | Claude Code 2.1.263 | Anthropic |
| Qwen3.8-27B | UD-Q4_K_M | OpenCode 1.18.29 | my RTX 5090 |
| Qwen3.8-27B | UD-Q4_K_M | Pi 0.73.1, via its SDK | my RTX 5090 |
| Qwen3.8-27B | UD-Q4_K_M | DeerFlow 2.1.0, two configurations | my RTX 5090 |
| Qwen3.8-27B | FP8, provider-declared | OpenCode 1.18.29 | OpenRouter, Reka |
| DeepSeek V4 Pro 0813 | FP8, provider-declared | OpenCode 1.18.29 | OpenRouter, CoreWeave |
| Qwen3.6-27B | Q4_K_M | OpenCode 1.18.29 | my RTX 5090 |
Every local agent ran inside a Docker container built for the test: a read-only root filesystem, 8 GB of RAM, two CPUs, a fresh home directory with none of my configuration, memories or MCP servers, and outbound networking dropped except for loopback and one port on the host. That port belongs to a logging proxy on the Mac. The proxy records every request and response, checks that the model, output cap, provider, quantisation and data-retention flags are what the run expects, and only then forwards to llama.cpp on the 5090 or to OpenRouter. Claude Code cannot run in that container, so it ran on the Mac in its own sandbox: safe mode, restricted, no session persistence, the home directory unreadable, no MCP servers, skills or browser, and an empty network allowlist for the commands it runs. Each attempt got a fresh workspace holding the task fixtures and README, snapshotted on every change, and an independent grader scored the final files afterwards.
One Mac drives everything: Claude Code in its own sandbox, and a fresh Docker container per attempt for each local agent, all through the logging proxy.
Claude Code ran with --model claude-opus-4-7 --effort high --permission-mode dontAsk. An earlier run in auto permission mode quietly made side calls to Sonnet 5 for classification, so I rejected it and kept the log. Budgets were 30 minutes for the app task and 45 minutes for each repository task, with no coaching from grader feedback. Each agent kept its own prompts and tool set, because that is part of what was being measured.
One agent-specific trap: Pi’s CLI caps output at 32,000 tokens and has no flag to raise it. The stock-CLI attempt spent its entire first response and never made a tool call. The scored Pi run went through Pi’s SDK with a 128,000-token ceiling, which the proxy verified on every request.
Part way through, llama.cpp crashed under agent load and I disabled CUDA graphs on the server. The Qwen3.8 app run through OpenCode and the app and queue runs through Pi happened before that change; every other local run happened after it. The crashed attempts were kept and their replacements labelled. Nothing was substituted silently.
Three tasks
| Task | Work | Scoring |
|---|---|---|
| Orbital laboratory | Build an interactive planets app in HTML with required behaviour, responsive layout and accessibility checks | 70 points, independent browser grader |
| Queue concurrency repair | Diagnose and fix cross-process leasing, crash-recovery and acknowledgement bugs in an existing repository; add regression tests | 85 automated, 15 review |
| Clock-change incident | Investigate code, logs, a database snapshot, configuration and git history; fix scheduling across a daylight-saving transition; add tests and an incident report | 70 automated, 30 review |
These are my tasks with frozen prompts and rubrics, not SWE-bench. The scores describe these requirements and nothing else.
Results
| Configuration | App /70 | Queue /100 | Incident /100 | Combined time | vs Opus |
|---|---|---|---|---|---|
| Opus 4.7, Claude Code | 70 | 100 | 99 | 14m41s | 1.0x |
| Qwen3.6-27B Q4, OpenCode, local | 51 | 100 | 99 | 16m51s | 1.1x |
| DeepSeek V4 Pro FP8, OpenCode, cloud | 66* | 100 | 99 | 19m43s | 1.3x |
| Qwen3.8-27B Q4, DeerFlow revised, local | 70 | 100 | 99 | 24m06s | 1.6x |
| Qwen3.8-27B Q4, OpenCode, local | 70 | 100 | 99 | 29m25s | 2.0x |
| Qwen3.8-27B Q4, DeerFlow first config, local | 63* | 93* | 82.5* | 30m36s | 2.1x |
| Qwen3.8-27B Q4, Pi, local | 70 | 100 | 100 | 42m05s | 2.9x |
| Qwen3.8-27B FP8, OpenCode, OpenRouter | 70* | 100 | 86* | 90m10s | 6.1x |
Times cover model responses and tool execution and exclude setup, grading and review. Combined time is not a quality ranking: a run that stops early looks fast.
The starred cells need a word each. DeepSeek lost four app points because touch input added a comet and an ordinary desktop click did not. Cloud Qwen earned every automated app point but hit the 30-minute limit while debugging malformed trajectory trails, and its incident submission did not meet the completion gate. The first DeerFlow configuration was stopped by its own limits on two of the three tasks, which I come back to below.
Three local Qwen3.8 configurations matched or beat Opus 4.7’s 70, 100 and 99 in the score column; Pi took the one extra incident review point. In the time column Opus finished in 14 minutes 41 seconds, and the same-score local runs took 24, 29 and 42 minutes.
What the agents built
The app task is the one with an artefact worth looking at. Every configuration had to build an interactive orbital laboratory from the same spec: bodies on a canvas, working physics, mouse and touch input, a responsive layout, accessibility checks. The grader drove each finished app in a browser at 1440 by 900 and at 375 by 812 and kept screenshots.

The eight apps at 1440 by 900 as the grader saw them, in the order of the results table.
Eight readings of one spec. Every app has a sun and a planet on a dark canvas, play, reset and add-comet controls, gravity and time-scale sliders and an energy-drift readout. Opus put the controls in a narrow right-hand panel with a shortcuts list. The others split between a top bar and a side panel, and some drew orbit rings. Nothing in the screenshots separates the runs that scored 70, which is the point: the spec was met eight ways, and the grader could only see whether it was met.
The finished apps are hosted as the agents left them, each opening in a new tab. Space plays and pauses, R resets, and a click or tap on the canvas adds a comet.
- Opus 4.7, Claude Code
- Qwen3.6-27B Q4, OpenCode
- DeepSeek V4 Pro FP8, OpenCode
- Qwen3.8-27B Q4, DeerFlow revised
- Qwen3.8-27B Q4, OpenCode
- Qwen3.8-27B Q4, DeerFlow first config
- Qwen3.8-27B Q4, Pi
- Qwen3.8-27B FP8, OpenCode via OpenRouter
The queue fix and the incident investigation produce a diff and a written report rather than a screen, so I have not embedded them. The incident reports are where judgement shows, and a blind re-read of those is the follow-up I want most.
Generating faster, finishing slower
The local model generated tokens faster than Opus and still finished later. The app task shows it most clearly.
| Configuration | Wall time | Model turns | Tool calls | Output tokens | Output t/s over API time |
|---|---|---|---|---|---|
| Opus 4.7, Claude Code | 7m16s | 30 | 29 | 33,597 | 78.5 |
| Qwen3.8 Q4, OpenCode | 9m02s | 25 | 26 | 54,477 | 103.2 |
| Qwen3.8 Q4, Pi | 11m14s | 36 | 36 | 64,848 | 103.8 |
| DeepSeek V4 Pro, OpenRouter | 5m22s | 32 | 33 | 32,312 | 106.4 |
| Qwen3.8 FP8, OpenRouter | 30m02s | 24 | 25 | 67,992 | 41.3 |
Local Qwen generated about 30 percent faster than Opus and finished 1m46s later through OpenCode and 3m58s later through Pi. It wrote 1.6 to 1.9 times as many tokens to reach the same result. Pi’s first response alone was 26,356 tokens before its first tool call. On the incident task the spread was wider: Opus 2m59s, revised DeerFlow 11m40s, OpenCode 15m45s, Pi 20m53s.
Task time is prompt processing plus reasoning plus output volume plus tool execution plus the number of attempts before something works. Tokens per second is one term in that sum. The timing boundaries also differ, since local timings exclude model load while Claude Code’s include the CLI round trip, so the direction holds and the precision does not. Stravica saw the same direction in August with the 27B at NVFP4 on a DGX Spark and single-call tasks: Opus 1.4 to 3.6 times faster per call.
The FP8 endpoint
I added Qwen3.8-27B at FP8 through OpenRouter to test whether my local quantisation was what held the model back. FP8 stores each weight as an eight-bit floating-point number; the local Q4_K_M file stores them at about four and a half bits. If the missing bits were the problem, the higher-precision weights should have shown it.
They did not, or at least this endpoint could not show it. The FP8 run was the slowest configuration by a wide margin, timed out on the app task and failed to complete the incident task, at 41.3 output tokens per second. Provider serving, sampling and stochastic decisions were not controlled, so this says nothing about Q4 against FP8 in general. It does say that paying for this endpoint bought no improvement over the local Q4 on these three tasks, at $0.39 for the app task alone.
Where the test stopped discriminating
Qwen3.6-27B, one generation older, scored 100 and 99 on the two repository tasks and 51 on the app. The repository tasks catch failures; they do not separate systems that can already pass them. Several passing submissions carried caveats the frozen rubric did not deduct: revised DeerFlow added a seven-day catch-up policy nobody asked for and an unsupported claim about tenant reactivation, and Pi left its verification scripts and screenshots in the submission. Each condition ran once, with no matched seeds, and serving conditions changed part way through when graphs were disabled. Nothing here tests days of accumulated context, large repositories or repeated compaction; the largest context any local run reached was 94,334 tokens.
DeerFlow deserves one paragraph because its first result would have been easy to misread. It scored 63, 93 and 82.5 with the same checkpoint the other agents used. The traces showed why: an internal graph-step limit of 1,000 stopped the app task after 71 model responses and 77 tool calls, well inside the time budget, and a repeated-tool-call guard stopped the incident task, partly because DeerFlow requires a re-read after every write and then counted those re-reads as repetition. With the guard off and the step limit raised, same checkpoint, same prompts, same graders, it scored 70, 100 and 99 in 24 minutes. Those were fresh attempts, so the configuration change is not proven to be the whole cause, but a low score from an agent needs its trace read before it is believed.
What I take from it
The card’s claim survives quantisation and a consumer GPU in the score column: on these three tasks, inside three different agents, a 27B at Q4 produced work the grader could not tell apart from Opus 4.7’s. The gap is wall-clock. Opus 4.7 through Claude Code reached the same outputs in 34 to 61 percent of the time.
That splits the decision cleanly for me. Structured implementation, code checks, extraction and anything with an objective test go local. Ambiguous judgement, causal diagnosis, long sessions and recovery from tool failures go to the frontier model. Matching a rubric is not the same as matching judgement, and this test was not built to measure the second.