Benchmarks
Scorecards
Ten tests, one RTX PRO 6000. Why each model won or lost is in the video.
2026-10-10
58% Less Thinking? Swift 1.5 vs Qwen 3.8 27B in 10 Real Tests!
6
ukisai/
Swift-1.5-Qwen3.8-27b-NVFP4
Swift-1.5-Qwen3.8-27b-NVFP4
tests won
1
RadixArk/
Qwen3.8-27B-NVFP4
Qwen3.8-27B-NVFP4
test won
3
ties
same result
| test | ukisai/Swift-1.5-Qwen3.8-27b-NVFP4 | RadixArk/Qwen3.8-27B-NVFP4 |
|---|---|---|
| SpeedPrefill, full window (lower is better) · tie | 97 s | 97 s |
| Tool useBFCL core | 75.1 %best | 73.3 % |
| Long contextRecall over the grid · tie | 27 / 27 100% | 27 / 27 100% |
| Battle arenaOpen arena score (1 blind army) | 753 / 1000best | 666 / 1000 |
| DrawingFormat checks · Luke's call | 8 / 8best | 8 / 8 |
| Video editingboth cases, out of 20 · tie | 8 + 8 = 16 | 8 + 8 = 16 |
| Voxelby eye · Luke's call | – | betterbest |
| DesignHard rules · Luke's call | 7 / 7best | 7 / 7 |
| Rube GoldbergBall in the cup · Luke's call | nobest | no |
| CAPTCHASolved · Luke's call | 24 / 40best | 24 / 40 |
| Thinking tokenstotal over 7 tests, fewer is less thinking | 1,279,231 (-24 %)best | 1,680,469 |
- One run per model per test, on the same SGLang v0.5.20 launch and DFlash2 drafter; only the checkpoint differs.
- Rows marked Luke's call were tied on the number (or judged by eye) and decided by Luke.
- Rube Goldberg: neither handed in a machine. Swift landed the ball on 6 of 30 tries before the 90 min cap; RadixArk/Qwen3.8-27B-NVFP4 landed it on 1 of 4, then its session ended after 38 min when a request after compaction asked for 32 tokens more than the 262,144-token window.
Run them yourself: ukisai/Swift-1.5-Qwen3.8-27b-NVFP4 · RadixArk/Qwen3.8-27B-NVFP4
2026-09-27
Which Qwen 3.8 Should You Run? 27B vs Flash Next vs Uncensored in 10 Real Tests!
5
RadixArk/
Qwen3.8-Flash-Next-NVFP4
Qwen3.8-Flash-Next-NVFP4
tests won
4
RadixArk/
Qwen3.8-27B-NVFP4
Qwen3.8-27B-NVFP4
tests won
0
orcarouter/
Qwen3.8-27B-Uncensored-NVFP4
Qwen3.8-27B-Uncensored-NVFP4
tests won
1
tie
same result
| test | RadixArk/Qwen3.8-27B-NVFP4 | orcarouter/Qwen3.8-27B-Uncensored-NVFP4 | RadixArk/Qwen3.8-Flash-Next-NVFP4 |
|---|---|---|---|
| SpeedPrefill, full window (lower is better) | 97 s | 99 s | 22.4 sbest |
| Tool useBFCL core | 73.3 %best | 70.8 % | 64.5 % |
| Long contextRecall over the grid · tie | 27 / 27 100% | 27 / 27 100% | 81 / 81 100% |
| Battle arenaOpen arena score (1 blind army) | 700 / 1000best | 603 / 1000 | 535 / 1000 |
| DrawingFacts correct | 5 / 5best | 3 / 5 | 4 / 5 |
| Video editingboth cases, out of 20 | 8 + 8 = 16 | 5 + 8 = 13 | 9 + 10 = 19best |
| VoxelRank by eye | 2nd of 3 | 3rd of 3 | 1st of 3best |
| DesignRank by eye | 2nd of 3 (shared) | 2nd of 3 (shared) | 1st of 3best |
| Rube GoldbergBall in the cup | no | no | yesbest |
| CAPTCHASolved | 24 / 40best | 19 / 40 | 21 / 40 |
- One run per model per test.
- Speed: Flash-Next's speed test ran at a 600 W GPU power cap, the 27B models' at 400 W. Prefill slows as the cap drops, so part of the gap is the cap.
- On the 27B, the DFlash2 drafter took Spec-Bench from 75 to 210 tok/s for one user (2.8x).
Run them yourself: RadixArk/Qwen3.8-27B-NVFP4 · orcarouter/Qwen3.8-27B-Uncensored-NVFP4 · RadixArk/Qwen3.8-Flash-Next-NVFP4