BottleCap AI has launched ThinkingCap-Qwen3.8-27B, the second mannequin in its ThinkingCap sequence. It’s a fine-tune of the Qwen workforce’s Qwen3.8-27B with one slim purpose: shorter reasoning traces. Throughout 12 benchmarks, it spends 37.2% fewer considering tokens on common. Macro-average accuracy strikes from 86.65% to 85.79%, a 0.86pp drop.
Deployable? Sure. It drops in for Qwen3.8-27B on vLLM or SGLang, with FP8, NVFP4, GGUF and MLX builds. The repo is gated, and industrial use past the small-business license wants a BottleCap settlement.
What Drawback Does ThinkingCap Goal?
Reasoning fashions typically spend extra considering tokens than a query wants. BottleCap’s place is that lots of these additional tokens don’t change the ultimate reply. The first release in the series utilized this concept to Qwen3.6-27B.
The target this time was intentionally conservative. BottleCap didn’t attempt to add data or change reply fashion. Reasoning capacity, instruction following and security behaviour have been meant to go by way of untouched. The analysis workforce additionally targeted tougher on math, reasoning, long-context and agentic benchmarks.
Benchmark Outcomes at xhigh Effort
All fundamental numbers use reasoning_effort=xhigh, the chat template default. Each benchmark will get shorter, with cuts starting from 10.7% to 65.5%.
Data and multilingual duties shrink essentially the most. MMMLU drops 65.5% (1,656 to 571 tokens) and MMLU-Professional drops 57.3%. GPQA-Diamond falls from 12,772 to 7,267 tokens, a 43.1% minimize. IFBench thinks 46.4% much less with accuracy practically flat (79.75% to 79.71%).
Lengthy-context retrieval improves. AA-LCR accuracy rises 2.25pp, from 81.75% to 84.00%, with 38.6% fewer considering tokens. LiveCodeBench v6 edges up 0.07pp whereas considering 20.3% much less.
Agentic outcomes maintain near the bottom. τ²-bench provides up 1.01pp for a 30.9% minimize. Terminal-Bench 2.1 loses 0.56pp, properly inside its ±4.26 interval, for a ten.7% minimize.
The costliest commerce is AIME 2026. Accuracy falls 3.85pp, from 98.13% to 94.27%, for 30.2% much less considering.
Please be aware that the 37.2% determine is the imply of the 12 per-benchmark reductions. Pooled imply considering tokens fall from 15,735 to 12,144.
BottleCap additionally experiences a price range curve. Underneath a 16K-token cap per response, ThinkingCap scores greater than the bottom mannequin. Truncated traces fall from 0.51% to 0.34%, and looping from 0.06% to 0.05%.

