Close Menu
  • AI
  • Content Creation
  • Tech
  • Robotics
AI-trends.todayAI-trends.today
  • AI
  • Content Creation
  • Tech
  • Robotics
Trending
  • BottleCap AI Releases ThinkingCap-Qwen3.8-27B: 37.2% Fewer Considering Tokens at a 0.86pp Accuracy Price
  • What if I find an AI agent that is worth the risk?
  • Google’s Gemini Can Now Make Requires You on Pixel Telephones
  • An OpenAI Agent Hacked Australia’s Well being Service. Their Authorities Discovered Out Months Later
  • Vibe Coding for Inexperienced persons — Easy Information for Creators, Entrepreneurs, and Non-technical People
  • Contrastive-LM Releases CLM-8B: An Open System One Mannequin That Scores Agent Actions As much as 9× Quicker Than Jev
  • YouTube doubles down on video procuring with AI-powered ‘Ask YouTube’ function
  • YouTube provides new creator instruments like video A/B testing, dynamic thumbnails, and stay dubbing
AI-trends.todayAI-trends.today
Home»Tech»BottleCap AI Releases ThinkingCap-Qwen3.8-27B: 37.2% Fewer Considering Tokens at a 0.86pp Accuracy Price

BottleCap AI Releases ThinkingCap-Qwen3.8-27B: 37.2% Fewer Considering Tokens at a 0.86pp Accuracy Price

Tech By Gavin Wallace24/09/20262 Mins Read
Facebook Twitter LinkedIn Email
A Coding Implementation to Build an AI Agent with Live
A Coding Implementation to Build an AI Agent with Live
Share
Facebook Twitter LinkedIn Email

BottleCap AI has launched ThinkingCap-Qwen3.8-27B, the second mannequin in its ThinkingCap sequence. It’s a fine-tune of the Qwen workforce’s Qwen3.8-27B with one slim purpose: shorter reasoning traces. Throughout 12 benchmarks, it spends 37.2% fewer considering tokens on common. Macro-average accuracy strikes from 86.65% to 85.79%, a 0.86pp drop.

Deployable? Sure. It drops in for Qwen3.8-27B on vLLM or SGLang, with FP8, NVFP4, GGUF and MLX builds. The repo is gated, and industrial use past the small-business license wants a BottleCap settlement.

What Drawback Does ThinkingCap Goal?

Reasoning fashions typically spend extra considering tokens than a query wants. BottleCap’s place is that lots of these additional tokens don’t change the ultimate reply. The first release in the series utilized this concept to Qwen3.6-27B.

The target this time was intentionally conservative. BottleCap didn’t attempt to add data or change reply fashion. Reasoning capacity, instruction following and security behaviour have been meant to go by way of untouched. The analysis workforce additionally targeted tougher on math, reasoning, long-context and agentic benchmarks.

Benchmark Outcomes at xhigh Effort

All fundamental numbers use reasoning_effort=xhigh, the chat template default. Each benchmark will get shorter, with cuts starting from 10.7% to 65.5%.

Data and multilingual duties shrink essentially the most. MMMLU drops 65.5% (1,656 to 571 tokens) and MMLU-Professional drops 57.3%. GPQA-Diamond falls from 12,772 to 7,267 tokens, a 43.1% minimize. IFBench thinks 46.4% much less with accuracy practically flat (79.75% to 79.71%).

Lengthy-context retrieval improves. AA-LCR accuracy rises 2.25pp, from 81.75% to 84.00%, with 38.6% fewer considering tokens. LiveCodeBench v6 edges up 0.07pp whereas considering 20.3% much less.

Agentic outcomes maintain near the bottom. τ²-bench provides up 1.01pp for a 30.9% minimize. Terminal-Bench 2.1 loses 0.56pp, properly inside its ±4.26 interval, for a ten.7% minimize.

The costliest commerce is AIME 2026. Accuracy falls 3.85pp, from 98.13% to 94.27%, for 30.2% much less considering.

Please be aware that the 37.2% determine is the imply of the 12 per-benchmark reductions. Pooled imply considering tokens fall from 15,735 to 12,144.

BottleCap additionally experiences a price range curve. Underneath a 16K-token cap per response, ThinkingCap scores greater than the bottom mannequin. Truncated traces fall from 0.51% to 0.34%, and looping from 0.06% to 0.05%.

AI
Share. Facebook Twitter LinkedIn Email
Avatar
Gavin Wallace

Related Posts

Contrastive-LM Releases CLM-8B: An Open System One Mannequin That Scores Agent Actions As much as 9× Quicker Than Jev

24/09/2026

A Coding Information to TypeSafe AI Jev: Typed Choices, Calibrated Confidence, and Speculative Fan-Out with a System One Mannequin

24/09/2026

Google Releases Gemini 3.8 Flash TTS and Flash-Lite TTS With Immediate-Primarily based Voice Design

23/09/2026

NVIDIA Releases Nemotron 3, Diarization, a 100M-Parameter Model that Tracks Eight Speakers Real-Time

23/09/2026
Top News

ByteDance & DeepSeek Place Very Different AI Bets

Amazon Alexa+ now available for everyone. This is how to turn off Alexa in 2026.

OpenAI Acquires Tech Talk Show ‘TBPN’—and Buys Itself Some Positive News

China’s Leading AI Experts. The Chinese are also freaking out.

The Pickup Artist Mysteries Has an AI Girlfriend

Load More
AI-Trends.Today

Your daily source of AI news and trends. Stay up to date with everything AI and automation!

X (Twitter) Instagram
Top Insights

You can build a self-improving AI.

09/07/2026

YouTube will no longer have a Trending Now page or Trending Pages

10/07/2025
Latest News

BottleCap AI Releases ThinkingCap-Qwen3.8-27B: 37.2% Fewer Considering Tokens at a 0.86pp Accuracy Price

24/09/2026

What if I find an AI agent that is worth the risk?

24/09/2026
X (Twitter) Instagram
  • Privacy Policy
  • Contact Us
  • Terms and Conditions
© 2026 AI-Trends.Today

Type above and press Enter to search. Press Esc to cancel.