Close Menu
  • AI
  • Content Creation
  • Tech
  • Robotics
AI-trends.todayAI-trends.today
  • AI
  • Content Creation
  • Tech
  • Robotics
Trending
  • Aikido Safety Releases Altar-1: An Open-Weight Safety Mannequin Pruned From GLM-5.3 to 328 GB
  • Black Forest Labs Releases FLUX 3 Motion: A 7B Open-Weights World Motion Mannequin That Tops RoboLab-120
  • Fastino Releases GLiNER2.5-Resolve: A 340M Open-Weight Determination Mannequin That Runs on CPU
  • BottleCap AI Releases ThinkingCap-Qwen3.8-27B: 37.2% Fewer Considering Tokens at a 0.86pp Accuracy Price
  • What if I find an AI agent that is worth the risk?
  • Google’s Gemini Can Now Make Requires You on Pixel Telephones
  • An OpenAI Agent Hacked Australia’s Well being Service. Their Authorities Discovered Out Months Later
  • Vibe Coding for Inexperienced persons — Easy Information for Creators, Entrepreneurs, and Non-technical People
AI-trends.todayAI-trends.today
Home»AI»AI Models’ Inner Thoughts are Revealed by a New Trick

AI Models’ Inner Thoughts are Revealed by a New Trick

AI By Gavin Wallace11/08/20263 Mins Read
Facebook Twitter LinkedIn Email
A United Arab Emirates Lab Announces Frontier AI Projects—and a
A United Arab Emirates Lab Announces Frontier AI Projects—and a
Share
Facebook Twitter LinkedIn Email

Recently, computer scientists have been able to The hidden data can be extracted. “thinking” The frontier AI models Perform as you work on complex problems.

The findings provide some evidence—although not conclusive proof—that certain Chinese models may have been trained by “distillingThe researchers were able to recover “reasoning information” from US models, which was hidden due to the similarity of some thinking patterns or reasoning styles. They also showed that they could recover sensitive information like API keys or passwords from the models’ inner reasoning. This vulnerability, however, has since been corrected.

“All major frontier model providers we tested share this vulnerability,” says Alexander Panfilov,⁩ a computer scientist at University of Tübingen in Germany who was involved with the work. “It can lead to personal information leakage, and it enables large-scale reasoning distillation attacks.”

Panfilov’s colleagues, from the University of Tubingen Max Planck Institute MATS Research AI Safety Institute and Snyk Security, found the same issues with Frontier models that were accessed through an application programming API (API) from OpenAI Anthropic or Google.

You can also find out more about the following: a paper The researchers, by outlining the research work show that an open-weight Chinese model or one which can be downloadable is the best. Kimi K3 from Moonshot AI produces a strikingly similar output to the hidden reasoning traces—the written-out reasoning steps involved in solving a problem—of Claude Opus 4.8 and GPT 5.6 Sol for certain prompts. The authors note that despite the similarities the work is different. “cannot causally establish distillation.” The researchers found that the reasoning of Claude Opus and two other models with open weights, DeepSeek from China, Inkling by Thinking Machines in the US, were not similar.

Moonshot AI or Z.ai had not yet responded to our request for a comment at the time this article was published.

Distillation The technique is widely known and used for copying existing model capabilities to new models.

Recently, however, distillation is a hot topic because Chinese AI firms are alleged to use it as a way of copying the best US models. OpenAI released its first version of OpenAI’s AI in February. told US lawmakers DeekSeek appeared to have used one of their models in order to create a reasoning system called R1. In June, Anthropic told lawmakers Alibaba has systematically refined its existing models to create its own Qwen.

It’s not known if Chinese AI firms used this technique specifically to extract AI models from the US. Panfilov, along with his collaborators, claim that their technique would allow for more information to be extracted from closed AI models.

Mini-Me Models

AI-based advanced models are capable of solving complex problems through a process that breaks them down into smaller parts, which can then be analyzed by artificial reasoning. “chain of thought.” The reasoning of proprietary models is often kept secret by companies to stop others using it to create new models. A user can also receive an encrypted version to reduce the amount of computation.

Researchers’ attacks are based on the fact most AI firms also offer related models in different sizes. The larger models may be more powerful, but they are also more costly to operate and to access. For certain tasks, the user may select smaller models that are weaker to reduce cost.

Panfilov’s colleagues discovered that feeding encryption traces of reasoning to a small version of the model could reveal hidden reasoning. Smaller models are more likely to be willing to share their innermost thoughts because they have less alignment training.

ai safety anthropic artificial intelligence china claude openai research
Share. Facebook Twitter LinkedIn Email
Avatar
Gavin Wallace

Related Posts

What if I find an AI agent that is worth the risk?

24/09/2026

Google’s Gemini Can Now Make Requires You on Pixel Telephones

24/09/2026

An OpenAI Agent Hacked Australia’s Well being Service. Their Authorities Discovered Out Months Later

24/09/2026

Meta VR Glasses, Ray-Ban Meta Audio, Ray-Ban Meta Gen 3: Specs, Options, Costs

24/09/2026
Top News

Google’s New Chrome ‘Auto Browse’ Agent Attempts to Roam the Web Without You

Shut These Laptops! Anthropic Places Its Claude Cowork Agent on Your Cellphone

America’s largest bitcoin miners are shifting to AI

AI-Designed drugs by a DeepMind spinoff are headed to human trials

Reid Hoffman believes doctors should consult AI to get a second opinion

Load More
AI-Trends.Today

Your daily source of AI news and trends. Stay up to date with everything AI and automation!

X (Twitter) Instagram
Top Insights

Sharing content on social media the right way

05/11/2025

UN AI Panel Invokes Precautionary Precept on Loss-of-Management Threat – Unite.AI

21/09/2026
Latest News

Aikido Safety Releases Altar-1: An Open-Weight Safety Mannequin Pruned From GLM-5.3 to 328 GB

25/09/2026

Black Forest Labs Releases FLUX 3 Motion: A 7B Open-Weights World Motion Mannequin That Tops RoboLab-120

25/09/2026
X (Twitter) Instagram
  • Privacy Policy
  • Contact Us
  • Terms and Conditions
© 2026 AI-Trends.Today

Type above and press Enter to search. Press Esc to cancel.