Close Menu
  • AI
  • Content Creation
  • Tech
  • Robotics
AI-trends.todayAI-trends.today
  • AI
  • Content Creation
  • Tech
  • Robotics
Trending
  • Spanberger Indicators Knowledge Heart Accountability Order, Creates AI Job Pressure – Unite.AI
  • GGUF vs GPTQ vs AWQ vs EXL2: LLM Mannequin Codecs Defined (2026)
  • Shoppers Sue Anthropic, OpenAI, SpaceXAI and Google Over Alleged AI Pact – Unite.AI
  • Ben Bernstein, Supervisor of Cybersecurity Advisors at Huntress – Interview Sequence – Unite.AI
  • Jina AI Releases jina-ocr-v1: A 3.4B MoE Doc Parser With Constructed-In Speculative Decoding for Low-Finances GPUs
  • Anthropic Faucets Accenture’s School for Embedded AI Mannequin Analysis – Unite.AI
  • Right here’s How an AI Slowdown May Really Be Enforced
  • PrismML Releases Ternary Bonnesai 2, 27B. A 5.9 GB Apache 2.2 Model that Retains 98.2% Qwen3.8 Performance 27B.
AI-trends.todayAI-trends.today
Home»Tech»ICM: An Unsupervised, Label-Free Training Framework for LLMs

ICM: An Unsupervised, Label-Free Training Framework for LLMs

Tech By Gavin Wallace14/06/20254 Mins Read
Facebook Twitter LinkedIn Email
NVIDIA Releases Llama Nemotron Nano 4B: An Efficient Open Reasoning
NVIDIA Releases Llama Nemotron Nano 4B: An Efficient Open Reasoning
Share
Facebook Twitter LinkedIn Email

Human supervision is required to define desired behaviors in post-training methods of language models that have been pre-trained. This can be done through either demonstrations or feedback on preferences. This approach is limited by the complexity of tasks and behaviors. The human supervisor is not reliable in such scenarios, as LMs can mimic errors in demos and exploit feedback system flaws. In order to overcome this challenge, LMs must be trained for tasks which are beyond the human ability in terms of reliability. Recently, research has revealed a variety of failure modes including the reward-hacking or human-designed signals for supervision and real humans.

LLM post-training: Limitations in Human Supervision

Researchers have investigated several ways to increase scale without human supervision. Standard methods include high-quality rewards that can be verified, like matching outputs of models with real-world solutions. While there is evidence to suggest that the pre-trained models are capable of performing downstream tasks with minimal post-training, the challenge remains in eliciting latent knowledge. Contrast Consistent Search is a method of unsupervised elicitation that relies on logical consistency in order to discover latent knowledge. CCS, however, underperforms the supervised approach and is often unable to identify latent knowledge as other prominent features satisfy consistency properties.

Introduction to Internal Coherence Maximization

Internal Coherence Maximization is a method that researchers from Anthropic have developed. It allows them to fine tune pre-trained model on labels they generate themselves, without any labels provided. ICM finds label sets that satisfy the requirements of both the pre-trained and logically consistent model. ICM’s simulated-annealing search algorithm approximates the maximum goal, since optimal label set recognition is computationally impossible. The method also matches training with golden labels for TruthfulQA or GSM8K. It even outperforms the crowdsourced labels used by Alpaca.

What is the ICM Algorithm?

ICM follows a three-step iterative process. (a) The system selects a sample of a newly unlabeled instance from the dataset to be considered for inclusion. (b) It determines an optimal label while also resolving logical inconsistencies. (c) Finally, the algorithm decides whether or not it accepts this newly labeled case based on a scoring function. ICM’s performance is assessed across three datasets, including TruthfulQA, which assesses truthfulness, GSM8K, which verifies mathematical accuracy, and Alpaca, which measures helpfulness and harmlessness. The researchers tested four baselines, including Zero-shot, Golden Label and Human Label. The Experiments also used Llama 3.1, 8B, 70B models as well as two proprietary models: Claude 3.5 Haiku, and Claude 3.4 Haiku.

Benchmark performance and model comparisons

ICM is more accurate than humans in superhuman capabilities elicitation. It matches the golden standard of supervision at 80%. Researchers successfully trained a chatbot assistant without the need for human supervision using ICM reward models. Unsupervised reward models achieve 75.0% accuracy in RewardBench compared with 72.2% when trained by humans using production data. Using both unsupervised and human supervised RMs, two policies were trained using RL in order to produce helpful, innocent, and honest Assistants. A policy that is trained using the unsupervised RM has a win rate of 60%. These policies are still behind the previously released Claude 3.5 Haiku which has a 92% success rate.

The Future Outlook

In this paper, we introduce Internal Coherence Maximization as an advance in unsupervised LM to fine-tune pre-trained models for self-generated labels. It consistently outperforms human supervision, including crowdsourced, in GSM8K Verification, TruthfulQA Tasks, and Alpaca Reward Modeling. ICM has limitations, including a dependency on the salience of concepts within pre-trained model and an inability to handle long inputs because context window restrictions. ICM is a promising alternative to traditional RLHF as LMs progress beyond the human evaluation capability. It ensures model alignment with intent, without any human supervision.


Click here to find out more Paper. The researchers are the sole owners of all credit. Also, feel free to follow us on Twitter Don’t forget about our 100k+ ML SubReddit Subscribe Now our Newsletter.


Sajjad is in his final year of undergraduate studies at IIT Kharagpur. Tech enthusiast Sajjad is interested in the applications of AI, with an emphasis on their impact and real-world implications. He strives to make complex AI ideas clear and understandable.

AI
Share. Facebook Twitter LinkedIn Email
Avatar
Gavin Wallace

Related Posts

GGUF vs GPTQ vs AWQ vs EXL2: LLM Mannequin Codecs Defined (2026)

19/09/2026

Jina AI Releases jina-ocr-v1: A 3.4B MoE Doc Parser With Constructed-In Speculative Decoding for Low-Finances GPUs

18/09/2026

PrismML Releases Ternary Bonnesai 2, 27B. A 5.9 GB Apache 2.2 Model that Retains 98.2% Qwen3.8 Performance 27B.

18/09/2026

Alibaba Qwen Releases Qwen3.8-Omni-Flash: A 1M-Context Omni-Modal Mannequin Constructed Round Agentic Audio-Video Understanding and Software Use

18/09/2026
Top News

Trump and Energy Industry are Eager to Use Fossil Energy for AI

Who are the real losers in the Musk v. Altman Trial?

Google DeepMind Workers vote for unionization over AI military deals

Senators Want to Know how Much Energy Data Centers Use

Gear News of the Weekly: Apple’s AI-powered wearable, as well as a smartphone that runs Android, Linux or Windows

Load More
AI-Trends.Today

Your daily source of AI news and trends. Stay up to date with everything AI and automation!

X (Twitter) Instagram
Top Insights

‘It’s a Modern-Day Draft’: Why Stanford Students Walked Out on Sundar Pichai’s Commencement Speech

21/07/2026

Napster Is Again, and It Desires to Digitally Clone Academics

18/09/2026
Latest News

Spanberger Indicators Knowledge Heart Accountability Order, Creates AI Job Pressure – Unite.AI

19/09/2026

GGUF vs GPTQ vs AWQ vs EXL2: LLM Mannequin Codecs Defined (2026)

19/09/2026
X (Twitter) Instagram
  • Privacy Policy
  • Contact Us
  • Terms and Conditions
© 2026 AI-Trends.Today

Type above and press Enter to search. Press Esc to cancel.