Close Menu
  • AI
  • Content Creation
  • Tech
  • Robotics
AI-trends.todayAI-trends.today
  • AI
  • Content Creation
  • Tech
  • Robotics
Trending
  • What if I find an AI agent that is worth the risk?
  • Google’s Gemini Can Now Make Requires You on Pixel Telephones
  • An OpenAI Agent Hacked Australia’s Well being Service. Their Authorities Discovered Out Months Later
  • Vibe Coding for Inexperienced persons — Easy Information for Creators, Entrepreneurs, and Non-technical People
  • Contrastive-LM Releases CLM-8B: An Open System One Mannequin That Scores Agent Actions As much as 9× Quicker Than Jev
  • YouTube doubles down on video procuring with AI-powered ‘Ask YouTube’ function
  • YouTube provides new creator instruments like video A/B testing, dynamic thumbnails, and stay dubbing
  • YouTube releases new AI options for creators inside its Studio app
AI-trends.todayAI-trends.today
Home»Tech»Microsoft Releases Phi-4-Reasoning-Vision-15B: A Compact Multimodal Model for Math, Science, and GUI Understanding

Microsoft Releases Phi-4-Reasoning-Vision-15B: A Compact Multimodal Model for Math, Science, and GUI Understanding

Tech By Gavin Wallace07/03/20265 Mins Read
Facebook Twitter LinkedIn Email
NVIDIA AI Releases Llama Nemotron Nano VL: A Compact Vision-Language
NVIDIA AI Releases Llama Nemotron Nano VL: A Compact Vision-Language
Share
Facebook Twitter LinkedIn Email

Microsoft released a new version of its software Phi-4-reasoning-vision-15BThe, Multimodal open weight reasoning model with 15 billion parameters Designed for text and image tasks that demand both selective reasoning and perception. This compact model is designed to balance the requirements for training data, computation efficiency and reasoning quality. Scientific and mathematical reasoning You can also find out more about the following: understanding user interfaces.

https://arxiv.org/pdf/2603.03975

On what is the model built?

Phi-4-reasoning-vision-15B combines the Phi-4-Reasoning Language backbone is the SigLIP-2 vision encoder using a The mid-fusion architectural style. The vision encoder converts the images first into visual tokens. These tokens are then projected into the language embedding space, and the pretrained language models process them. This is an effective trade-off, as it maintains strong crossmodal reasoning and keeps training costs low compared to heavier early-fusion design.

https://arxiv.org/pdf/2603.03975

Microsoft’s decision to go smaller?

The number of parameters and the tokens used in many recent models for vision-language have increased, increasing both deployment and latency. Phi-4-reasoning-vision-15B was built as a smaller alternative that still handles common multimodal workloads without relying on extremely large training datasets or excessive inference-time token generation. The model was developed using 200 billion multimodal tokensThe building of a nation Phi-4-ReasoningIt was taught on 16 billion tokensThe final word is on the Phi-4 The base model on which the training was conducted 400 billion unique tokens. Microsoft compares this with Over 1 trillion tokens Several recent multimodal models, such as Qwen 2,5 VL, Qwen 3 VL, Kimi-VL” Gemma 3.

https://arxiv.org/pdf/2603.03975

Designing for high-resolution was the core choice

Microsoft team explained one of their most valuable technical lessons in a technical report, that multimodal reason often fails due to perception failing first. The models can be wrong not because of a lack of reasoning, but rather because they are unable to discern the visual detail in dense images, such as documents or small interfaces.

Phi-4-reasoning-vision-15B uses a Dynamic resolution visual encoder up to 3,600 tokensThis is a high-resolution tool that supports tasks such as The GUI Grounding You can also find out more about the following: Fine-grained Document Analysis. Microsoft states that High-resolution and dynamic-resolution encoders provide consistent improvementsIt is important to note that For high-quality reasoning, accurate perception is necessary.

Mixing reasoning is better than forcing it everywhere

Second, the design of the model is important. The mixed reasoning strategy and the non-reasoning approach to training. Rather than forcing chain-of-thought-style reasoning for all tasks, Microsoft team trained the model to switch between two modes. Samples of reasoning include ... Non-reasoning samples start with These tasks are primarily perception-oriented and include The VQA is a simple method of OCR and captioning.. This data is made up of reasoning. Around 20% Overall training is a mixture of all the above.

This hybrid set-up is designed to allow the model to respond more quickly on certain tasks, where a longer explanation would add latency and reduce accuracy. However, structured reasoning can still be used on other tasks like math or science. Microsoft also points out an important limitation. The boundary between the two modes are learned implicitly so switching may not be optimal. The default behavior can be overridden by users through an explicit prompt. The following are some examples of how to use tokens.

Which areas have the strongest performance?

Microsoft highlights two main areas of application. Microsoft team highlights two main application areas. Scientific and mathematical reasoning using visual inputsThe second is the quantitative document. This includes handwritten documents, such as equations and diagrams. Second is computer-use agent tasksModels that interpret content on the screen, localize GUI components, and support interaction with desktop interfaces, mobile interfaces, or both.

https://arxiv.org/pdf/2603.03975

Benchmark results

Microsoft team reports the following benchmark scores for Phi-4-reasoning-vision-15B: AI2DTEST 84.8, ChartQATEST, MathVerseMINI 44.9, MathVisionMINI, version 36.2, MathVistaMINI Version 75.2, 54.3 on MMMUVAL, 64.5 on MMStar, 76.0 on OCRBench” ScreenSpotv2: 88.2 on ScreenSpot. This report notes also that the results were obtained using Eureka ML Insights You can also find out more about the following: VLMEvalKitMicrosoft presents the results as comparisons, not leaderboard claims.

The Key Takeaways

  • Phi-4-reasoning-vision-15B is a 15B open-weight multimodal model Built by Combining Phi-4-Reasoning With the SigLIP-2 vision encoder in a The mid-fusion architectural style.
  • Microsoft creates compact model of multimodal reasoningThe focus is on Math, Science, Document Understanding, and GUI GroundingScaling to an even larger number of parameters is preferable.
  • The system is built around high-resolution vision.With support from Encoding dynamic resolution and up to 3600 visual tokensThis is useful for tasks that require a lot of screenshots, documentation, or interfaces.
  • This model combines mixed reasoning with non-reasoning.The switch allows the user to choose between You can also find out more about the following: Modes are based on the task, whether it requires explicit reasoning or a direct output based on perception.
  • Microsoft’s benchmark results show high performance for its sizeResults on ChartQATEST AI2DTEST MathVistaMINI OCRBench ScreenSpotv2This supports the model’s positioning as an efficient but compact vision-language reasoning system.

Click here to find out more Paper, Repo You can also find out more about the following: Model Weights. Also, feel free to follow us on Twitter Don’t forget about our 120k+ ML SubReddit Subscribe now our Newsletter. Wait! What? now you can join us on telegram as well.


ATH math microsoft Science
Share. Facebook Twitter LinkedIn Email
Avatar
Gavin Wallace

Related Posts

Contrastive-LM Releases CLM-8B: An Open System One Mannequin That Scores Agent Actions As much as 9× Quicker Than Jev

24/09/2026

A Coding Information to TypeSafe AI Jev: Typed Choices, Calibrated Confidence, and Speculative Fan-Out with a System One Mannequin

24/09/2026

Google Releases Gemini 3.8 Flash TTS and Flash-Lite TTS With Immediate-Primarily based Voice Design

23/09/2026

NVIDIA Releases Nemotron 3, Diarization, a 100M-Parameter Model that Tracks Eight Speakers Real-Time

23/09/2026
Top News

Gemini Spark in Action: My boyfriend was friend-zoned by Gemini after I gave it access to my life.

OpenClaw is banned by Meta and other tech companies over cyber security concerns

OpenAI loses another high-profile researcher to Meta

Before layoffs, Meta employees scramble to use up their benefits

Mark Zuckerberg is Offering AI Talent Top Paying Jobs

Load More
AI-Trends.Today

Your daily source of AI news and trends. Stay up to date with everything AI and automation!

X (Twitter) Instagram
Top Insights

AI is being used to falsely identify the federal agent who shot Renee Good

08/01/2026

Genie Envisioner: A Unified Video-Generative Platform for Scalable, Instruction-Driven Robotic Manipulation

12/08/2025
Latest News

What if I find an AI agent that is worth the risk?

24/09/2026

Google’s Gemini Can Now Make Requires You on Pixel Telephones

24/09/2026
X (Twitter) Instagram
  • Privacy Policy
  • Contact Us
  • Terms and Conditions
© 2026 AI-Trends.Today

Type above and press Enter to search. Press Esc to cancel.