Close Menu
  • AI
  • Content Creation
  • Tech
  • Robotics
AI-trends.todayAI-trends.today
  • AI
  • Content Creation
  • Tech
  • Robotics
Trending
  • Nokia Open-Sources AnyJev: A Coaching-Free Layer That Turns Any Open LLM Right into a Calibrated Resolution Mannequin
  • Find out how to Use AI With Your Privateness Intact
  • OpenAI Releases GPT-6 Sol and Luna: 50% Cheaper API Pricing and Benchmarks
  • A New Chatbot Needs to Unlock the Secrets and techniques in Tattered Historical Greek Information
  • SpeakON Ships a MagSafe AI Voice Button With Its Personal Microphone: Turning Your Voice into Polished Communication, and Motion throughout Apps
  • Patti Harrison Had Goals of a Tech Utopia. Silicon Valley Smashed Them
  • A New Instrument Discovered Malware That’s Guided by an AI Hive Thoughts—No People in Sight
  • Tips on how to Declare Your Reduce of Apple’s $250 Million Siri Settlement
AI-trends.todayAI-trends.today
Home»Tech»Liquid AI Open Sources: Pipette is a reproducible benchmarking tool that can measure on-device model, quantity, runtime and hardware all together.

Liquid AI Open Sources: Pipette is a reproducible benchmarking tool that can measure on-device model, quantity, runtime and hardware all together.

Tech By Gavin Wallace26/08/20265 Mins Read
Facebook Twitter LinkedIn Email
A Coding Implementation to Build an Interactive Transcript and PDF
A Coding Implementation to Build an Interactive Transcript and PDF
Share
Facebook Twitter LinkedIn Email

Model cards describe the performance of the device under full-precision, server-class conditions. The numbers don’t always predict the behavior of a model on a smartphone. Liquid AI launched this week. Pipette. This is a platform that was built by partnering with. Artificial Analysis Validator of independent methodologies. Pipette considers on-device behaviors as properties of the deployed system and not the isolated model. Its measurement unit is a complete configuration. model + quantization + runtime + device. The launch dataset covers five on-device performance metrics across more than 1,000 model × quantization × runtime × device × context configurations, spanning 30+ models, llama.cpp builds for macOS, iOS, Windows and Android, and context lengths from 256 to 8,192 tokens. The MacBook Pro M5 Max with an iPhone 17 Pro, a Galaxy S26 Ultra and the iPhone 17 Pro are used to verify initial results. Tests can be done to verify the practical claim: Two 350M models on the same device, at the same quantumization retain 78.4% of the decoded throughput.

Is it deployable?

No,, Pipette Apache 2.0 is shipped as a standard infrastructure.pipette-mgmt, pipette-clients, pipette-scores), a public results dataset, a hosted dashboardNative Americans and. iOS The following are some examples of how to get started: Android benchmark apps. There is no waitlist. Publication of the results from community submissions is still in beta.

  • Which Companies?Any team that ships a model on hardware they do not own. The dashboards and apps can be used by solo developers or seed stage startups without any infrastructure. The mid-market product team can deploy the client across their internal devices. The whole process can be controlled by the large chip manufacturers, OEMs and enterprises.
  • Industries: Consumer electronics and smartphone OEMs, automotive, industrial and robotics, healthcare devices, financial services, defense — anywhere latency, privacy or connectivity forces inference onto the device.
  • AppsModel and quantization choice before committing to a sprint; validation of the SoCs and hardware; regression tests when an OS, driver or runtime is updated; context-length planning for capacity; independent confirmation of performance claims from vendors.

Liquid AI is shipped

Liquid AI releases Pipette as a partnership with Artificial AnalysisAn independent validater reviewed and confirmed the methodology. The assumption is simple and effective: On-device behaviour is a property not only of the model, but also of the deployed system.

The launch dataset covers five on-device performance metrics across more than 1,000 model × quantization × runtime × device × context configurations. This dataset includes 30+ models and multiple quantization formats. It also contains llama.cpp build for macOS iOS Windows Android and Android. And contexts ranging from 256 tokens to 8,192. The initial results are from MacBook Pros with the M5 Max and an iPhone 17 Pro, and Galaxy S26 Ultras. AMD Ryzen AI max+ 395, and Radeon 8060S, results will be published soon.

Pipette’s unit of measurement in a deployment is configuration. model + quantization + runtime + device. The benchmark defines the token and metric, producing latency, memory or throughput results. Separate quality metrics are tracked. IFBench, GPQA Diamond The following are some examples of how to get started: MATH-500. Those quality scores currently come from llama.cpp evaluation runs on NVIDIA H100 80GB reference systems, then get matched to on-device runs sharing the same model and quantization — a quality number shown next to phone throughput was not produced on the phone.

The context of deployment changes the answer

Four comparisons published show the impact a configuration has on a decision.:

  • The context scaling may diverge even when the parameter count is identical. The Q4_K_M parameter on Galaxy S26 Ultra. Granite-4.0-H-350M The decoder retains approximately 78.4% its throughput when 256 input tokens are used. Granite-4.0-350M Retains only 33.8%
  • Memory is not bought by sparse activation, but rather speed. With 2,048 inputs tokens, LFM2.5-8B-A1B Decodes up to 2.4x faster Qwen3.5-4B The 2.6x speedier than Ministral-3-3B-Instruct-2512. The memory usage is still 5.29 GiB, even though it activates 1,5B out of the 8.5B parameters.
  • Quality and speed are not synonymous. On iPhone 17 Pro at Q4_K_M, MiniCPM5-1B The workload of 2,048 in / 256 out is completed in just 3.47 seconds, compared to 4.12 seconds when using LFM2.5-1.2B-InstructA 15.8% decrease in time. LFM scored 9.0 more points on the MATH 500 for these same artifacts.
  • Nearly identical profiles of a system can conceal task-level reversals. The Q4_K_M system profile has 2,048 tokens for M5 Max. Granite-4.1-8B The following are some examples of how to get started: Ministral-3-8B-Instruct-2512 There is a difference of 2.4% between decode throughput (throughput) and peak RAM (1.2%). Granite is ahead of IFBench, while Ministral has a 14.0 point lead over GPQA Diamond.

What are the measurement methods?

Follow the a for performance runs published methodologyFive measured repetitions, readiness gate, and fixed-token shapes. Prior to each repeated timed test, the platform is checked for thermal conditions and loads. The evaluations are conducted using a protocol that is model blind and deterministic. Pipette scores never know the generation source. Each submission includes information on the benchmark version, token artifacts, model settings and versions, device OS and hardware, as well as runtime version.

Interactive explainer

AI ar Benchmark ces ETH hardware open source war
Share. Facebook Twitter LinkedIn Email
Avatar
Gavin Wallace

Related Posts

Nokia Open-Sources AnyJev: A Coaching-Free Layer That Turns Any Open LLM Right into a Calibrated Resolution Mannequin

23/09/2026

OpenAI Releases GPT-6 Sol and Luna: 50% Cheaper API Pricing and Benchmarks

23/09/2026

SpeakON Ships a MagSafe AI Voice Button With Its Personal Microphone: Turning Your Voice into Polished Communication, and Motion throughout Apps

23/09/2026

Anthropic releases Claude Opus: Performance of Fable 5.1 at 40% less running costs than Opus 5

22/09/2026
Top News

There is Only One AI Company. Blob Welcome!

Truth Social’s AI chatbot, Donald Trump’s Media Diet Incarnate is the new AI chatbot from Truth Social

OpenAI’s Chief Communication Officer is Leaving the Company

SpaceX Listed Grok’s ‘Spicy’ Mode as a Risk in Its IPO Filing

The Meta AI App Lets You ‘Discover’ People’s Bizarrely Personal Chats

Load More
AI-Trends.Today

Your daily source of AI news and trends. Stay up to date with everything AI and automation!

X (Twitter) Instagram
Top Insights

A new audio foundation model from Liquid AI, LFM2-Audio-1.50B with response times of under 100 milliseconds.

01/10/2025

Cerebras Runs OpenAI’s GPT-5.6 Sol at 750 Tokens Per Second in New Ultrafast Tier – Unite.AI

16/08/2026
Latest News

Nokia Open-Sources AnyJev: A Coaching-Free Layer That Turns Any Open LLM Right into a Calibrated Resolution Mannequin

23/09/2026

Find out how to Use AI With Your Privateness Intact

23/09/2026
X (Twitter) Instagram
  • Privacy Policy
  • Contact Us
  • Terms and Conditions
© 2026 AI-Trends.Today

Type above and press Enter to search. Press Esc to cancel.