Liquid AI Points Small Multimodal d1 Models at Edge Devices

A Hugging Face post highlights compact LiquidAI image-text models, including d1-omni-600M, plus a demo Space with camera, drawing and text inputs.

October 7, 2026

A Hugging Face blog post titled “Multimodal open d1 decision models for the edge” points to three LiquidAI models. LFM2.5-Encoder-350M is a fill-mask model of about 0.4B parameters. LFM2.5-VL-3B is a 3B image-text-to-text model. d1-omni-600M is a 0.6B image-text-to-text model that was updated two days before the page was captured.

The post also links a Space called System One Arcade, which runs on an L4 GPU and offers camera, drawing and text demos of d1-3B. It lists two earlier posts from the same author on faster inference: one on accelerating vision-language models with LFM2.5-VL-DSpark (September 24, 2026), and one on up to 3.2x faster inference with LFM2.5-DSpark (August 20, 2026). The captured text contains little of the article body, so it gives no technical details on how the d1 models work or how they perform.

Why it matters

  • The models are small, from 0.4B to 3B parameters, and the post frames them for the edge. That suggests multimodal AI aimed at lighter hardware rather than only large servers.
  • The interactive demo takes camera, drawing and text input, so people can try the d1 models in a browser without setting anything up.
  • The linked posts on faster inference suggest speed is a focus for this model family. The captured text gives no benchmarks for the d1 models themselves, so those claims need checking against the full article.

Source

Summary written by FoxaMind with AI assistance from the source above. Check the original for full details.