TOP 5 Free AI Models That Understand and Read Video

37 views
Skip to first unread message

Sergei Grichine

unread,
Sep 2, 2026, 12:02:49 PM (9 days ago) Sep 2
to hbrob...@googlegroups.com
I came across an interesting list of models, just want to share it. No personal experience with any of them.

I'd guess that if one could feed a camera stream to a local model and read its output through a "perception adapter," the perceived objects, distances, etc. could feed into behaviors, scripts, and other high-level logic.

Meanwhile, this is what the Overlords say about it: https://chatgpt.com/s/t_6a9846d23174819193276d90079510ec  - note the remark about "temporal perception"

TOP 5 Free AI Models That Understand and Read Video.

1⃣Qwen3.8-Flash-Next
The absolute leader among open models, outperforming Claude Opus 4.6 on a number of benchmarks.
125 billion parameters, but activates only 6 billion per token, providing high efficiency. Supports video, images, and text, with a native context window of 262,144 tokens (expandable to 1 million). 
HuggingFace

2⃣VideoChat3-4B
The best fully open model for video, outperforming larger open-source alternatives.
A 4-billion-parameter model that understands video over time—from subtle movements to stories lasting up to half an hour.
HuggingFace

3⃣LLaVA-OneVision-2-8B
The next generation of one of the most influential open multimodal AI models.
An 8-billion-parameter model from Glint Lab, built on the Qwen3-8B language foundation. It combines understanding of images, long videos, and spatial scenes.
HuggingFace

4⃣JoyAI-VL-Interaction
A new model outperforming Qwen3-VL-8B-Instruct on 26 benchmarks.
The model achieves an average score of 57.53% across 26 standard benchmarks, which is 3.37% higher than Qwen3-VL-8B. It is available in quantized form (INT4), making it easier to run.
HuggingFace

5⃣SmolVLM2-2.2B
The most efficient model for resource-constrained environments, running even on phones.
A model from Hugging Face with 2.2 billion parameters, designed to operate under severe memory constraints. It requires only 5.2 GB of VRAM and runs on an RTX 3060, MacBook Pro M2, and even the free Google Colab T4.
HuggingFace

Best Regards,
-- Sergei

Charles de Montaigu

unread,
Sep 8, 2026, 8:44:23 AM (3 days ago) Sep 8
to hbrob...@googlegroups.com
Thanks for Sharing Sergei

Google is playing hard ball in MultiModal Hidden Space Embedding : Gemma  should get an honorable mention on that list.

Qwen is barnstorming its open weight Qwen3.8-Flash-Next

 Chaz

--
You received this message because you are subscribed to the Google Groups "HomeBrew Robotics Club" group.
To unsubscribe from this group and stop receiving emails from it, send an email to hbrobotics+...@googlegroups.com.
To view this discussion visit https://groups.google.com/d/msgid/hbrobotics/CA%2BKVXVPQvpVmRp2X%3DGreMxfghYqxQCB%2BQNCKdciRm9Ss0yvZMA%40mail.gmail.com.

core...@coreyfro.com

unread,
Sep 8, 2026, 11:44:22 PM (3 days ago) Sep 8
to HomeBrew Robotics Club
Qwen 3.8 Flash Next is not ready for people to use.  Two reasons:
  • It's a tech demo for features coming out in Qwen 4.0, and these features aren't fully baked.
  • Even where the model is fine, the software packages that interface with it (Llama.cpp, vLLM, etc) do not have the necessary features to take advantage of these new advances.
Use Qwen 3.8-27B. It's the gold standard for local LLMs at the moment, and the software around it is mature.

Don't get me wrong, Qwen 3.8 Flash Next is doing lots of amazing things. I just want to make sure that people have a fruitful experience with local LLMs.



Charles de Montaigu

unread,
Sep 9, 2026, 2:41:24 PM (2 days ago) Sep 9
to hbrob...@googlegroups.com
Things are moving Fast ! Yes its leading Edge Qwen 4 preview but there is vLLM Recipe and Docker image available for local.

For production it is Qwen Cloud for now and it's accessible on OpenRouter.

HuggingFace Card with some Benchmark vs 27B

Obviously changing LLM in your harness/control plane will impact your solution behaviors






--
You received this message because you are subscribed to the Google Groups "HomeBrew Robotics Club" group.
To unsubscribe from this group and stop receiving emails from it, send an email to hbrobotics+...@googlegroups.com.
Reply all
Reply to author
Forward
0 new messages