What launched

The Allen Institute for AI put out Molmo 2 this week, an open model that watches video the way the previous Molmo looked at still images. Three variants shipped: an 8B model built on Qwen-3 that Ai2 calls its best overall for video grounding and Q&A, a 4B for cheaper deployments, and a 7B built on Ai2's own Olmo. They handle single images, multiple images, and video clips, and they can point at things in a frame, track objects, and answer questions about what happened.

The numbers, read carefully

Ai2 says the Molmo 2 models beat Gemini 3 Pro and other open-weight rivals on video tracking benchmarks, and that the 8B leads all open-weight models on image and multi-image reasoning. It is honest about the ceiling: on video grounding, nothing on the market scores above 40% accuracy yet, open or closed. Larger proprietary models still lead on human preference ratings. So the honest read is narrower than "open beats closed." It is that on the specific, hard task of following objects through video, a small open model now trades punches with the biggest names.

Why it is worth your attention

Molmo 2 is not a video generator. Veo and Sora make footage; Molmo 2 watches it. That is exactly why it matters for the video workflow: the missing piece in most AI video pipelines is analysis, not generation. An open model that can ground objects and answer questions about clips means creators and enterprises can build search, moderation, and editing tools on top of video without renting a frontier API. Open weights also mean it runs where the footage already lives, instead of uploading everything to someone else's cloud.

Sources

  1. [1] VentureBeat — “Ai2’s Molmo 2 shows open source models can rival proprietary giants in video”Read source