Tiny multimodal vision-language model — 1.8B params, runs on laptops.
Tiny multimodal — laptop-class image understanding.
Thirty seconds keeps the numbers honest.
Thanks for the ping — we'll take a look.
Check your connection and try again, or email [email protected].