This model makes calibrated, typed decisions about an image plus optional text. It answers choice, score and noul (yes/no probability) questions in one forward pass, with no text generation. It adds image input to Laya by replacing Laya's ModernBERT encoder with SmolVLM-256M-Instruct.
Source: [Hacker News](https://huggingface.co/thaitea/laya-vision-smolvlm-256m)