AI glossary Models & Architecture
What is VLM?
VLM (Vision-Language Model)
← All glossary termsModel that jointly understands images and text for captioning, OCR, or visual Q&A.
Explained
Vision-language models fuse encoders for pixels and tokens, enabling document AI, visual inspection, and multimodal assistants. They extend RAG to charts, slides, and scans.