PaliGemma: A lightweight open vision-language model (VLM)
Google's new PaliGemma vision-language model combines SigLIP image encoding with Gemma text processing for tasks like image captioning and visual question answering. Released in May 2024, it offers open-source pretrained and fine-tuned checkpoints in various resolutions, enabling powerful multimodal AI capabilities for researchers and developers.