Evaluate vision-language model perception by isolating failures in visual encoders and language heads using standardized tasks and probes.
-
Updated
Jul 26, 2026 - Python
Evaluate vision-language model perception by isolating failures in visual encoders and language heads using standardized tasks and probes.
This research project investigates whether Vision-Language Models (VLMs) like CLIP/BLIP truly "understand" compositions or simply memorize patterns. It features a sophisticated diagnostic suite—including CKA, Concept Probing, and Negation Consistency—to quantify generalization gaps and visualize model attention via Grad-CAM.
Add a description, image, and links to the concept-probing topic page so that developers can more easily learn about it.
To associate your repository with the concept-probing topic, visit your repo's landing page and select "manage topics."