Household-object VLM benchmark

Compare submitted classification, captioning, and open-vocabulary detection outputs across model families.

Hosted demo: select a benchmark example and compare the exact saved outputs from the submitted experiments. The repository also contains app.py, the full Gradio live-inference implementation for arbitrary uploaded images.

Selected benchmark example
Models to compare (select one or more)
Tip: select multiple models to compare their submitted outputs for the same benchmark image.
Model outputs
ModelTaskOutputDetails
Select an example and run a comparison.