How the vision list works
The page groups active text-generation offers that list image input and text output into conservative model identities.
Inclusion criteria
- The provider offer has explicit typed text-generation support.
- The input modality set includes image and the output modality set includes text.
Exclusions
- The list excludes image-generation-only records that do not return text.
- The list excludes catalog-only, deprecated, retired, and disallowed offers.
Data rules
- Vision meaning: Vision means that the catalog lists image as an input modality. It does not mean image generation.
- Grouping: Provider offers are grouped by the conservative model ID and database alias rules used by the LLM models directory.
Limits of the data
- Image input metadata does not measure visual accuracy, supported image size, or document quality.
- A missing modality can mean that the source does not have a confirmed value.
What vision means here
This page uses the catalog modality fields. A model is present when an active text offer accepts images and returns text. The list does not claim that every image model has the same level of visual reasoning.
Sources
- llm_db — Source for typed execution, modality, provider, context, and price data. Checked 2026-07-30.