Direct answer
Vision capability depends on the model, endpoint, and input protocol. Confirm the target model’s endpoint in the catalog, then submit images, video, PDFs, or files through an OpenAI-compatible, Responses, Gemini-native, or provider-native structure. Image support does not imply video, PDF, OCR, tool, or every MIME type is supported.
This page was last verified on September 2, 2026.
Choose a protocol
Minimal image-understanding request
This shows only a common OpenAI Chat Completions structure:
Base64, file IDs, video, and PDF inputs can use different structures and must follow the model-specific page.
Production acceptance
- supported MIME, size, resolution, page count, duration, and file source;
- multi-image ordering, rotation, transparency, and EXIF;
- OCR, table, chart, and small-text accuracy;
- video frames, audio track, time localization, and processing state;
- errors for inaccessible URLs, corrupt files, and unsupported formats;
- cross-border processing, upstream retention, and sensitive information;
- usage, billing, and call logs.
Vision output can omit or misread content. Legal, medical, financial, identity, and other high-impact use cases require human review and must not use the model result as the sole conclusion.