Skip to main content

Direct answer

Vision capability depends on the model, endpoint, and input protocol. Confirm the target model’s endpoint in the catalog, then submit images, video, PDFs, or files through an OpenAI-compatible, Responses, Gemini-native, or provider-native structure. Image support does not imply video, PDF, OCR, tool, or every MIME type is supported. This page was last verified on September 2, 2026.

Choose a protocol

Minimal image-understanding request

This shows only a common OpenAI Chat Completions structure:
Base64, file IDs, video, and PDF inputs can use different structures and must follow the model-specific page.

Production acceptance

  • supported MIME, size, resolution, page count, duration, and file source;
  • multi-image ordering, rotation, transparency, and EXIF;
  • OCR, table, chart, and small-text accuracy;
  • video frames, audio track, time localization, and processing state;
  • errors for inaccessible URLs, corrupt files, and unsupported formats;
  • cross-border processing, upstream retention, and sensitive information;
  • usage, billing, and call logs.
Vision output can omit or misread content. Legal, medical, financial, identity, and other high-impact use cases require human review and must not use the model result as the sole conclusion.