Skip to main content
Perceptron Mk1.5 accepts text, images, video, and audio. It generates text, spatial and temporal annotations, structured responses, and function calls for your application to execute.

Specifications

Audio shares the context window with the rest of the input and the output budget. The audio limit is specific to the current Mk1.5 model; a request can reach the context limit before reaching the per-item audio limit. Video soundtracks require vision_config.enable_audio_in_video: true. Responses contain text and structured content; the model does not generate audio. See Audio for input forms and limits.

Pricing

Multilook reuses a shared context within one request. Cached tokens are a subset of prompt tokens; do not add them to the prompt-token total a second time. Use the tokenization guide to budget media, reasoning, and conversation history, and the scaling guide to manage throughput.

Choose the output for your task

track is a container in generated markup, not an annotation_format value. The optional asset_idx attribute identifies a media occurrence among the assets available at that point in the conversation; it is not a persistent file identifier. If no explicit or inherited selector is present, it defaults to the last asset available when the annotation was produced. See the asset resolution rules.

Integration notes

  • Use n: 1 for chat completions. Multilook has separate sampling controls.
  • Inspect finish_reason before using an answer. length means the output may be incomplete; tool calls require a caller-managed round trip.
  • Use reasoning_effort for new reasoning integrations.
  • Function tools cannot be declared in the same request as JSON Schema or regex constrained output. Request a constrained final answer after the tool loop.