Specifications
Audio shares the context window with the rest of the input and the output budget. The audio limit is specific to the current Mk1.5 model; a request can reach the context limit before reaching the per-item audio limit. Video soundtracks require
vision_config.enable_audio_in_video: true. Responses contain text and structured content; the model does not generate audio. See Audio for input forms and limits.
Pricing
Multilook reuses a shared context within one request. Cached tokens are a subset of prompt tokens; do not add them to the prompt-token total a second time.
Use the tokenization guide to budget media, reasoning, and conversation history, and the scaling guide to manage throughput.
Choose the output for your task
track is a container in generated markup, not an annotation_format value. The optional asset_idx attribute identifies a media occurrence among the assets available at that point in the conversation; it is not a persistent file identifier. If no explicit or inherited selector is present, it defaults to the last asset available when the annotation was produced. See the asset resolution rules.
Integration notes
- Use
n: 1for chat completions. Multilook has separate sampling controls. - Inspect
finish_reasonbefore using an answer.lengthmeans the output may be incomplete; tool calls require a caller-managed round trip. - Use
reasoning_effortfor new reasoning integrations. - Function tools cannot be declared in the same request as JSON Schema or regex constrained output. Request a constrained final answer after the tool loop.