Multilook: many prompts, one context
Multilook sends one sharedcontext — an image, a video, a few reference frames, or any message sequence you would pass to /v1/chat/completions — together with up to 16 independent prompts. The context is prefilled once and reused across every prompt, and the reused tokens are billed at the reduced cache-read rate. Each prompt gets its own completions, in request order, without seeing any other prompt.
Multilook Chat Completions
Full field reference, response shape, and request limits.
Pricing
Input, output, and cached-input rates for Perceptron Mk1.
When to use it
Reach for Multilook whenever several questions share the same media:- Checklists and rubrics — run every item of an inspection checklist over one clip in a single call.
- Unrelated questions about one input — count people, describe the setting, and flag an action, without re-sending the video three times.
- Sampling several looks — set
n > 1to draw multiple completions per prompt for self-consistency or majority voting. - Comparing prompt phrasings — send the same question worded several ways and compare the answers side by side.
/v1/chat/completions calls, one Multilook request bills the shared context at the cache-read rate for every prompt after the first, and counts as a single request against its own rate-limit bucket. Compared with client-side fan-out from the Batch guide, it removes the per-call media upload and the coordination code.
Multilook is not the right tool when:
- You need
stream,response_format, orstop— none are supported on this endpoint. - Prompts depend on one another. Prompts are fully isolated: no prompt observes another prompt’s text or completions, so multi-turn conversations still belong on
/v1/chat/completions.
How it works
contextis prefilled once. It uses the same message format as/v1/chat/completions, including media parts (image_url,video_url,video_frames,image_file_id,video_file_id).- Each prompt is an implicit final
userturn that extends the context. A prompt is either a bare string or{ "content": [ ...parts ] }with the same part types as a user message. - Each prompt produces
ncompletions (1–8), sampled with the sametemperature,top_p,top_k, penalties, andvision_configfor every prompt. results[]comes back in request order. Each entry carriesprompt_indexplus eithercompletions(with per-promptusage.completion_tokens) or anerror.usageis reported once per call.prompt_tokens_details.cached_tokensis the subset ofprompt_tokensserved from the in-request prefill.
Quickstart
Per-prompt media
A prompt can carry its own content parts, including images or video. Use this when every prompt shares a reference but also brings its own input — for example, comparing a batch of parts against one approved sample:context is prefilled once and shows up in cached_tokens. The per-prompt images (part-001.jpg, part-002.jpg) sit outside the shared prefix: they are fetched, preprocessed, and prefilled per prompt, and are not reported in cached_tokens. Media you intend to reuse across prompts belongs in context.
File ids work in both places. See Working with files for uploading media once and referencing it by id.
Sampling multiple looks (n > 1)
Set n (1–8) to draw several completions per prompt. Two constraints apply:
n > 1requirestemperature > 0.temperaturedefaults to0.0, so set it explicitly.prompts × nmust not exceed 64 completions per call.
n > 1 do not re-bill prompt tokens — you pay only for the extra output. A simple majority vote over the sampled looks:
Handling partial failures
A prompt that fails returns anerror object in place of its completions. The other prompts are unaffected, and the call still returns 200 as long as at least one prompt succeeded. If every prompt fails, the whole request returns the first error’s status.
The
error.type and error.code values above are illustrative. Branch on the presence of error rather than on specific codes.- Request budget: each call has 300 seconds to finish. On a timeout, retry with fewer prompts, a lower
n, or a smallermax_completion_tokens. - Rate limits: Multilook has its own 150 requests/min bucket, separate from
/v1/chat/completions. Handle429with backoff as described in the Scaling guide.
Billing
Each call produces one usage event, billed at Perceptron Mk1’s standard rates:prompt_tokens − cached_tokens— $0.15 / M tokens (input rate)cached_tokens— $0.0375 / M tokens (cache-read rate)completion_tokens— $1.50 / M tokens (output rate)
prompt_tokens counts the shared context once per prompt, matching /v1/chat/completions accounting. cached_tokens is the part of that total served from the in-request prefill, and it approaches (P − 1) / P of prompt_tokens for P prompts as each prompt’s own text becomes small relative to the context. Reuse is scoped to the call: a subsequent identical call reports cached_tokens: 0.
Worked example. Four prompts over one ~60-frame video (~8,900 context tokens each) bill 35,612 prompt tokens, of which 25,872 are at the cache-read rate, plus 94 completion tokens:
The same four questions as four separate
/v1/chat/completions calls would bill all 35,612 prompt tokens at the input rate — about $0.0055 in total, roughly twice the cost. Figures are illustrative; actual token counts depend on your media and prompts.