> ## Documentation Index
> Fetch the complete documentation index at: https://docs.perceptron.inc/llms.txt
> Use this file to discover all available pages before exploring further.

# Perceptron Mk1.5

> Image, video, and audio understanding with visual grounding, reasoning, and function calling.

Perceptron Mk1.5 accepts text, images, video, and audio. It generates text, spatial and temporal annotations, structured responses, and function calls for your application to execute.

## Specifications

| Property | Value |
| - | - |
| Model ID | `perceptron-mk1.5` |
| Context window | 36,864 tokens |
| Maximum output | 8,192 tokens |
| Input modalities | Text, images, video, audio |
| Audio formats | WAV, MP3, FLAC |
| Audio limit | 16,384 audio tokens per item; about 21.8 minutes at approximately 750 tokens per minute |
| Reasoning | Supported; configure with `reasoning_effort` |
| Function calling | Supported on chat completions |
| Constrained responses | JSON Schema and regex |
| Multilook | Supported for independent prompts over shared context |

Audio shares the context window with the rest of the input and the output budget. The audio limit is specific to the current Mk1.5 model; a request can reach the context limit before reaching the per-item audio limit. Video soundtracks require `vision_config.enable_audio_in_video: true`. Responses contain text and structured content; the model does not generate audio. See [Audio](/perceptron-mk1.5/capabilities/audio) for input forms and limits.

## Pricing

| Token category | Price per million tokens |
| - | - |
| Input | \$0.15 |
| Output | \$1.50 |
| Cached input | \$0.0375 |

[Multilook](/perceptron-mk1.5/guides/multilook) reuses a shared context within one request. Cached tokens are a subset of prompt tokens; do not add them to the prompt-token total a second time.

Use the [tokenization guide](/perceptron-mk1.5/guides/tokenization) to budget media, reasoning, and conversation history, and the [scaling guide](/perceptron-mk1.5/guides/scaling) to manage throughput.

## Choose the output for your task

| Task | API controls | Guide |
| - | - | - |
| Describe an image or answer a question | Natural-language prompt | [Image understanding](/perceptron-mk1.5/capabilities/image-understanding) |
| Summarize a recording or describe sounds | `input_audio`, `audio_url`, or `audio_file_id` | [Audio understanding](/perceptron-mk1.5/capabilities/audio) |
| Ask questions about a recording | Audio input and a question | [Audio Q\&A](/perceptron-mk1.5/capabilities/audio-qa) |
| Transcribe speech | Audio input and a transcription instruction | [Audio transcription](/perceptron-mk1.5/capabilities/audio-transcription) |
| Locate audible events or speech in time | Audio input and `vision_config.annotation_format: "clip"` | [Audio clipping](/perceptron-mk1.5/capabilities/audio-clipping) |
| Combine video frames with their soundtrack | `vision_config.enable_audio_in_video: true` | [Video soundtracks](/perceptron-mk1.5/capabilities/video-understanding#analyze-video-soundtracks) |
| Locate objects | `vision_config.annotation_format`: `point`, `box`, or `polygon` | [Annotation format](/perceptron-mk1.5/concepts/annotations) |
| Locate events in video | `vision_config.annotation_format`: `clip` | [Video understanding](/perceptron-mk1.5/capabilities/video-understanding) |
| Track an object | Spatial annotation format plus a prompt requesting `<track>` | [Video tracking](/perceptron-mk1.5/capabilities/video-tracking) |
| Return a constrained final answer | `response_format` with JSON Schema, or `regex` | [Structured outputs](/perceptron-mk1.5/capabilities/structured-outputs) |
| Ask your application to run a function | `tools` and `tool_choice: "auto"` | [Tool calling](/perceptron-mk1.5/capabilities/tool-calling) |

`track` is a container in generated markup, not an `annotation_format` value. The optional `asset_idx` attribute identifies a media occurrence among the assets available at that point in the conversation; it is not a persistent file identifier. If no explicit or inherited selector is present, it defaults to the last asset available when the annotation was produced. See the [asset resolution rules](/perceptron-mk1.5/guides/multiple-assets#resolve-selectors-before-drawing).

## Integration notes

* Use `n: 1` for chat completions. Multilook has separate sampling controls.
* Inspect `finish_reason` before using an answer. `length` means the output may be incomplete; tool calls require a caller-managed round trip.
* Use [`reasoning_effort`](/perceptron-mk1.5/capabilities/thinking) for new reasoning integrations.
* Function tools cannot be declared in the same request as JSON Schema or regex constrained output. Request a constrained final answer after the tool loop.
