Choose the caption style
Installperceptron>=0.4.0 and set PERCEPTRON_API_KEY. This example sends the same public street image in two independent requests, one for each style.
concise and detailed are labels in this example application.
Ground the caption in regions
Request boxes for the subjects mentioned in the description. Run this after the setup above:mention describes the cited region. Boxes use two normalized 0–1000 coordinates: top-left, then bottom-right. Convert using the dimensions of the asset selected by asset_idx; do not treat the numbers as source pixels.
Use the shared rendering guide to turn response.txt into an overlay. The annotation reference also covers points, polygons, and collections. For captions comparing several images, request explicit selectors and follow multiple-asset ordering.
Write captions for their destination
For metadata that must be parsed, define a JSON Schema with structured outputs. Asking for JSON in a caption prompt alone does not enforce a schema. Use image Q&A when you need an answer to a specific question rather than an overall description.