Skip to main content
Use a reference image to show Perceptron Mk1.5 what to look for in a video. The request combines the example, an explanation of which features matter, and a query video. This can help with product matching, finding a familiar object, or reviewing footage for a visually defined condition. The example guides the current request. Keep it in the supplied conversation when asking follow-up questions that depend on it.

Find a reference product on a shelf

This example supplies an image of a cereal package, then asks whether a matching package is visible in a shelf video. It asks for temporal evidence and permits an uncertain answer when packaging details cannot be read clearly. Visibility in this recording does not establish a store’s current inventory. Install the perceptron>=0.4.0 Python package and set PERCEPTRON_API_KEY before running:
The reference image is asset 0; the query video is asset 1. Temporal annotations must therefore select asset 1. This is an illustrative matching interval, not measured output for the sample video:
Validate the selected asset before using a timestamp in a player. If you add more references or retain earlier media in the conversation, update the selectors according to multiple-asset ordering.

Make the comparison precise

  • State whether you want the same class, the same product design, or the same individual object. Those are different matching tasks.
  • Use a reference where the distinguishing features are visible. If it contains several objects, identify the intended one with a reviewed box, as in image in-context learning.
  • For several reference concepts, label each one and request separate evidence for each. Do not assume every reference has a match in the query video.
  • Treat a missed or obscured match as uncertainty about the supplied footage. Selected frames do not show every instant of the source video.

Choose the output you need

Request clip annotations when the result should seek to supporting evidence. Request video tracks when you need the matching object’s position over time; the track belongs to the video asset, not the reference image. If a track has an uncertain interval, request new model observations around the gap before continuing local tracking. Keep crop coordinates and source timestamps associated with their own assets when combining the results.