Perceptron released Mk1.5 on September 25, adding native audio and video tracking to its perception model. This September 28 assessment looks at what those capabilities mean for an application that must identify objects and follow them through a recording. Perceptron describes outputs including points, boxes, polygons and object tracks, alongside text. Release announcement
A useful distinction is whether a system can connect observations over time. A box around a package in one frame answers where it appears. A track can help an application follow that package through later frames, where another object may overlap it or the camera may move. That makes identity changes and missing observations part of the evaluation, alongside detection accuracy.
Tracking creates a different evaluation problem
Perceptron positions Mk1.5 for embodied applications and describes timestamped geometry across video. Those are vendor claims, not evidence that a particular robot or inspection workflow will perform reliably. We have not independently tested the model. Perceptron’s launch description
For a recorded inspection workflow, start with clips whose relevant objects, positions and event times have been labeled by a person. Include occlusion, poor lighting and objects leaving the frame. Check whether the same identity survives those transitions and whether timestamps align with the recording. A fluent summary can look convincing even when a track jumps between similar objects.
The application also needs a policy for uncertain observations. A model output should not silently become a physical action. For an initial evaluation, keep outputs in a review queue and measure how often a person must correct them before considering automated downstream steps.
Budget audio, video and output together
The current model documentation lists a 36,864-token context window and an 8,192-token maximum output. It accepts text, images, video and audio and returns text. Standard prices per million tokens are $0.15 input, $1.50 output and $0.0375 cached input. Mk1.5 model documentation
Audio shares the context budget with other inputs and the response. Video soundtracks require vision_config.enable_audio_in_video: true; supplying video alone does not establish that its audio was processed. The documentation also says function tools cannot be declared in the same request as JSON Schema or regex constrained output. It recommends producing a constrained final answer after the tool loop. Input and output constraints
These details affect how to divide a recording into requests. Test enough overlap to retain an object’s identity at clip boundaries, then measure the extra input cost. Keep space for the returned tracks instead of filling the entire context with media.
Start with the hosted integration
The quickstart exposes Mk1.5 through a Chat Completions API using the identifier perceptron-mk1.5. API quickstart
The hosted API is the concrete starting point documented here. Custom infrastructure discussions do not establish that downloadable weights or a self-hosting license are available. Confirm those terms separately if local deployment is a requirement.
For teams comparing visual models, keep dated inputs and outputs. Our GPT-6 Sol and Luna update describes a recent image-processing fix that warrants rerunning earlier evaluations. The same discipline applies here: compare verified tracks on representative recordings, with cost and correction effort recorded together.
