Skip to content

Perceptron Mk1.5 and Mk1

Perceptron models are vision-language models that answer from visual evidence. Mk1.5 understands video, images and audio together; it can tell you in which interval of a video an event happens, mark objects with coordinates, and follow them through a video. Mk1 is the previous generation; it accepts images and video, and supports neither audio nor tool calling. Both are called through /v1/chat/completions.

Model Context Input / Output ($/1M)
perceptron/perceptron-mk1.5 36,864 0.15 / 1.50
perceptron/perceptron-mk1 32,768 0.15 / 1.50

On both models a single response is at most 8,192 tokens; a request with a larger max_tokens is refused with 400.

Terminal window
curl "$LLMTR_BASE_URL/v1/chat/completions" \
-H "Authorization: Bearer llmtr-your_key" \
-H "Content-Type: application/json" \
-d '{
"model": "perceptron/perceptron-mk1.5",
"messages": [
{
"role": "user",
"content": [
{ "type": "video_url", "video_url": { "url": "https://example.com/match.mp4" } },
{ "type": "text", "text": "What happens in this video, in order?" }
]
}
],
"max_tokens": 500
}'

Send the video as a video_url part. The address can be a public URL starting with https:// or a base64 data URL (data:video/mp4;base64,...). Send large files by https address rather than base64.

The model samples frames across the whole video; the longer the recording, the sparser the frames, so a brief event can be missed. Video can fill much of the 36,864-token context window. Send only the relevant part of long recordings and track actual use in the response's usage.prompt_tokens field.

To ask when an event happens, set vision_config.annotation_format to "clip":

{
"model": "perceptron/perceptron-mk1.5",
"messages": [
{
"role": "user",
"content": [
{ "type": "video_url", "video_url": { "url": "https://example.com/surf.mp4" } },
{ "type": "text", "text": "Find the intervals where the surfer is riding a wave." }
]
}
],
"vision_config": { "annotation_format": "clip" }
}

Intervals come back as markup inside the response text:

<clip mention="the surfer is riding a wave" t="0 seconds 7.9 seconds" />

t can carry a single moment ("3.2 seconds") or a start and end ("1.0 seconds 3.2 seconds").

By default Mk1.5 processes only the picture. To have it take speech and sounds into account as well, send vision_config.enable_audio_in_video as true:

"vision_config": { "enable_audio_in_video": true }

The soundtrack adds roughly 750 tokens per minute. These tokens are counted in usage.prompt_tokens and billed at the input rate. Mk1 does not process a video's soundtrack; enable_audio_in_video: true returns 400 on Mk1.

Choosing the frames yourself: video_frames

Section titled “Choosing the frames yourself: video_frames”

If you split the video into frames yourself, you can send the frames with timestamps as one video (Mk1.5 only):

{
"type": "video_frames",
"video_frames": {
"frames": [
{ "image_url": { "url": "https://example.com/frame-000.jpg" }, "timestamp_ms": 0 },
{ "image_url": { "url": "https://example.com/frame-001.jpg" }, "timestamp_ms": 500 },
{ "image_url": { "url": "https://example.com/frame-002.jpg" }, "timestamp_ms": 1000 }
]
}
}
  • A group carries at least 2 frames. A request can carry at most 256 media items, earlier messages included; every frame and every image, video and audio recording counts as one.
  • timestamp_ms values must not decrease.
  • Every frame address follows the same rules as image_url: a public https address or a data:image/... base64 URL.
  • A frame group has no soundtrack.

Too few frames can miss a brief event; more frames increase the cost.

Send the image as an image_url part. To ask where objects are, set vision_config.annotation_format to "point", "box" or "polygon" and say in the prompt what to mark:

{
"model": "perceptron/perceptron-mk1.5",
"messages": [
{
"role": "user",
"content": [
{ "type": "image_url", "image_url": { "url": "https://example.com/shelf.jpg" } },
{ "type": "text", "text": "Draw a box around each package on the shelf." }
]
}
],
"vision_config": { "annotation_format": "box" }
}

Annotations come back inside the response text. Coordinates are scaled to 0 to 1000 regardless of the image's size:

<point mention="button center"> (120,200) </point>
<point_box mention="package"> (100,150) (300,350) </point_box>
<polygon mention="surface"> (20,40) (60,40) (60,80) (20,80) </polygon>

To have an object followed through a video, set annotation_format to "box" and ask for tracking in the prompt. The answer can come back as <track> markup containing timestamped boxes; waypoints can be sparse, and a box for every frame is not guaranteed.

Mk1.5 also accepts an audio recording on its own. Use either form:

{ "type": "input_audio", "input_audio": { "data": "<base64>", "format": "wav" } }
{ "type": "audio_url", "audio_url": { "url": "https://example.com/recording.mp3" } }
  • Supported formats are WAV, MP3 and FLAC. Give wav, mp3 or flac for input_audio.format; use audio/wav, audio/mpeg or audio/flac in audio_url data URLs.
  • Audio takes roughly 750 tokens per minute, and the limit per recording is 16,384 audio tokens (about 21 minutes). Split longer recordings.

Mk1 does not accept audio.

Mk1.5 answers without reasoning unless you ask for it. Turn reasoning on with the reasoning_effort field; the valid values are none, minimal, low, medium and high. xhigh and max return 400 on this model.

perceptron/perceptron-mk1.5 -> no reasoning
perceptron/perceptron-mk1.5:low -> low effort
perceptron/perceptron-mk1.5:high -> high effort

Reasoning tokens are counted in completion_tokens and billed at the output rate; they also use up the max_tokens budget. When you use a high effort, keep max_tokens well above the answer you expect. See reasoning effort for the general behaviour.

On Mk1, reasoning is switched on or off rather than set by level: the :think / :fast suffix, or "reasoning": true / false.

Mk1.5 supports function calling. Only "auto" (the default) and "none" can be used for tool_choice; "required" or a named function choice returns 400. If the model must not call a function in a turn, leave tools out of that request.

Terminal window
curl "$LLMTR_BASE_URL/v1/chat/completions" \
-H "Authorization: Bearer llmtr-your_key" \
-H "Content-Type: application/json" \
-d '{
"model": "perceptron/perceptron-mk1.5",
"messages": [{ "role": "user", "content": "What is the weather in Ankara?" }],
"tools": [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Gets the current weather for a city",
"parameters": {
"type": "object",
"properties": { "city": { "type": "string" } },
"required": ["city"]
}
}
}
]
}'

Mk1 does not support tool calling; a request with tools returns 400.

Mk1.5 produces constrained output in two ways:

  • json_schema in response_format: the answer is JSON that matches your schema.
  • A top-level regex field: the answer is text that matches the expression, for example "regex": "[0-9]{4}".

response_format: { "type": "json_object" } is not supported and returns 400; give a json_schema instead. tools cannot be combined with json_schema or regex in the same request: finish the tool loop first, then get the constrained final answer in a separate request without tools.

  • n can only be 1.
  • Files uploaded to LLMTR (input_file) cannot be sent to these models. Send images as image_url, video as video_url, and audio as input_audio or audio_url.
  • When streaming, send stream_options: { "include_usage": true } to get usage in the last chunk.