Skip to solution
mediumAI Engineering

How do multimodal models handle images, audio and other non-text inputs?

872 views
01

Understand the problem

Vision encoders, audio tokens, and what multimodality means for app design.

multimodalvisionaudiofundamentals
02

Attempt it yourself

Sketch your approach before reading the solution — that's what interviews test.

Nudge consolestandby

Stuck? Beam a request up — the console returns a conceptual nudge that guides your logic without spoiling the implementation.

03

Study the solution

Multimodal models encode non-text inputs into the same token/embedding space the language model reasons over — an image is split into patches and projected into visual tokens, audio into acoustic tokens. This enables OCR-free document understanding, screenshot reasoning, and voice pipelines without separate transcripti

Solution ready — 2 min read

Classified // press E to declassify

04

Read the code

Vision request with cost control
const resized = await sharp(buf).resize({ width: 1024 }).toBuffer(); // cap tokens

await client.messages.create({
  model,
  max_tokens: 500,
  messages: [{
    role: "user",
    content: [
      { type: "image", source: { type: "base64", media_type: "image/png",
        data: resized.toString("base64") } },
      { type: "text", text: "From this invoice, extract vendor, date, total as JSON." },
    ],
  }],
});
05

Join the discussion

Discussion (0)

Sign in to join the discussion.

No responses yet. Be the first to share what you think.

Transmission complete // awaiting log

KEEP THE
STREAK ALIVE.

Dossier 24 of 80 decoded in the AI Engineering track. One more won't hurt.

Back to track